BackendCUDA
Decode160 tok/s
VRAM24 GB
Memory bandwidth1008 GB/s
ModelQwen 7B
Parameters7B
QuantizationQ4
Context8K
KV quantQ4
OffloadFully on GPU
About this result
Estimated from the GPU AI Performance Calculator. Fully resident.
Status: Estimated from the GPU AI Performance Calculator – expect variation on real hardware.