BackendCUDA
Decode160 tok/s
VRAM24 GB
Memory bandwidth1008 GB/s
ModelQwen 7B
Parameters7B
QuantizationQ4
Context8K
KV quantQ4
OffloadFully on GPU
About this result
Estimated from the GPU AI Performance Calculator. Fully resident.
This result is estimated from the GPU AI Performance Calculator – expect variation on real hardware.