BackendCUDA
Decode42 tok/s
VRAM16 GB
Memory bandwidth288 GB/s
ModelQwen 7B
Parameters7B
QuantizationQ4
Context8K
KV quantQ4
OffloadFully on GPU

About this result

Estimated. Low bandwidth limits decode.

This result is estimated from the GPU AI Performance Calculator – expect variation on real hardware.

Related