BackendCUDA
Decode160 tok/s
VRAM24 GB
Memory bandwidth1008 GB/s
ModelQwen 7B
Parameters7B
QuantizationQ4
Context8K
KV quantQ4
OffloadFully on GPU

About this result

Estimated from the GPU AI Performance Calculator. Fully resident.

This result is estimated from the GPU AI Performance Calculator – expect variation on real hardware.

Related