BackendCUDA
Decode60 tok/s
VRAM24 GB
Memory bandwidth1008 GB/s
ModelQwen 27B
Parameters27B
QuantizationQ4
Context8K
KV quantQ4
OffloadFully on GPU

About this result

Estimated. Tight but resident on 24 GB.

This result is estimated from the GPU AI Performance Calculator – expect variation on real hardware.

Related