BackendCUDA
Decode120 tok/s
VRAM24 GB
Memory bandwidth1008 GB/s
ModelQwen 14B
Parameters14B
QuantizationQ4
Context8K
KV quantQ4
OffloadFully on GPU

About this result

Estimated from the GPU AI Performance Calculator. Fully resident.

This result is estimated from the GPU AI Performance Calculator – expect variation on real hardware.

Related