BackendROCm
Decode68 tok/s
VRAM16 GB
Memory bandwidth644 GB/s
ModelQwen 14B
Parameters14B
QuantizationQ4
Context8K
KV quantQ4
OffloadFully on GPU
About this result
Estimated. Fully resident on 16 GB.
This result is estimated from the GPU AI Performance Calculator – expect variation on real hardware.