BackendVulkan
Decode19.66 tok/s
Prompt eval83.1 tok/s
VRAM12 GB
Memory bandwidth~384 GB/s
ModelOrnith 35B-A3B MoE
Parameters35B (3B active)
QuantizationQ5_K/Q4_K
Context262K
KV quantq8_0
OffloadPartial (some layers in RAM)

About this result

28 CPU experts, MoE-sparse + q8 KV + CPU offload. ROCm wins decode (+30%). Decodes ~3B active per token, so speed holds even at 262K. Needs --load-mode none, no -ngl 999.

This result was tested in the VelsTech Lab on real hardware. Read the full Lab report →

Related