BackendROCm (MTP)
Decode29.09 tok/s
Prompt eval111.4 tok/s
VRAM12 GB
Memory bandwidth~384 GB/s
ModelTiel-Coder 35B-A3B
Parameters35B (3B active)
QuantizationQ4_K_XL
Context262K
KV quantq8_0
OffloadPartial (some layers in RAM)

About this result

MTP speculative decoding on MoE (28 vs 32 CPU experts). MTP +14.6% decode (25.39 → 29.09 tok/s), prompt -6.4%, 54.7% acceptance (2.64 mean). 35B Q4 ~17.5GB + q8 KV 262K needs --n-cpu-moe offload on 12GB.

This result was tested in the VelsTech Lab on real hardware. Read the full Lab report →

Related