BackendROCm (PTQ1_0, gfx1031 build)
Decode18.8 tok/s
Prompt eval22.5 tok/s
VRAM12 GB
Memory bandwidth~384 GB/s
ModelTernary-Bonsai-2-27B
Parameters27B (ternary hybrid)
QuantizationPQ2_0 / PTQ1_0
Context32K
KV quantf16
OffloadFully on GPU
About this result
Stock llama.cpp rejects type 142/143; Prism fork b10709 built with -DCMAKE_HIP_ARCHITECTURES=gfx1031 required (prebuilt ROCm 7.2 gives 'device kernel image is invalid' on gfx1031). PQ2_0 wins prompt +185% and decode +74% over PTQ1_0 on RDNA2; PTQ1_0 saves 1.3GB (5.95 vs 7.21GB).
Status: Measured in the VelsTech Lab on real hardware. Read the full Lab report →