About this result
Server timings at full 262K ctx, q8_0 KV, flash-attn on. Decode falls as more experts move to CPU — 24.3 (28 experts) → 22.7 (30) → 20.9 (32) tok/s — while warm prompt eval holds ~89-105 tok/s across runs (the WebUI's first-run 26.3 tok/s was cold-start; expect ~10% run variance, WebUI cross-check read 22.5/100.9). VRAM follows the same curve: 12.77 GB (99%) → 12.06 GB (94%) → 11.08 GB (86%). That headroom decides vision: at 99% an image OOM-crashes the server inside clip image encode (ROCm out of memory, abort), but at 94% and 86% the same VT-logo grounding succeeds with accurate descriptions. Second-turn ~153-160 tok/s prompt figures are KV-cache reuse (LCP slot match), not raw eval. The 21.7 GB Q4_K_M file plus mmproj F16 cannot fully reside on 12 GB. Loader ignores unused blk.40 tensors; Qwen-VL grounding wants --image-min-tokens 1024. Caveat: the model introduces itself as Qwen3.8-FP8 on a cloud endpoint — identity is baked into the template, never introspect self-reported quant or host.
Status: Measured in the VelsTech Lab on real hardware.