Every number on this page comes from real testing on this site's hardware (RX 6800M, ROCm vs Vulkan) plus estimates from the GPU AI Performance Calculator for the big desktop cards. Filter by GPU, model, backend or tested/estimated status, then click any column header to sort. The fastest decode and prompt speed in the current view are highlighted in blue.

Loading…

GPU Model Backend Status Decode tok/s Prompt tok/s Context Quant VRAM Offload
Notes & methodology

Tested rows were measured directly on this site's hardware and are described in detail in their source articles (linked below). Estimated rows come from the GPU AI Performance Calculator assuming a fully resident model and are approximations, not lab measurements.

Decode (generation) speed is the number that matters most for interactive chat – it's bandwidth-bound, so lower-bandwidth cards like the RTX 4060 Ti drop off hard. Prompt (prefill) speed is compute-bound. Where multiple backends were tested, each backend is listed as its own row so you can compare ROCm vs Vulkan directly.

Related guides