Search, filter, and compare how fast different LLMs run on different GPUs. Tested results come from the VelsTech Lab on real hardware. Estimated results come from the GPU AI Performance Calculator. Every row links to a dedicated benchmark page.

GPU Model Quant Context Backend Decode

Decode = tokens per second while generating. Higher is better for chat. Offloaded results run partly in system RAM and are slower.

How to read these numbers

Decode speed is the number that matters most for chatting – it's how fast the model streams its answer. Prompt eval (not shown per row) matters for long documents. Numbers are only comparable within the same backend, context, and quantization.

Want to know what fits your card? Check the LLM VRAM Calculator, or read the best GPU for local LLMs guide.