llamaperf

How open-weight LLMs run on your hardware

81 performance reports, crowdsourced from the community.

VRAM Calculator

Pick your GPU - see which models fit, at which quant, and how fast they run.

Open calculator →

Qwen3.6 35B (3B active)

RTX 3090 · llama.cpp · 65,536 ctx

throughput:
97.7 t/s gen · 1330.0 t/s pp
quant:
Q6_K (gguf)
kv:
Q8

Benchmark of Qwen3.6-35B-A3B Q6_K on RTX 3090 with llama.cpp. Tuned config offloads 8 MoE expert layers to CPU, increasing batch sizes, improving PP from 564.5 to 1330.0 tok/s (2.36x) while TG unchanged (~97.7 tok/s). Context 64K, KV cache Q8. CPU: Threadripper PRO 3955WX, ~100GB DDR4. Two-repetition measurements, ~1.6% drift.

User is setting up a 16x DGX Spark cluster to run frontier models locally. Mentions DeepSeek V4 pro, Kimi K3, GLM 5.5, and Minimax M4 as future models. No benchmark numbers provided.

Tone: positive
quant:
WinterMix58 (mlx)

Post describes a new MLX quantization method (WinterMix) for Qwen3.5-122B-A10B, with two builds: WinterMix58 (82 GiB) and WinterMix48 (68 GiB). Benchmarks show perplexity improvements over existing MLX quants. The author is enthusiastic about the results and the method's advantages for agentic workflows on Apple Silicon.

Tone: positive
throughput:
5.5 t/s gen

Custom Swift/Metal engine. Also tested on M5 MacBook Pro: 31-35 tok/s. OpenAI-compatible server with streaming and tool-call support.

Gemma 4

M5 32GB · MLX · 130,173 ctx

throughput:
3029.0 t/s pp

User built custom w8a8 kernels for M5 MacBook Air. Baseline prefill 2193 tps, improved to 3029 tps for 130k tokens. Also mentions llama.cpp for Macs.

Qwen3.6 27B

RTX 5060 Ti 16GB · 256,000 ctx

Tone: positive
throughput:
52.2 t/s gen · 608.0 t/s pp
quant:
Q8
kv:
F16
mtp (multi-token prediction):
on
coding

Benchmark on Vast AI instance with 4x RTX 5060 Ti 16GB. Q8 quant, FP16 KV cache, MTP enabled. 256K context. Cold prefill 608 t/s, decode 52.2 t/s. User considers this excellent for $2K hardware.

Kimi K2.6

RTX 5090 · llama.cpp

throughput:
471.4 t/s pp
quant:
IQ3_M (gguf)

LLM prompt processing benchmark with Kimi K2.5 IQ3_M (80GB offload) at 500W. RTX 5090 achieved 471.40 t/s PP. Also tested GLM 5.1 IQ4_NL (70GB offload) at 574.98 t/s PP. Comparison with RTX 6000 PRO MaxQ shunt modded.

Qwen2.5 27B

RTX 3090 · llama.cpp

Tone: positive
throughput:
70.0 t/s gen · 1850.0 t/s pp
quant:
Q6_K_XL (gguf)
coding

Multi-token prediction enabled. 96GB total VRAM (24+24? but user says 96GB system). Reliable for code generation and codebase ingestion.

Qwen3.6 27B

RTX 3090 Ti · llama.cpp · 196,608 ctx

Tone: positive
throughput:
100.0 t/s gen
quant:
Q8_0 (gguf)

Tensor split-mode improved t/s from 70+ to 100+. Peak 130 t/s reported. Power draw 750W+.

Tone: positive
quant:
Q4_K_M (gguf)
vision

Champion model in vision benchmark. Best quality and stability with thinking disabled. 90/90 successful runs. Speed: 70 s/img. Tested on Apple M2 Max 96GB with llama.cpp b9690.

Qwen3.6 27B

RTX 5060 Ti 16GB · llama.cpp · 131,072 ctx

Tone: positive
throughput:
19.0 t/s gen
quant:
IQ4_XS (gguf)
kv:
F16

User reports that offloading KV cache to RAM (with -nkvo) allows fitting the whole model on GPU with f16 KV cache, achieving 19 tps peak and 14 tps during long generation at 65k context. With 128k context and 63 layers on GPU, speed remained similar. KV cache quant to RAM didn't improve performance.

Gemma 4

RX 7900 XTX · llama-swap

Tone: positive

Benchmark of Gemma 4 QAT vs regular quants on AMD 7900 XTX. No token/s reported, but wall clock times show significant speedups (e.g., 12B QAT 45% faster, 83% throughput increase). Quality reported identical. Models tested: 12B, 26B, 31B, E4B.

Benchmark of abliteration tools (Apostate, Huihui, Heretic) on Qwen 2.5 7B. Evaluated with lm-evaluation-harness via vLLM 0.19.0, bf16 on RTX 5090 32GB. Reports MMLU, GSM8K, HellaSwag, ARC Challenge, WinoGrande, TruthfulQA MC2, PiQA, LAMBADA ppl, HarmBench ASR, KL divergence. No tokens/sec reported.

Tone: positive
throughput:
138.0 t/s gen
coding

Benchmark of Gemma 4 26B-A4B vs 12B on RTX 4090. 26B-A4B used 15GB VRAM, 138 tok/s; 12B used 9GB, 80 tok/s. 26B-A4B won every scene and ran ~1.7x faster. 12B ideal for 16GB laptop.

Qwen3.6 27B

RTX 3090 · Ollama · 32,000 ctx

Tone: mixed
quant:
Q6_K (gguf)
codingagentic

User replaced Claude with Qwen3.6-27B in multi-agent orchestrator for 2 weeks. Plan generation good, tool-call reliability poor (12% format error), long-context drift past ~14k tokens, cascade-failure handling weak. Viable as reasoning layer but not execution layer.

quant:
4bit
agenticcoding

User currently runs Qwen3.6-35B-A3B-4bit on M3 Max 128GB for production sub-agent delegations. Also mentions GLM-5.1 for orchestration. Considering building a 5090 rig.

Showing 120 of 96
Page 1 of 5

Community benchmarks snapshot

96 records · 26 GPUs · 9 model families · 5 engines

Records by GPU

RTX 3090 13 RTX 5090 12 M5 Max 128GB 6 RTX 3060 12GB 5 RTX 5060 Ti 16GB 4 AMD Strix Halo 128GB 4 RX 7900 XTX 4 RTX 4090 4 M5 Max 64GB 3 M2 Max 96GB 3

Records by model

90 total
Qwen3.657
Gemma 420
DeepSeek V33
Qwen2.53
Kimi K2.62
Qwen3.52
other3

Records by engine

63 total
llama.cpp33
vLLM13
Ollama9
MLX6
LM Studio2

Use cases

coding 36agentic 11text-generation 9tool-use 7summarization 6long-context 5vision 4creative-writing 3math 3multilingual 1
coding 36 agentic 11 text-generation 9 tool-use 7 summarization 6 long-context 5 vision 4 creative-writing 3

Avg gen t/s by GPU

Pro 6000 3500 5090 648 4070 Ti Super 110 3090 Ti 100 4090 98 H100 80GB 85 3090 65 M5 Max 64GB 64 5080 56 4070 55

Avg gen t/s by model

Qwen3.6 197 Gemma 4 96 Qwen2.5 49 DeepSeek V3 28 Kimi K2.6 10 Llama 3.1 1

Quants

Q4_K_M 16 IQ4_XS 6 Q8 5 NVFP4 4 Q6_K 3 Q8_0 3 Q4 3 Q5_K_S 2 4bit 1 UD-Q5_K_XL 1