- quant:
- Q3 (gguf)
User expresses amazement at running a frontier model on a home PC with 24GB VRAM, but notes it's slow.
DeepSeek · 3 reports
User expresses amazement at running a frontier model on a home PC with 24GB VRAM, but notes it's slow.
User is setting up a 16x DGX Spark cluster to run frontier models locally. Mentions DeepSeek V4 pro, Kimi K3, GLM 5.5, and Minimax M4 as future models. No benchmark numbers provided.
Prefill performance mentioned but no numbers given. Decode speeds at various depths: 28 t/s at start, 23.5 t/s at 45k, 18 t/s at 192k. Maintained with 8k token output.