Best local LLMs for 8 to 96 GB of VRAM
On one 24 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is Muse Glimmer 30B (29.8B, 19.5 GB, about 29–40 tokens/s on one RTX 4090), the largest MoE is a tie between Ornith 1.5 35B-A3B and Qwen3.6 35B-A3B (36.0B total, 3.0B active, 23.5 GB, about 103–184 tokens/s), among models that take 12 GB or more, the longest context goes to Nemotron 3 Nano 30B-A3B (its full 256K in 21.7 GB), and the largest at Q8_0 is Gemma 4 12B (12.0B, 14.6 GB).
The table lists, for each VRAM size, the largest models that fit one GPU with 32,768 tokens of context, one request and an FP16 KV cache; each size links to its full list, newest models and questions. This compares memory, not quality. Total = weights + KV cache + 0.5 GB + 10% overhead; GB = GiB. Speeds are bandwidth estimates on one typical card of that size. Updated .
Not sure how much VRAM you have? Let the browser detect your GPU and see what your PC can run.
Top picks by VRAM size (Q4_K_M, 32K context)
| VRAM | Largest dense | Largest MoE | Longest context (models using half the VRAM or more) | Largest at Q8_0 | Models that fit |
|---|---|---|---|---|---|
| 8 GB RTX 4060 8GB | MiMo V2.6 Distill Qwen 9B 7.4 GB, 21–30 tok/s | LFM2.5 8B-A1B (MoE) 6.2 GB, 57–99 tok/s | Ling 3.0 Tiny 7.9B-A1.3B (MoE) 128K, 6.3 GB | Spark-X2.5 4B 6.3 GB | 8 / 55 |
| 12 GB RTX 4070 12GB | Gemma 4 12B 9.0 GB, 32–45 tok/s | LFM2.5 8B-A1B (MoE) 6.2 GB, 98–175 tok/s | Gemma 4 12B 178K, 11.5 GB | LFM2.5 8B-A1B (MoE) 10.1 GB | 14 / 55 |
| 16 GB RTX 5060 Ti 16GB | Gemma 4 12B 9.0 GB, 29–40 tok/s | gpt-oss-20b MXFP4, 15.4 GB, 40–68 tok/s | Gemma 4 12B 256K, 12.8 GB | Gemma 4 12B 14.6 GB | 15 / 55 |
| 24 GB RTX 4090 | Muse Glimmer 30B 19.5 GB, 29–40 tok/s | Ornith 1.5 35B-A3B (MoE) 23.5 GB, 103–184 tok/s | Nemotron 3 Nano 30B-A3B (MoE) 256K, 21.7 GB | Gemma 4 12B 14.6 GB | 27 / 55 |
| 32 GB RTX 5090 | Gemma 4 31B 23.9 GB, 40–57 tok/s | K2-Horizon MoVA 36B-A4B (MoE) 30.3 GB, 56–96 tok/s | Ornith 1.5 35B-A3B (MoE) 256K, 28.3 GB | Gemma 4 26B-A4B (MoE) 29.6 GB | 29 / 55 |
| 48 GB L40S | Gemma 4 31B 23.9 GB, 20–28 tok/s | K2-Horizon MoVA 36B-A4B (MoE) 30.3 GB, 28–48 tok/s | K2-Horizon MoVA 36B-A4B (MoE) 115K, 47.5 GB | Ornith 1.5 35B-A3B (MoE) 40.3 GB | 29 / 55 |
| 96 GB RTX PRO 6000 Blackwell | Mistral Medium 3.5 128B 91.8 GB, 11–15 tok/s | Qwen3.5 122B-A10B (MoE) 78.8 GB, 70–123 tok/s | Qwen3.5 122B-A10B (MoE) 256K, 84.6 GB | AliceAI Foundation 80B-A3B (MoE) 89.8 GB | 36 / 55 |
The short answer for each size
- Best local LLMs for 8 GB of VRAM: On one 8 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is MiMo V2.6 Distill Qwen 9B (9.4B, 7.4 GB, about 21–30 tokens/s on one RTX 4060 8GB), the largest MoE is LFM2.5 8B-A1B (8.5B total, 1.5B active, 6.2 GB, about 57–99 tokens/s), among models that take 4 GB or more, the longest context goes to Ling 3.0 Tiny 7.9B-A1.3B (its full 128K in 6.3 GB), and the largest at Q8_0 is Spark-X2.5 4B (4.1B, 6.3 GB).
- Best local LLMs for 12 GB of VRAM: On one 12 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is Gemma 4 12B (12.0B, 9.0 GB, about 32–45 tokens/s on one RTX 4070 12GB), the largest MoE is LFM2.5 8B-A1B (8.5B total, 1.5B active, 6.2 GB, about 98–175 tokens/s), among models that take 6 GB or more, the longest context goes to Gemma 4 12B (up to 178K tokens in 11.5 GB), and the largest at Q8_0 is LFM2.5 8B-A1B (MoE, 8.5B total, 1.5B active, 10.1 GB).
- Best local LLMs for 16 GB of VRAM: On one 16 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is Gemma 4 12B (12.0B, 9.0 GB, about 29–40 tokens/s on one RTX 5060 Ti 16GB), the largest MoE is gpt-oss-20b (20.9B total, 3.6B active, MXFP4, 15.4 GB, about 40–68 tokens/s), among models that take 8 GB or more, the longest context goes to Gemma 4 12B (its full 256K in 12.8 GB), and the largest at Q8_0 is Gemma 4 12B (12.0B, 14.6 GB).
- Best local LLMs for 24 GB of VRAM: On one 24 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is Muse Glimmer 30B (29.8B, 19.5 GB, about 29–40 tokens/s on one RTX 4090), the largest MoE is a tie between Ornith 1.5 35B-A3B and Qwen3.6 35B-A3B (36.0B total, 3.0B active, 23.5 GB, about 103–184 tokens/s), among models that take 12 GB or more, the longest context goes to Nemotron 3 Nano 30B-A3B (its full 256K in 21.7 GB), and the largest at Q8_0 is Gemma 4 12B (12.0B, 14.6 GB).
- Best local LLMs for 32 GB of VRAM: On one 32 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is Gemma 4 31B (31.3B, 23.9 GB, about 40–57 tokens/s on one RTX 5090), the largest MoE is K2-Horizon MoVA 36B-A4B (37.4B total, 4.0B active, 30.3 GB, about 56–96 tokens/s), among models that take 16 GB or more, the longest context goes to a tie between Ornith 1.5 35B-A3B and Qwen3.6 35B-A3B (the full 256K in 28.3 GB), and the largest at Q8_0 is Gemma 4 26B-A4B (MoE, 25.8B total, 4.0B active, 29.6 GB).
- Best local LLMs for 48 GB of VRAM: On one 48 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is Gemma 4 31B (31.3B, 23.9 GB, about 20–28 tokens/s on one L40S), the largest MoE is K2-Horizon MoVA 36B-A4B (37.4B total, 4.0B active, 30.3 GB, about 28–48 tokens/s), among models that take 24 GB or more, the longest context goes to K2-Horizon MoVA 36B-A4B (up to 115K tokens in 47.5 GB), and the largest at Q8_0 is a tie between Ornith 1.5 35B-A3B and Qwen3.6 35B-A3B (MoE, 36.0B total, 3.0B active, 40.3 GB).
- Best local LLMs for 96 GB of VRAM: On one 96 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is Mistral Medium 3.5 128B (127.7B, 91.8 GB, about 11–15 tokens/s on one RTX PRO 6000 Blackwell), the largest MoE is Qwen3.5 122B-A10B (125.1B total, 10.0B active, 78.8 GB, about 70–123 tokens/s), among models that take 48 GB or more, the longest context goes to Qwen3.5 122B-A10B (its full 256K in 84.6 GB), and the largest at Q8_0 is AliceAI Foundation 80B-A3B (MoE, 81.3B total, 3.0B active, 89.8 GB).
All numbers, including 8K and 128K context, FP16 and 80 GB cards, are in the open dataset and the model and GPU report; summary figures across all models are on the LLM VRAM statistics page. Each model page shows every step of its calculation.