Best local LLMs for 8 to 96 GB of VRAM

On one 24 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is Muse Glimmer 30B (29.8B, 19.5 GB, about 29–40 tokens/s on one RTX 4090), the largest MoE is a tie between Ornith 1.5 35B-A3B and Qwen3.6 35B-A3B (36.0B total, 3.0B active, 23.5 GB, about 103–184 tokens/s), among models that take 12 GB or more, the longest context goes to Nemotron 3 Nano 30B-A3B (its full 256K in 21.7 GB), and the largest at Q8_0 is Gemma 4 12B (12.0B, 14.6 GB).

The table lists, for each VRAM size, the largest models that fit one GPU with 32,768 tokens of context, one request and an FP16 KV cache; each size links to its full list, newest models and questions. This compares memory, not quality. Total = weights + KV cache + 0.5 GB + 10% overhead; GB = GiB. Speeds are bandwidth estimates on one typical card of that size. Updated .

Not sure how much VRAM you have? Let the browser detect your GPU and see what your PC can run.

Top picks by VRAM size (Q4_K_M, 32K context)

VRAM Largest dense Largest MoE Longest context (models using half the VRAM or more) Largest at Q8_0 Models that fit
8 GB
RTX 4060 8GB
MiMo V2.6 Distill Qwen 9B
7.4 GB, 21–30 tok/s
LFM2.5 8B-A1B (MoE)
6.2 GB, 57–99 tok/s
Ling 3.0 Tiny 7.9B-A1.3B (MoE)
128K, 6.3 GB
Spark-X2.5 4B
6.3 GB
8 / 55
12 GB
RTX 4070 12GB
Gemma 4 12B
9.0 GB, 32–45 tok/s
LFM2.5 8B-A1B (MoE)
6.2 GB, 98–175 tok/s
Gemma 4 12B
178K, 11.5 GB
LFM2.5 8B-A1B (MoE)
10.1 GB
14 / 55
16 GB
RTX 5060 Ti 16GB
Gemma 4 12B
9.0 GB, 29–40 tok/s
gpt-oss-20b
MXFP4, 15.4 GB, 40–68 tok/s
Gemma 4 12B
256K, 12.8 GB
Gemma 4 12B
14.6 GB
15 / 55
24 GB
RTX 4090
Muse Glimmer 30B
19.5 GB, 29–40 tok/s
Ornith 1.5 35B-A3B (MoE)
23.5 GB, 103–184 tok/s
Nemotron 3 Nano 30B-A3B (MoE)
256K, 21.7 GB
Gemma 4 12B
14.6 GB
27 / 55
32 GB
RTX 5090
Gemma 4 31B
23.9 GB, 40–57 tok/s
K2-Horizon MoVA 36B-A4B (MoE)
30.3 GB, 56–96 tok/s
Ornith 1.5 35B-A3B (MoE)
256K, 28.3 GB
Gemma 4 26B-A4B (MoE)
29.6 GB
29 / 55
48 GB
L40S
Gemma 4 31B
23.9 GB, 20–28 tok/s
K2-Horizon MoVA 36B-A4B (MoE)
30.3 GB, 28–48 tok/s
K2-Horizon MoVA 36B-A4B (MoE)
115K, 47.5 GB
Ornith 1.5 35B-A3B (MoE)
40.3 GB
29 / 55
96 GB
RTX PRO 6000 Blackwell
Mistral Medium 3.5 128B
91.8 GB, 11–15 tok/s
Qwen3.5 122B-A10B (MoE)
78.8 GB, 70–123 tok/s
Qwen3.5 122B-A10B (MoE)
256K, 84.6 GB
AliceAI Foundation 80B-A3B (MoE)
89.8 GB
36 / 55

The short answer for each size

All numbers, including 8K and 128K context, FP16 and 80 GB cards, are in the open dataset and the model and GPU report; summary figures across all models are on the LLM VRAM statistics page. Each model page shows every step of its calculation.