Best local LLMs for 8 GB of VRAM
On one 8 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is MiMo V2.6 Distill Qwen 9B (9.4B, 7.4 GB, about 21–30 tokens/s on one RTX 4060 8GB), the largest MoE is LFM2.5 8B-A1B (8.5B total, 1.5B active, 6.2 GB, about 57–99 tokens/s), among models that take 4 GB or more, the longest context goes to Ling 3.0 Tiny 7.9B-A1.3B (its full 128K in 6.3 GB), and the largest at Q8_0 is Spark-X2.5 4B (4.1B, 6.3 GB).
Updated , the latest date the data of a model listed here was checked; 55 models in total. Ordered by memory and size, not quality. Total = weights + FP16 KV cache for one request + 0.5 GB + 10% overhead; GB = GiB. A model counts as fitting when it leaves at least 0.5 GB free, as on the Can I run it? pages. Speeds are bandwidth estimates for one request with the 32K context full, on one RTX 4060 8GB (272 GB/s), not measurements. Not sure how much VRAM you have? Detect your GPU.
Top picks for 8 GB
| Pick | Model | Parameters | Active | VRAM | Tokens/s (RTX 4060 8GB) |
|---|---|---|---|---|---|
| Largest dense (Q4_K_M) | MiMo V2.6 Distill Qwen 9B | 9.4B | dense | 7.4 GB | 21–30 |
| Largest MoE (Q4_K_M) | LFM2.5 8B-A1B (MoE) | 8.5B | 1.5B | 6.2 GB | 57–99 |
| Longest context, 4 GB+ models (Q4_K_M) | Ling 3.0 Tiny 7.9B-A1.3B (MoE) | 7.9B | 1.3B | 128K (full), 6.3 GB | 72–126 |
| Largest at Q8_0 | Spark-X2.5 4B | 4.1B | dense | 6.3 GB | 25–35 |
What is the best local LLM for 8 GB VRAM?
By size, the largest model that fits one 8 GB GPU at Q4_K_M with 32K context is MiMo V2.6 Distill Qwen 9B (9.4B, 7.4 GB, 0.6 GB spare); 8 of the 55 models we track fit. This page ranks by memory and size, not by benchmark scores, which our data does not include.
Can 8 GB run a 70B model?
Not on one card: Llama 3.1 70B (70.6B) needs 55.2 GB at Q4_K_M with 32K context, and still 41.3 GB at IQ3_XXS.
Every step is on the Llama 3.1 70B VRAM page.
Is Q8_0 or a bigger model better on 8 GB?
On one 8 GB GPU with 32K context, the largest model at Q8_0 is Spark-X2.5 4B (4.1B, 6.3 GB, about 25–35 tokens/s) and at Q4_K_M it is MiMo V2.6 Distill Qwen 9B (9.4B, 7.4 GB, about 21–30 tokens/s), 2.3× the total parameters. Q8_0 stores about 8.5 bits per weight and Q4_K_M 4.84; our data covers memory and speed, not output quality, so this page does not score the trade-off.
Bits per weight for every GGUF type are in GGUF quantization explained.
Newest models that fit 8 GB
Sorted by the date each model's data was added or checked, newest first (Q4_K_M, 32K context). New models appear here as they are added.
| Model | Added / checked | Parameters | Active | VRAM (GB) | Tokens/s |
|---|---|---|---|---|---|
| Spark-X2.5 4B | 4.1B | dense | 4.40 | 37–52 | |
| Ling 3.0 Tiny 7.9B-A1.3B (MoE) | 7.9B | 1.3B | 5.62 | 72–126 | |
| LFM2.5 8B-A1B (MoE) | 8.5B | 1.5B | 6.16 | 57–99 | |
| Limite 1B Violetto | 1.0B | dense | 1.62 | 113–170 | |
| MiMo V2.6 Distill Qwen 9B | 9.4B | dense | 7.43 | 21–30 | |
| Gemma 4 E4B | 8.0B | dense | 6.05 | 27–37 | |
| Nemotron 3 Nano 4B | 4.0B | dense | 3.51 | 47–67 | |
| MiniCPM5 2B | 2.5B | dense | 3.50 | 47–67 |
All 8 models that fit 8 GB at Q4_K_M
| Model | Parameters | Active | VRAM (GB) | Spare (GB) | Tokens/s | Calculator |
|---|---|---|---|---|---|---|
| MiMo V2.6 Distill Qwen 9B | 9.4B | dense | 7.43 | 0.57 | 21–30 | Open |
| LFM2.5 8B-A1B (MoE) | 8.5B | 1.5B | 6.16 | 1.84 | 57–99 | Open |
| Gemma 4 E4B | 8.0B | dense | 6.05 | 1.95 | 27–37 | Open |
| Ling 3.0 Tiny 7.9B-A1.3B (MoE) | 7.9B | 1.3B | 5.62 | 2.38 | 72–126 | Open |
| Spark-X2.5 4B | 4.1B | dense | 4.40 | 3.60 | 37–52 | Open |
| Nemotron 3 Nano 4B | 4.0B | dense | 3.51 | 4.49 | 47–67 | Open |
| MiniCPM5 2B | 2.5B | dense | 3.50 | 4.50 | 47–67 | Open |
| Limite 1B Violetto | 1.0B | dense | 1.62 | 6.38 | 113–170 | Open |
All 4 models that fit 8 GB at Q8_0
| Model | Parameters | Active | VRAM (GB) | Spare (GB) | Tokens/s | Calculator |
|---|---|---|---|---|---|---|
| Spark-X2.5 4B | 4.1B | dense | 6.33 | 1.67 | 25–35 | Open |
| Nemotron 3 Nano 4B | 4.0B | dense | 5.38 | 2.62 | 30–42 | Open |
| MiniCPM5 2B | 2.5B | dense | 4.68 | 3.32 | 35–49 | Open |
| Limite 1B Violetto | 1.0B | dense | 2.11 | 5.89 | 83–122 | Open |
Macs and CPU-only PCs with about 8 GB to use
These machines leave a model between 8 GB and the next size up: on a Mac, the unified memory macOS lets the GPU use by default (about two thirds up to 32 GB, 75% above; sudo sysctl iogpu.wired_limit_mb raises it), on a CPU-only PC the RAM minus about 6 GB for the system. The lists above apply to them too; speed follows each one's bandwidth, and a CPU is also much slower at reading long prompts.
| Machine | Usable memory | Bandwidth | Models at Q4_K_M (8K) | Biggest at 32K | Tokens/s |
|---|---|---|---|---|---|
| M1 Mac (16 GB) | 10.7 GB | 68.25 GB/s | 14 | Gemma 4 12B | 4.5–6.2 |
| M2 or M3 Mac (16 GB) | 10.7 GB | 100 GB/s | 14 | Gemma 4 12B | 6.6–9.0 |
| M4 Mac (16 GB) | 10.7 GB | 120 GB/s | 14 | Gemma 4 12B | 7.9–11 |
| M5 Mac (16 GB) | 10.7 GB | 153 GB/s | 14 | Gemma 4 12B | 10–14 |
| M1 Pro or M2 Pro Mac (16 GB) | 10.7 GB | 200 GB/s | 14 | Gemma 4 12B | 13–18 |
| CPU only, DDR4-3200 dual channel (16 GB) | 10 GB | 51.2 GB/s | 14 | Gemma 4 12B | 3.1–4.9 |
Other VRAM sizes
All numbers, including 8K and 128K context, are in the open dataset, with summary figures on LLM VRAM statistics; any other setting can be worked out in the LLM VRAM Calculator.