Best local LLMs for 8 GB of VRAM

On one 8 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is MiMo V2.6 Distill Qwen 9B (9.4B, 7.4 GB, about 21–30 tokens/s on one RTX 4060 8GB), the largest MoE is LFM2.5 8B-A1B (8.5B total, 1.5B active, 6.2 GB, about 57–99 tokens/s), among models that take 4 GB or more, the longest context goes to Ling 3.0 Tiny 7.9B-A1.3B (its full 128K in 6.3 GB), and the largest at Q8_0 is Spark-X2.5 4B (4.1B, 6.3 GB).

Updated , the latest date the data of a model listed here was checked; 55 models in total. Ordered by memory and size, not quality. Total = weights + FP16 KV cache for one request + 0.5 GB + 10% overhead; GB = GiB. A model counts as fitting when it leaves at least 0.5 GB free, as on the Can I run it? pages. Speeds are bandwidth estimates for one request with the 32K context full, on one RTX 4060 8GB (272 GB/s), not measurements. Not sure how much VRAM you have? Detect your GPU.

Top picks for 8 GB

Pick Model Parameters Active VRAM Tokens/s (RTX 4060 8GB)
Largest dense (Q4_K_M) MiMo V2.6 Distill Qwen 9B 9.4B dense 7.4 GB 21–30
Largest MoE (Q4_K_M) LFM2.5 8B-A1B (MoE) 8.5B 1.5B 6.2 GB 57–99
Longest context, 4 GB+ models (Q4_K_M) Ling 3.0 Tiny 7.9B-A1.3B (MoE) 7.9B 1.3B 128K (full), 6.3 GB 72–126
Largest at Q8_0 Spark-X2.5 4B 4.1B dense 6.3 GB 25–35

What is the best local LLM for 8 GB VRAM?

By size, the largest model that fits one 8 GB GPU at Q4_K_M with 32K context is MiMo V2.6 Distill Qwen 9B (9.4B, 7.4 GB, 0.6 GB spare); 8 of the 55 models we track fit. This page ranks by memory and size, not by benchmark scores, which our data does not include.

Can 8 GB run a 70B model?

Not on one card: Llama 3.1 70B (70.6B) needs 55.2 GB at Q4_K_M with 32K context, and still 41.3 GB at IQ3_XXS.

Every step is on the Llama 3.1 70B VRAM page.

Is Q8_0 or a bigger model better on 8 GB?

On one 8 GB GPU with 32K context, the largest model at Q8_0 is Spark-X2.5 4B (4.1B, 6.3 GB, about 25–35 tokens/s) and at Q4_K_M it is MiMo V2.6 Distill Qwen 9B (9.4B, 7.4 GB, about 21–30 tokens/s), 2.3× the total parameters. Q8_0 stores about 8.5 bits per weight and Q4_K_M 4.84; our data covers memory and speed, not output quality, so this page does not score the trade-off.

Bits per weight for every GGUF type are in GGUF quantization explained.

Newest models that fit 8 GB

Sorted by the date each model's data was added or checked, newest first (Q4_K_M, 32K context). New models appear here as they are added.

Model Added / checked Parameters Active VRAM (GB) Tokens/s
Spark-X2.5 4B 4.1B dense 4.40 37–52
Ling 3.0 Tiny 7.9B-A1.3B (MoE) 7.9B 1.3B 5.62 72–126
LFM2.5 8B-A1B (MoE) 8.5B 1.5B 6.16 57–99
Limite 1B Violetto 1.0B dense 1.62 113–170
MiMo V2.6 Distill Qwen 9B 9.4B dense 7.43 21–30
Gemma 4 E4B 8.0B dense 6.05 27–37
Nemotron 3 Nano 4B 4.0B dense 3.51 47–67
MiniCPM5 2B 2.5B dense 3.50 47–67

All 8 models that fit 8 GB at Q4_K_M

Model Parameters Active VRAM (GB) Spare (GB) Tokens/s Calculator
MiMo V2.6 Distill Qwen 9B 9.4B dense 7.43 0.57 21–30 Open
LFM2.5 8B-A1B (MoE) 8.5B 1.5B 6.16 1.84 57–99 Open
Gemma 4 E4B 8.0B dense 6.05 1.95 27–37 Open
Ling 3.0 Tiny 7.9B-A1.3B (MoE) 7.9B 1.3B 5.62 2.38 72–126 Open
Spark-X2.5 4B 4.1B dense 4.40 3.60 37–52 Open
Nemotron 3 Nano 4B 4.0B dense 3.51 4.49 47–67 Open
MiniCPM5 2B 2.5B dense 3.50 4.50 47–67 Open
Limite 1B Violetto 1.0B dense 1.62 6.38 113–170 Open

All 4 models that fit 8 GB at Q8_0

Model Parameters Active VRAM (GB) Spare (GB) Tokens/s Calculator
Spark-X2.5 4B 4.1B dense 6.33 1.67 25–35 Open
Nemotron 3 Nano 4B 4.0B dense 5.38 2.62 30–42 Open
MiniCPM5 2B 2.5B dense 4.68 3.32 35–49 Open
Limite 1B Violetto 1.0B dense 2.11 5.89 83–122 Open

Macs and CPU-only PCs with about 8 GB to use

These machines leave a model between 8 GB and the next size up: on a Mac, the unified memory macOS lets the GPU use by default (about two thirds up to 32 GB, 75% above; sudo sysctl iogpu.wired_limit_mb raises it), on a CPU-only PC the RAM minus about 6 GB for the system. The lists above apply to them too; speed follows each one's bandwidth, and a CPU is also much slower at reading long prompts.

Machine Usable memory Bandwidth Models at Q4_K_M (8K) Biggest at 32K Tokens/s
M1 Mac (16 GB) 10.7 GB 68.25 GB/s 14 Gemma 4 12B 4.5–6.2
M2 or M3 Mac (16 GB) 10.7 GB 100 GB/s 14 Gemma 4 12B 6.6–9.0
M4 Mac (16 GB) 10.7 GB 120 GB/s 14 Gemma 4 12B 7.9–11
M5 Mac (16 GB) 10.7 GB 153 GB/s 14 Gemma 4 12B 10–14
M1 Pro or M2 Pro Mac (16 GB) 10.7 GB 200 GB/s 14 Gemma 4 12B 13–18
CPU only, DDR4-3200 dual channel (16 GB) 10 GB 51.2 GB/s 14 Gemma 4 12B 3.1–4.9

Other VRAM sizes

All numbers, including 8K and 128K context, are in the open dataset, with summary figures on LLM VRAM statistics; any other setting can be worked out in the LLM VRAM Calculator.