Best local LLMs for 12 GB of VRAM

On one 12 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is Gemma 4 12B (12.0B, 9.0 GB, about 32–45 tokens/s on one RTX 4070 12GB), the largest MoE is LFM2.5 8B-A1B (8.5B total, 1.5B active, 6.2 GB, about 98–175 tokens/s), among models that take 6 GB or more, the longest context goes to Gemma 4 12B (up to 178K tokens in 11.5 GB), and the largest at Q8_0 is LFM2.5 8B-A1B (MoE, 8.5B total, 1.5B active, 10.1 GB).

Updated , the latest date the data of a model listed here was checked; 55 models in total. Ordered by memory and size, not quality. Total = weights + FP16 KV cache for one request + 0.5 GB + 10% overhead; GB = GiB. A model counts as fitting when it leaves at least 0.5 GB free, as on the Can I run it? pages. Speeds are bandwidth estimates for one request with the 32K context full, on one RTX 4070 12GB (504 GB/s), not measurements. Not sure how much VRAM you have? Detect your GPU.

Top picks for 12 GB

Pick Model Parameters Active VRAM Tokens/s (RTX 4070 12GB)
Largest dense (Q4_K_M) Gemma 4 12B 12.0B dense 9.0 GB 32–45
Largest MoE (Q4_K_M) LFM2.5 8B-A1B (MoE) 8.5B 1.5B 6.2 GB 98–175
Longest context, 6 GB+ models (Q4_K_M) Gemma 4 12B 12.0B dense 178K, 11.5 GB 32–45
Largest at Q8_0 LFM2.5 8B-A1B (MoE) 8.5B 1.5B 10.1 GB 68–119

What is the best local LLM for 12 GB VRAM?

By size, the largest model that fits one 12 GB GPU at Q4_K_M with 32K context is Gemma 4 12B (12.0B, 9.0 GB, 3.0 GB spare); 14 of the 55 models we track fit. This page ranks by memory and size, not by benchmark scores, which our data does not include.

Can 12 GB run a 70B model?

Not on one card: Llama 3.1 70B (70.6B) needs 55.2 GB at Q4_K_M with 32K context, and still 41.3 GB at IQ3_XXS.

Every step is on the Llama 3.1 70B VRAM page.

Is Q8_0 or a bigger model better on 12 GB?

On one 12 GB GPU with 32K context, the largest model at Q8_0 is LFM2.5 8B-A1B (MoE, 8.5B total, 1.5B active, 10.1 GB, about 68–119 tokens/s) and at Q4_K_M it is Gemma 4 12B (12.0B, 9.0 GB, about 32–45 tokens/s), 1.4× the total parameters. Q8_0 stores about 8.5 bits per weight and Q4_K_M 4.84; our data covers memory and speed, not output quality, so this page does not score the trade-off.

Bits per weight for every GGUF type are in GGUF quantization explained.

Newest models that fit 12 GB

Sorted by the date each model's data was added or checked, newest first (Q4_K_M, 32K context). New models appear here as they are added.

Model Added / checked Parameters Active VRAM (GB) Tokens/s
Spark-X2.5 4B 4.1B dense 4.40 66–95
Ling 3.0 Tiny 7.9B-A1.3B (MoE) 7.9B 1.3B 5.62 122–221
LFM2.5 8B-A1B (MoE) 8.5B 1.5B 6.16 98–175
Limite 1B Violetto 1.0B dense 1.62 183–294
ZDTaichu 5.0 9B 9.8B dense 7.67 37–53
MiMo V2.6 Distill Qwen 9B 9.4B dense 7.43 39–54
Ornith 1.5 9B 9.7B dense 7.58 38–53
Qwen3.5 9B 9.7B dense 7.58 38–53

All 14 models that fit 12 GB at Q4_K_M

Model Parameters Active VRAM (GB) Spare (GB) Tokens/s Calculator
Gemma 4 12B 12.0B dense 8.98 3.02 32–45 Open
ZDTaichu 5.0 9B 9.8B dense 7.67 4.33 37–53 Open
Ornith 1.5 9B 9.7B dense 7.58 4.42 38–53 Open
Qwen3.5 9B 9.7B dense 7.58 4.42 38–53 Open
MiMo V2.6 Distill Qwen 9B 9.4B dense 7.43 4.57 39–54 Open
LFM2.5 8B-A1B (MoE) 8.5B 1.5B 6.16 5.84 98–175 Open
Qwen3 8B 8.2B dense 10.53 1.47 27–38 Open
Llama 3.1 8B 8.0B dense 9.88 2.12 29–40 Open
Gemma 4 E4B 8.0B dense 6.05 5.95 48–67 Open
Ling 3.0 Tiny 7.9B-A1.3B (MoE) 7.9B 1.3B 5.62 6.38 122–221 Open
Spark-X2.5 4B 4.1B dense 4.40 7.60 66–95 Open
Nemotron 3 Nano 4B 4.0B dense 3.51 8.49 83–121 Open
MiniCPM5 2B 2.5B dense 3.50 8.50 83–121 Open
Limite 1B Violetto 1.0B dense 1.62 10.38 183–294 Open

All 7 models that fit 12 GB at Q8_0

Model Parameters Active VRAM (GB) Spare (GB) Tokens/s Calculator
LFM2.5 8B-A1B (MoE) 8.5B 1.5B 10.13 1.87 68–119 Open
Gemma 4 E4B 8.0B dense 9.80 2.20 29–41 Open
Ling 3.0 Tiny 7.9B-A1.3B (MoE) 7.9B 1.3B 9.32 2.68 82–145 Open
Spark-X2.5 4B 4.1B dense 6.33 5.67 45–64 Open
Nemotron 3 Nano 4B 4.0B dense 5.38 6.62 54–76 Open
MiniCPM5 2B 2.5B dense 4.68 7.32 62–88 Open
Limite 1B Violetto 1.0B dense 2.11 9.89 140–215 Open

Macs and CPU-only PCs with about 12 GB to use

These machines leave a model between 12 GB and the next size up: on a Mac, the unified memory macOS lets the GPU use by default (about two thirds up to 32 GB, 75% above; sudo sysctl iogpu.wired_limit_mb raises it), on a CPU-only PC the RAM minus about 6 GB for the system. The lists above apply to them too; speed follows each one's bandwidth, and a CPU is also much slower at reading long prompts.

Machine Usable memory Bandwidth Models at Q4_K_M (8K) Biggest at 32K Tokens/s
M3 Pro Mac (18 GB) 12 GB 150 GB/s 14 Gemma 4 12B 9.8–14

Other VRAM sizes

All numbers, including 8K and 128K context, are in the open dataset, with summary figures on LLM VRAM statistics; any other setting can be worked out in the LLM VRAM Calculator.