2026 Local LLM VRAM Report: 64 Open Models on 8–80 GB GPUs and Macs

Every number in this report is computed at build time from the site's 64 open models with the same functions as the VRAM calculator; none is typed in. Setting: Q4_K_M weights (gpt-oss at its published MXFP4), 32K context, one request, FP16 KV cache, and a model fits only with 0.5 GB left free. Data version 1.2026-09-30, model data checked . GB means GiB.

Key findings you can quote

  1. At 32K context, the FP16 KV cache of one request ranges from 0.19 GB (Nemotron 3 Nano 30B-A3B) to 11.0 GB (Mistral Medium 3.5 128B) across 63 open models, a 58.7-fold difference.
  2. Parameter count does not predict the cache: Nemotron 3 Nano 30B-A3B (31.6B parameters) needs 0.19 GB of KV cache at 32K, while Granite 4.2 30B (29.3B) needs 8.0 GB, 43 times more.
  3. 34 of 64 open models fit a 24 GB GPU at Q4_K_M with 32K context; the largest is Ornith 1.5 35B-A3B at 23.5 GB.
  4. A 48 GB GPU holds no larger model than a 32 GB one at Q4_K_M with 32K: both top out at K2-Horizon MoVA 36B-A4B, which gains context (37K to 115K tokens) rather than parameters; Llama 3.1 70B needs 55.2 GB.
  5. A 128 GB Mac, with 96 GB usable by the GPU, runs 44 of the 64 models at Q4_K_M with 32K, up to Mistral Medium 3.5 128B (127.7B, 91.8 GB).

How to cite

Plain text: Data: ModelVRAM (modelvram.com), "2026 Local LLM VRAM Report", https://modelvram.com/reports/local-llm-vram-2026/, data version 1.2026-09-30 (2026-09-30), CC BY 4.0.

APA: ModelVRAM. (2026). 2026 Local LLM VRAM Report (Version 1.2026-09-30) [Data set]. https://modelvram.com/reports/local-llm-vram-2026/

The data is under CC BY 4.0: copy, adapt and use it commercially, crediting "Data: ModelVRAM" with a link to this page.

The largest model each memory size holds (Q4_K_M, 32K)

Largest by total parameters; "longest context" is how far that model's context goes on the tier with 0.5 GB still free. Macs count the unified memory macOS lets the GPU use.

MemoryModels that fit Largest modelSetting VRAM at 32KLongest context
8 GB GPU 11 / 64 LensVLM 9B Q4_K_M 7.43 GB 33K
12 GB GPU 18 / 64 Gemma 4 12B Q4_K_M 8.98 GB 178K
16 GB GPU 19 / 64 gpt-oss-20b MXFP4 15.44 GB 34K
24 GB GPU 34 / 64 Ornith 1.5 35B-A3B (MoE) Q4_K_M 23.47 GB 33K
32 GB GPU 37 / 64 K2-Horizon MoVA 36B-A4B (MoE) Q4_K_M 30.31 GB 37K
48 GB GPU 37 / 64 K2-Horizon MoVA 36B-A4B (MoE)same as the smaller tier Q4_K_M 30.31 GB 115K
80 GB GPU 43 / 64 Qwen3.5 122B-A10B (MoE) Q4_K_M 78.85 GB 57K
64 GB Mac (48 GB usable) 37 / 64 K2-Horizon MoVA 36B-A4B (MoE)same as the smaller tier Q4_K_M 30.31 GB 115K
128 GB Mac (96 GB usable) 44 / 64 Mistral Medium 3.5 128B Q4_K_M 91.75 GB 41K

The 48 GB gap

A 48 GB card (and a 64 GB Mac) holds no larger model than a 32 GB card: both top out at K2-Horizon MoVA 36B-A4B (MoE), and the extra memory only takes its context from 37K to 115K tokens. The dataset has 1 model between 40B and 75B parameters (Llama 3.1 70B, 55.2 GB at Q4_K_M with 32K), and none fits 48 GB. This is a gap in the dataset's coverage, not proof that no model of that size runs on 48 GB; no figures have been made up to fill it.

KV cache at 32K context, smallest to largest

One request, FP16. The smallest, Nemotron 3 Nano 30B-A3B (MoE), needs 0.19 GB; the largest, Mistral Medium 3.5 128B, 11.0 GB: 58.7 times more. The attention layout decides it: hybrid models that keep full keys and values in only a few layers cache far less than models with full attention in every layer.

#ModelParametersKV cache at 32KAttention layout
1 Nemotron 3 Nano 30B-A3B (MoE) 31.6B3.5B active 0.19 GB 6 of its 52 layers use full attention and 46 are Mamba or feed-forward layers with no growing cache
2 Ling 3.0 Tiny 7.9B-A1.3B (MoE) 7.9B1.3B active 0.21 GB 6 of its 24 layers use multi-head latent attention and 18 are KDA recurrent layers with no growing cache
3 Nemotron 3 Super 120B-A12B (MoE) 123.6B12.0B active 0.25 GB 8 of its 88 layers use full attention and 80 are Mamba or feed-forward layers with no growing cache
4 LFM2.5 8B-A1B (MoE) 8.5B1.5B active 0.38 GB 6 of its 24 layers use full attention and 18 are short-convolution layers with no growing cache
5 GLM-5.3 Flash 321.3B18.0B active 0.39 GB 11 of its 45 layers use multi-head latent attention and 34 are linear-attention layers with no growing cache
6 Limite 1B Violetto 1.0B 0.44 GB 12 of its 48 layers use full attention and 36 keep a sliding window of 1,025 tokens
7 Nemotron 3 Nano 4B 4.0B 0.50 GB 4 of its 42 layers use full attention and 38 are Mamba or feed-forward layers with no growing cache
8 Muse Glimmer 30B 29.8B 0.50 GB 13 of its 52 layers use full attention and 39 keep a sliding window of 2,048 tokens
9 Gemma 4 E4B 8.0B 0.54 GB 4 of its 42 layers use full attention, 20 keep a sliding window of 512 tokens and 18 reuse the cache of earlier layers
10 Ornith 1.0 35B (MoE) 35.1B3.0B active 0.63 GB 10 of its 40 layers use full attention and 30 are linear-attention layers with no growing cache
11 Ornith 1.5 35B-A3B (MoE) 36.0B3.0B active 0.63 GB 10 of its 40 layers use full attention and 30 are linear-attention layers with no growing cache
12 Qwen3.6 35B-A3B (MoE) 36.0B3.0B active 0.63 GB 10 of its 40 layers use full attention and 30 are linear-attention layers with no growing cache
13 AliceAI Foundation 80B-A3B (MoE) 81.3B3.0B active 0.75 GB 12 of its 48 layers use full attention and 36 are KDA recurrent layers with no growing cache
14 Qwen3-Coder-Next (80B MoE) 79.7B3.0B active 0.75 GB 12 of its 48 layers use full attention and 36 are linear-attention layers with no growing cache
15 Qwen3.5 122B-A10B (MoE) 125.1B10.0B active 0.75 GB 12 of its 48 layers use full attention and 36 are linear-attention layers with no growing cache
16 Qwen3.8 Flash Next (180B MoE) 180.0B6.0B active 0.75 GB 12 of its 48 layers use full attention and 36 are linear-attention layers with no growing cache
17 gpt-oss-20b 20.9B3.6B active 0.77 GB 12 of its 24 layers use full attention and 12 keep a sliding window of 128 tokens
18 Kimi K3 2.78T104.0B active 0.84 GB 24 of its 93 layers use multi-head latent attention and 69 are linear-attention layers with no growing cache
19 MiMo V2.6 Flash 310.8B15.0B active 0.85 GB 9 of its 48 layers use full attention and 39 keep a sliding window of 128 tokens
20 Gemma 4 26B-A4B (MoE) 25.8B4.0B active 0.92 GB 5 of its 30 layers use full attention and 25 keep a sliding window of 1,024 tokens
21 Gemma 4 12B 12.0B 0.97 GB 8 of its 48 layers use full attention and 40 keep a sliding window of 1,024 tokens
22 LensVLM 9B 9.4B 1.00 GB 8 of its 32 layers use full attention and 24 are linear-attention layers with no growing cache
23 MiMo V2.6 Distill Qwen 9B 9.4B 1.00 GB 8 of its 32 layers use full attention and 24 are linear-attention layers with no growing cache
24 Ornith 1.0 9B 9.4B 1.00 GB 8 of its 32 layers use full attention and 24 are linear-attention layers with no growing cache
25 Ornith 1.5 9B 9.7B 1.00 GB 8 of its 32 layers use full attention and 24 are linear-attention layers with no growing cache
26 Qwen3.5 9B 9.7B 1.00 GB 8 of its 32 layers use full attention and 24 are linear-attention layers with no growing cache
27 ZDTaichu 5.0 9B 9.8B 1.00 GB 8 of its 32 layers use full attention and 24 are linear-attention layers with no growing cache
28 gpt-oss-120b 116.8B5.1B active 1.15 GB 18 of its 36 layers use full attention and 18 keep a sliding window of 128 tokens
29 Spark-X2.5 4B 4.1B 1.23 GB 9 of its 36 layers use full attention and 27 keep a sliding window of 512 tokens
30 MiniCPM5 2B 2.5B 1.31 GB All 42 layers use full attention
31 Xing 4.0 29B-A4B (MoE) 31.2B4.0B active 1.41 GB All 40 layers use multi-head latent attention
32 Step 3.7 Flash 196B-A11B (MoE) 201.4B11.0B active 1.63 GB 12 of its 45 layers use full attention and 33 keep a sliding window of 512 tokens
33 GLM-4.7 Flash 31.2B3.0B active 1.65 GB All 47 layers use multi-head latent attention
34 MiMo V2.6 Pro 1.02T42.0B active 1.78 GB 10 of its 70 layers use full attention and 60 keep a sliding window of 128 tokens
35 Hemmingway-1 27B 27.3B 2.00 GB 16 of its 64 layers use full attention and 48 are linear-attention layers with no growing cache
36 LLM-jp-4.1 32B-A3B Thinking (MoE) 32.1B3.8B active 2.00 GB All 32 layers use full attention
37 Qwen3.6 27B 27.8B 2.00 GB 16 of its 64 layers use full attention and 48 are linear-attention layers with no growing cache
38 Qwen3.8 27B 27.8B 2.00 GB 16 of its 64 layers use full attention and 48 are linear-attention layers with no growing cache
39 ThinkingCap Qwen3.8 27B 27.8B 2.00 GB 16 of its 64 layers use full attention and 48 are linear-attention layers with no growing cache
40 DeepSeek V3 / R1 (671B) 684.5B37.0B active 2.14 GB All 61 layers use multi-head latent attention
41 DeepSeek V3.2 685.4B37.0B active 2.38 GB All 61 layers use multi-head latent attention
42 DeepSeek V4.1 Flash 763.2B16.0B active 2.50 GB All 40 layers use full attention
43 Granite 4.2 3B 3.7B 2.50 GB All 40 layers use full attention
44 DeepSeek V4 Flash 290.9B13.0B active 2.69 GB All 43 layers use full attention
45 DeepSeek V4 Flash 0731 304.2B13.0B active 2.69 GB All 43 layers use full attention
46 GLM-5.2 753.3B40.0B active 2.82 GB All 78 layers use multi-head latent attention
47 GLM-5.3 753.3B40.0B active 2.82 GB All 78 layers use multi-head latent attention
48 Hy4 Preview 770B-A49B (MoE) 780.0B49.0B active 2.82 GB All 78 layers use multi-head latent attention
49 Qwen3.8 2.4T-A95B (MoE) 2.45T95.0B active 2.88 GB 23 of its 92 layers use full attention and 69 are linear-attention layers with no growing cache
50 Qwen3 30B-A3B (MoE) 30.5B3.3B active 3.00 GB All 48 layers use full attention
51 Qwen3-Coder 30B-A3B (MoE) 30.5B3.3B active 3.00 GB All 48 layers use full attention
52 Gemma 4 31B 31.3B 3.67 GB 10 of its 60 layers use full attention and 50 keep a sliding window of 1,024 tokens
53 MiniMax M3 427.0B23.0B active 3.75 GB All 60 layers use full attention
54 DeepSeek V4 Pro 1.60T49.0B active 3.81 GB All 61 layers use full attention
55 Llama 3.1 8B 8.0B 4.00 GB All 32 layers use full attention
56 IQuest-Q1 320B-A15B (MoE) 320.3B15.0B active 4.23 GB 25 of its 88 layers use full attention and 63 keep a sliding window of 4,096 tokens
57 Qwen3 8B 8.2B 4.50 GB All 36 layers use full attention
58 Granite 4.2 8B 8.8B 5.00 GB All 40 layers use full attention
59 K2-Horizon MoVA 36B-A4B (MoE) 37.4B4.0B active 6.00 GB All 48 layers use full attention
60 MiniMax M2.7 228.7B10.0B active 7.75 GB All 62 layers use full attention
61 Granite 4.2 30B 29.3B 8.00 GB All 64 layers use full attention
62 Llama 3.1 70B 70.6B 10.00 GB All 80 layers use full attention
63 Mistral Medium 3.5 128B 127.7B 11.00 GB All 88 layers use full attention

Download and embed

To put the chart in your article, copy this code (the image links back to this report, with its credit):

<a href="https://modelvram.com/reports/local-llm-vram-2026/"><img src="https://modelvram.com/reports/local-llm-vram-2026/kv-cache-32k.svg" width="760" height="1226" alt="KV cache at 32K context for 63 open LLMs, from 0.19 GB to 11.0 GB" loading="lazy"></a>
<p>Source: <a href="https://modelvram.com/reports/local-llm-vram-2026/">ModelVRAM 2026 Local LLM VRAM Report</a> (CC BY 4.0)</p>

KV cache at 32K context for 63 open LLMs, from 0.19 GB to 11.0 GB

Method and limits

Each model's layers, attention layout and parameter count are read from its repository's config.json and file sizes on Hugging Face. Total = weights + FP16 KV cache + 0.5 GB + 10% of weights and cache. The KV cache follows the real layout: full-attention layers grow with context, sliding-window layers keep only their window, MLA stores a compressed latent, and linear-attention or state-space layers do not grow.

Limits: one request only; vLLM's gpu_memory_utilization reservation, CUDA graphs and vision encoders are left out; these are estimates, not measured peaks. See predicted vs measured for how they compare with 20 public llama.cpp and vLLM measurements, and the LLM VRAM Calculator for any other setting.

Updates: this page is generated with the open dataset and changes when the data does; cite the data version above so a reader can match your numbers. First published .