LLM VRAM statistics
15 figures, all computed at build time from the open dataset of 55 open models with the same formulas as the calculator. Each comes with its method and a line to cite it with. Data version 1.2026-09-29b; model data checked . GB means GiB.
How many open LLMs fit on a 24 GB GPU?
28 of 55
28 of 55 open models (51%) fit one 24 GB GPU at Q4_K_M with 32K tokens of context.
Method: Weights at 4.84 bits per weight, FP16 KV cache for 32,768 tokens and one request, plus 0.5 GB and 10% overhead, against 24 GiB; models with a shorter maximum context do not count as fitting.
Cite this: Source: ModelVRAM, modelvram.com/llm-vram-stats/, data version 1.2026-09-29b
How many open LLMs fit on a 16 GB GPU?
15 of 55
15 of 55 open models (27%) fit one 16 GB GPU at Q4_K_M with 32K tokens of context.
Method: The same rule as the 24 GB figure, against 16 GiB.
Cite this: Source: ModelVRAM, modelvram.com/llm-vram-stats/, data version 1.2026-09-29b
How much VRAM do you need to run half of the open LLMs?
24 GB
A GPU needs 24 GB of memory to hold at least half of the 55 models at Q4_K_M with 32K tokens of context (28 fit), as on an RTX 3090, RTX 4090 or RX 7900 XTX.
Method: The smallest memory size among the GPUs and Macs with a page on this site at which the 24 GB rule above holds for half the models or more.
Cite this: Source: ModelVRAM, modelvram.com/llm-vram-stats/, data version 1.2026-09-29b
How much VRAM does a typical open LLM need?
23.9 GB
The median open model needs 23.9 GB of VRAM at Q4_K_M with 32K tokens of context.
Method: Median of the q4_k_m_32k_total_gib column over the 55 models whose context reaches 32,768 tokens.
Cite this: Source: ModelVRAM, modelvram.com/llm-vram-stats/, data version 1.2026-09-29b
How much of a MoE model is used per token?
5.4%
The median mixture-of-experts model uses 5.4% of its parameters per token, yet all of them have to be in memory.
Method: Active parameters from the model cards divided by total parameters, median over the 36 MoE models.
Cite this: Source: ModelVRAM, modelvram.com/llm-vram-stats/, data version 1.2026-09-29b
How much less KV cache do hybrid-attention models need?
24 KB vs 121 KB
Hybrid models add a median 24 KB of FP16 KV cache per token of context, against 121 KB for models with full attention in every layer.
Method: Hybrid: 34 models with linear-attention, state-space, sliding-window or cache-sharing layers. Full attention: 14 models where every layer caches keys and values; the 7 models with MLA in every layer are in neither group. KB = 1,024 bytes, once sliding windows are full.
Cite this: Source: ModelVRAM, modelvram.com/llm-vram-stats/, data version 1.2026-09-29b
How much VRAM does 128K context add?
+5.4 GB
Going from 8K to 128K tokens of context adds a median 5.4 GB of VRAM at Q4_K_M.
Method: Median of q4_k_m_128k_total_gib minus q4_k_m_8k_total_gib over the 51 models whose context reaches 131,072 tokens; FP16 cache, one request, overhead included.
Cite this: Source: ModelVRAM, modelvram.com/llm-vram-stats/, data version 1.2026-09-29b
When is the KV cache bigger than the model weights?
6 of 51
For 6 of 51 models (12%), the FP16 KV cache of one 128K-token request is larger than the Q4_K_M weights.
Method: kv_cache_128k_gib compared with q4_k_m_weights_gib, over the models whose context reaches 131,072 tokens.
Cite this: Source: ModelVRAM, modelvram.com/llm-vram-stats/, data version 1.2026-09-29b
How many open LLMs support 256K context?
67%
37 of 55 open models (67%) accept 256K tokens of context or more.
Method: The max_context column (the longest context in config.json) at least 262,144 tokens.
Cite this: Source: ModelVRAM, modelvram.com/llm-vram-stats/, data version 1.2026-09-29b
What is the largest LLM that fits on a 24 GB GPU?
36.0B
The largest open models that fit one 24 GB GPU at Q4_K_M with 32K tokens of context are Ornith 1.5 35B-A3B and Qwen3.6 35B-A3B, tied at 36.0B parameters and 23.5 GB. Model page.
Method: Largest total parameter count among the models counted in the 24 GB figure.
Cite this: Source: ModelVRAM, modelvram.com/llm-vram-stats/, data version 1.2026-09-29b
How many open LLMs need more than one 80 GB GPU?
20 of 55
20 of 55 open models (36%) need more than one 80 GB GPU even at Q4_K_M with 32K tokens of context.
Method: The 24 GB rule against 80 GiB, at 32,768 tokens or the model's whole context if shorter.
Cite this: Source: ModelVRAM, modelvram.com/llm-vram-stats/, data version 1.2026-09-29b
How many open LLMs are too big for 80 GB at FP16?
26 of 55
26 of 55 open models (47%) have FP16 weights larger than 80 GB, before any KV cache.
Method: fp16_weights_gib above 80: total parameters × 2 bytes.
Cite this: Source: ModelVRAM, modelvram.com/llm-vram-stats/, data version 1.2026-09-29b
How big is the median open LLM?
36.0B
The median open model tracked here has 36.0B parameters, every expert counted.
Method: Median of total_params over all 55 models.
Cite this: Source: ModelVRAM, modelvram.com/llm-vram-stats/, data version 1.2026-09-29b
Where do these numbers come from?
Every model's layers, attention layout and parameter count are read from its config.json and file sizes on Hugging Face; VRAM = weights + FP16 KV cache for one request + 0.5 GB + 10% overhead. These are estimates, not measured peaks. The figures here count every model, gpt-oss included, at Q4_K_M with no headroom beyond the estimate; the VRAM tier and "Can I run" pages count gpt-oss at MXFP4 and leave 0.5 GB free, so their counts can differ slightly. The per-model numbers are in the open dataset (CSV, JSON), the models by VRAM tier on best local LLMs for 8 to 96 GB, and any other setting in the LLM VRAM Calculator. The data is released under CC BY 4.0.
Which models fit my VRAM?
Top picks, newest models and common questions for each size: best local LLMs for 8 GB, best local LLMs for 12 GB, best local LLMs for 16 GB, best local LLMs for 24 GB, best local LLMs for 32 GB, best local LLMs for 48 GB, best local LLMs for 96 GB.