Open LLM VRAM dataset
VRAM estimates for 55 open models, generated at build time by the same functions as the calculator, the model pages and the report.
- modelvram-vram.csv (one row per model)
- modelvram-vram.json (the same rows, plus column notes, units and assumptions)
Version 1.2026-09-29b; model data last checked .
Quotable figures computed from these rows, such as how many models fit a 24 GB GPU, are on the LLM VRAM statistics page.
Each model also has a static VRAM badge for its Hugging Face README.
License and attribution
The data is released under CC BY 4.0: copy, adapt and use it commercially, with attribution:
Data: ModelVRAM (modelvram.com), linking to https://modelvram.com/
Units
Columns ending in _gib are in GiB (1024³ bytes); the site writes that as GB. Parameters are counts, contexts are tokens and kv_bytes_per_token_fp16 is bytes. Values a model does not have, such as totals past its longest context, are empty.
Method
Weight memory is total parameters × bits per parameter ÷ 8 (Q4_K_M 4.84, Q8_0 8.5, FP16 / BF16 16 bits; mixture-of-experts models count every expert). The KV cache follows each model's attention layout, at FP16 for one request: full attention grows with context, sliding-window layers keep only their window, MLA stores a compressed latent, and linear or state-space layers add no context-growing cache. Total = weights + KV cache + 0.5 GB + 10% of weights and cache. The "fits" columns say whether that total at 32,768 tokens fits one GPU of the given size. Contexts are 8,192, 32,768, 131,072 tokens. gpt-oss is at each column's bits per parameter too; it is published in MXFP4 and its GGUF quants keep the experts in MXFP4, so its real files are all about the MXFP4 size, and the "Can I run", GPU and VRAM tier pages list it at MXFP4 only, with at least 0.5 GB left free.
These are estimates, not measured peaks; vLLM reservations, CUDA graphs and vision encoders can need more. See the KV cache explanation, and predicted vs measured for how they compare with public measurements. Measured a model yourself? Submit a measurement.
Sources
Model shapes and parameter counts come from each repository on Hugging Face, its config.json and file sizes, read on the date in source_checked. The calculation code and tests are on GitHub; the data is also on Hugging Face Datasets.
JSON API
The same numbers are served as static JSON, built with the site: no key, any origin (Access-Control-Allow-Origin: *), CC BY 4.0 like the files above. Every response carries version (the data version above), a checked date, license and source. It covers the 55 models and 34 GPU setups on the site; for any other Hugging Face model, use the calculator.
/api/v1/models.json: every model: parameters, architecture, KV bytes per token at f16 and q8_0, weights per precision (GiB) and links/api/v1/models/<id>.json: one model: totals per precision at 4K, 8K, 32K, 128K and its longest context with f16 and q8_0 KV, checked GGUF repos, the smallest setup and a "Can I run" verdict per single GPU/api/v1/gpus.json: every GPU and multi-GPU setup/api/v1/gpus/<id>.json: the models one setup holds (best precision at 8K, total, longest context, speed) and the ones it does not
curl -s https://modelvram.com/api/v1/models/qwen3.6-35b-a3b.json | jq '.precisions[] | select(.id=="q4_k_m") | .totals_gib.f16'
curl -s https://modelvram.com/api/v1/gpus/rtx-4090.json | jq '.fits[] | {name, best_precision, total_gib}' IDs are the page slugs, such as qwen3.6-35b-a3b, rtx-4090 or 2x-rtx-3090. A site overview for AI assistants is at /llms.txt, and the calculators' URL parameters are in /agent.md. Within v1, fields may be added but never removed or changed in meaning.
Columns
| Column | Unit | Description |
|---|---|---|
model | Model name as shown on the site | |
hf_id | Hugging Face repository ID | |
hf_url | URL | Source repository on Hugging Face |
model_url | URL | ModelVRAM page with the full calculations |
source_checked | date | When config.json and file sizes were read |
total_params | parameters | All parameters, every expert included |
active_params | parameters | Parameters used per token (MoE); empty for dense models |
experts | count | Routed experts (MoE); empty for dense models |
layers | count | Transformer layers |
max_context | tokens | Longest context in the model config |
kv_bytes_per_token_fp16 | bytes | KV cache added per token of one request at FP16, once sliding windows are full |
attention_layout | Layers that decide the KV cache | |
q4_k_m_weights_gib | GiB | Weight memory at Q4_K_M |
q8_0_weights_gib | GiB | Weight memory at Q8_0 |
fp16_weights_gib | GiB | Weight memory at FP16 / BF16 |
q4_k_m_8k_total_gib | GiB | Total VRAM at Q4_K_M, 8,192 tokens: weights + KV cache + overhead; empty past max_context |
q4_k_m_32k_total_gib | GiB | Total VRAM at Q4_K_M, 32,768 tokens: weights + KV cache + overhead; empty past max_context |
q4_k_m_128k_total_gib | GiB | Total VRAM at Q4_K_M, 131,072 tokens: weights + KV cache + overhead; empty past max_context |
q8_0_8k_total_gib | GiB | Total VRAM at Q8_0, 8,192 tokens: weights + KV cache + overhead; empty past max_context |
q8_0_32k_total_gib | GiB | Total VRAM at Q8_0, 32,768 tokens: weights + KV cache + overhead; empty past max_context |
q8_0_128k_total_gib | GiB | Total VRAM at Q8_0, 131,072 tokens: weights + KV cache + overhead; empty past max_context |
fp16_8k_total_gib | GiB | Total VRAM at FP16 / BF16, 8,192 tokens: weights + KV cache + overhead; empty past max_context |
fp16_32k_total_gib | GiB | Total VRAM at FP16 / BF16, 32,768 tokens: weights + KV cache + overhead; empty past max_context |
fp16_128k_total_gib | GiB | Total VRAM at FP16 / BF16, 131,072 tokens: weights + KV cache + overhead; empty past max_context |
kv_cache_8k_gib | GiB | FP16 KV cache of one request at 8,192 tokens |
kv_cache_32k_gib | GiB | FP16 KV cache of one request at 32,768 tokens |
kv_cache_128k_gib | GiB | FP16 KV cache of one request at 131,072 tokens |
q4_k_m_32k_fits_16gib | boolean | Q4_K_M with 32,768 tokens fits on one 16 GB GPU |
q4_k_m_32k_fits_24gib | boolean | Q4_K_M with 32,768 tokens fits on one 24 GB GPU |
q4_k_m_32k_fits_32gib | boolean | Q4_K_M with 32,768 tokens fits on one 32 GB GPU |
q4_k_m_32k_fits_48gib | boolean | Q4_K_M with 32,768 tokens fits on one 48 GB GPU |
q4_k_m_32k_fits_80gib | boolean | Q4_K_M with 32,768 tokens fits on one 80 GB GPU |
q8_0_32k_fits_16gib | boolean | Q8_0 with 32,768 tokens fits on one 16 GB GPU |
q8_0_32k_fits_24gib | boolean | Q8_0 with 32,768 tokens fits on one 24 GB GPU |
q8_0_32k_fits_32gib | boolean | Q8_0 with 32,768 tokens fits on one 32 GB GPU |
q8_0_32k_fits_48gib | boolean | Q8_0 with 32,768 tokens fits on one 48 GB GPU |
q8_0_32k_fits_80gib | boolean | Q8_0 with 32,768 tokens fits on one 80 GB GPU |
fp16_32k_fits_16gib | boolean | FP16 / BF16 with 32,768 tokens fits on one 16 GB GPU |
fp16_32k_fits_24gib | boolean | FP16 / BF16 with 32,768 tokens fits on one 24 GB GPU |
fp16_32k_fits_32gib | boolean | FP16 / BF16 with 32,768 tokens fits on one 32 GB GPU |
fp16_32k_fits_48gib | boolean | FP16 / BF16 with 32,768 tokens fits on one 48 GB GPU |
fp16_32k_fits_80gib | boolean | FP16 / BF16 with 32,768 tokens fits on one 80 GB GPU |
note | Known cache optimization the estimate leaves out |