Open LLM VRAM dataset

VRAM estimates for 55 open models, generated at build time by the same functions as the calculator, the model pages and the report.

Version 1.2026-09-29b; model data last checked .

Quotable figures computed from these rows, such as how many models fit a 24 GB GPU, are on the LLM VRAM statistics page.

Each model also has a static VRAM badge for its Hugging Face README.

License and attribution

The data is released under CC BY 4.0: copy, adapt and use it commercially, with attribution:

Data: ModelVRAM (modelvram.com), linking to https://modelvram.com/data/.

Units

Columns ending in _gib are in GiB (1024³ bytes); the site writes that as GB. Parameters are counts, contexts are tokens and kv_bytes_per_token_fp16 is bytes. Values a model does not have, such as totals past its longest context, are empty.

Method

Weight memory is total parameters × bits per parameter ÷ 8 (Q4_K_M 4.84, Q8_0 8.5, FP16 / BF16 16 bits; mixture-of-experts models count every expert). The KV cache follows each model's attention layout, at FP16 for one request: full attention grows with context, sliding-window layers keep only their window, MLA stores a compressed latent, and linear or state-space layers add no context-growing cache. Total = weights + KV cache + 0.5 GB + 10% of weights and cache. The "fits" columns say whether that total at 32,768 tokens fits one GPU of the given size. Contexts are 8,192, 32,768, 131,072 tokens. gpt-oss is at each column's bits per parameter too; it is published in MXFP4 and its GGUF quants keep the experts in MXFP4, so its real files are all about the MXFP4 size, and the "Can I run", GPU and VRAM tier pages list it at MXFP4 only, with at least 0.5 GB left free.

These are estimates, not measured peaks; vLLM reservations, CUDA graphs and vision encoders can need more. See the KV cache explanation, and predicted vs measured for how they compare with public measurements. Measured a model yourself? Submit a measurement.

Sources

Model shapes and parameter counts come from each repository on Hugging Face, its config.json and file sizes, read on the date in source_checked. The calculation code and tests are on GitHub; the data is also on Hugging Face Datasets.

JSON API

The same numbers are served as static JSON, built with the site: no key, any origin (Access-Control-Allow-Origin: *), CC BY 4.0 like the files above. Every response carries version (the data version above), a checked date, license and source. It covers the 55 models and 34 GPU setups on the site; for any other Hugging Face model, use the calculator.

curl -s https://modelvram.com/api/v1/models/qwen3.6-35b-a3b.json | jq '.precisions[] | select(.id=="q4_k_m") | .totals_gib.f16'
curl -s https://modelvram.com/api/v1/gpus/rtx-4090.json | jq '.fits[] | {name, best_precision, total_gib}'

IDs are the page slugs, such as qwen3.6-35b-a3b, rtx-4090 or 2x-rtx-3090. A site overview for AI assistants is at /llms.txt, and the calculators' URL parameters are in /agent.md. Within v1, fields may be added but never removed or changed in meaning.

Columns

ColumnUnitDescription
modelModel name as shown on the site
hf_idHugging Face repository ID
hf_urlURLSource repository on Hugging Face
model_urlURLModelVRAM page with the full calculations
source_checkeddateWhen config.json and file sizes were read
total_paramsparametersAll parameters, every expert included
active_paramsparametersParameters used per token (MoE); empty for dense models
expertscountRouted experts (MoE); empty for dense models
layerscountTransformer layers
max_contexttokensLongest context in the model config
kv_bytes_per_token_fp16bytesKV cache added per token of one request at FP16, once sliding windows are full
attention_layoutLayers that decide the KV cache
q4_k_m_weights_gibGiBWeight memory at Q4_K_M
q8_0_weights_gibGiBWeight memory at Q8_0
fp16_weights_gibGiBWeight memory at FP16 / BF16
q4_k_m_8k_total_gibGiBTotal VRAM at Q4_K_M, 8,192 tokens: weights + KV cache + overhead; empty past max_context
q4_k_m_32k_total_gibGiBTotal VRAM at Q4_K_M, 32,768 tokens: weights + KV cache + overhead; empty past max_context
q4_k_m_128k_total_gibGiBTotal VRAM at Q4_K_M, 131,072 tokens: weights + KV cache + overhead; empty past max_context
q8_0_8k_total_gibGiBTotal VRAM at Q8_0, 8,192 tokens: weights + KV cache + overhead; empty past max_context
q8_0_32k_total_gibGiBTotal VRAM at Q8_0, 32,768 tokens: weights + KV cache + overhead; empty past max_context
q8_0_128k_total_gibGiBTotal VRAM at Q8_0, 131,072 tokens: weights + KV cache + overhead; empty past max_context
fp16_8k_total_gibGiBTotal VRAM at FP16 / BF16, 8,192 tokens: weights + KV cache + overhead; empty past max_context
fp16_32k_total_gibGiBTotal VRAM at FP16 / BF16, 32,768 tokens: weights + KV cache + overhead; empty past max_context
fp16_128k_total_gibGiBTotal VRAM at FP16 / BF16, 131,072 tokens: weights + KV cache + overhead; empty past max_context
kv_cache_8k_gibGiBFP16 KV cache of one request at 8,192 tokens
kv_cache_32k_gibGiBFP16 KV cache of one request at 32,768 tokens
kv_cache_128k_gibGiBFP16 KV cache of one request at 131,072 tokens
q4_k_m_32k_fits_16gibbooleanQ4_K_M with 32,768 tokens fits on one 16 GB GPU
q4_k_m_32k_fits_24gibbooleanQ4_K_M with 32,768 tokens fits on one 24 GB GPU
q4_k_m_32k_fits_32gibbooleanQ4_K_M with 32,768 tokens fits on one 32 GB GPU
q4_k_m_32k_fits_48gibbooleanQ4_K_M with 32,768 tokens fits on one 48 GB GPU
q4_k_m_32k_fits_80gibbooleanQ4_K_M with 32,768 tokens fits on one 80 GB GPU
q8_0_32k_fits_16gibbooleanQ8_0 with 32,768 tokens fits on one 16 GB GPU
q8_0_32k_fits_24gibbooleanQ8_0 with 32,768 tokens fits on one 24 GB GPU
q8_0_32k_fits_32gibbooleanQ8_0 with 32,768 tokens fits on one 32 GB GPU
q8_0_32k_fits_48gibbooleanQ8_0 with 32,768 tokens fits on one 48 GB GPU
q8_0_32k_fits_80gibbooleanQ8_0 with 32,768 tokens fits on one 80 GB GPU
fp16_32k_fits_16gibbooleanFP16 / BF16 with 32,768 tokens fits on one 16 GB GPU
fp16_32k_fits_24gibbooleanFP16 / BF16 with 32,768 tokens fits on one 24 GB GPU
fp16_32k_fits_32gibbooleanFP16 / BF16 with 32,768 tokens fits on one 32 GB GPU
fp16_32k_fits_48gibbooleanFP16 / BF16 with 32,768 tokens fits on one 48 GB GPU
fp16_32k_fits_80gibbooleanFP16 / BF16 with 32,768 tokens fits on one 80 GB GPU
noteKnown cache optimization the estimate leaves out