How accurate are ModelVRAM's estimates? Predicted vs measured

Across 19 public measurements the median error is 1.2%. The KV cache formula matched llama.cpp's logs exactly in 8 of 8 cases, weights were 1.5% off (median), and whole-card totals came out 5.6%–8.1% high, because 0.5 GB + 10% overhead is conservative for one llama.cpp request.

Error summary

Compared Rows Median |error| Over Under Within 0.5%
All rows 19 1.2% 6 4 9
KV cache 8 0.0% 0 0 8
Weights 7 1.5% 2 4 1
Whole card 4 6.4% 4 0 0

Error = (predicted − measured) ÷ measured; positive means we predict more. GB and GiB both mean 1024³ bytes and MiB 1024² bytes, as in llama.cpp's logs.

Every comparison

Predictions are computed when the site is built, with the calculators' own functions: kvCacheBytes for the KV cache, weightBytes for weights and estimate (0.5 GB + 10% overhead) for the whole card, one request (the Gemma 4 26B row uses the 4 slots in its log). The context is the number of cache cells llama.cpp allocated, which is the context rounded up to a multiple of 256.

Model Part Setup Predicted Measured Error Why Source
Mistral Small 3.1 24B KV cache Q4_K_M, 20,736 cells, f16 KV
koboldcpp 1.88, RTX 4090 24GB
3,240 MiB 3,240 MiB
CUDA0 KV buffer size = 3240.00 MiB
0.0% Plain grouped-query attention: 2 × 40 layers × 8 heads × 128 values × 2 bytes per token. github.com
2025-04-13
Mistral Small 3.1 24B KV cache Q4_K_M, 57,600 cells, q4_0 KV
koboldcpp 1.88, RTX 4090 24GB
2,531.3 MiB 2,531.3 MiB
CUDA0 KV buffer size = 2531.25 MiB
0.0% q4_0 stores 32 values in 18 bytes, the 4.5 bits per value used here. github.com
2025-04-13
Gemma 2 9B KV cache 3,072 tokens, f16 KV, GTX 1070
llama.cpp (Oct 2024), GTX 1070 8GB
1,008 MiB 1,008 MiB
CUDA0 KV buffer size = 1008.00 MiB
0.0% Shorter than Gemma 2’s 4,096-token window, so every layer keeps every token. github.com
2024-10-28
Llama 3.1 8B KV cache Q4_K_M, -c 132000 (132,096 cells), f16 KV
llama.cpp b7583, DGX Spark
16,512 MiB 16,512 MiB
CUDA0 KV buffer size = 16512.00 MiB
0.0% 128 KiB per token, exactly as the formula gives. github.com
2025-12-30
Llama 3.1 70B KV cache Q4_K_L, 131,072 tokens, q4_0 KV
llama.cpp b3889 (ROCm), AMD multi-GPU (ROCm)
11,520 MiB 11,520 MiB
KV self size = 11520.00 MiB, K (q4_0): 5760.00 MiB, V (q4_0): 5760.00 MiB
0.0% Same formula at 4.5 bits per value. github.com
2024-10-06
gpt-oss-20b KV cache MXFP4, 32,768 tokens, f16 KV
llama.cpp b6360, NVIDIA (CUDA)
786 MiB 786 MiB
KV buffer size = 768.00 MiB (32,768 cells) + 18.00 MiB (SWA, 768 cells)
0.0% The 12 full-attention layers always matched. The sliding-window layers get 768 cells (128-token window plus a 512-token batch, rounded up to 256), and we have counted them that way since 2026-09-29; before that we counted 128 cells and came out 15 MiB (1.9%) low. github.com
2025-09-04
Gemma 4 12B KV cache QAT q4_0, 4,096 tokens, q8_0 KV
Ollama 0.30.6, RTX 4060 Laptop 8GB
289 MiB 289 MiB
non-SWA 34.00 MiB (4,096 cells) + SWA 255.00 MiB (1,536 cells), q8_0
0.0% Exact since 2026-09-29. Before that we undercounted by 35%, in two places: llama.cpp stores the values of Gemma 4’s global layers separately, where we counted the keys once, and it gives each sliding layer 1,536 cells (1,024-token window plus a 512-token batch), where we counted 1,024. github.com
2026-06-07
Gemma 4 26B-A4B KV cache Q4_K_M, 248,320 tokens shared by 4 slots, f16 KV
llama.cpp b8648, RTX 4090 24GB
5,750 MiB 5,750 MiB
non-SWA 4850.00 MiB (248,320 cells) + SWA 900.00 MiB (4,608 cells)
0.0% Exact since 2026-09-29, counted as 4 slots of 62,080 cells in one unified cache. Before that we undercounted by 54%: the global layers took exactly twice our figure, and 4 parallel slots give the sliding layers 4 × 1,024 + 512 cells, where we counted 1,024. github.com
2026-04-04
Llama 3.1 8B Weights bartowski Q4_K_M GGUF
llama.cpp b7583, DGX Spark
4,633.2 MiB 4,689.9 MiB
file size = 4.58 GiB (4.89 BPW)
−1.2% The file averages 4.89 bits per weight, a little above the 4.84 we use. github.com
2025-12-30
Llama 3.1 70B Weights Q4_K_L GGUF (Q4_K_M with Q8_0 embeddings)
llama.cpp b3889 (ROCm), AMD multi-GPU (ROCm)
40,707.6 MiB 41,287.7 MiB
model size = 40.32 GiB (4.91 BPW)
−1.4% Compared with our Q4_K_M; the Q8_0 embedding and output layers add about 1%. github.com
2024-10-06
Mistral Small 3.1 24B Weights openfree Q4_K_M GGUF
koboldcpp 1.88, RTX 4090 24GB
13,600.6 MiB 13,662.4 MiB
CUDA0 model buffer size = 13302.36 MiB + CPU_Mapped 360.00 MiB
−0.5% Within half a percent. github.com
2025-04-13
gpt-oss-20b Weights gpt-oss-20b-mxfp4.gguf
llama.cpp b6360, NVIDIA (CUDA)
13,123.8 MiB 11,540.5 MiB
file size = 11.27 GiB (4.63 BPW)
+13.7% We count the published checkpoint: MXFP4 experts plus 1.8B BF16 weights. The GGUF stores part of the BF16 in smaller types; its 579M-parameter embedding table, for one, is Q8_0 (the 586.82 MiB CPU buffer in the same log). github.com
2025-09-04
gpt-oss-120b Weights openai/gpt-oss-120b, MXFP4
vLLM 0.14.1, H100 80GB
62,226.1 MiB 65,925.1 MiB
Model loading took 64.38 GiB memory
−5.6% vLLM reports the GPU memory in use after loading, which is more than the checkpoint’s bytes; we count the checkpoint. github.com
2026-01-27
Gemma 4 26B-A4B Weights ggml-org Q4_K_M GGUF
llama.cpp b8648, RTX 4090 24GB
14,889.3 MiB 16,005.1 MiB
file size = 15.63 GiB (5.32 BPW)
−7.0% Q4_K_M keeps some tensors in higher-precision types, and in this model they are a larger share: the file averages 5.32 bits per weight against the 4.84 we use. github.com
2026-04-04
Gemma 4 31B Weights Q5_K_M GGUF
llama.cpp b9031, 2 NVIDIA GPUs (CUDA)
21,138 MiB 20,817.9 MiB
file size = 20.33 GiB (5.69 BPW)
+1.5% Bits per weight agree (5.69 vs 5.67); our parameter count includes the vision encoder, which this GGUF leaves out. github.com
2026-05-06
Mistral Small 3.1 24B Whole card Q4_K_M, f16 KV, 20,736 cells, flash attention
koboldcpp 1.88, RTX 4090 24GB
18.59 GB 17.60 GB
19.2 GB used − 1.6 GB desktop
+5.6% Weights and cache match the logs; the gap is our 0.5 GB + 10% overhead, about 2–3 times what llama.cpp added on top of them for one request. github.com
2025-04-13
Mistral Small 3.1 24B Whole card Q4_K_M, q4_0 KV, 20,736 cells, flash attention
koboldcpp 1.88, RTX 4090 24GB
16.09 GB 15.20 GB
16.8 GB used − 1.6 GB desktop
+5.8% Weights and cache match the logs; the gap is our 0.5 GB + 10% overhead, about 2–3 times what llama.cpp added on top of them for one request. github.com
2025-04-13
Mistral Small 3.1 24B Whole card Q4_K_M, f16 KV, 41,216 cells, flash attention
koboldcpp 1.88, RTX 4090 24GB
22.03 GB 20.60 GB
22.2 GB used − 1.6 GB desktop
+6.9% Weights and cache match the logs; the gap is our 0.5 GB + 10% overhead, about 2–3 times what llama.cpp added on top of them for one request. github.com
2025-04-13
Mistral Small 3.1 24B Whole card Q4_K_M, q4_0 KV, 57,600 cells, flash attention
koboldcpp 1.88, RTX 4090 24GB
17.83 GB 16.50 GB
18.1 GB used − 1.6 GB desktop
+8.1% Weights and cache match the logs; the gap is our 0.5 GB + 10% overhead, about 2–3 times what llama.cpp added on top of them for one request. github.com
2025-04-13

The four whole-card rows are the VRAM in use the reporter read, minus the 1.6 GB their desktop already took. Mistral Small 3.1 is counted at 23.57B parameters without its vision encoder, like the text-only GGUF in the report; Gemma 2 9B's shape is from its config.json. The other models use the calculator's presets.

Image models

The Qwen-Image-2.1 calculator gives a range: exact file sizes plus the working memory seen in real runs.

GPU Setup Predicted range Measured Result Source
RTX 3090 24GB Qwen-Image-2.1: DiT GGUF Q4_K_M + encoder INT8 ConvRot + VAE BF16, all on the GPU, 1024×1024
ComfyUI
15.35 GB–17.95 GB 15.39 GB Inside github.com
2026-09-21
Apple M5 Max 48GB Qwen-Image-2.1: DiT GGUF Q4_K_M; encoder, VAE and image size not stated, compared as BF16 encoder + FP32 VAE, all in unified memory, 1 MP
Unsloth Desktop (macOS)
23.60 GB–26.20 GB 42.00 GB Outside (measured higher) github.com
2026-09-29

RTX 3090 24GB: Inside the range, at its low end. This run is one of those the range was set from, so it checks consistency more than accuracy.

Apple M5 Max 48GB: Outside the range: the measurement is far above it, even with the largest encoder and VAE. The ~34 GB after loading is close to the files with the DiT held at BF16 (32.4 decimal GB with a BF16 encoder and VAE) rather than the Q4_K_M file (22.4), so this path likely dequantizes the DiT, which the calculator does not model. The resolution is unknown and the engine has not confirmed the dtype; read as decimal GB, 42 GB is still far above the range.

More runs, and low-VRAM setups for 8–16 GB, are on the Qwen-Image 2.1 VRAM calculator. MiniMax H3's public runs mostly stream weights, so the card being full says how big the card is, not what the job needs; they are not compared here.

vLLM's KV cache pool is counted in tokens rather than bytes, so its checks are on the vLLM KV cache calculator's log table: 6 public startup logs.

Where we overestimate and underestimate

What we'd change

On 2026-09-29 the KV cache formula was fixed: Gemma 4's global layers count K and V separately, and sliding-window layers get llama.cpp's cell count, window × slots + 512 rounded up to 256, and never more than the context (src/llama-kv-cache-iswa.cpp). The table above uses the fixed formula; before it we predicted 187 MiB for Gemma 4 12B, 2,625 MiB for 26B and 771 MiB for gpt-oss-20b. The rest has not changed yet: this page shows how the current formulas do, and when they change the table is recomputed with them at build time.

Reports we could not use

Two 2026 reports of Qwen3.8 27B on 16 GB cards (autodidacts.io, jrell's IQ4_XS mix) give the quant, context and KV type but no memory reading, only that it runs. They do not contradict the estimates, but there is no error to compute. Runs with an nvidia-smi reading or llama.cpp buffer lines are welcome; see About for contact.

Submit a measurement

Ran a model yourself? Submit a measurement on GitHub: the form asks for the model, file or quant, engine and version, GPU, -c, parallel slots and KV type, and the log lines as printed (llama.cpp's KV buffer size and model buffer size, an nvidia-smi reading, or vLLM's KV cache lines). Each one is checked by hand, then added to the table above next to the build-time prediction.

Frequently asked questions

How accurate is the ModelVRAM VRAM calculator?

Across 19 public measurements the median error is 1.2%. The KV cache formula matched llama.cpp's logs exactly in 8 of 8 cases, weights were 1.5% off (median), and whole-card totals came out 5.6%–8.1% high, because 0.5 GB + 10% overhead is conservative for one llama.cpp request.

Does the calculator overestimate or underestimate?

Whole-card totals: every one over, by 5.6% to 8.1%. Under: GGUF files that keep more tensors at high precision (Gemma 4 26B Q4_K_M −7.0%) and vLLM's loaded gpt-oss-120b (−5.6%). No KV row is under now. Before 2026-09-29 we undercounted Gemma 4's KV cache in llama.cpp (−35.3% and −54.3%) and gpt-oss's sliding-window cache (15 MiB short); both were fixed that day.

Why does the estimate differ from nvidia-smi?

nvidia-smi shows the whole card: the desktop, the CUDA context and whatever the engine reserves. vLLM fills memory up to gpu_memory_utilization, and llama.cpp allocates the KV cache for the full context at startup. This page compares only what lines up: the KV and model buffers in the logs, and whole-card readings with the desktop's share taken off.

How were the measurements chosen?

Only public reports with a source link and enough settings to reproduce them (model, quantized file, context, KV type, engine). Every source was reopened and re-read on 2026-09-29. Reports that give settings but no memory reading are left out.

Sources checked .