How accurate are ModelVRAM's estimates? Predicted vs measured
Across 19 public measurements the median error is 1.2%. The KV cache formula matched llama.cpp's logs exactly in 8 of 8 cases, weights were 1.5% off (median), and whole-card totals came out 5.6%–8.1% high, because 0.5 GB + 10% overhead is conservative for one llama.cpp request.
Error summary
| Compared | Rows | Median |error| | Over | Under | Within 0.5% |
|---|---|---|---|---|---|
| All rows | 19 | 1.2% | 6 | 4 | 9 |
| KV cache | 8 | 0.0% | 0 | 0 | 8 |
| Weights | 7 | 1.5% | 2 | 4 | 1 |
| Whole card | 4 | 6.4% | 4 | 0 | 0 |
Error = (predicted − measured) ÷ measured; positive means we predict more. GB and GiB both mean 1024³ bytes and MiB 1024² bytes, as in llama.cpp's logs.
Every comparison
Predictions are computed when the site is built, with the calculators' own functions: kvCacheBytes for the KV cache, weightBytes for weights and estimate (0.5 GB + 10% overhead) for the whole card, one request (the Gemma 4 26B row uses the 4 slots in its log). The context is the number of cache cells llama.cpp allocated, which is the context rounded up to a multiple of 256.
| Model | Part | Setup | Predicted | Measured | Error | Why | Source |
|---|---|---|---|---|---|---|---|
| Mistral Small 3.1 24B | KV cache | Q4_K_M, 20,736 cells, f16 KV koboldcpp 1.88, RTX 4090 24GB | 3,240 MiB | 3,240 MiB CUDA0 KV buffer size = 3240.00 MiB | 0.0% | Plain grouped-query attention: 2 × 40 layers × 8 heads × 128 values × 2 bytes per token. | github.com 2025-04-13 |
| Mistral Small 3.1 24B | KV cache | Q4_K_M, 57,600 cells, q4_0 KV koboldcpp 1.88, RTX 4090 24GB | 2,531.3 MiB | 2,531.3 MiB CUDA0 KV buffer size = 2531.25 MiB | 0.0% | q4_0 stores 32 values in 18 bytes, the 4.5 bits per value used here. | github.com 2025-04-13 |
| Gemma 2 9B | KV cache | 3,072 tokens, f16 KV, GTX 1070 llama.cpp (Oct 2024), GTX 1070 8GB | 1,008 MiB | 1,008 MiB CUDA0 KV buffer size = 1008.00 MiB | 0.0% | Shorter than Gemma 2’s 4,096-token window, so every layer keeps every token. | github.com 2024-10-28 |
| Llama 3.1 8B | KV cache | Q4_K_M, -c 132000 (132,096 cells), f16 KV llama.cpp b7583, DGX Spark | 16,512 MiB | 16,512 MiB CUDA0 KV buffer size = 16512.00 MiB | 0.0% | 128 KiB per token, exactly as the formula gives. | github.com 2025-12-30 |
| Llama 3.1 70B | KV cache | Q4_K_L, 131,072 tokens, q4_0 KV llama.cpp b3889 (ROCm), AMD multi-GPU (ROCm) | 11,520 MiB | 11,520 MiB KV self size = 11520.00 MiB, K (q4_0): 5760.00 MiB, V (q4_0): 5760.00 MiB | 0.0% | Same formula at 4.5 bits per value. | github.com 2024-10-06 |
| gpt-oss-20b | KV cache | MXFP4, 32,768 tokens, f16 KV llama.cpp b6360, NVIDIA (CUDA) | 786 MiB | 786 MiB KV buffer size = 768.00 MiB (32,768 cells) + 18.00 MiB (SWA, 768 cells) | 0.0% | The 12 full-attention layers always matched. The sliding-window layers get 768 cells (128-token window plus a 512-token batch, rounded up to 256), and we have counted them that way since 2026-09-29; before that we counted 128 cells and came out 15 MiB (1.9%) low. | github.com 2025-09-04 |
| Gemma 4 12B | KV cache | QAT q4_0, 4,096 tokens, q8_0 KV Ollama 0.30.6, RTX 4060 Laptop 8GB | 289 MiB | 289 MiB non-SWA 34.00 MiB (4,096 cells) + SWA 255.00 MiB (1,536 cells), q8_0 | 0.0% | Exact since 2026-09-29. Before that we undercounted by 35%, in two places: llama.cpp stores the values of Gemma 4’s global layers separately, where we counted the keys once, and it gives each sliding layer 1,536 cells (1,024-token window plus a 512-token batch), where we counted 1,024. | github.com 2026-06-07 |
| Gemma 4 26B-A4B | KV cache | Q4_K_M, 248,320 tokens shared by 4 slots, f16 KV llama.cpp b8648, RTX 4090 24GB | 5,750 MiB | 5,750 MiB non-SWA 4850.00 MiB (248,320 cells) + SWA 900.00 MiB (4,608 cells) | 0.0% | Exact since 2026-09-29, counted as 4 slots of 62,080 cells in one unified cache. Before that we undercounted by 54%: the global layers took exactly twice our figure, and 4 parallel slots give the sliding layers 4 × 1,024 + 512 cells, where we counted 1,024. | github.com 2026-04-04 |
| Llama 3.1 8B | Weights | bartowski Q4_K_M GGUF llama.cpp b7583, DGX Spark | 4,633.2 MiB | 4,689.9 MiB file size = 4.58 GiB (4.89 BPW) | −1.2% | The file averages 4.89 bits per weight, a little above the 4.84 we use. | github.com 2025-12-30 |
| Llama 3.1 70B | Weights | Q4_K_L GGUF (Q4_K_M with Q8_0 embeddings) llama.cpp b3889 (ROCm), AMD multi-GPU (ROCm) | 40,707.6 MiB | 41,287.7 MiB model size = 40.32 GiB (4.91 BPW) | −1.4% | Compared with our Q4_K_M; the Q8_0 embedding and output layers add about 1%. | github.com 2024-10-06 |
| Mistral Small 3.1 24B | Weights | openfree Q4_K_M GGUF koboldcpp 1.88, RTX 4090 24GB | 13,600.6 MiB | 13,662.4 MiB CUDA0 model buffer size = 13302.36 MiB + CPU_Mapped 360.00 MiB | −0.5% | Within half a percent. | github.com 2025-04-13 |
| gpt-oss-20b | Weights | gpt-oss-20b-mxfp4.gguf llama.cpp b6360, NVIDIA (CUDA) | 13,123.8 MiB | 11,540.5 MiB file size = 11.27 GiB (4.63 BPW) | +13.7% | We count the published checkpoint: MXFP4 experts plus 1.8B BF16 weights. The GGUF stores part of the BF16 in smaller types; its 579M-parameter embedding table, for one, is Q8_0 (the 586.82 MiB CPU buffer in the same log). | github.com 2025-09-04 |
| gpt-oss-120b | Weights | openai/gpt-oss-120b, MXFP4 vLLM 0.14.1, H100 80GB | 62,226.1 MiB | 65,925.1 MiB Model loading took 64.38 GiB memory | −5.6% | vLLM reports the GPU memory in use after loading, which is more than the checkpoint’s bytes; we count the checkpoint. | github.com 2026-01-27 |
| Gemma 4 26B-A4B | Weights | ggml-org Q4_K_M GGUF llama.cpp b8648, RTX 4090 24GB | 14,889.3 MiB | 16,005.1 MiB file size = 15.63 GiB (5.32 BPW) | −7.0% | Q4_K_M keeps some tensors in higher-precision types, and in this model they are a larger share: the file averages 5.32 bits per weight against the 4.84 we use. | github.com 2026-04-04 |
| Gemma 4 31B | Weights | Q5_K_M GGUF llama.cpp b9031, 2 NVIDIA GPUs (CUDA) | 21,138 MiB | 20,817.9 MiB file size = 20.33 GiB (5.69 BPW) | +1.5% | Bits per weight agree (5.69 vs 5.67); our parameter count includes the vision encoder, which this GGUF leaves out. | github.com 2026-05-06 |
| Mistral Small 3.1 24B | Whole card | Q4_K_M, f16 KV, 20,736 cells, flash attention koboldcpp 1.88, RTX 4090 24GB | 18.59 GB | 17.60 GB 19.2 GB used − 1.6 GB desktop | +5.6% | Weights and cache match the logs; the gap is our 0.5 GB + 10% overhead, about 2–3 times what llama.cpp added on top of them for one request. | github.com 2025-04-13 |
| Mistral Small 3.1 24B | Whole card | Q4_K_M, q4_0 KV, 20,736 cells, flash attention koboldcpp 1.88, RTX 4090 24GB | 16.09 GB | 15.20 GB 16.8 GB used − 1.6 GB desktop | +5.8% | Weights and cache match the logs; the gap is our 0.5 GB + 10% overhead, about 2–3 times what llama.cpp added on top of them for one request. | github.com 2025-04-13 |
| Mistral Small 3.1 24B | Whole card | Q4_K_M, f16 KV, 41,216 cells, flash attention koboldcpp 1.88, RTX 4090 24GB | 22.03 GB | 20.60 GB 22.2 GB used − 1.6 GB desktop | +6.9% | Weights and cache match the logs; the gap is our 0.5 GB + 10% overhead, about 2–3 times what llama.cpp added on top of them for one request. | github.com 2025-04-13 |
| Mistral Small 3.1 24B | Whole card | Q4_K_M, q4_0 KV, 57,600 cells, flash attention koboldcpp 1.88, RTX 4090 24GB | 17.83 GB | 16.50 GB 18.1 GB used − 1.6 GB desktop | +8.1% | Weights and cache match the logs; the gap is our 0.5 GB + 10% overhead, about 2–3 times what llama.cpp added on top of them for one request. | github.com 2025-04-13 |
The four whole-card rows are the VRAM in use the reporter read, minus the 1.6 GB their desktop already took. Mistral Small 3.1 is counted at 23.57B parameters without its vision encoder, like the text-only GGUF in the report; Gemma 2 9B's shape is from its config.json. The other models use the calculator's presets.
Image models
The Qwen-Image-2.1 calculator gives a range: exact file sizes plus the working memory seen in real runs.
| GPU | Setup | Predicted range | Measured | Result | Source |
|---|---|---|---|---|---|
| RTX 3090 24GB | Qwen-Image-2.1: DiT GGUF Q4_K_M + encoder INT8 ConvRot + VAE BF16, all on the GPU, 1024×1024 ComfyUI | 15.35 GB–17.95 GB | 15.39 GB | Inside | github.com 2026-09-21 |
| Apple M5 Max 48GB | Qwen-Image-2.1: DiT GGUF Q4_K_M; encoder, VAE and image size not stated, compared as BF16 encoder + FP32 VAE, all in unified memory, 1 MP Unsloth Desktop (macOS) | 23.60 GB–26.20 GB | 42.00 GB | Outside (measured higher) | github.com 2026-09-29 |
RTX 3090 24GB: Inside the range, at its low end. This run is one of those the range was set from, so it checks consistency more than accuracy.
Apple M5 Max 48GB: Outside the range: the measurement is far above it, even with the largest encoder and VAE. The ~34 GB after loading is close to the files with the DiT held at BF16 (32.4 decimal GB with a BF16 encoder and VAE) rather than the Q4_K_M file (22.4), so this path likely dequantizes the DiT, which the calculator does not model. The resolution is unknown and the engine has not confirmed the dtype; read as decimal GB, 42 GB is still far above the range.
More runs, and low-VRAM setups for 8–16 GB, are on the Qwen-Image 2.1 VRAM calculator. MiniMax H3's public runs mostly stream weights, so the card being full says how big the card is, not what the job needs; they are not compared here.
vLLM's KV cache pool is counted in tokens rather than bytes, so its checks are on the vLLM KV cache calculator's log table: 6 public startup logs.
Where we overestimate and underestimate
- The KV cache formula matches llama.cpp to the byte in 8 of 8 logs. Models with plain attention (Llama 3.1, Mistral Small, Gemma 2) always did, at f16 and at q4_0; Gemma 4 and gpt-oss do since 2026-09-29.
- Whole-card totals are 5.6%–8.1% high. Weights and cache line up; the gap is all overhead. 0.5 GB + 10% comes to 1.9–2.5 GB here, while llama.cpp used 0.7–1.1 GB beyond its weights and cache for one request (compute buffers, the CUDA context and the like).
- Before 2026-09-29, Gemma 4's KV cache was 35.3% and 54.3% low. We counted the global layers' keys once, as values equal to keys; llama.cpp still allocates the values separately (it ropes K and RMS-normalizes V). Its sliding-window layers also get one extra batch (512 tokens), more with parallel slots. That made some "fits" verdicts too optimistic. It is fixed, and the model, GPU and "Can I run" pages were recomputed with the new formula.
- GGUF file sizes can be 7.0% low to 13.7% high. Bits per weight work for Llama and Mistral, but Gemma 4's Q4_K_M averages 5.32 bits, and gpt-oss's GGUF stores the embedding table in Q8_0. When you have the file, its download size is the exact figure.
- vLLM reported 5.6% more than the checkpoint for gpt-oss-120b: its figure is the memory in use after loading, not the file.
What we'd change
- A llama.cpp overhead of its own: about 1 GB flat rather than 0.5 GB + 10%. The conservative default stays for now, because vLLM and CUDA graphs do need more.
Count Gemma 4's global-layer values separately, and size sliding windows as window × slots + one batch.Done on 2026-09-29, in every calculator on the site.- Use download sizes for known GGUF files, and closer bits per weight for models like Gemma 4.
- Let K and V use different types (such as
-ctk q8_0 -ctv q4_0), which the calculator cannot express yet.
On 2026-09-29 the KV cache formula was fixed: Gemma 4's global layers count K and V separately, and sliding-window layers get llama.cpp's cell count, window × slots + 512 rounded up to 256, and never more than the context (src/llama-kv-cache-iswa.cpp). The table above uses the fixed formula; before it we predicted 187 MiB for Gemma 4 12B, 2,625 MiB for 26B and 771 MiB for gpt-oss-20b. The rest has not changed yet: this page shows how the current formulas do, and when they change the table is recomputed with them at build time.
Reports we could not use
Two 2026 reports of Qwen3.8 27B on 16 GB cards (autodidacts.io, jrell's IQ4_XS mix) give the quant, context and KV type but no memory reading, only that it runs. They do not contradict the estimates, but there is no error to compute. Runs with an nvidia-smi reading or llama.cpp buffer lines are welcome; see About for contact.
Submit a measurement
Ran a model yourself? Submit a measurement on GitHub: the form asks for the model, file or quant, engine and version, GPU, -c, parallel slots and KV type, and the log lines as printed (llama.cpp's KV buffer size and model buffer size, an nvidia-smi reading, or vLLM's KV cache lines). Each one is checked by hand, then added to the table above next to the build-time prediction.
Frequently asked questions
How accurate is the ModelVRAM VRAM calculator?
Across 19 public measurements the median error is 1.2%. The KV cache formula matched llama.cpp's logs exactly in 8 of 8 cases, weights were 1.5% off (median), and whole-card totals came out 5.6%–8.1% high, because 0.5 GB + 10% overhead is conservative for one llama.cpp request.
Does the calculator overestimate or underestimate?
Whole-card totals: every one over, by 5.6% to 8.1%. Under: GGUF files that keep more tensors at high precision (Gemma 4 26B Q4_K_M −7.0%) and vLLM's loaded gpt-oss-120b (−5.6%). No KV row is under now. Before 2026-09-29 we undercounted Gemma 4's KV cache in llama.cpp (−35.3% and −54.3%) and gpt-oss's sliding-window cache (15 MiB short); both were fixed that day.
Why does the estimate differ from nvidia-smi?
nvidia-smi shows the whole card: the desktop, the CUDA context and whatever the engine reserves. vLLM fills memory up to gpu_memory_utilization, and llama.cpp allocates the KV cache for the full context at startup. This page compares only what lines up: the KV and model buffers in the logs, and whole-card readings with the desktop's share taken off.
How were the measurements chosen?
Only public reports with a source link and enough settings to reproduce them (model, quantized file, context, KV type, engine). Every source was reopened and re-read on 2026-09-29. Reports that give settings but no memory reading are left out.
Sources checked .