LLM VRAM guides
9 guides, each answering one question people ask before or after they load a model. The number under each is taken from that guide's own tables, so it always matches the page. For an answer about your own model and card, start with the LLM VRAM calculator.
One GPU: weights and usable memory
- GGUF quantization explained
What Q8_0, Q5_K_M, Q4_K_M and Q2_K mean in bits per weight, and how big a model is at each.
Llama 3.1 8B: 15.0 GB of weights in BF16, 7.95 GB at Q8_0, 4.52 GB at Q4_K_M.
- Why a 24 GB GPU does not give an LLM 24 GB
How much of a card is free once the CUDA context, the desktop and Windows have taken their share, and how the calculator budgets for it.
An RTX 3090 had 23,863 of its 24,126 MiB free in a public llama.cpp log; the calculator budgets 20.9 GB of weights and cache on a 24 GB card.
Macs and unified memory
- How much of a Mac's unified memory the GPU can use
The default GPU limit for each Mac memory size, how to raise it with iogpu.wired_limit_mb, and the largest model each size runs.
A 64 GB Mac lets the GPU use 48 GB by default, a 128 GB one 96 GB.
Several GPUs
- Splitting a model across two or more GPUs
How llama.cpp and vLLM divide weights and KV cache between cards, and how much each card needs.
Llama 3.1 70B at Q4_K_M, 8K context, on two RTX 3090s: 23.7 GB + 23.2 GB per card.
KV cache and context
- KV cache explained
Why long context costs memory, the formula behind it, and why some architectures need far less.
Llama 3.1 8B's FP16 cache for a 128K-token context: 16.0 GB.
- KV cache quantization: q8_0 vs q4_0 vs f16
How much memory a quantized cache saves, how to turn it on in llama.cpp, vLLM and Ollama, and what it costs in quality.
Llama 3.1 70B at 128K: 40.0 GB of cache at f16, 21.3 GB at q8_0, 11.3 GB at q4_0.
Mixture-of-experts models
- Big MoE models on one GPU: --n-cpu-moe and -ot
Which flag moves experts to system RAM, how to pick its value, and how much VRAM and RAM five MoE models then need.
gpt-oss-120b on an RTX 3090 at 32K: --n-cpu-moe 24, 22.7 GB of VRAM and 38.5 GB of RAM.
Troubleshooting and accuracy
- The llama-server --fit context regression
Why some llama.cpp builds picked huge contexts on their own, which builds, and which flags to set instead.
Builds b10999 to b11200 gave a 128 GB Mac a 828,160-token context; b11201 reverted the change.
- How accurate the estimates are
Every estimate on the site next to public measurements, with each one's error.
VRAM: 20 public measurements, median error 0.8%; prompt speed: 32 runs, all within ±28%.
Data behind the guides: the 2026 local LLM VRAM report, the model and GPU report and the open dataset.