KV cache quantization: q8_0 vs q4_0 vs f16, how much VRAM it saves and what it costs
A q8_0 KV cache stores 32 values in 34 bytes (8.5 bits, 53% of f16) and q4_0 in 18 bytes (4.5 bits, 28%), so Llama 3.1 70B's cache for a 128K-token context drops from 40.0 GB at f16 to 21.3 GB at q8_0 and 11.3 GB at q4_0; in the only public perplexity test found, q8_0 cost +0.03% perplexity and q4_0 +3.31%.
Why the cache is as big as it is in the first place (layers, KV heads, sliding windows, latent attention) is on KV cache explained; this page is only about storing it in fewer bits.
Bytes per value
llama.cpp's cache types are its block formats: 32 values share one 16-bit scale (and, for the _1 types, a 16-bit minimum).
Bits per value = block bytes × 8 ÷ 32. Read at llama.cpp 19e28a2
(ggml/src/ggml-common.h).
| Type | Block | Bytes per 32 values | Bits per value | Share of f16 |
|---|---|---|---|---|
| f16 | one 16-bit float | 64 | 16 | 100% |
| q8_0 | fp16 scale + 32 × int8 | 34 | 8.5 | 53% |
| q5_1 | fp16 scale + fp16 min + 32 high bits + 32 × 4 bits | 24 | 6 | 38% |
| q5_0 | fp16 scale + 32 high bits + 32 × 4 bits | 22 | 5.5 | 34% |
| q4_1 | fp16 scale + fp16 min + 32 × 4 bits | 20 | 5 | 31% |
| q4_0 | fp16 scale + 32 × 4 bits | 18 | 4.5 | 28% |
vLLM's FP8 cache is exactly 8 bits per value (50% of BF16); its scales are per tensor or per head, not per value.
KV cache of five models at f16, q8_0 and q4_0
One request, the site's own cache formula (the same as the calculator's KV cache setting). Gemma 4 and gpt-oss keep most layers at a sliding window, so their per-token figure is what each token adds once the windows are full. The “q8_0 K + q4_0 V” column is the mixed setting the quality test below favours (6.5 bits on average). GB means GiB.
| Model | Context | f16 | q8_0 | q8_0 K + q4_0 V | q4_0 |
|---|---|---|---|---|---|
| Qwen3.8 27B | per token | 64 KB | 34 KB | 26 KB | 18 KB |
| 32K | 2.00 GB | 1.06 GB | 0.81 GB | 0.56 GB | |
| 128K | 8.00 GB | 4.25 GB | 3.25 GB | 2.25 GB | |
| Gemma 4 31B | per token | 80 KB | 42.5 KB | 32.5 KB | 22.5 KB |
| 32K | 3.67 GB | 1.95 GB | 1.49 GB | 1.03 GB | |
| 128K | 11.2 GB | 5.94 GB | 4.54 GB | 3.14 GB | |
| Llama 3.1 8B | per token | 128 KB | 68 KB | 52 KB | 36 KB |
| 32K | 4.00 GB | 2.13 GB | 1.63 GB | 1.13 GB | |
| 128K | 16.0 GB | 8.50 GB | 6.50 GB | 4.50 GB | |
| Llama 3.1 70B | per token | 320 KB | 170 KB | 130 KB | 90 KB |
| 32K | 10.0 GB | 5.31 GB | 4.06 GB | 2.81 GB | |
| 128K | 40.0 GB | 21.3 GB | 16.3 GB | 11.3 GB | |
| gpt-oss-120b | per token | 36 KB | 19.1 KB | 14.6 KB | 10.1 KB |
| 32K | 1.15 GB | 0.61 GB | 0.47 GB | 0.32 GB | |
| 128K | 4.53 GB | 2.40 GB | 1.84 GB | 1.27 GB |
Whole-model total at Q4_K_M weights with 32K context, with the site's 0.5 GB + 10% overhead; each figure opens the calculator with that cache type:
| Model | f16 cache | q8_0 cache | q4_0 cache |
|---|---|---|---|
| Qwen3.8 27B | 19.9 GB | 18.9 GB | 18.3 GB |
| Gemma 4 31B | 23.9 GB | 22.0 GB | 21.0 GB |
| Llama 3.1 8B | 9.88 GB | 7.81 GB | 6.71 GB |
| Llama 3.1 70B | 55.2 GB | 50.1 GB | 47.3 GB |
| gpt-oss-120b | 74.2 GB | 73.6 GB | 73.3 GB |
Where it matters most is context on a fixed card. On a 24 GB GPU at Q4_K_M, with 0.5 GB left free, the longest context goes for Qwen3.8 27B from 83K tokens at f16 to 158K at q8_0, and for Gemma 4 31B from 26K to 64K (134K at q4_0). A smaller cache matters less for gpt-oss-120b, whose 128-token windows keep the cache small next to its weights.
Checked against a llama.cpp log: in issue #22748 (2026-05-06; Gemma 4 31B Q5_K_M, -c 128000 -ctk q8_0 -ctv q4_0 -ub 256, on an RTX 5060 Ti + RTX 3060) the
full-attention cache is 4062.50 MiB and the sliding-window cache 406.25 MiB
across the two GPUs; at 6.5 bits per value the formula gives 4062.50 and 406.25 MiB
(1280 sliding cells, because that run set -ub 256).
llama.cpp: -ctk, -ctv and flash attention
-
-ctk/--cache-type-kand-ctv/--cache-type-vset the K and V types separately, orLLAMA_ARG_CACHE_TYPE_K/_V(common/arg.cpp:2433-2458 ). Allowed: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1 (304-314). Both default to f16 (common/common.h:343-344 ). - A quantized V cache needs flash attention. Without it llama.cpp stops with “quantized V cache was requested, but
this requires Flash Attention” (src/
llama-context.cpp:464-468 ). A quantized K cache has no such check.-fatakes on, off or auto (arg.cpp:1751-1765) and defaults to auto (llama-context.cpp:3724), which uses flash attention when the backend supports it. -
Example:
llama-server -m model.gguf -c 65536 -fa on -ctk q8_0 -ctv q8_0. The log'sllama_kv_cache: CUDA0 KV buffer sizeline shows the smaller cache. -
If you leave out
-c, llama-server's--fituses the freed memory for more context rather than leaving it free; the --fit preview shows what it would pick for a given KV type.
vLLM: --kv-cache-dtype
-
Values include auto (the model's dtype), fp8 (= fp8_e4m3) and fp8_e5m2; “CUDA 11.8+ supports fp8 (=fp8_e4m3) and fp8_e5m2. ROCm
(AMD GPU) supports fp8 (=fp8_e4m3)” (vllm/
config/ ; list at 39-57). Read at ac7f3e1 (v0.30.0 logic).cache.py:123-132 -
vLLM's quantized KV cache docs: without calibration all scales are 1.0; calibrated scales from
llm-compressor are the recommended path;
--kv-cache-dtype-skip-layers sliding_windowkeeps sliding-window layers, which it calls more sensitive, at the model's dtype. -
Memory: vLLM fills what
--gpu-memory-utilizationleaves after the weights with cache blocks, so FP8 does not lower memory use but about doubles the tokens that fit. The vLLM memory calculator has an FP8 option.
Ollama: OLLAMA_KV_CACHE_TYPE
From the Ollama FAQ (main b3f78b7; the variable itself in
envconfig/OLLAMA_KV_CACHE_TYPE to f16
(default), q8_0 or q4_0 for the server; it is global to all models, and the cache “can be quantized … when Flash
Attention is enabled”. Ollama's own description: q8_0 has “a very small loss in precision” and is recommended if not using f16; q4_0 a
“small-medium loss in precision that may be more noticeable at higher context sizes”, and models with a high GQA ratio may lose more.
What it costs in quality
The one public side-by-side found is a perplexity and KL-divergence table posted in llama.cpp PR #7412 (2024-05-20), a research build that was not merged; the comment does not name the model or text. Rows are weights / K / V:
| Weights / K / V | KV bits per value | Perplexity | vs f16 cache | KL divergence |
|---|---|---|---|---|
| f16/f16/f16 | 16 | 6.2322 | — | 0.000189 |
| f16/q8_0/q8_0 | 8.5 | 6.2344 | +0.03% | 0.000980 |
| f16/q8_0/q4_0 | 6.5 | 6.2541 | +0.35% | 0.005079 |
| f16/q4_0/f16 | 10.25 | 6.4207 | +3.03% | 0.032916 |
| f16/q4_0/q4_0 | 4.5 | 6.4384 | +3.31% | 0.036509 |
| q4_K_M/f16/f16 | 16 | 6.4064 | +2.80% | 0.031280 |
The author's reading: “The K cache seems to be much more sensitive to quantization than the V cache.” and “There seems to be no significant quality loss from using q8_0 instead of FP16 for the KV cache.” A q4_0 cache cost slightly more than Q4_K_M weights did in the same test, and a q8_0 K cache with a q4_0 V cache kept most of the q4_0 saving at a fraction of the loss. No public comparison was found for Qwen3.8 27B, Gemma 4, gpt-oss or Llama 3.1, or for vLLM's FP8 cache on them, so treat these as one data point rather than a rule.
Related
- The LLM VRAM calculator has a KV cache type setting (FP16, FP8, Q8_0, Q4_0) for any model and context.
- On several GPUs the cache is split with the layers: splitting a model across GPUs.
Questions
How much VRAM does a q8_0 KV cache save?
Almost half: q8_0 keeps 34 bytes per 32 values, 53% of f16. For Llama 3.1 8B at 32K context the cache goes from 4.00 GB to 2.13 GB, and its Q4_K_M total from 9.88 GB to 7.81 GB. q4_0 keeps 18 bytes per 32 values, 28% of f16.
Does a quantized KV cache need flash attention in llama.cpp?
A quantized V cache does: llama.cpp refuses to start with "quantized V cache was requested, but this requires Flash Attention" when flash attention is off. A quantized K cache alone does not. Flash attention defaults to auto, which turns it on where the backend supports it; pass -fa on to require it.
Is q8_0 or q4_0 KV cache worse for quality?
q4_0 is. In the public llama.cpp test (PR #7412, model not named) q8_0 for K and V raised perplexity from 6.232 to 6.234, and q4_0 to 6.438, slightly more than going from f16 to Q4_K_M weights (6.406). The K cache was the sensitive half: q8_0 K with q4_0 V gave 6.254. No public comparison was found for Qwen3.8, Gemma 4 or gpt-oss.
How do I quantize the KV cache in Ollama?
Set OLLAMA_KV_CACHE_TYPE to q8_0 or q4_0 (default f16) in the environment of the Ollama server. It is global, so every model uses it, and it only applies when flash attention is on; Ollama turns flash attention on automatically where the backend supports it, and OLLAMA_FLASH_ATTENTION=1 forces it.
What does --kv-cache-dtype fp8 do in vLLM?
It stores keys and values as 8-bit floats (fp8 means fp8_e4m3), half of BF16, so the pool holds about twice the tokens. Without calibration every scale is 1.0; vLLM's docs recommend scales calibrated with llm-compressor, and --kv-cache-dtype-skip-layers sliding_window keeps the sliding-window layers, which the docs call more sensitive, at full precision.
Sources read on .