KV cache quantization: q8_0 vs q4_0 vs f16, how much VRAM it saves and what it costs

A q8_0 KV cache stores 32 values in 34 bytes (8.5 bits, 53% of f16) and q4_0 in 18 bytes (4.5 bits, 28%), so Llama 3.1 70B's cache for a 128K-token context drops from 40.0 GB at f16 to 21.3 GB at q8_0 and 11.3 GB at q4_0; in the only public perplexity test found, q8_0 cost +0.03% perplexity and q4_0 +3.31%.

Why the cache is as big as it is in the first place (layers, KV heads, sliding windows, latent attention) is on KV cache explained; this page is only about storing it in fewer bits.

Bytes per value

llama.cpp's cache types are its block formats: 32 values share one 16-bit scale (and, for the _1 types, a 16-bit minimum). Bits per value = block bytes × 8 ÷ 32. Read at llama.cpp 19e28a2 (ggml/src/ggml-common.h).

TypeBlockBytes per 32 valuesBits per valueShare of f16
f16 one 16-bit float 64 16 100%
q8_0 fp16 scale + 32 × int8 34 8.5 53%
q5_1 fp16 scale + fp16 min + 32 high bits + 32 × 4 bits 24 6 38%
q5_0 fp16 scale + 32 high bits + 32 × 4 bits 22 5.5 34%
q4_1 fp16 scale + fp16 min + 32 × 4 bits 20 5 31%
q4_0 fp16 scale + 32 × 4 bits 18 4.5 28%

vLLM's FP8 cache is exactly 8 bits per value (50% of BF16); its scales are per tensor or per head, not per value.

KV cache of five models at f16, q8_0 and q4_0

One request, the site's own cache formula (the same as the calculator's KV cache setting). Gemma 4 and gpt-oss keep most layers at a sliding window, so their per-token figure is what each token adds once the windows are full. The “q8_0 K + q4_0 V” column is the mixed setting the quality test below favours (6.5 bits on average). GB means GiB.

Model Context f16q8_0q8_0 K + q4_0 Vq4_0
Qwen3.8 27B per token 64 KB34 KB26 KB18 KB
32K 2.00 GB1.06 GB0.81 GB0.56 GB
128K 8.00 GB4.25 GB3.25 GB2.25 GB
Gemma 4 31B per token 80 KB42.5 KB32.5 KB22.5 KB
32K 3.67 GB1.95 GB1.49 GB1.03 GB
128K 11.2 GB5.94 GB4.54 GB3.14 GB
Llama 3.1 8B per token 128 KB68 KB52 KB36 KB
32K 4.00 GB2.13 GB1.63 GB1.13 GB
128K 16.0 GB8.50 GB6.50 GB4.50 GB
Llama 3.1 70B per token 320 KB170 KB130 KB90 KB
32K 10.0 GB5.31 GB4.06 GB2.81 GB
128K 40.0 GB21.3 GB16.3 GB11.3 GB
gpt-oss-120b per token 36 KB19.1 KB14.6 KB10.1 KB
32K 1.15 GB0.61 GB0.47 GB0.32 GB
128K 4.53 GB2.40 GB1.84 GB1.27 GB

Whole-model total at Q4_K_M weights with 32K context, with the site's 0.5 GB + 10% overhead; each figure opens the calculator with that cache type:

Modelf16 cacheq8_0 cacheq4_0 cache
Qwen3.8 27B 19.9 GB 18.9 GB 18.3 GB
Gemma 4 31B 23.9 GB 22.0 GB 21.0 GB
Llama 3.1 8B 9.88 GB 7.81 GB 6.71 GB
Llama 3.1 70B 55.2 GB 50.1 GB 47.3 GB
gpt-oss-120b 74.2 GB 73.6 GB 73.3 GB

Where it matters most is context on a fixed card. On a 24 GB GPU at Q4_K_M, with 0.5 GB left free, the longest context goes for Qwen3.8 27B from 83K tokens at f16 to 158K at q8_0, and for Gemma 4 31B from 26K to 64K (134K at q4_0). A smaller cache matters less for gpt-oss-120b, whose 128-token windows keep the cache small next to its weights.

Checked against a llama.cpp log: in issue #22748 (2026-05-06; Gemma 4 31B Q5_K_M, -c 128000 -ctk q8_0 -ctv q4_0 -ub 256, on an RTX 5060 Ti + RTX 3060) the full-attention cache is 4062.50 MiB and the sliding-window cache 406.25 MiB across the two GPUs; at 6.5 bits per value the formula gives 4062.50 and 406.25 MiB (1280 sliding cells, because that run set -ub 256).

llama.cpp: -ctk, -ctv and flash attention

vLLM: --kv-cache-dtype

Ollama: OLLAMA_KV_CACHE_TYPE

From the Ollama FAQ (main b3f78b7; the variable itself in envconfig/config.go:221-222): set OLLAMA_KV_CACHE_TYPE to f16 (default), q8_0 or q4_0 for the server; it is global to all models, and the cache “can be quantized … when Flash Attention is enabled”. Ollama's own description: q8_0 has “a very small loss in precision” and is recommended if not using f16; q4_0 a “small-medium loss in precision that may be more noticeable at higher context sizes”, and models with a high GQA ratio may lose more.

What it costs in quality

The one public side-by-side found is a perplexity and KL-divergence table posted in llama.cpp PR #7412 (2024-05-20), a research build that was not merged; the comment does not name the model or text. Rows are weights / K / V:

Weights / K / VKV bits per valuePerplexityvs f16 cacheKL divergence
f16/f16/f16 16 6.2322 — 0.000189
f16/q8_0/q8_0 8.5 6.2344 +0.03% 0.000980
f16/q8_0/q4_0 6.5 6.2541 +0.35% 0.005079
f16/q4_0/f16 10.25 6.4207 +3.03% 0.032916
f16/q4_0/q4_0 4.5 6.4384 +3.31% 0.036509
q4_K_M/f16/f16 16 6.4064 +2.80% 0.031280

The author's reading: “The K cache seems to be much more sensitive to quantization than the V cache.” and “There seems to be no significant quality loss from using q8_0 instead of FP16 for the KV cache.” A q4_0 cache cost slightly more than Q4_K_M weights did in the same test, and a q8_0 K cache with a q4_0 V cache kept most of the q4_0 saving at a fraction of the loss. No public comparison was found for Qwen3.8 27B, Gemma 4, gpt-oss or Llama 3.1, or for vLLM's FP8 cache on them, so treat these as one data point rather than a rule.

Related

Questions

How much VRAM does a q8_0 KV cache save?

Almost half: q8_0 keeps 34 bytes per 32 values, 53% of f16. For Llama 3.1 8B at 32K context the cache goes from 4.00 GB to 2.13 GB, and its Q4_K_M total from 9.88 GB to 7.81 GB. q4_0 keeps 18 bytes per 32 values, 28% of f16.

Does a quantized KV cache need flash attention in llama.cpp?

A quantized V cache does: llama.cpp refuses to start with "quantized V cache was requested, but this requires Flash Attention" when flash attention is off. A quantized K cache alone does not. Flash attention defaults to auto, which turns it on where the backend supports it; pass -fa on to require it.

Is q8_0 or q4_0 KV cache worse for quality?

q4_0 is. In the public llama.cpp test (PR #7412, model not named) q8_0 for K and V raised perplexity from 6.232 to 6.234, and q4_0 to 6.438, slightly more than going from f16 to Q4_K_M weights (6.406). The K cache was the sensitive half: q8_0 K with q4_0 V gave 6.254. No public comparison was found for Qwen3.8, Gemma 4 or gpt-oss.

How do I quantize the KV cache in Ollama?

Set OLLAMA_KV_CACHE_TYPE to q8_0 or q4_0 (default f16) in the environment of the Ollama server. It is global, so every model uses it, and it only applies when flash attention is on; Ollama turns flash attention on automatically where the backend supports it, and OLLAMA_FLASH_ATTENTION=1 forces it.

What does --kv-cache-dtype fp8 do in vLLM?

It stores keys and values as 8-bit floats (fp8 means fp8_e4m3), half of BF16, so the pool holds about twice the tokens. Without calibration every scale is 1.0; vLLM's docs recommend scales calibrated with llm-compressor, and --kv-cache-dtype-skip-layers sliding_window keeps the sliding-window layers, which the docs call more sensitive, at full precision.

Sources read on .