vLLM KV Cache & Concurrency Calculator

vLLM gives the KV cache whatever --gpu-memory-utilization leaves after the weights and its warm-up pass, then cuts it into 16-token blocks. Pick a model, GPU and flags to see the pool it will print at startup, how many requests fit at your context length, and the vllm serve command.

– tokens of KV cache ("GPU KV cache size")

Weights per GPU
–
Held back: activations, CUDA graphs, non-torch
–
KV cache memory per GPU
–
KV cache per token, per GPU
–
Max concurrency at max model length
–
Requests at the average length
–

Command

What vLLM should print at startup

Checked against vLLM logs

Each row is a vLLM startup log posted in a GitHub issue. "Ours from logged memory" takes the Available KV cache memory the log prints and applies this page's formulas, which tests the bytes per token, the grouping and the block rounding. "Ours from scratch" uses only the GPU, the settings and the reserve estimate, nothing from the log. Before v0.24.0 the log printed blocks ÷ groups × 16 for hybrid models, so those rows are compared in that format.

Log Model GPUs and settings Logged KV memory Logged pool Ours from logged memory Ours from scratch
#42593 Llama 3.1 8B H100 80GB HBM3, TP 1, util 0.9, auto KV, 131,072 tokens, vLLM 0.14.1 51.39 GiB 420,944 420,976 +0.01% 432,048 +2.6%; 52.74 GiB
#27604 Llama 3.1 8B H100 80GB HBM3, TP 1, util 0.9, fp8 KV, 131,072 tokens, vLLM 0.11.0 54.32 GiB 890,000 889,968 0.00% 864,112 −2.9%; 52.74 GiB
#38107 Qwen3 8B 2× Radeon PRO W6800 32GB (ROCm), TP 2, util 0.9, auto KV, 32,768 tokens, vLLM 0.1.dev15181 (2026-03) 20.23 GiB 294,624 294,608 −0.01% 280,592 −4.8%; 19.27 GiB
#26480 gpt-oss-120b 2× A100 80GB PCIe, TP 2, util 0.85, auto KV, 131,072 tokens, vLLM 0.11.0 31.78 GiB 925,696 925,648 −0.01% 977,808 +5.6%; 33.57 GiB
#36849 gpt-oss-20b 2× RTX 5090, TP 2, util 0.9, fp8 KV, 100,000 tokens, vLLM 0.17.1 16.62 GiB 1,452,000 1,452,272 +0.02% 1,525,536 +5.1%; 17.46 GiB
#58580 Gemma 4 26B-A4B H200, TP 1, util 0.9, auto KV, 131,072 tokens, vLLM 0.28.0 73.43 GiB 1,652,749 1,652,749 0.00% 1,671,267 +1.1%; 74.25 GiB

#26480: A100 has no native MXFP4 kernels; vLLM repacks the experts, so the weights come out 6% above our figure.

#36849: Run with --max-num-batched-tokens 98304, twelve times the usual batch, so the activation reserve was large.

#58580: The weights include the MTP drafter (gemma-4-26B-A4B-it-assistant) this run loaded; its draft layers reuse the main model's cache. The log also breaks the rest down: 1.17 GiB non-torch, 1.89 GiB peak activation, 0.16 GiB CUDA graphs.

How vLLM sizes the KV cache

The formulas follow vLLM's main branch (ac7f3e11, the same as v0.30.0), checked .

  1. Memory (gpu_worker.py): on each GPU, total memory × gpu_memory_utilization, minus the weights, the peak activation of a profiling pass of max_num_batched_tokens tokens, non-torch memory (CUDA context, NCCL) and the CUDA graph estimate (subtracted by default since v0.20). What is left is the Available KV cache memory in the log. The default utilization is 0.92 since v0.20.0, 0.9 before.
  2. Pages (kv_cache_interface.py): one block of one layer = 16 tokens × KV heads per GPU × (K width + V width) × bytes per value. Tensor parallelism splits the KV heads over the GPUs; with fewer heads than GPUs each GPU keeps one (replicated). MLA stores one latent vector per token, which every GPU keeps whole.
  3. Blocks: blocks = memory ÷ bytes per block, rounded down. A model with one attention type is one group of all its layers. Sliding-window plus full-attention models are split into groups of equal layer count, and when the pages differ the smaller one gets a longer block (Gemma 4's 512-wide global layers use 32-token blocks).
  4. Capacity: a request takes blocks in every group: full-attention layers for max_model_len tokens, sliding-window layers for at most window − 1 + 2 × max_num_batched_tokens tokens (async scheduling, the default, keeps two batches in flight) plus one block. Max concurrency = blocks ÷ blocks per request, and the log's GPU KV cache size is max concurrency × max_model_len.

Step 1's non-weight part cannot be measured in a browser, so this page uses an estimate by the max_num_batched_tokens vLLM picks for the card: about 1.5 GB at 2,048 (0.4–2.5), 3.5 GB at 8,192 (2–5) and 4 GB at 16,384 (2.5–6), from the logs above. Prefix caching does not change the pool, but requests that hit it share blocks, so real concurrency can be higher.

For one request's memory and which cards fit, see the LLM VRAM calculator; for generation speed, the LLM speed calculator.

How to use

  1. Pick a model, or load any Hugging Face id, and the GPU and how many of them (tensor parallel).
  2. Set --gpu-memory-utilization, the weight precision and the KV cache dtype.
  3. Set --max-model-len and the average tokens per request (prompt plus output).
  4. Read the pool, the concurrency and the command. If you already have a startup log, compare its "Available KV cache memory" with the estimate here.

Frequently asked questions

How does vLLM decide how big the KV cache is?

At startup it loads the weights, runs one forward pass of max_num_batched_tokens tokens to measure the peak activation memory, and estimates the CUDA graph memory. The KV cache gets total GPU memory × --gpu-memory-utilization, minus the weights, that peak and the memory outside PyTorch (CUDA context, NCCL), cut into blocks of 16 tokens. For Qwen3 8B on one H100 with the defaults that is about 54.0 GB, or 393,392 tokens.

What do "GPU KV cache size" and "Maximum concurrency" in the startup log mean?

The first is how many tokens of context the pool holds; the second is how many requests of exactly --max-model-len tokens fit in it at once (12.01x for Qwen3 8B on an H100 at 32,768). Real requests are shorter: at 4,096 tokens each the same pool holds about 96. vLLM does not reject the rest; it queues them, or preempts running requests when blocks run out.

Does a second GPU double the KV cache?

More than that when the weights are large. With --tensor-parallel-size 2 each GPU keeps half the weights and half the KV heads. Llama 3.1 70B in FP8 on one H100 leaves 3.6 GB for cache, 11,696 tokens, less than one 32K request; on two H100s it leaves 36.4 GB on each and the pool holds 238,720 tokens (7.29x at 32K). With fewer KV heads than GPUs, each GPU keeps a whole copy of one head, so the cache per token stops shrinking.

Should I use --kv-cache-dtype fp8?

It halves the bytes per token, so the pool holds twice the tokens: Qwen3 8B on an H100 goes from 393,392 to 786,784. The quality cost is usually small but depends on the model and the attention backend, so check it on your own task.

Why does vLLM report fewer tokens for gpt-oss and Gemma than the bytes per token suggest?

They mix sliding-window and full-attention layers. vLLM reserves the window − 1 plus the tokens of two batches in flight (2 × max_num_batched_tokens with async scheduling) and one block per request for the sliding layers, and since v0.24 the logged token count is concurrency × --max-model-len, which counts every layer group's blocks. gpt-oss-120b on one H100 at 131,072 tokens leaves about 4.9 GB for cache and 125,885 tokens: borderline for one full-length request.

Why is the "Available KV cache memory" in my log different?

The part vLLM holds back besides the weights (activation peak, CUDA graphs, NCCL, the CUDA context) depends on the version, the attention backend, the vocabulary and max_num_batched_tokens. In the 5 public logs on this page that ran vLLM's default batch sizes it was 0.5–4.8 GB. Once you have your own log, that line is exact, and the pool follows from it: from the logged memory this page's formulas matched every log's token count within 0.02%.

Should I raise --gpu-memory-utilization?

It is the share of total GPU memory this vLLM instance may use, weights included: 0.92 by default since v0.20.0, 0.9 before. Whatever the weights and activations leave inside it becomes KV cache, so 0.95 on an H100 adds 2.4 GB of cache for Qwen3 8B (393,392 to 410,672 tokens). Leave room for anything else on the GPU; a desktop or a second process can make vLLM fail at startup.

Is anything I enter uploaded?

No. The calculation runs in your browser; the only network request is to Hugging Face when you load a model by its id.

More calculators

Updated