Speculative Decoding VRAM Calculator (MTP & draft models)
Speculative decoding adds a second model or head to llama.cpp, and llama.cpp's automatic fitting does not always count it. Pick the main file and the draft to see the exact extra bytes: draft weights on the GPU, the draft KV cache at the draft context (always as long as the main one), the extra recurrent-state copies hybrid models keep, and an estimate of the draft compute buffer.
– extra VRAM for speculative decoding
- Draft / MTP weights on the GPU
- –
- Draft KV cache at the draft context
- –
- Extra recurrent-state copies of the main model
- –
- Draft model's own recurrent state
- –
- Draft compute buffer (estimate)
- –
- Model alone (file + KV + state + 1 GB buffers)
- –
- With speculative decoding
- –
llama.cpp's automatic fitting
| GPU memory | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB |
|---|---|---|---|---|---|
| Model alone | – | – | – | – | – |
| With the draft | – | – | – | – | – |
“Tight”: fits only with the low end of the compute estimate. Weights, KV and state are exact byte counts; the compute buffer is the range two Qwen3.8 27B MTP logs show (126.77–272.28 MiB), plus the F16 copy CUDA flash attention makes of a quantized draft cache. The model-alone figure counts the whole file and 1 GB of buffers, as on the other pages here. Leave about 0.5–1 GB for the desktop if the GPU drives your screen.
GB = GiB (1024³ bytes) and MB = MiB; logged values are shown as llama.cpp printed them.
llama-server command
Every draft compared
With the Q4_K_M file of each model, f16 caches, one slot and --spec-draft-n-max 3. Extra = weights + draft KV + recurrent copies + the compute estimate. “-fit counts it”: whether llama.cpp's automatic fitting includes the draft.
| Draft | --spec-type | File size | On the GPU | Draft KV, 32K | Draft KV, 128K | Extra, 32K | Extra, 128K | -fit counts it |
|---|---|---|---|---|---|---|---|---|
| MTP layer in the model file Qwen3.8 27B | draft-mtp | – | 335 MB | 128 MB | 512 MB | 1.01 GB–1.16 GB | 1.39 GB–1.53 GB | yes |
| Unsloth MTP file (mtp-Qwen3.8-27B-Q4_0) Qwen3.8 27B | draft-mtp | 1.28 GB | 1.08 GB | 128 MB | 512 MB | 1.77 GB–1.91 GB | 2.15 GB–2.29 GB | yes |
| DFlash2 Q4_K_M (z-lab) Qwen3.8 27B | draft-dflash | 1.06 GB | 1.05 GB | 50.0 MB | 50.0 MB | 1.67 GB–1.81 GB | 1.67 GB–1.81 GB | no |
| DFlash2 Q8_0 (z-lab) Qwen3.8 27B | draft-dflash | 1.92 GB | 1.90 GB | 50.0 MB | 50.0 MB | 2.52 GB–2.66 GB | 2.52 GB–2.66 GB | no |
| DFlash2 BF16 (z-lab) Qwen3.8 27B | draft-dflash | 3.60 GB | 3.58 GB | 50.0 MB | 50.0 MB | 4.20 GB–4.34 GB | 4.20 GB–4.34 GB | no |
| Qwen3.5 0.8B Q4_K_M Qwen3.8 27B | draft-simple | 508 MB | 497 MB | 384 MB | 1.50 GB | 1.00 GB–1.15 GB | 2.13 GB–2.27 GB | yes |
| Qwen3.5 0.8B Q8_0 Qwen3.8 27B | draft-simple | 774 MB | 764 MB | 384 MB | 1.50 GB | 1.26 GB–1.41 GB | 2.39 GB–2.53 GB | yes |
| Qwen3.5 2B Q4_K_M Qwen3.8 27B | draft-simple | 1.19 GB | 1.18 GB | 384 MB | 1.50 GB | 1.70 GB–1.84 GB | 2.83 GB–2.97 GB | yes |
| Qwen3.5 2B Q8_0 Qwen3.8 27B | draft-simple | 1.87 GB | 1.86 GB | 384 MB | 1.50 GB | 2.38 GB–2.52 GB | 3.51 GB–3.65 GB | yes |
| Unsloth MTP assistant Q8_0 Gemma 4 31B | draft-mtp | 491 MB | 476 MB | 0 | 0 | 603 MB–748 MB | 603 MB–748 MB | no |
| Unsloth MTP assistant F16 Gemma 4 31B | draft-mtp | 911 MB | 896 MB | 0 | 0 | 1.00 GB–1.14 GB | 1.00 GB–1.14 GB | no |
| DFlash Q4_K_M (Anbeeld) Gemma 4 31B | draft-dflash | 870 MB | 855 MB | 168 MB | 552 MB | 1.12 GB–1.27 GB | 1.50 GB–1.64 GB | no |
| DFlash Q8_0 (Anbeeld) Gemma 4 31B | draft-dflash | 1.53 GB | 1.52 GB | 168 MB | 552 MB | 1.81 GB–1.95 GB | 2.18 GB–2.33 GB | no |
The llama.cpp flags involved
Read from common/arg.cpp and common/common.h on llama.cpp master at commit 4364bf7 (2026-09-28).
--spec-type- draft-mtp (built-in MTP layer, Qwen MTP file or Gemma 4 assistant), draft-dflash, draft-dspark, draft-eagle3, draft-simple (a small model), plus the n-gram types that need no model. With -md and no --spec-type, llama.cpp reads the type from the draft file.
--spec-draft-n-max / --spec-draft-n-min- Most and fewest tokens drafted per step, default 3 and 0. On hybrid models (Qwen3.8) every extra draft token keeps one more copy of the recurrent state: 150 MB each for Qwen3.8 27B.
-md, --spec-draft-model- The draft, MTP or DFlash file. Not needed for an MTP layer inside the model file.
-ngld, --spec-draft-ngl- Draft layers on the GPU: a number, auto (the default) or all.
-ctkd / -ctvd- Draft KV cache types, f16 by default. -ctk / -ctv do not apply to the draft.
-devd, --spec-draft-device- Which devices hold the draft, e.g. a second GPU; none keeps it on the CPU.
There is no flag for the draft context size: the draft context always holds as many tokens as the main one
llama.cpp logs the formulas reproduce
| What was logged | Logged | Source |
|---|---|---|
| Recurrent state, no speculation (Qwen3.8 27B, 1 slot) | 149.62 MiB | #27814 2026-08-27 |
| Recurrent state, MTP --spec-draft-n-max 3 (Qwen3.6 27B, same shape) | 598.50 MiB | #28112 2026-08-31 |
| Recurrent state, MTP --spec-draft-n-max 4 (Qwen3.6 27B) | 748.12 MiB | #23577 2026-05-23 |
| MTP draft KV cache, 4 × 131,072 tokens, -ctk q8_0 (draft stays f16) | 2,048.00 MiB | #28115 2026-08-31 |
| MTP draft compute buffer, -c 64000, f16 draft cache | 126.77 MiB | #28378 2026-09-04 |
| Same with -ctkd q8_0 -ctvd q8_0 | 374.03 MiB | #28378 2026-09-04 |
| MTP draft compute buffer requested, -c 196608, q4_0 draft cache (OOM) | 1,040.28 MiB | #27282 2026-08-17 |
File sizes and tensor bytes read from the GGUF headers and the Hugging Face API on : unsloth/
How to use
- Pick the main model file. Qwen3.8 27B files from Unsloth include an MTP layer; the GSQ-RCO IQ3_XXS file has none.
- Pick the draft: the built-in MTP layer, a separate MTP file, a DFlash drafter or a small Qwen3.5 model; for Gemma 4 31B, Unsloth's MTP assistant or a DFlash drafter.
- Set the context, the main and draft KV cache types and --spec-draft-n-max as you run llama-server.
- Read the extra VRAM, the total with and without the draft on 12–48 GB cards, and copy the llama-server command.
Frequently asked questions
How much extra VRAM does MTP take on Qwen3.8 27B?
With UD-Q4_K_M at 64K context, f16 caches and the default --spec-draft-n-max 3: 335 MB for the MTP layer, 256 MB for its KV cache, 449 MB for three more copies of the recurrent state and about 127–272 MB of compute buffer, 1.14–1.28 GB in all. At 256K the draft cache grows to 1.00 GB and the total to 1.89–2.03 GB.
Does the MTP layer in the file use VRAM when MTP is off?
No. llama.cpp creates the blk.64 tensors with TENSOR_SKIP unless --spec-type draft-mtp is set, so a file with the layer costs 335 MB more download (430 MB for Q8_0), not more VRAM; the load log lists them as unused. In a Swift-1.5 discussion on Hugging Face a user running a DFlash2 drafter saw about 1 GB more VRAM and suspected the MTP layer; 1 GB is what the drafter itself takes (DFlash2 Q4_K_M puts 1.05 GB on the GPU), while the MTP layer would be 335 MB.
DFlash2 or MTP: which needs more memory?
DFlash2 Q4_K_M has larger weights (1.05 GB against 335 MB) but a small cache: its five layers use a 2,048-token window, so it keeps 2,560 cells, 50 MB in f16 at any context. The MTP cache grows with the context: 256 MB at 64K, 1.00 GB at 256K. Both keep the same extra recurrent-state copies (150 MB per draft token). At 64K, DFlash2 Q4_K_M adds 1.67–1.81 GB and MTP 1.14–1.28 GB; at 256K they are 1.67–1.81 GB and 1.89–2.03 GB.
Why does Gemma 4 31B with an MTP draft run out of memory when it works without?
In llama.cpp issue #29521 (Gemma 4 31B with --model-draft mtp-gemma-4-31B-it-F16.gguf), common/fit.cpp logs "failed to measure the memory of the extra model, fitting without it": the assistant needs the main context to be created, so the fitter counts it as zero and fills the memory with context. The draft context is always given the same n_ctx as the main one. Budget the draft yourself: the F16 assistant puts 896 MB on the GPU (476 MB for Q8_0) plus about 127–272 MB of compute buffer, and adds no KV cache because it reads the main model's. Set -c and -np 1 yourself, or raise --fit-target by about 1.2 GB.
Is there a flag for the draft context size?
Not on llama.cpp master (commit 4364bf7, 2026-09-28): common/
Does -ctk q8_0 also shrink the draft cache?
No. The draft cache has its own types, -ctkd and -ctvd, and they default to f16 (a log in issue #28115 shows a 2,048 MiB f16 MTP cache next to a q8_0 main cache). Quantizing it can cost memory on CUDA: flash attention keeps an F16 copy in the compute buffer, which went from 126.77 to 374.03 MiB at -c 64000 in PR #28378, more than the cache saved.
Is GB here GB or GiB?
GiB, and MB is MiB, the units llama.cpp and nvidia-smi print, as everywhere on this site.
More calculators
- Fine-Tuning VRAM Calculator GPU memory for full fine-tuning, LoRA and QLoRA in Transformers or Unsloth, checked against published runs. Open →
- vLLM KV Cache & Concurrency Calculator The KV cache pool vLLM allocates, in tokens, and how many requests fit at once, with the vllm serve command. Open →
- What Can My PC Run? Detects your GPU in the browser and lists the local LLMs it runs, with the best quantization and speed. Open →
- MiniMax H3 VRAM Calculator
(ComfyUI) VRAM and system RAM for MiniMax H3 video in ComfyUI: pruned, INT8, NVFP4 and GGUF files on 8–96 GB GPUs. Open → - MoE Offload Calculator
(--n-cpu-moe) The smallest llama.cpp --n-cpu-moe that fits your GPU, from real GGUF tensor sizes. Open → - Qwen3.8 27B GGUF Quants: Bonsai 2 vs GSQ-RCO vs UD Ternary Bonsai 2, GSQ-RCO and Unsloth UD files of Qwen3.8 27B: VRAM, quality, speed and the engine each needs. Open →
Updated