MoE Offload Calculator (--n-cpu-moe)
Pick a MoE GGUF, your GPU, RAM and context length. The planner adds up the real size of every layer’s expert tensors, read from the file’s header, and finds the smallest --n-cpu-moe that fits your VRAM.
Qwen3.6 35B-A3B · UD-Q4_K_M · 20.6 GB
–
- VRAM used
- –
- RAM for offloaded experts
- –
- KV cache
- –
- Writing speed, one request
- –
Every --n-cpu-moe value
| --n-cpu-moe | VRAM | RAM | Tokens/s | Status |
|---|
Smallest --n-cpu-moe at 32K context
FP16 KV cache, one request, 1 GB for buffers. RAM is what the offloaded experts take, plus the input embedding. Qwen3.8 Flash Next also reads a 26.8 GB n-gram table from the mapped file, not counted here.
| File | Size | 8 GB GPU | 12 GB GPU | 16 GB GPU | 24 GB GPU |
|---|---|---|---|---|---|
| Qwen3.6 35B-A3B UD-Q4_K_M | 20.6 GB | 31 · 14.6 GB RAM | 22 · 10.5 GB RAM | 13 · 6.39 GB RAM | fits whole |
| Qwen3.6 35B-A3B UD-IQ4_XS | 16.5 GB | 28 · 10.2 GB RAM | 17 · 6.41 GB RAM | 5 · 2.24 GB RAM | fits whole |
| Qwen3.6 35B-A3B Q8_0 | 34.4 GB | 35 · 28.4 GB RAM | 30 · 24.4 GB RAM | 25 · 20.4 GB RAM | 15 · 12.5 GB RAM |
| Ornith 1.5 35B-A3B Q4_K_M | 20.2 GB | 31 · 14.2 GB RAM | 22 · 10.2 GB RAM | 13 · 6.20 GB RAM | fits whole |
| Ornith 1.5 35B-A3B Q8_0 | 35.2 GB | 36 · 29.2 GB RAM | 31 · 25.2 GB RAM | 26 · 21.2 GB RAM | 16 · 13.3 GB RAM |
| Qwen3 30B-A3B Q4_K_M | 17.3 GB | 39 · 13.3 GB RAM | 27 · 9.33 GB RAM | 15 · 5.34 GB RAM | fits whole |
| Qwen3 30B-A3B Q8_0 | 30.2 GB | 44 · 26.6 GB RAM | 37 · 22.4 GB RAM | 31 · 18.8 GB RAM | 17 · 10.5 GB RAM |
| Gemma 4 26B-A4B UD-Q4_K_M | 15.8 GB | 21 · 10.0 GB RAM | 12 · 6.05 GB RAM | 3 · 2.06 GB RAM | fits whole |
| GLM-4.7 Flash Q4_K_M | 17.0 GB | 36 · 11.9 GB RAM | 24 · 7.93 GB RAM | 12 · 3.94 GB RAM | fits whole |
| GLM-4.7 Flash UD-Q4_K_XL | 16.3 GB | 35 · 11.1 GB RAM | 23 · 7.24 GB RAM | 10 · 3.13 GB RAM | fits whole |
| Nemotron 3 Nano 30B-A3B Q4_K_M | 22.9 GB | 41 · 16.5 GB RAM | 30 · 12.2 GB RAM | 21 · 8.70 GB RAM | fits whole |
| gpt-oss-20b MXFP4 | 11.3 GB | 12 · 5.31 GB RAM | 2 · 1.36 GB RAM | fits whole | fits whole |
| gpt-oss-120b MXFP4 | 59.0 GB | 34 · 54.3 GB RAM | 31 · 49.6 GB RAM | 29 · 46.4 GB RAM | 24 · 38.5 GB RAM |
| Qwen3-Coder-Next Q4_K_M | 45.2 GB | 44 · 39.9 GB RAM | 39 · 35.3 GB RAM | 35 · 31.6 GB RAM | 26 · 23.6 GB RAM |
| Qwen3-Coder-Next IQ4_XS | 39.7 GB | 42 · 33.6 GB RAM | 37 · 29.6 GB RAM | 32 · 25.7 GB RAM | 22 · 17.7 GB RAM |
| Qwen3-Coder-Next UD-Q4_K_XL | 46.2 GB | 44 · 40.2 GB RAM | 40 · 36.6 GB RAM | 35 · 32.0 GB RAM | 27 · 24.8 GB RAM |
| Qwen3.5 122B-A10B Q4_K_M | 71.3 GB | 48 · 66.0 GB RAM | 45 · 61.9 GB RAM | 42 · 57.8 GB RAM | 36 · 49.7 GB RAM |
| Qwen3.8 Flash Next UD-Q4_K_XL | 103.7 GB | 47 · 70.6 GB RAM | 45 · 67.5 GB RAM | 42 · 63.1 GB RAM | 37 · 55.8 GB RAM |
| Qwen3.8 Flash Next UD-IQ4_XS | 87.2 GB | 47 · 54.6 GB RAM | 44 · 50.8 GB RAM | 40 · 46.4 GB RAM | 33 · 38.6 GB RAM |
| Qwen3.8 Flash Next UD-IQ3_XXS | 76.3 GB | 46 · 43.9 GB RAM | 42 · 40.1 GB RAM | 37 · 35.4 GB RAM | 29 · 27.9 GB RAM |
| Qwen3.8 Flash Next UD-Q2_K_XL | 73.4 GB | 45 · 40.7 GB RAM | 41 · 37.1 GB RAM | 36 · 32.6 GB RAM | 27 · 24.6 GB RAM |
| Qwen3.8 Flash Next UD-Q3_K_XL | 83.8 GB | 47 · 51.2 GB RAM | 44 · 47.7 GB RAM | 40 · 43.5 GB RAM | 32 · 35.2 GB RAM |
| Qwen3.8 Flash Next UD-Q5_K_XL | 147.4 GB | 48 · 92.2 GB RAM | 45 · 86.5 GB RAM | 43 · 82.7 GB RAM | 39 · 75.1 GB RAM |
| Qwen3.8 Flash Next GSQ-RCO IQ3_S | 77.9 GB | 46 · 44.8 GB RAM | 43 · 41.6 GB RAM | 39 · 37.5 GB RAM | 31 · 29.6 GB RAM |
Tensor sizes read from the GGUF headers on Hugging Face on : Qwen3.6-35B-A3B-GGUF UD-Q4_K_M, Qwen3.6-35B-A3B-GGUF UD-IQ4_XS, Qwen3.6-35B-A3B-GGUF Q8_0, Ornith-1.5-35B-A3B-GGUF Q4_K_M, Ornith-1.5-35B-A3B-GGUF Q8_0, Qwen3-30B-A3B-GGUF Q4_K_M, Qwen3-30B-A3B-GGUF Q8_0, gemma-4-26B-A4B-it-GGUF UD-Q4_K_M, GLM-4.7-Flash-GGUF Q4_K_M, GLM-4.7-Flash-GGUF UD-Q4_K_XL, Nemotron-3-Nano-30B-A3B-GGUF Q4_K_M, gpt-oss-20b-GGUF MXFP4, gpt-oss-120b-GGUF MXFP4, Qwen3-Coder-Next-GGUF Q4_K_M, Qwen3-Coder-Next-GGUF IQ4_XS, Qwen3-Coder-Next-GGUF UD-Q4_K_XL, Qwen3.5-122B-A10B-GGUF Q4_K_M, Qwen3.8-Flash-Next-GGUF UD-Q4_K_XL, Qwen3.8-Flash-Next-GGUF UD-IQ4_XS, Qwen3.8-Flash-Next-GGUF UD-IQ3_XXS, Qwen3.8-Flash-Next-GGUF UD-Q2_K_XL, Qwen3.8-Flash-Next-GGUF UD-Q3_K_XL, Qwen3.8-Flash-Next-GGUF UD-Q5_K_XL, Qwen3.8-Flash-Next-GSQ-RCO-GGUF GSQ-RCO IQ3_S. For whole-model VRAM see the LLM VRAM calculator; for GPU-only speed, the LLM speed calculator.
How to use
- Pick the model file. Every layer’s expert tensors were measured from its GGUF header, so mixed quants such as unsloth UD are exact.
- Choose your GPU, your RAM speed, the context length and the KV cache type
(-ctk/-ctv). - Read the suggested --n-cpu-moe, the VRAM and RAM it uses and the speed range. The table shows every other value.
- Paste the command, then check llama.cpp’s load log; --fit (on by default) can also find a split for you.
Frequently asked questions
How do I choose --n-cpu-moe for a 16 GB GPU?
Use the smallest N that fits: for Qwen3.6-35B-A3B UD-Q4_K_M with 32,768 tokens, an FP16 KV cache and 1 GB of buffers, that is --n-cpu-moe 13, which moves 5.89 GB of experts to system RAM and uses 15.8 GB of VRAM. The file has 18.2 GB of expert tensors and 2.38 GB of everything else, so a 24 GB card holds it whole. A larger N frees more VRAM for context but reads more from slower system RAM.
What does --n-cpu-moe N do?
It keeps the routed-expert weights (the ffn_*_exps tensors) of the first N layers in system RAM and runs them on the CPU. Attention, the router, shared experts and the output head stay on the GPU. It was added to llama.cpp in PR #15077 (August 2025); --cpu-moe does the same for every layer.
Why offload experts rather than whole layers?
Each token uses only a few experts, for example 8 of 256 in Qwen3.6-35B-A3B, so the CPU reads only about 3% of each offloaded layer’s expert bytes per token. The attention and shared weights that every token needs stay on the fast GPU memory.
Does N count layers from the start?
Yes: layers 0 to N−1, by index, and layers without experts move nothing. GLM-4.7 Flash’s first layer is dense, and in Nemotron 3 Nano 29 of the 52 layers are Mamba or attention layers with no experts, so N counts more layers there than it moves.
How accurate is the --n-cpu-moe estimate?
The weight bytes are exact. The KV cache follows each model’s config, including hybrid and sliding-window layers. Compute buffers and the CUDA context are a setting (1 GB by default), and speeds are estimates from memory bandwidth: 30–50% of the limit, as llama.cpp reaches with MoE models.
More calculators
- Fine-Tuning VRAM Calculator GPU memory for full fine-tuning, LoRA and QLoRA in Transformers or Unsloth, checked against published runs. Open →
- vLLM KV Cache & Concurrency Calculator The KV cache pool vLLM allocates, in tokens, and how many requests fit at once, with the vllm serve command. Open →
- What Can My PC Run? Detects your GPU in the browser and lists the local LLMs it runs, with the best quantization and speed. Open →
- MiniMax H3 VRAM Calculator
(ComfyUI) VRAM and system RAM for MiniMax H3 video in ComfyUI: pruned, INT8, NVFP4 and GGUF files on 8–96 GB GPUs. Open → - Qwen3.8 27B GGUF Quants: Bonsai 2 vs GSQ-RCO vs UD Ternary Bonsai 2, GSQ-RCO and Unsloth UD files of Qwen3.8 27B: VRAM, quality, speed and the engine each needs. Open →
- Qwen-Image-2.1 VRAM Calculator Peak VRAM for every DiT, text encoder and VAE combination, with real measurements. Open →
Updated