MoE Offload Calculator (--n-cpu-moe)

Pick a MoE GGUF, your GPU, RAM and context length. The planner adds up the real size of every layer’s expert tensors, read from the file’s header, and finds the smallest --n-cpu-moe that fits your VRAM.

Qwen3.6 35B-A3B · UD-Q4_K_M · 20.6 GB

–

VRAM used
–
RAM for offloaded experts
–
KV cache
–
Writing speed, one request
–

Every --n-cpu-moe value

--n-cpu-moeVRAMRAMTokens/sStatus

Smallest --n-cpu-moe at 32K context

FP16 KV cache, one request, 1 GB for buffers. RAM is what the offloaded experts take, plus the input embedding. Qwen3.8 Flash Next also reads a 26.8 GB n-gram table from the mapped file, not counted here.

FileSize8 GB GPU12 GB GPU16 GB GPU24 GB GPU
Qwen3.6 35B-A3B UD-Q4_K_M20.6 GB31 · 14.6 GB RAM22 · 10.5 GB RAM13 · 6.39 GB RAMfits whole
Qwen3.6 35B-A3B UD-IQ4_XS16.5 GB28 · 10.2 GB RAM17 · 6.41 GB RAM5 · 2.24 GB RAMfits whole
Qwen3.6 35B-A3B Q8_034.4 GB35 · 28.4 GB RAM30 · 24.4 GB RAM25 · 20.4 GB RAM15 · 12.5 GB RAM
Ornith 1.5 35B-A3B Q4_K_M20.2 GB31 · 14.2 GB RAM22 · 10.2 GB RAM13 · 6.20 GB RAMfits whole
Ornith 1.5 35B-A3B Q8_035.2 GB36 · 29.2 GB RAM31 · 25.2 GB RAM26 · 21.2 GB RAM16 · 13.3 GB RAM
Qwen3 30B-A3B Q4_K_M17.3 GB39 · 13.3 GB RAM27 · 9.33 GB RAM15 · 5.34 GB RAMfits whole
Qwen3 30B-A3B Q8_030.2 GB44 · 26.6 GB RAM37 · 22.4 GB RAM31 · 18.8 GB RAM17 · 10.5 GB RAM
Gemma 4 26B-A4B UD-Q4_K_M15.8 GB21 · 10.0 GB RAM12 · 6.05 GB RAM3 · 2.06 GB RAMfits whole
GLM-4.7 Flash Q4_K_M17.0 GB36 · 11.9 GB RAM24 · 7.93 GB RAM12 · 3.94 GB RAMfits whole
GLM-4.7 Flash UD-Q4_K_XL16.3 GB35 · 11.1 GB RAM23 · 7.24 GB RAM10 · 3.13 GB RAMfits whole
Nemotron 3 Nano 30B-A3B Q4_K_M22.9 GB41 · 16.5 GB RAM30 · 12.2 GB RAM21 · 8.70 GB RAMfits whole
gpt-oss-20b MXFP411.3 GB12 · 5.31 GB RAM2 · 1.36 GB RAMfits wholefits whole
gpt-oss-120b MXFP459.0 GB34 · 54.3 GB RAM31 · 49.6 GB RAM29 · 46.4 GB RAM24 · 38.5 GB RAM
Qwen3-Coder-Next Q4_K_M45.2 GB44 · 39.9 GB RAM39 · 35.3 GB RAM35 · 31.6 GB RAM26 · 23.6 GB RAM
Qwen3-Coder-Next IQ4_XS39.7 GB42 · 33.6 GB RAM37 · 29.6 GB RAM32 · 25.7 GB RAM22 · 17.7 GB RAM
Qwen3-Coder-Next UD-Q4_K_XL46.2 GB44 · 40.2 GB RAM40 · 36.6 GB RAM35 · 32.0 GB RAM27 · 24.8 GB RAM
Qwen3.5 122B-A10B Q4_K_M71.3 GB48 · 66.0 GB RAM45 · 61.9 GB RAM42 · 57.8 GB RAM36 · 49.7 GB RAM
Qwen3.8 Flash Next UD-Q4_K_XL103.7 GB47 · 70.6 GB RAM45 · 67.5 GB RAM42 · 63.1 GB RAM37 · 55.8 GB RAM
Qwen3.8 Flash Next UD-IQ4_XS87.2 GB47 · 54.6 GB RAM44 · 50.8 GB RAM40 · 46.4 GB RAM33 · 38.6 GB RAM
Qwen3.8 Flash Next UD-IQ3_XXS76.3 GB46 · 43.9 GB RAM42 · 40.1 GB RAM37 · 35.4 GB RAM29 · 27.9 GB RAM
Qwen3.8 Flash Next UD-Q2_K_XL73.4 GB45 · 40.7 GB RAM41 · 37.1 GB RAM36 · 32.6 GB RAM27 · 24.6 GB RAM
Qwen3.8 Flash Next UD-Q3_K_XL83.8 GB47 · 51.2 GB RAM44 · 47.7 GB RAM40 · 43.5 GB RAM32 · 35.2 GB RAM
Qwen3.8 Flash Next UD-Q5_K_XL147.4 GB48 · 92.2 GB RAM45 · 86.5 GB RAM43 · 82.7 GB RAM39 · 75.1 GB RAM
Qwen3.8 Flash Next GSQ-RCO IQ3_S77.9 GB46 · 44.8 GB RAM43 · 41.6 GB RAM39 · 37.5 GB RAM31 · 29.6 GB RAM

Tensor sizes read from the GGUF headers on Hugging Face on : Qwen3.6-35B-A3B-GGUF UD-Q4_K_M, Qwen3.6-35B-A3B-GGUF UD-IQ4_XS, Qwen3.6-35B-A3B-GGUF Q8_0, Ornith-1.5-35B-A3B-GGUF Q4_K_M, Ornith-1.5-35B-A3B-GGUF Q8_0, Qwen3-30B-A3B-GGUF Q4_K_M, Qwen3-30B-A3B-GGUF Q8_0, gemma-4-26B-A4B-it-GGUF UD-Q4_K_M, GLM-4.7-Flash-GGUF Q4_K_M, GLM-4.7-Flash-GGUF UD-Q4_K_XL, Nemotron-3-Nano-30B-A3B-GGUF Q4_K_M, gpt-oss-20b-GGUF MXFP4, gpt-oss-120b-GGUF MXFP4, Qwen3-Coder-Next-GGUF Q4_K_M, Qwen3-Coder-Next-GGUF IQ4_XS, Qwen3-Coder-Next-GGUF UD-Q4_K_XL, Qwen3.5-122B-A10B-GGUF Q4_K_M, Qwen3.8-Flash-Next-GGUF UD-Q4_K_XL, Qwen3.8-Flash-Next-GGUF UD-IQ4_XS, Qwen3.8-Flash-Next-GGUF UD-IQ3_XXS, Qwen3.8-Flash-Next-GGUF UD-Q2_K_XL, Qwen3.8-Flash-Next-GGUF UD-Q3_K_XL, Qwen3.8-Flash-Next-GGUF UD-Q5_K_XL, Qwen3.8-Flash-Next-GSQ-RCO-GGUF GSQ-RCO IQ3_S. For whole-model VRAM see the LLM VRAM calculator; for GPU-only speed, the LLM speed calculator.

How to use

  1. Pick the model file. Every layer’s expert tensors were measured from its GGUF header, so mixed quants such as unsloth UD are exact.
  2. Choose your GPU, your RAM speed, the context length and the KV cache type (-ctk/-ctv).
  3. Read the suggested --n-cpu-moe, the VRAM and RAM it uses and the speed range. The table shows every other value.
  4. Paste the command, then check llama.cpp’s load log; --fit (on by default) can also find a split for you.

Frequently asked questions

How do I choose --n-cpu-moe for a 16 GB GPU?

Use the smallest N that fits: for Qwen3.6-35B-A3B UD-Q4_K_M with 32,768 tokens, an FP16 KV cache and 1 GB of buffers, that is --n-cpu-moe 13, which moves 5.89 GB of experts to system RAM and uses 15.8 GB of VRAM. The file has 18.2 GB of expert tensors and 2.38 GB of everything else, so a 24 GB card holds it whole. A larger N frees more VRAM for context but reads more from slower system RAM.

What does --n-cpu-moe N do?

It keeps the routed-expert weights (the ffn_*_exps tensors) of the first N layers in system RAM and runs them on the CPU. Attention, the router, shared experts and the output head stay on the GPU. It was added to llama.cpp in PR #15077 (August 2025); --cpu-moe does the same for every layer.

Why offload experts rather than whole layers?

Each token uses only a few experts, for example 8 of 256 in Qwen3.6-35B-A3B, so the CPU reads only about 3% of each offloaded layer’s expert bytes per token. The attention and shared weights that every token needs stay on the fast GPU memory.

Does N count layers from the start?

Yes: layers 0 to N−1, by index, and layers without experts move nothing. GLM-4.7 Flash’s first layer is dense, and in Nemotron 3 Nano 29 of the 52 layers are Mamba or attention layers with no experts, so N counts more layers there than it moves.

How accurate is the --n-cpu-moe estimate?

The weight bytes are exact. The KV cache follows each model’s config, including hybrid and sliding-window layers. Compute buffers and the CUDA context are a setting (1 GB by default), and speeds are estimates from memory bandwidth: 30–50% of the limit, as llama.cpp reaches with MoE models.

More calculators

Updated