Running big MoE models on one GPU: --n-cpu-moe, -ot and how much VRAM and RAM you need

gpt-oss-120b's 59.0 GB MXFP4 file is 96% routed-expert weights, so on one 24 GB RTX 3090, --n-cpu-moe 24 keeps the experts of its first 24 layers in system RAM and the model needs about 22.7 GB of VRAM and 38.5 GB of RAM at 32K context, for an estimated 16–27 tokens/s with dual-channel DDR5-5600; a 16 GB card needs --n-cpu-moe 29 and 46.4 GB of RAM.

The numbers here come from the MoE offload calculator, which reads every layer's expert tensors from the GGUF header; use it for your own file, card and context. This page is about which flag to use and how to choose its value.

Where the bytes of a MoE model are

In a mixture-of-experts model most of the file is routed experts, and each token uses only a few of them. The attention, the router, shared experts and the output head are small and used by every token, so they belong on the GPU; the input embedding stays in RAM either way.

FileTotalRouted expertsEverything elseExperts read per token
gpt-oss-120b MXFP4, 36 layers 59.0 GB 56.9 GB (96%) 2.14 GB 4 of 128 (3.1%)
Qwen3.6 35B-A3B UD-Q4_K_M, 40 layers 20.6 GB 18.2 GB (88%) 2.38 GB 8 of 256 (3.1%)
GLM-4.7 Flash Q4_K_M, 47 layers 17.0 GB 15.6 GB (92%) 1.43 GB 4 of 64 (6.3%)
Qwen3.5 122B-A10B Q4_K_M, 48 layers 71.3 GB 65.3 GB (92%) 6.02 GB 8 of 256 (3.1%)
Qwen3.8 Flash Next UD-Q4_K_XL, 48 layers 104 GB 71.7 GB (69%) 31.9 GBof which 26.8 GB is an n-gram table kept in RAM 10 of 512 (2.0%)

The flags

Read at llama.cpp 19e28a2 (2026-09-29).

N, VRAM and RAM for five models on 16, 24 and 32 GB cards

The smallest --n-cpu-moe that fits, at the calculator's defaults: 32,768 tokens of context, FP16 KV cache, 1.00 GB for compute buffers and the CUDA context, the file memory-mapped. RAM is the moved experts plus the input embedding; “64 GB / 128 GB” says whether that leaves 6 GB for the OS. Speeds are the calculator's estimate with DDR5-5600, 2 channels (90 GB/s). Each N opens the calculator with that file and card.

ModelRTX 5060 Ti 16GB (16 GB)RTX 3090 (24 GB)RTX 5090 (32 GB)
gpt-oss-120b MXFP4 --n-cpu-moe 29 VRAM 14.8 GB, RAM 46.4 GB; 64 GB yes, 128 GB yes; 12–20 tokens/s --n-cpu-moe 24 VRAM 22.7 GB, RAM 38.5 GB; 64 GB yes, 128 GB yes; 16–27 tokens/s --n-cpu-moe 19 VRAM 30.6 GB, RAM 30.6 GB; 64 GB yes, 128 GB yes; 22–37 tokens/s
Qwen3.6 35B-A3B UD-Q4_K_M --n-cpu-moe 13 VRAM 15.8 GB, RAM 6.39 GB; 64 GB yes, 128 GB yes; 31–53 tokens/s whole on GPU VRAM 21.7 GB, RAM 515 MB; 64 GB yes, 128 GB yes; 76–133 tokens/s whole on GPU VRAM 21.7 GB, RAM 515 MB; 64 GB yes, 128 GB yes; 131–239 tokens/s
GLM-4.7 Flash Q4_K_M --n-cpu-moe 12 VRAM 15.8 GB, RAM 3.94 GB; 64 GB yes, 128 GB yes; 25–42 tokens/s whole on GPU VRAM 19.5 GB, RAM 170 MB; 64 GB yes, 128 GB yes; 61–106 tokens/s whole on GPU VRAM 19.5 GB, RAM 170 MB; 64 GB yes, 128 GB yes; 108–194 tokens/s
Qwen3.5 122B-A10B Q4_K_M --n-cpu-moe 42 VRAM 15.2 GB, RAM 57.8 GB; 64 GB yes, 128 GB yes; 8–14 tokens/s --n-cpu-moe 36 VRAM 23.3 GB, RAM 49.7 GB; 64 GB yes, 128 GB yes; 11–19 tokens/s --n-cpu-moe 30 VRAM 31.5 GB, RAM 41.5 GB; 64 GB yes, 128 GB yes; 15–26 tokens/s
Qwen3.8 Flash Next UD-Q4_K_XL --n-cpu-moe 42 VRAM 15.5 GB, RAM 63.1 GB; 64 GB no, 128 GB yes; 11–18 tokens/s --n-cpu-moe 37 VRAM 22.8 GB, RAM 55.8 GB; 64 GB yes, 128 GB yes; 15–26 tokens/s --n-cpu-moe 31 VRAM 31.6 GB, RAM 47.0 GB; 64 GB yes, 128 GB yes; 20–34 tokens/s

Qwen3.8 Flash Next's 26.8 GB n-gram table stays in the mapped file and is read as needed; with --load-mode none (no mmap) it is copied into RAM as well. The 1 GB for buffers and context is a setting in the calculator: on Windows or with a desktop on the card, raise it (why a card does not give a model all its memory). A q8_0 KV cache frees context memory for fewer moved layers (KV cache quantization).

The trade-off: VRAM for speed

gpt-oss-120b on an RTX 3090 with DDR5-5600, 2 channels (90 GB/s), every sixth N (✗ = over the card's 24 GB):

--n-cpu-moeVRAMRAMEstimated tokens/sUpper limit
0 ✗ 60.6 GB 587 MB 53–92 194
6 ✗ 51.1 GB 10.1 GB 34–58 119
12 ✗ 41.6 GB 19.5 GB 25–42 86
18 ✗ 32.2 GB 29.0 GB 20–33 68
24 22.7 GB 38.5 GB 16–27 56
30 13.2 GB 48.0 GB 14–23 47
36 3.72 GB 57.5 GB 12–20 41

How the speed is estimated: per token, the GPU reads its non-expert weights, the routed share of its own experts and the KV cache at the card's bandwidth, and the CPU reads the routed share of the moved experts at RAM bandwidth; the upper limit is one over that time, and the range is 30–50% of it plus a fixed cost per token, what llama.cpp reaches with MoE models. It assumes one request, the CPU keeping up with its memory, and no prompt processing. Once most experts are in RAM, the RAM term dominates, so more memory channels matter more than a faster GPU.

Checked against public reports

Questions

How do I pick N for --n-cpu-moe?

Take the smallest N that fits your card with the context you want: every layer moved to RAM makes each token slower. At 32K context with an FP16 cache and 1 GB for buffers, gpt-oss-120b needs --n-cpu-moe 24 on 24 GB and 29 on 16 GB, Qwen3.6 35B-A3B at UD-Q4_K_M needs 13 on 16 GB and fits a 24 GB card whole. The MoE offload calculator gives N for any file, card and context.

What is the difference between --cpu-moe, --n-cpu-moe and -ot?

--cpu-moe keeps every routed-expert tensor in system RAM. --n-cpu-moe N does that for layers 0 to N−1 only. Both are shortcuts for -ot, which takes "regex=buffer type" pairs: --n-cpu-moe N adds one pattern per layer, blk.i.ffn_(up|down|gate|gate_up)_(ch|)exps=CPU. Use -ot directly to pick other tensors or other layers.

Do I need --n-cpu-moe if llama-server has --fit?

Not always. With nothing set, --fit first tries to fit by shrinking the context, then puts the MoE tensors in system memory and moves experts back to the GPU while there is room, which is what --n-cpu-moe does by hand. Once you pass -ngl, --n-cpu-moe or -ot, --fit leaves the tensor placement to you and logs "failed to fit params to free device memory" if it would have had to change it.

How much system RAM do I need?

The experts you move plus the input embedding, plus room for the OS: this guide counts a machine as big enough when that total leaves 6 GB free. Qwen3.8 Flash Next at UD-Q4_K_XL on a 16 GB card needs 63.1 GB of RAM, more than a 64 GB machine has to spare; everything else here fits in 64 GB.

How fast is a MoE model with experts in RAM?

Each token reads only the experts it routes to, a few percent of each layer, but the part in RAM is read at RAM speed. The estimate here is the time to read those bytes over GPU and RAM bandwidth, at the 30–50% of that limit llama.cpp reaches with MoE models: 16–27 tokens/s for gpt-oss-120b on an RTX 3090 with dual-channel DDR5-5600. More memory channels help directly.

Sources read on .