Running big MoE models on one GPU: --n-cpu-moe, -ot and how much VRAM and RAM you need
gpt-oss-120b's 59.0 GB MXFP4 file is 96% routed-expert weights, so on one 24 GB RTX 3090, --n-cpu-moe 24 keeps the experts of its first 24 layers in system RAM and the model needs about 22.7 GB of VRAM and 38.5 GB of RAM at 32K context, for an estimated 16–27 tokens/s with dual-channel DDR5-5600; a 16 GB card needs --n-cpu-moe 29 and 46.4 GB of RAM.
The numbers here come from the MoE offload calculator, which reads every layer's expert tensors from the GGUF header; use it for your own file, card and context. This page is about which flag to use and how to choose its value.
Where the bytes of a MoE model are
In a mixture-of-experts model most of the file is routed experts, and each token uses only a few of them. The attention, the router, shared experts and the output head are small and used by every token, so they belong on the GPU; the input embedding stays in RAM either way.
| File | Total | Routed experts | Everything else | Experts read per token |
|---|---|---|---|---|
| gpt-oss-120b MXFP4, 36 layers | 59.0 GB | 56.9 GB (96%) | 2.14 GB | 4 of 128 (3.1%) |
| Qwen3.6 35B-A3B UD-Q4_K_M, 40 layers | 20.6 GB | 18.2 GB (88%) | 2.38 GB | 8 of 256 (3.1%) |
| GLM-4.7 Flash Q4_K_M, 47 layers | 17.0 GB | 15.6 GB (92%) | 1.43 GB | 4 of 64 (6.3%) |
| Qwen3.5 122B-A10B Q4_K_M, 48 layers | 71.3 GB | 65.3 GB (92%) | 6.02 GB | 8 of 256 (3.1%) |
| Qwen3.8 Flash Next UD-Q4_K_XL, 48 layers | 104 GB | 71.7 GB (69%) | 31.9 GBof which 26.8 GB is an n-gram table kept in RAM | 10 of 512 (2.0%) |
The flags
Read at llama.cpp 19e28a2 (2026-09-29).
-
--cpu-moe(-cmoe): “keep all Mixture of Experts (MoE) weights in the CPU” (common/arg.cpp:2756-2762 ). -
--n-cpu-moe N(-ncmoe): “keep the Mixture of Experts (MoE) weights of the first N layers in the CPU” (arg.cpp:2763-2772). It adds one override per layer 0 to N−1 for the tensors matching\.ffn_(up|down|gate|gate_up)_(ch|)exps(common/common.h:1199-1218 ). Layers without experts count towards N but move nothing. -
-ot/--override-tensortakes comma-separatedpattern=buffer typepairs (arg.cpp:2750-2755, 252-284); an unknown buffer type makes llama.cpp print the ones available. Each tensor name is regex-searched against the patterns and the first match decides (src/llama-model-loader.cpp:1234-1260 ).--n-cpu-moe 3is the same as-ot "blk\.(0|1|2)\.ffn_(up|down|gate|gate_up)_(ch|)exps=CPU". - With
--fit(on by default): if you set nothing, it puts the MoE tensors in system memory and then moves experts back to the GPU front to back while they fit (common/fit.cpp:488-490 , 735-738). If you set-ngl,--n-cpu-moe,--cpu-moeor-otand the model still does not fit after shrinking the context, it stops with “tensor_buft_overrides already set by user” (463-486) and logs “failed to fit params to free device memory” (894-896): the placement is then yours. The --fit preview replays what it would pick. - Memory mapping: with CPU overrides and mmap on, llama.cpp warns “consider using --load-mode none for better performance” (model-loader.cpp:1246). One report found that without mmap the OS could swap the experts out under sustained load when RAM was short (llama.cpp #26110).
- Several GPUs: any tensor override turns llama.cpp's pipeline parallelism off
(src/
llama-context.cpp:429-434 ); how layers split between cards is in splitting a model across GPUs.
N, VRAM and RAM for five models on 16, 24 and 32 GB cards
The smallest --n-cpu-moe that fits, at the calculator's defaults: 32,768 tokens of context, FP16 KV
cache, 1.00 GB for compute buffers and the CUDA context, the file memory-mapped. RAM is the moved experts plus the input
embedding; “64 GB / 128 GB” says whether that leaves 6 GB for the OS. Speeds are the calculator's
estimate with DDR5-5600, 2 channels (90 GB/s). Each N opens the calculator with that file and card.
| Model | RTX 5060 Ti 16GB (16 GB) | RTX 3090 (24 GB) | RTX 5090 (32 GB) |
|---|---|---|---|
| gpt-oss-120b MXFP4 | --n-cpu-moe 29 VRAM 14.8 GB, RAM 46.4 GB; 64 GB yes, 128 GB yes; 12–20 tokens/s | --n-cpu-moe 24 VRAM 22.7 GB, RAM 38.5 GB; 64 GB yes, 128 GB yes; 16–27 tokens/s | --n-cpu-moe 19 VRAM 30.6 GB, RAM 30.6 GB; 64 GB yes, 128 GB yes; 22–37 tokens/s |
| Qwen3.6 35B-A3B UD-Q4_K_M | --n-cpu-moe 13 VRAM 15.8 GB, RAM 6.39 GB; 64 GB yes, 128 GB yes; 31–53 tokens/s | whole on GPU VRAM 21.7 GB, RAM 515 MB; 64 GB yes, 128 GB yes; 76–133 tokens/s | whole on GPU VRAM 21.7 GB, RAM 515 MB; 64 GB yes, 128 GB yes; 131–239 tokens/s |
| GLM-4.7 Flash Q4_K_M | --n-cpu-moe 12 VRAM 15.8 GB, RAM 3.94 GB; 64 GB yes, 128 GB yes; 25–42 tokens/s | whole on GPU VRAM 19.5 GB, RAM 170 MB; 64 GB yes, 128 GB yes; 61–106 tokens/s | whole on GPU VRAM 19.5 GB, RAM 170 MB; 64 GB yes, 128 GB yes; 108–194 tokens/s |
| Qwen3.5 122B-A10B Q4_K_M | --n-cpu-moe 42 VRAM 15.2 GB, RAM 57.8 GB; 64 GB yes, 128 GB yes; 8–14 tokens/s | --n-cpu-moe 36 VRAM 23.3 GB, RAM 49.7 GB; 64 GB yes, 128 GB yes; 11–19 tokens/s | --n-cpu-moe 30 VRAM 31.5 GB, RAM 41.5 GB; 64 GB yes, 128 GB yes; 15–26 tokens/s |
| Qwen3.8 Flash Next UD-Q4_K_XL | --n-cpu-moe 42 VRAM 15.5 GB, RAM 63.1 GB; 64 GB no, 128 GB yes; 11–18 tokens/s | --n-cpu-moe 37 VRAM 22.8 GB, RAM 55.8 GB; 64 GB yes, 128 GB yes; 15–26 tokens/s | --n-cpu-moe 31 VRAM 31.6 GB, RAM 47.0 GB; 64 GB yes, 128 GB yes; 20–34 tokens/s |
Qwen3.8 Flash Next's 26.8 GB n-gram table stays in the mapped file and is read as needed; with
--load-mode none (no mmap) it is copied into RAM as well. The 1 GB for buffers and context is a setting in the calculator:
on Windows or with a desktop on the card, raise it (why a card does not give a model all its memory). A q8_0
KV cache frees context memory for fewer moved layers (KV cache quantization).
The trade-off: VRAM for speed
gpt-oss-120b on an RTX 3090 with DDR5-5600, 2 channels (90 GB/s), every sixth N (✗ = over the card's 24 GB):
| --n-cpu-moe | VRAM | RAM | Estimated tokens/s | Upper limit |
|---|---|---|---|---|
| 0 ✗ | 60.6 GB | 587 MB | 53–92 | 194 |
| 6 ✗ | 51.1 GB | 10.1 GB | 34–58 | 119 |
| 12 ✗ | 41.6 GB | 19.5 GB | 25–42 | 86 |
| 18 ✗ | 32.2 GB | 29.0 GB | 20–33 | 68 |
| 24 | 22.7 GB | 38.5 GB | 16–27 | 56 |
| 30 | 13.2 GB | 48.0 GB | 14–23 | 47 |
| 36 | 3.72 GB | 57.5 GB | 12–20 | 41 |
How the speed is estimated: per token, the GPU reads its non-expert weights, the routed share of its own experts and the KV cache at the card's bandwidth, and the CPU reads the routed share of the moved experts at RAM bandwidth; the upper limit is one over that time, and the range is 30–50% of it plus a fixed cost per token, what llama.cpp reaches with MoE models. It assumes one request, the CPU keeping up with its memory, and no prompt processing. Once most experts are in RAM, the RAM term dominates, so more memory channels matter more than a faster GPU.
Checked against public reports
- llama.cpp #15253 (2025-08-11): gpt-oss-120b MXFP4 (lmstudio-community's file) on an RTX 3080 Ti with
--n-cpu-moe 33loggedCUDA0 model buffer size = 6461.28 MiB; the calculator's GPU weights for that N, from ggml-org's MXFP4 file, are 6461.25 MiB. - Qwen3.8-Flash-Next UD-IQ4_XS on a Strix Halo (Vulkan), every layer on the GPU: Vulkan0 model buffer 61,222.07 MiB, CPU 27,465.95 MiB (the n-gram table), Vulkan_Host 644.14 MiB (the input embedding), each equal to the GGUF header to 0.01 MiB.
- RTX 4080 16 GB + EPYC 7B13 (memory configuration not stated), unsloth gpt-oss-120b "F16" GGUF, --n-cpu-moe 30 (2025-08-25): 39.0–40.8 tokens/s generating. Not the same hardware or file as the table, so not a check of its speeds.
- RTX 3090 + RTX 5060 Ti, Ryzen 9 9900X, 64 GB dual-channel DDR5, ggml-org MXFP4, --tensor-split 19,7 --n-cpu-moe 17 (2025-12-07): 24.72 tokens/s generating. Not the same hardware or file as the table, so not a check of its speeds.
- llama.cpp #19035 (2026-01-23): with an RTX 4090 and 64 GB RAM, a build that found no GPU put all 60,438.47 MiB of gpt-oss-120b in RAM and the process was killed for running out of memory: the whole file does not fit in 64 GB next to the OS, which is what moving only some layers avoids.
Questions
How do I pick N for --n-cpu-moe?
Take the smallest N that fits your card with the context you want: every layer moved to RAM makes each token slower. At 32K context with an FP16 cache and 1 GB for buffers, gpt-oss-120b needs --n-cpu-moe 24 on 24 GB and 29 on 16 GB, Qwen3.6 35B-A3B at UD-Q4_K_M needs 13 on 16 GB and fits a 24 GB card whole. The MoE offload calculator gives N for any file, card and context.
What is the difference between --cpu-moe, --n-cpu-moe and -ot?
--cpu-moe keeps every routed-expert tensor in system RAM. --n-cpu-moe N does that for layers 0 to N−1 only. Both are shortcuts for -ot, which takes "regex=buffer type" pairs: --n-cpu-moe N adds one pattern per layer, blk.i.ffn_(up|down|gate|gate_up)_(ch|)exps=CPU. Use -ot directly to pick other tensors or other layers.
Do I need --n-cpu-moe if llama-server has --fit?
Not always. With nothing set, --fit first tries to fit by shrinking the context, then puts the MoE tensors in system memory and moves experts back to the GPU while there is room, which is what --n-cpu-moe does by hand. Once you pass -ngl, --n-cpu-moe or -ot, --fit leaves the tensor placement to you and logs "failed to fit params to free device memory" if it would have had to change it.
How much system RAM do I need?
The experts you move plus the input embedding, plus room for the OS: this guide counts a machine as big enough when that total leaves 6 GB free. Qwen3.8 Flash Next at UD-Q4_K_XL on a 16 GB card needs 63.1 GB of RAM, more than a 64 GB machine has to spare; everything else here fits in 64 GB.
How fast is a MoE model with experts in RAM?
Each token reads only the experts it routes to, a few percent of each layer, but the part in RAM is read at RAM speed. The estimate here is the time to read those bytes over GPU and RAM bandwidth, at the 30–50% of that limit llama.cpp reaches with MoE models: 16–27 tokens/s for gpt-oss-120b on an RTX 3090 with dual-channel DDR5-5600. More memory channels help directly.
Sources read on .