Can I run gpt-oss-20b on an RTX 4070 12GB?

With --n-cpu-moe 2: the MXFP4 GGUF keeps 11.7 GB on the RTX 4070 and 1.36 GB in system RAM, with the experts of 2 of its 24 layers moved there. Whole, gpt-oss-20b needs about 15.4 GB at MXFP4 with 32K tokens of context, 3.44 GB more than the RTX 4070 12GB holds. The smallest setup here that holds gpt-oss-20b at MXFP4 with 32K is RTX 4060 Ti 16GB.

Why 15.4 GB with vLLM, but 11.7 GB on the RTX 4070 with llama.cpp

The 15.4 GB counts the published checkpoint as vLLM loads it: 12.8 GB of weights, the 32K KV cache, 0.5 GB and 10% overhead. llama.cpp's measured MXFP4 GGUF is 11.3 GB, and its 587 MB of token embeddings stay in RAM. With every layer on the RTX 4070 it takes 12.5 GB at 32K with 1 GB of buffers, 471 MB more than it holds, so the experts of 2 layers (809 MB) move to system RAM. That leaves 338 MB of the 12 GB free, and the PC needs 1.36 GB of free RAM for those experts and the embeddings.

With offload the MXFP4 GGUF with --n-cpu-moe 2 and 32K tokens of context

On the card
11.7 GB
RTX 4070
12 GB, 504 GB/s
In RAM
1.36 GB
Tokens/s
37–63 tokens/s

With the MXFP4 GGUF, --n-cpu-moe 2 and 32K tokens of context it writes about 37–63 tokens/s for one request on an RTX 4070 12GB and dual-channel DDR5-5600.

gpt-oss-20b on the RTX 4070 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; the MXFP4 GGUF columns are the measured llama.cpp file with every layer on the GPU and 1 GB of buffers; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context MXFP4FreeTokens/sMXFP4 GGUFFreeTokens/s
4K 14.7 GB −2.72 GB — 11.8 GB 201 MB (tight) 52–89
8K 14.8 GB −2.82 GB — 11.9 GB 105 MB (tight) 50–86
16K 15.0 GB −3.03 GB — 12.1 GB −87 MB —
32K 15.4 GB −3.44 GB — 12.5 GB −471 MB —
64K 16.3 GB −4.27 GB — 13.2 GB −1.21 GB —
128K (full) 17.9 GB −5.92 GB — 14.7 GB −2.71 GB —

--n-cpu-moe for gpt-oss-20b on the RTX 4070

With --n-cpu-moe 2 the MXFP4 GGUF keeps 11.7 GB on the RTX 4070 and 1.36 GB in RAM, about 37–63 tokens/s with DDR5-5600. Measured MXFP4 file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 0 11.9 GB 587 MB 50–8650–8650–86
32K 2 11.7 GB 1.36 GB 33–5637–6337–64
64K 4 11.6 GB 2.15 GB 24–4128–4729–49
128K (full) 7 11.9 GB 3.34 GB 16–2719–3320–34

gpt-oss-20b on more than one RTX 4070

CardsBest at 32KTokens/sMXFP4 longestMXFP4 tokens/s
2× (24 GB) MXFP4 71–131 128K (full) 71–131
4× (48 GB) MXFP4 114–224 128K (full) 114–224

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run gpt-oss-20b on the RTX 4070 with llama-server

llama-server -hf ggml-org/gpt-oss-20b-GGUF:MXFP4 -c 32768 --n-cpu-moe 2 -np 1

gpt-oss-20b-MXFP4.gguf, 12.1 GB, from ggml-org/gpt-oss-20b-GGUF (checked 2026-09-29); --n-cpu-moe 2 with -c 32768 is the plan above. -np 1: one slot, one sliding window.

Questions

Can I run gpt-oss-20b on an RTX 4070 12GB?

With --n-cpu-moe 2: the MXFP4 GGUF keeps 11.7 GB on the RTX 4070 and 1.36 GB in system RAM, with the experts of 2 of its 24 layers moved there. Whole, gpt-oss-20b needs about 15.4 GB at MXFP4 with 32K tokens of context, 3.44 GB more than the RTX 4070 12GB holds. The smallest setup here that holds gpt-oss-20b at MXFP4 with 32K is RTX 4060 Ti 16GB.

How fast is gpt-oss-20b on an RTX 4070 12GB?

With the MXFP4 GGUF, --n-cpu-moe 2 and 32K tokens of context it writes about 37–63 tokens/s for one request on an RTX 4070 12GB and dual-channel DDR5-5600.

Does gpt-oss-20b need --n-cpu-moe on an RTX 4070 12GB?

With --n-cpu-moe 2 the MXFP4 GGUF keeps 11.7 GB on the RTX 4070 and 1.36 GB in RAM, about 37–63 tokens/s with DDR5-5600.

What does a second RTX 4070 12GB change for gpt-oss-20b?

Two RTX 4070 12GB cards (24 GB in one tensor-parallel group) hold gpt-oss-20b at MXFP4 with 32K, and MXFP4 up to 128K (full) tokens, at about 71–131 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also gpt-oss-20b VRAM requirements, what LLMs an RTX 4070 12GB can run and every pair, or detect your own GPU. Model data checked .