Can I run gpt-oss-20b on an RTX 4070 12GB?
With --n-cpu-moe 2: the MXFP4 GGUF keeps 11.7 GB on the RTX 4070 and 1.36 GB in system RAM, with the experts of 2 of its 24 layers moved there. Whole, gpt-oss-20b needs about 15.4 GB at MXFP4 with 32K tokens of context, 3.44 GB more than the RTX 4070 12GB holds. The smallest setup here that holds gpt-oss-20b at MXFP4 with 32K is RTX 4060 Ti 16GB.
Why 15.4 GB with vLLM, but 11.7 GB on the RTX 4070 with llama.cpp
The 15.4 GB counts the published checkpoint as vLLM loads it: 12.8 GB of weights, the 32K KV cache, 0.5 GB and 10% overhead. llama.cpp's measured MXFP4 GGUF is 11.3 GB, and its 587 MB of token embeddings stay in RAM. With every layer on the RTX 4070 it takes 12.5 GB at 32K with 1 GB of buffers, 471 MB more than it holds, so the experts of 2 layers (809 MB) move to system RAM. That leaves 338 MB of the 12 GB free, and the PC needs 1.36 GB of free RAM for those experts and the embeddings.
With offload the MXFP4 GGUF with --n-cpu-moe 2 and 32K tokens of context
- On the card
- 11.7 GB
- RTX 4070
- 12 GB, 504 GB/s
- In RAM
- 1.36 GB
- Tokens/s
- 37–63 tokens/s
With the MXFP4 GGUF, --n-cpu-moe 2 and 32K tokens of context it writes about 37–63 tokens/s for one request on an RTX 4070 12GB and dual-channel DDR5-5600.
gpt-oss-20b on the RTX 4070 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; the MXFP4 GGUF columns are the measured llama.cpp file with every layer on the GPU and 1 GB of buffers; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | MXFP4 | Free | Tokens/s | MXFP4 GGUF | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 14.7 GB | −2.72 GB | — | 11.8 GB | 201 MB (tight) | 52–89 |
| 8K | 14.8 GB | −2.82 GB | — | 11.9 GB | 105 MB (tight) | 50–86 |
| 16K | 15.0 GB | −3.03 GB | — | 12.1 GB | −87 MB | — |
| 32K | 15.4 GB | −3.44 GB | — | 12.5 GB | −471 MB | — |
| 64K | 16.3 GB | −4.27 GB | — | 13.2 GB | −1.21 GB | — |
| 128K (full) | 17.9 GB | −5.92 GB | — | 14.7 GB | −2.71 GB | — |
--n-cpu-moe for gpt-oss-20b on the RTX 4070
With --n-cpu-moe 2 the MXFP4 GGUF keeps 11.7 GB on the RTX 4070 and 1.36 GB in RAM, about 37–63 tokens/s with DDR5-5600. Measured MXFP4 file, 1 GB of buffers; tokens/s by system RAM speed.
gpt-oss-20b on more than one RTX 4070
| Cards | Best at 32K | Tokens/s | MXFP4 longest | MXFP4 tokens/s |
|---|---|---|---|---|
| 2× (24 GB) | MXFP4 | 71–131 | 128K (full) | 71–131 |
| 4× (48 GB) | MXFP4 | 114–224 | 128K (full) | 114–224 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- Smallest setup for MXFP4 at 32K: RTX 4060 Ti 16GB.
Run gpt-oss-20b on the RTX 4070 with llama-server
llama-server -hf ggml-org/gpt-oss-20b-GGUF:MXFP4 -c 32768 --n-cpu-moe 2 -np 1 gpt-oss-20b-MXFP4.gguf, 12.1 GB, from ggml-org/
Questions
Can I run gpt-oss-20b on an RTX 4070 12GB?
With --n-cpu-moe 2: the MXFP4 GGUF keeps 11.7 GB on the RTX 4070 and 1.36 GB in system RAM, with the experts of 2 of its 24 layers moved there. Whole, gpt-oss-20b needs about 15.4 GB at MXFP4 with 32K tokens of context, 3.44 GB more than the RTX 4070 12GB holds. The smallest setup here that holds gpt-oss-20b at MXFP4 with 32K is RTX 4060 Ti 16GB.
How fast is gpt-oss-20b on an RTX 4070 12GB?
With the MXFP4 GGUF, --n-cpu-moe 2 and 32K tokens of context it writes about 37–63 tokens/s for one request on an RTX 4070 12GB and dual-channel DDR5-5600.
Does gpt-oss-20b need --n-cpu-moe on an RTX 4070 12GB?
With --n-cpu-moe 2 the MXFP4 GGUF keeps 11.7 GB on the RTX 4070 and 1.36 GB in RAM, about 37–63 tokens/s with DDR5-5600.
What does a second RTX 4070 12GB change for gpt-oss-20b?
Two RTX 4070 12GB cards (24 GB in one tensor-parallel group) hold gpt-oss-20b at MXFP4 with 32K, and MXFP4 up to 128K (full) tokens, at about 71–131 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also gpt-oss-20b VRAM requirements, what LLMs an RTX 4070 12GB can run and every pair, or detect your own GPU. Model data checked .