Can I run gpt-oss-20b on an RTX 5090?
Yes: gpt-oss-20b needs about 15.4 GB at MXFP4 with 32K tokens of context, which fits the 32 GB RTX 5090 with 16.6 GB to spare. MXFP4 is the format gpt-oss-20b is published in, and on the RTX 5090 it runs up to 128K (full) tokens with 0.5 GB to spare.
Why 15.4 GB with vLLM, but 12.5 GB on the RTX 5090 with llama.cpp
The 15.4 GB counts the published checkpoint as vLLM loads it: 12.8 GB of weights, the 32K KV cache, 0.5 GB and 10% overhead. llama.cpp's measured MXFP4 GGUF is 11.3 GB, and its 587 MB of token embeddings stay in RAM. With every layer on the RTX 5090 it takes 12.5 GB at 32K with 1 GB of buffers, 19.5 GB under its 32 GB.
Yes MXFP4 with 32K tokens of context
- MXFP4, 32K
- 15.4 GB
- RTX 5090
- 32 GB, 1,792 GB/s
- To spare
- 16.6 GB
- Tokens/s
- 134–246 tokens/s
At MXFP4 with 32K tokens of context it writes about 134–246 tokens/s for one request on an RTX 5090.
Best precision for gpt-oss-20b on an RTX 5090
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | MXFP4 | 14.8 GB | 17.2 GB | 158–295 |
| 32K | MXFP4 | 15.4 GB | 16.6 GB | 134–246 |
| 128K (full) | MXFP4 | 17.9 GB | 14.1 GB | 84–148 |
gpt-oss-20b on the RTX 5090 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; the MXFP4 GGUF columns are the measured llama.cpp file with every layer on the GPU and 1 GB of buffers; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | MXFP4 | Free | Tokens/s | MXFP4 GGUF | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 14.7 GB | 17.3 GB | 163–305 | 11.8 GB | 20.2 GB | 154–285 |
| 8K | 14.8 GB | 17.2 GB | 158–295 | 11.9 GB | 20.1 GB | 149–276 |
| 16K | 15.0 GB | 17.0 GB | 149–277 | 12.1 GB | 19.9 GB | 141–260 |
| 32K | 15.4 GB | 16.6 GB | 134–246 | 12.5 GB | 19.5 GB | 128–233 |
| 64K | 16.3 GB | 15.7 GB | 112–202 | 13.2 GB | 18.8 GB | 107–193 |
| 128K (full) | 17.9 GB | 14.1 GB | 84–148 | 14.7 GB | 17.3 GB | 81–143 |
--n-cpu-moe for gpt-oss-20b on the RTX 5090
The measured MXFP4 GGUF fits the RTX 5090 whole at 32K (12.5 GB), so --n-cpu-moe is not needed there. Measured MXFP4 file, 1 GB of buffers; tokens/s by system RAM speed.
gpt-oss-20b on more than one RTX 5090
| Cards | Best at 32K | Tokens/s | MXFP4 longest | MXFP4 tokens/s | Rent per hour |
|---|---|---|---|---|---|
| 2× (64 GB) | MXFP4 | 155–324 | 128K (full) | 155–324 | $1.38 |
| 4× (128 GB) | MXFP4 | 201–456 | 128K (full) | 201–456 | $2.76 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- Renting an RTX 5090 costs about $0.69 an hour (median on getdeploying.com, 2026-09-29): $0.78–1.4 per million tokens at 134–246 tokens/s.
- Smallest setup for MXFP4 at 32K: RTX 4060 Ti 16GB.
Run gpt-oss-20b on the RTX 5090 with llama-server
llama-server -hf ggml-org/gpt-oss-20b-GGUF:MXFP4 -c 131072 -np 1 gpt-oss-20b-MXFP4.gguf, 12.1 GB, from ggml-org/
Questions
Can I run gpt-oss-20b on an RTX 5090?
Yes: gpt-oss-20b needs about 15.4 GB at MXFP4 with 32K tokens of context, which fits the 32 GB RTX 5090 with 16.6 GB to spare. MXFP4 is the format gpt-oss-20b is published in, and on the RTX 5090 it runs up to 128K (full) tokens with 0.5 GB to spare.
How fast is gpt-oss-20b on an RTX 5090?
At MXFP4 with 32K tokens of context it writes about 134–246 tokens/s for one request on an RTX 5090.
Does gpt-oss-20b need --n-cpu-moe on an RTX 5090?
The measured MXFP4 GGUF fits the RTX 5090 whole at 32K (12.5 GB), so --n-cpu-moe is not needed there.
What does a second RTX 5090 change for gpt-oss-20b?
Two RTX 5090 cards (64 GB in one tensor-parallel group) hold gpt-oss-20b at MXFP4 with 32K, and MXFP4 up to 128K (full) tokens, at about 155–324 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also gpt-oss-20b VRAM requirements, what LLMs an RTX 5090 can run and every pair, or detect your own GPU. Model data checked .