Can I run gpt-oss-20b on an RX 7900 XTX?

Yes: gpt-oss-20b needs about 15.4 GB at MXFP4 with 32K tokens of context, which fits the 24 GB RX 7900 XTX with 8.56 GB to spare. MXFP4 is the format gpt-oss-20b is published in, and on the RX 7900 XTX it runs up to 128K (full) tokens with 0.5 GB to spare.

Why 15.4 GB with vLLM, but 12.5 GB on the RX 7900 XTX with llama.cpp

The 15.4 GB counts the published checkpoint as vLLM loads it: 12.8 GB of weights, the 32K KV cache, 0.5 GB and 10% overhead. llama.cpp's measured MXFP4 GGUF is 11.3 GB, and its 587 MB of token embeddings stay in RAM. With every layer on the RX 7900 XTX it takes 12.5 GB at 32K with 1 GB of buffers, 11.5 GB under its 24 GB.

Yes MXFP4 with 32K tokens of context

MXFP4, 32K
15.4 GB
RX 7900 XTX
24 GB, 960 GB/s
To spare
8.56 GB
Tokens/s
79–140 tokens/s

At MXFP4 with 32K tokens of context it writes about 79–140 tokens/s for one request on an RX 7900 XTX.

Best precision for gpt-oss-20b on an RX 7900 XTX

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K MXFP4 14.8 GB 9.18 GB 95–170
32K MXFP4 15.4 GB 8.56 GB 79–140
128K (full) MXFP4 17.9 GB 6.08 GB 48–82

gpt-oss-20b on the RX 7900 XTX as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; the MXFP4 GGUF columns are the measured llama.cpp file with every layer on the GPU and 1 GB of buffers; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context MXFP4FreeTokens/sMXFP4 GGUFFreeTokens/s
4K 14.7 GB 9.28 GB 99–176 11.8 GB 12.2 GB 92–164
8K 14.8 GB 9.18 GB 95–170 11.9 GB 12.1 GB 89–158
16K 15.0 GB 8.97 GB 89–158 12.1 GB 11.9 GB 84–148
32K 15.4 GB 8.56 GB 79–140 12.5 GB 11.5 GB 75–132
64K 16.3 GB 7.73 GB 65–113 13.2 GB 10.8 GB 62–108
128K (full) 17.9 GB 6.08 GB 48–82 14.7 GB 9.29 GB 46–79

--n-cpu-moe for gpt-oss-20b on the RX 7900 XTX

The measured MXFP4 GGUF fits the RX 7900 XTX whole at 32K (12.5 GB), so --n-cpu-moe is not needed there. Measured MXFP4 file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 0 11.9 GB 587 MB 89–15889–15889–158
32K 0 12.5 GB 587 MB 75–13275–13275–132
64K 0 13.2 GB 587 MB 62–10862–10862–108
128K (full) 0 14.7 GB 587 MB 46–7946–7946–79

gpt-oss-20b on more than one RX 7900 XTX

CardsBest at 32KTokens/sMXFP4 longestMXFP4 tokens/s
2× (48 GB) MXFP4 111–216 128K (full) 111–216
4× (96 GB) MXFP4 159–338 128K (full) 159–338

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run gpt-oss-20b on the RX 7900 XTX with llama-server

llama-server -hf ggml-org/gpt-oss-20b-GGUF:MXFP4 -c 131072 -ngl 99 -np 1

gpt-oss-20b-MXFP4.gguf, 12.1 GB, from ggml-org/gpt-oss-20b-GGUF (checked 2026-09-29). At -c 131072 on the RX 7900 XTX: 17.9 GB of 24 GB, 6.08 GB free. -np 1: one slot, one sliding window.

Questions

Can I run gpt-oss-20b on an RX 7900 XTX?

Yes: gpt-oss-20b needs about 15.4 GB at MXFP4 with 32K tokens of context, which fits the 24 GB RX 7900 XTX with 8.56 GB to spare. MXFP4 is the format gpt-oss-20b is published in, and on the RX 7900 XTX it runs up to 128K (full) tokens with 0.5 GB to spare.

How fast is gpt-oss-20b on an RX 7900 XTX?

At MXFP4 with 32K tokens of context it writes about 79–140 tokens/s for one request on an RX 7900 XTX.

Does gpt-oss-20b need --n-cpu-moe on an RX 7900 XTX?

The measured MXFP4 GGUF fits the RX 7900 XTX whole at 32K (12.5 GB), so --n-cpu-moe is not needed there.

What does a second RX 7900 XTX change for gpt-oss-20b?

Two RX 7900 XTX cards (48 GB in one tensor-parallel group) hold gpt-oss-20b at MXFP4 with 32K, and MXFP4 up to 128K (full) tokens, at about 111–216 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also gpt-oss-20b VRAM requirements, what LLMs an RX 7900 XTX can run and every pair, or detect your own GPU. Model data checked .