Can I run gpt-oss-120b on an RX 7900 XTX?
With --n-cpu-moe 24: the MXFP4 GGUF keeps 22.7 GB on the RX 7900 XTX and 38.5 GB in system RAM, with the experts of 24 of its 36 layers moved there. Whole, gpt-oss-120b needs about 68.6 GB at MXFP4 with 32K tokens of context, 44.6 GB more than the 24 GB RX 7900 XTX holds. The smallest setup here that holds gpt-oss-120b at MXFP4 with 32K is A100 80GB.
Why 68.6 GB with vLLM, but 22.7 GB on the RX 7900 XTX with llama.cpp
The 68.6 GB counts the published checkpoint as vLLM loads it: 60.8 GB of weights, the 32K KV cache, 0.5 GB and 10% overhead. llama.cpp's measured MXFP4 GGUF is 59.0 GB, and its 587 MB of token embeddings stay in RAM. With every layer on the RX 7900 XTX it takes 60.6 GB at 32K with 1 GB of buffers, 36.6 GB more than it holds, so the experts of 24 layers (37.9 GB) move to system RAM. That leaves 1.32 GB of the 24 GB free, and the PC needs 38.5 GB of free RAM for those experts and the embeddings.
With offload the MXFP4 GGUF with --n-cpu-moe 24 and 32K tokens of context
- On the card
- 22.7 GB
- RX 7900 XTX
- 24 GB, 960 GB/s
- In RAM
- 38.5 GB
- Tokens/s
- 16–28 tokens/s
With the MXFP4 GGUF, --n-cpu-moe 24 and 32K tokens of context it writes about 16–28 tokens/s for one request on an RX 7900 XTX and dual-channel DDR5-5600.
gpt-oss-120b on the RX 7900 XTX as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; the MXFP4 GGUF columns are the measured llama.cpp file with every layer on the GPU and 1 GB of buffers; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | MXFP4 | Free | Tokens/s | MXFP4 GGUF | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 67.5 GB | −43.5 GB | — | 59.6 GB | −35.6 GB | — |
| 8K | 67.7 GB | −43.7 GB | — | 59.8 GB | −35.8 GB | — |
| 16K | 68.0 GB | −44.0 GB | — | 60.0 GB | −36.0 GB | — |
| 32K | 68.6 GB | −44.6 GB | — | 60.6 GB | −36.6 GB | — |
| 64K | 69.8 GB | −45.8 GB | — | 61.7 GB | −37.7 GB | — |
| 128K (full) | 72.3 GB | −48.3 GB | — | 64.0 GB | −40.0 GB | — |
--n-cpu-moe for gpt-oss-120b on the RX 7900 XTX
With --n-cpu-moe 24 the MXFP4 GGUF keeps 22.7 GB on the RX 7900 XTX and 38.5 GB in RAM, about 16–28 tokens/s with DDR5-5600. Measured MXFP4 file, 1 GB of buffers; tokens/s by system RAM speed.
gpt-oss-120b on more than one RX 7900 XTX
| Cards | Best at 32K | Tokens/s | MXFP4 longest | MXFP4 tokens/s |
|---|---|---|---|---|
| 2× (48 GB) | Nothing fits | — | Does not fit | — |
| 4× (96 GB) | MXFP4 | 142–292 | 128K (full) | 142–292 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- Smallest setup for MXFP4 at 32K: A100 80GB.
Run gpt-oss-120b on the RX 7900 XTX with llama-server
llama-server -hf ggml-org/gpt-oss-120b-GGUF:MXFP4 -c 32768 --n-cpu-moe 24 -np 1 gpt-oss-120b-MXFP4.gguf, 63.4 GB, from ggml-org/
Questions
Can I run gpt-oss-120b on an RX 7900 XTX?
With --n-cpu-moe 24: the MXFP4 GGUF keeps 22.7 GB on the RX 7900 XTX and 38.5 GB in system RAM, with the experts of 24 of its 36 layers moved there. Whole, gpt-oss-120b needs about 68.6 GB at MXFP4 with 32K tokens of context, 44.6 GB more than the 24 GB RX 7900 XTX holds. The smallest setup here that holds gpt-oss-120b at MXFP4 with 32K is A100 80GB.
How fast is gpt-oss-120b on an RX 7900 XTX?
With the MXFP4 GGUF, --n-cpu-moe 24 and 32K tokens of context it writes about 16–28 tokens/s for one request on an RX 7900 XTX and dual-channel DDR5-5600.
Does gpt-oss-120b need --n-cpu-moe on an RX 7900 XTX?
With --n-cpu-moe 24 the MXFP4 GGUF keeps 22.7 GB on the RX 7900 XTX and 38.5 GB in RAM, about 16–28 tokens/s with DDR5-5600.
What does a second RX 7900 XTX change for gpt-oss-120b?
Even two RX 7900 XTX cards (48 GB) do not hold gpt-oss-120b at 32K.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also gpt-oss-120b VRAM requirements, what LLMs an RX 7900 XTX can run and every pair, or detect your own GPU. Model data checked .