Can I run Nemotron 3 Nano 30B-A3B on an RX 7900 XTX?

Yes: Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RX 7900 XTX with 3.72 GB to spare. Nothing more precise fits at 32K with 0.5 GB to spare; Q4_K_M runs up to 256K (full) tokens on the RX 7900 XTX.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
20.3 GB
RX 7900 XTX
24 GB, 960 GB/s
To spare
3.72 GB
Tokens/s
105–188 tokens/s

At Q4_K_M with 32K tokens of context it writes about 105–188 tokens/s for one request on an RX 7900 XTX.

Best precision for Nemotron 3 Nano 30B-A3B on an RX 7900 XTX

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q5_K_M 23.5 GB 533 MB 97–173
32K Q4_K_M 20.3 GB 3.72 GB 105–188
128K Q4_K_M 20.9 GB 3.10 GB 86–152
256K (full) Q4_K_M 21.7 GB 2.28 GB 69–121

Nemotron 3 Nano 30B-A3B on the RX 7900 XTX as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 20.1 GB 3.90 GB 112–201 34.9 GB −10.9 GB —
8K 20.1 GB 3.88 GB 111–199 34.9 GB −10.9 GB —
16K 20.2 GB 3.83 GB 109–195 35.0 GB −11.0 GB —
32K 20.3 GB 3.72 GB 105–188 35.1 GB −11.1 GB —
64K 20.5 GB 3.52 GB 98–174 35.3 GB −11.3 GB —
128K 20.9 GB 3.10 GB 86–152 35.7 GB −11.7 GB —
256K (full) 21.7 GB 2.28 GB 69–121 36.5 GB −12.5 GB —

--n-cpu-moe for Nemotron 3 Nano 30B-A3B on the RX 7900 XTX

The measured Q4_K_M GGUF fits the RX 7900 XTX whole at 32K (23.8 GB), so --n-cpu-moe is not needed there. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 0 23.7 GB 231 MB 95–16895–16895–168
32K 0 23.8 GB 231 MB 90–16090–16090–160
64K 2 23.0 GB 1.27 GB 67–11674–12975–132
128K 2 23.4 GB 1.27 GB 61–10667–11768–119
256K (full) 4 23.1 GB 2.31 GB 45–7651–8853–91

Nemotron 3 Nano 30B-A3B on more than one RX 7900 XTX

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/s
2× (48 GB) Q8_0 97–186 256K (full) 133–269
4× (96 GB) BF16 103–198 256K (full) 181–399

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Nemotron 3 Nano 30B-A3B on the RX 7900 XTX with llama-server

llama-server -hf ggml-org/NVIDIA-Nemotron-3-Nano-30B-A3B-GGUF:Q4_K_M -c 4096

NVIDIA-Nemotron-3-Nano-30B-A3B-Q4_K_M.gguf, 22.4 GB, from ggml-org/NVIDIA-Nemotron-3-Nano-30B-A3B-GGUF (checked 2026-09-29). At -c 4096 on the RX 7900 XTX: 23.5 GB of 24 GB, 516 MB free. The file is 3.09 GB over the estimate above, so -c counts the file.

Questions

Can I run Nemotron 3 Nano 30B-A3B on an RX 7900 XTX?

Yes: Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RX 7900 XTX with 3.72 GB to spare. Nothing more precise fits at 32K with 0.5 GB to spare; Q4_K_M runs up to 256K (full) tokens on the RX 7900 XTX.

How fast is Nemotron 3 Nano 30B-A3B on an RX 7900 XTX?

At Q4_K_M with 32K tokens of context it writes about 105–188 tokens/s for one request on an RX 7900 XTX.

Does Nemotron 3 Nano 30B-A3B need --n-cpu-moe on an RX 7900 XTX?

The measured Q4_K_M GGUF fits the RX 7900 XTX whole at 32K (23.8 GB), so --n-cpu-moe is not needed there.

What does a second RX 7900 XTX change for Nemotron 3 Nano 30B-A3B?

Two RX 7900 XTX cards (48 GB in one tensor-parallel group) hold Nemotron 3 Nano 30B-A3B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 133–269 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Nemotron 3 Nano 30B-A3B VRAM requirements, what LLMs an RX 7900 XTX can run and every pair, or detect your own GPU. Model data checked .