Can I run Qwen3-Coder-Next on an RX 7900 XTX?

With --n-cpu-moe 26: the Q4_K_M GGUF keeps 23.3 GB on the RX 7900 XTX and 23.6 GB in system RAM, with the experts of 26 of its 48 layers moved there. Whole, Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, 26.7 GB more than the 24 GB RX 7900 XTX holds. The smallest setup here that holds Qwen3-Coder-Next at Q4_K_M with 32K is 2× RTX 5090 (64 GB).

With offload the Q4_K_M GGUF with --n-cpu-moe 26 and 32K tokens of context

On the card
23.3 GB
RX 7900 XTX
24 GB, 960 GB/s
In RAM
23.6 GB
Tokens/s
34–58 tokens/s

With the Q4_K_M GGUF, --n-cpu-moe 26 and 32K tokens of context it writes about 34–58 tokens/s for one request on an RX 7900 XTX and dual-channel DDR5-5600.

Qwen3-Coder-Next on the RX 7900 XTX as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 50.0 GB −26.0 GB — 87.3 GB −63.3 GB —
8K 50.1 GB −26.1 GB — 87.4 GB −63.4 GB —
16K 50.3 GB −26.3 GB — 87.6 GB −63.6 GB —
32K 50.7 GB −26.7 GB — 88.0 GB −64.0 GB —
64K 51.5 GB −27.5 GB — 88.9 GB −64.9 GB —
128K 53.2 GB −29.2 GB — 90.5 GB −66.5 GB —
256K (full) 56.5 GB −32.5 GB — 93.8 GB −69.8 GB —

--n-cpu-moe for Qwen3-Coder-Next on the RX 7900 XTX

With --n-cpu-moe 26 the Q4_K_M GGUF keeps 23.3 GB on the RX 7900 XTX and 23.6 GB in RAM, about 34–58 tokens/s with DDR5-5600. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 25 23.6 GB 22.8 GB 25–4237–6441–70
32K 26 23.3 GB 23.6 GB 23–3934–5837–63
64K 27 23.1 GB 24.6 GB 21–3630–5233–56
128K 28 23.7 GB 25.5 GB 19–3126–4327–46
256K (full) 32 23.2 GB 29.0 GB 14–2419–3220–34

Qwen3-Coder-Next on more than one RX 7900 XTX

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/s
2× (48 GB) Q3_K_M 134–273 Does not fit —
4× (96 GB) Q8_0 144–296 256K (full) 173–375

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Qwen3-Coder-Next on the RX 7900 XTX with llama-server

llama-server -hf unsloth/Qwen3-Coder-Next-GGUF:Q4_K_M -c 32768 -ngl 99 --n-cpu-moe 26

Qwen3-Coder-Next-Q4_K_M.gguf, 48.5 GB, from unsloth/Qwen3-Coder-Next-GGUF (checked 2026-09-29); --n-cpu-moe 26 with -c 32768 is the plan above.

Questions

Can I run Qwen3-Coder-Next on an RX 7900 XTX?

With --n-cpu-moe 26: the Q4_K_M GGUF keeps 23.3 GB on the RX 7900 XTX and 23.6 GB in system RAM, with the experts of 26 of its 48 layers moved there. Whole, Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, 26.7 GB more than the 24 GB RX 7900 XTX holds. The smallest setup here that holds Qwen3-Coder-Next at Q4_K_M with 32K is 2× RTX 5090 (64 GB).

How fast is Qwen3-Coder-Next on an RX 7900 XTX?

With the Q4_K_M GGUF, --n-cpu-moe 26 and 32K tokens of context it writes about 34–58 tokens/s for one request on an RX 7900 XTX and dual-channel DDR5-5600.

Does Qwen3-Coder-Next need --n-cpu-moe on an RX 7900 XTX?

With --n-cpu-moe 26 the Q4_K_M GGUF keeps 23.3 GB on the RX 7900 XTX and 23.6 GB in RAM, about 34–58 tokens/s with DDR5-5600.

What does a second RX 7900 XTX change for Qwen3-Coder-Next?

Two RX 7900 XTX cards (48 GB in one tensor-parallel group) hold Qwen3-Coder-Next at Q3_K_M with 32K, and Q4_K_M up to No tokens.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3-Coder-Next VRAM requirements, what LLMs an RX 7900 XTX can run and every pair, or detect your own GPU. Model data checked .