Can I run Qwen3.8 Flash Next on an RX 7900 XTX?
With --n-cpu-moe 37: the UD-Q4_K_XL GGUF keeps 22.8 GB on the RX 7900 XTX and 55.8 GB in system RAM, with the experts of 37 of its 48 layers moved there. Whole, Qwen3.8 Flash Next needs about 113 GB at Q4_K_M with 32K tokens of context, 88.9 GB more than the 24 GB RX 7900 XTX holds. The smallest setup here that holds Qwen3.8 Flash Next at Q4_K_M with 32K is DGX Spark (128 GB, 120 GB usable).
With offload the UD-Q4_K_XL GGUF with --n-cpu-moe 37 and 32K tokens of context
- On the card
- 22.8 GB
- RX 7900 XTX
- 24 GB, 960 GB/s
- In RAM
- 55.8 GB
- Tokens/s
- 15–26 tokens/s
With the UD-Q4_K_XL GGUF, --n-cpu-moe 37 and 32K tokens of context it writes about 15–26 tokens/s for one request on an RX 7900 XTX and dual-channel DDR5-5600.
Qwen3.8 Flash Next on the RX 7900 XTX as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 112 GB | −88.2 GB | — | 197 GB | −173 GB | — |
| 8K | 112 GB | −88.3 GB | — | 197 GB | −173 GB | — |
| 16K | 112 GB | −88.5 GB | — | 197 GB | −173 GB | — |
| 32K | 113 GB | −88.9 GB | — | 197 GB | −173 GB | — |
| 64K | 114 GB | −89.7 GB | — | 198 GB | −174 GB | — |
| 128K | 115 GB | −91.4 GB | — | 200 GB | −176 GB | — |
| 256K (full) | 119 GB | −94.7 GB | — | 203 GB | −179 GB | — |
--n-cpu-moe for Qwen3.8 Flash Next on the RX 7900 XTX
With --n-cpu-moe 37 the UD-Q4_K_XL GGUF keeps 22.8 GB on the RX 7900 XTX and 55.8 GB in RAM, about 15–26 tokens/s with DDR5-5600. Measured UD-Q4_K_XL file, 1 GB of buffers; tokens/s by system RAM speed.
Qwen3.8 Flash Next on more than one RX 7900 XTX
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s |
|---|---|---|---|---|
| 2× (48 GB) | Nothing fits | — | Does not fit | — |
| 4× (96 GB) | Q3_K_M | 148–308 | Does not fit | — |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 91.5 GB at 32K, 67.5 GB over the RX 7900 XTX; it does not fit with 0.5 GB to spare even with 1K tokens.
- Smallest setup for Q4_K_M at 32K: DGX Spark (128 GB, 120 GB usable).
Run Qwen3.8 Flash Next on the RX 7900 XTX with llama-server
llama-server -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL -c 32768 -ngl 99 --n-cpu-moe 37 UD-Q4_K_XL/
Questions
Can I run Qwen3.8 Flash Next on an RX 7900 XTX?
With --n-cpu-moe 37: the UD-Q4_K_XL GGUF keeps 22.8 GB on the RX 7900 XTX and 55.8 GB in system RAM, with the experts of 37 of its 48 layers moved there. Whole, Qwen3.8 Flash Next needs about 113 GB at Q4_K_M with 32K tokens of context, 88.9 GB more than the 24 GB RX 7900 XTX holds. The smallest setup here that holds Qwen3.8 Flash Next at Q4_K_M with 32K is DGX Spark (128 GB, 120 GB usable).
How fast is Qwen3.8 Flash Next on an RX 7900 XTX?
With the UD-Q4_K_XL GGUF, --n-cpu-moe 37 and 32K tokens of context it writes about 15–26 tokens/s for one request on an RX 7900 XTX and dual-channel DDR5-5600.
Does Qwen3.8 Flash Next need --n-cpu-moe on an RX 7900 XTX?
With --n-cpu-moe 37 the UD-Q4_K_XL GGUF keeps 22.8 GB on the RX 7900 XTX and 55.8 GB in RAM, about 15–26 tokens/s with DDR5-5600.
What does a second RX 7900 XTX change for Qwen3.8 Flash Next?
Even two RX 7900 XTX cards (48 GB) do not hold Qwen3.8 Flash Next at 32K.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3.8 Flash Next VRAM requirements, what LLMs an RX 7900 XTX can run and every pair, or detect your own GPU. Model data checked .