Can I run Qwen3-Coder-Next on an RTX 5090?
With --n-cpu-moe 17: the Q4_K_M GGUF keeps 31.3 GB on the RTX 5090 and 15.7 GB in system RAM, with the experts of 17 of its 48 layers moved there. Whole, Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, 18.7 GB more than the 32 GB RTX 5090 holds. The smallest setup here that holds Qwen3-Coder-Next at Q4_K_M with 32K is 2× RTX 5090 (64 GB).
With offload the Q4_K_M GGUF with --n-cpu-moe 17 and 32K tokens of context
- On the card
- 31.3 GB
- RTX 5090
- 32 GB, 1,792 GB/s
- In RAM
- 15.7 GB
- Tokens/s
- 52–91 tokens/s
With the Q4_K_M GGUF, --n-cpu-moe 17 and 32K tokens of context it writes about 52–91 tokens/s for one request on an RTX 5090 and dual-channel DDR5-5600.
Qwen3-Coder-Next on the RTX 5090 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 50.0 GB | −18.0 GB | — | 87.3 GB | −55.3 GB | — |
| 8K | 50.1 GB | −18.1 GB | — | 87.4 GB | −55.4 GB | — |
| 16K | 50.3 GB | −18.3 GB | — | 87.6 GB | −55.6 GB | — |
| 32K | 50.7 GB | −18.7 GB | — | 88.0 GB | −56.0 GB | — |
| 64K | 51.5 GB | −19.5 GB | — | 88.9 GB | −56.9 GB | — |
| 128K | 53.2 GB | −21.2 GB | — | 90.5 GB | −58.5 GB | — |
| 256K (full) | 56.5 GB | −24.5 GB | — | 93.8 GB | −61.8 GB | — |
--n-cpu-moe for Qwen3-Coder-Next on the RTX 5090
With --n-cpu-moe 17 the Q4_K_M GGUF keeps 31.3 GB on the RTX 5090 and 15.7 GB in RAM, about 52–91 tokens/s with DDR5-5600. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.
Qwen3-Coder-Next on more than one RTX 5090
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s | Rent per hour |
|---|---|---|---|---|---|
| 2× (64 GB) | Q5_K_M | 161–341 | 256K (full) | 168–362 | $1.38 |
| 4× (128 GB) | Q8_0 | 187–414 | 256K (full) | 212–492 | $2.76 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 41.2 GB at 32K, 9.22 GB over the RTX 5090; it does not fit with 0.5 GB to spare even with 1K tokens.
- Smallest setup for Q4_K_M at 32K: 2× RTX 5090 (64 GB).
Run Qwen3-Coder-Next on the RTX 5090 with llama-server
llama-server -hf unsloth/Qwen3-Coder-Next-GGUF:Q4_K_M -c 32768 --n-cpu-moe 17 Qwen3-Coder-Next-Q4_K_M.gguf, 48.5 GB, from unsloth/
Questions
Can I run Qwen3-Coder-Next on an RTX 5090?
With --n-cpu-moe 17: the Q4_K_M GGUF keeps 31.3 GB on the RTX 5090 and 15.7 GB in system RAM, with the experts of 17 of its 48 layers moved there. Whole, Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, 18.7 GB more than the 32 GB RTX 5090 holds. The smallest setup here that holds Qwen3-Coder-Next at Q4_K_M with 32K is 2× RTX 5090 (64 GB).
How fast is Qwen3-Coder-Next on an RTX 5090?
With the Q4_K_M GGUF, --n-cpu-moe 17 and 32K tokens of context it writes about 52–91 tokens/s for one request on an RTX 5090 and dual-channel DDR5-5600.
Does Qwen3-Coder-Next need --n-cpu-moe on an RTX 5090?
With --n-cpu-moe 17 the Q4_K_M GGUF keeps 31.3 GB on the RTX 5090 and 15.7 GB in RAM, about 52–91 tokens/s with DDR5-5600.
What does a second RTX 5090 change for Qwen3-Coder-Next?
Two RTX 5090 cards (64 GB in one tensor-parallel group) hold Qwen3-Coder-Next at Q5_K_M with 32K, and Q4_K_M up to 256K (full) tokens, at about 168–362 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3-Coder-Next VRAM requirements, what LLMs an RTX 5090 can run and every pair, or detect your own GPU. Model data checked .