Can I run Qwen3-Coder-Next on an RTX 5090?

With --n-cpu-moe 17: the Q4_K_M GGUF keeps 31.3 GB on the RTX 5090 and 15.7 GB in system RAM, with the experts of 17 of its 48 layers moved there. Whole, Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, 18.7 GB more than the 32 GB RTX 5090 holds. The smallest setup here that holds Qwen3-Coder-Next at Q4_K_M with 32K is 2× RTX 5090 (64 GB).

With offload the Q4_K_M GGUF with --n-cpu-moe 17 and 32K tokens of context

On the card
31.3 GB
RTX 5090
32 GB, 1,792 GB/s
In RAM
15.7 GB
Tokens/s
52–91 tokens/s

With the Q4_K_M GGUF, --n-cpu-moe 17 and 32K tokens of context it writes about 52–91 tokens/s for one request on an RTX 5090 and dual-channel DDR5-5600.

Qwen3-Coder-Next on the RTX 5090 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 50.0 GB −18.0 GB — 87.3 GB −55.3 GB —
8K 50.1 GB −18.1 GB — 87.4 GB −55.4 GB —
16K 50.3 GB −18.3 GB — 87.6 GB −55.6 GB —
32K 50.7 GB −18.7 GB — 88.0 GB −56.0 GB —
64K 51.5 GB −19.5 GB — 88.9 GB −56.9 GB —
128K 53.2 GB −21.2 GB — 90.5 GB −58.5 GB —
256K (full) 56.5 GB −24.5 GB — 93.8 GB −61.8 GB —

--n-cpu-moe for Qwen3-Coder-Next on the RTX 5090

With --n-cpu-moe 17 the Q4_K_M GGUF keeps 31.3 GB on the RTX 5090 and 15.7 GB in RAM, about 52–91 tokens/s with DDR5-5600. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 16 31.6 GB 14.8 GB 39–6658–10063–110
32K 17 31.3 GB 15.7 GB 36–6152–9157–99
64K 18 31.1 GB 16.6 GB 32–5547–8151–88
128K 19 31.7 GB 17.5 GB 29–4940–6943–74
256K (full) 23 31.2 GB 21.0 GB 22–3730–5132–54

Qwen3-Coder-Next on more than one RTX 5090

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/sRent per hour
2× (64 GB) Q5_K_M 161–341 256K (full) 168–362 $1.38
4× (128 GB) Q8_0 187–414 256K (full) 212–492 $2.76

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Qwen3-Coder-Next on the RTX 5090 with llama-server

llama-server -hf unsloth/Qwen3-Coder-Next-GGUF:Q4_K_M -c 32768 --n-cpu-moe 17

Qwen3-Coder-Next-Q4_K_M.gguf, 48.5 GB, from unsloth/Qwen3-Coder-Next-GGUF (checked 2026-09-29); --n-cpu-moe 17 with -c 32768 is the plan above.

Questions

Can I run Qwen3-Coder-Next on an RTX 5090?

With --n-cpu-moe 17: the Q4_K_M GGUF keeps 31.3 GB on the RTX 5090 and 15.7 GB in system RAM, with the experts of 17 of its 48 layers moved there. Whole, Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, 18.7 GB more than the 32 GB RTX 5090 holds. The smallest setup here that holds Qwen3-Coder-Next at Q4_K_M with 32K is 2× RTX 5090 (64 GB).

How fast is Qwen3-Coder-Next on an RTX 5090?

With the Q4_K_M GGUF, --n-cpu-moe 17 and 32K tokens of context it writes about 52–91 tokens/s for one request on an RTX 5090 and dual-channel DDR5-5600.

Does Qwen3-Coder-Next need --n-cpu-moe on an RTX 5090?

With --n-cpu-moe 17 the Q4_K_M GGUF keeps 31.3 GB on the RTX 5090 and 15.7 GB in RAM, about 52–91 tokens/s with DDR5-5600.

What does a second RTX 5090 change for Qwen3-Coder-Next?

Two RTX 5090 cards (64 GB in one tensor-parallel group) hold Qwen3-Coder-Next at Q5_K_M with 32K, and Q4_K_M up to 256K (full) tokens, at about 168–362 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3-Coder-Next VRAM requirements, what LLMs an RTX 5090 can run and every pair, or detect your own GPU. Model data checked .