Can I run Qwen3.8 Flash Next on an RTX 5090?

With --n-cpu-moe 31: the UD-Q4_K_XL GGUF keeps 31.6 GB on the RTX 5090 and 47.0 GB in system RAM, with the experts of 31 of its 48 layers moved there. Whole, Qwen3.8 Flash Next needs about 113 GB at Q4_K_M with 32K tokens of context, 80.9 GB more than the 32 GB RTX 5090 holds. The smallest setup here that holds Qwen3.8 Flash Next at Q4_K_M with 32K is DGX Spark (128 GB, 120 GB usable).

With offload the UD-Q4_K_XL GGUF with --n-cpu-moe 31 and 32K tokens of context

On the card
31.6 GB
RTX 5090
32 GB, 1,792 GB/s
In RAM
47.0 GB
Tokens/s
20–34 tokens/s

With the UD-Q4_K_XL GGUF, --n-cpu-moe 31 and 32K tokens of context it writes about 20–34 tokens/s for one request on an RTX 5090 and dual-channel DDR5-5600.

Qwen3.8 Flash Next on the RTX 5090 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 112 GB −80.2 GB — 197 GB −165 GB —
8K 112 GB −80.3 GB — 197 GB −165 GB —
16K 112 GB −80.5 GB — 197 GB −165 GB —
32K 113 GB −80.9 GB — 197 GB −165 GB —
64K 114 GB −81.7 GB — 198 GB −166 GB —
128K 115 GB −83.4 GB — 200 GB −168 GB —
256K (full) 119 GB −86.7 GB — 203 GB −171 GB —

--n-cpu-moe for Qwen3.8 Flash Next on the RTX 5090

With --n-cpu-moe 31 the UD-Q4_K_XL GGUF keeps 31.6 GB on the RTX 5090 and 47.0 GB in RAM, about 20–34 tokens/s with DDR5-5600. Measured UD-Q4_K_XL file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 31 31.1 GB 47.0 GB 13–2221–3523–39
32K 31 31.6 GB 47.0 GB 13–2220–3422–38
64K 32 30.9 GB 48.4 GB 13–2119–3321–36
128K 33 31.0 GB 49.9 GB 12–2018–3020–33
256K (full) 35 31.0 GB 52.8 GB 11–1816–2617–29

Qwen3.8 Flash Next on more than one RTX 5090

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/sRent per hour
2× (64 GB) Nothing fits — Does not fit — $1.38
4× (128 GB) Q4_K_M 180–394 256K (full) 180–394 $2.76

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Qwen3.8 Flash Next on the RTX 5090 with llama-server

llama-server -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL -c 32768 -ngl 99 --n-cpu-moe 31

UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf, 111.3 GB, from unsloth/Qwen3.8-Flash-Next-GGUF (checked 2026-09-29); --n-cpu-moe 31 with -c 32768 is the plan above.

Questions

Can I run Qwen3.8 Flash Next on an RTX 5090?

With --n-cpu-moe 31: the UD-Q4_K_XL GGUF keeps 31.6 GB on the RTX 5090 and 47.0 GB in system RAM, with the experts of 31 of its 48 layers moved there. Whole, Qwen3.8 Flash Next needs about 113 GB at Q4_K_M with 32K tokens of context, 80.9 GB more than the 32 GB RTX 5090 holds. The smallest setup here that holds Qwen3.8 Flash Next at Q4_K_M with 32K is DGX Spark (128 GB, 120 GB usable).

How fast is Qwen3.8 Flash Next on an RTX 5090?

With the UD-Q4_K_XL GGUF, --n-cpu-moe 31 and 32K tokens of context it writes about 20–34 tokens/s for one request on an RTX 5090 and dual-channel DDR5-5600.

Does Qwen3.8 Flash Next need --n-cpu-moe on an RTX 5090?

With --n-cpu-moe 31 the UD-Q4_K_XL GGUF keeps 31.6 GB on the RTX 5090 and 47.0 GB in RAM, about 20–34 tokens/s with DDR5-5600.

What does a second RTX 5090 change for Qwen3.8 Flash Next?

Even two RTX 5090 cards (64 GB) do not hold Qwen3.8 Flash Next at 32K.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3.8 Flash Next VRAM requirements, what LLMs an RTX 5090 can run and every pair, or detect your own GPU. Model data checked .