Can I run Qwen3.6 35B-A3B on an RTX 5090?

Yes: Qwen3.6 35B-A3B needs about 23.5 GB at Q4_K_M with 32K tokens of context, which fits the 32 GB RTX 5090 with 8.53 GB to spare. At 32K the RTX 5090 holds up to Q6_K (31.4 GB), and Q4_K_M runs up to 256K (full) tokens.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
23.5 GB
RTX 5090
32 GB, 1,792 GB/s
To spare
8.53 GB
Tokens/s
163–305 tokens/s

At Q4_K_M with 32K tokens of context it writes about 163–305 tokens/s for one request on an RTX 5090.

Best precision for Qwen3.6 35B-A3B on an RTX 5090

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q6_K 30.9 GB 1.13 GB 157–291
32K Q6_K 31.4 GB 626 MB 137–250
128K Q5_K_M 29.4 GB 2.65 GB 96–170
256K (full) Q4_K_M 28.3 GB 3.72 GB 67–117

Qwen3.6 35B-A3B on the RTX 5090 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 22.9 GB 9.13 GB 199–382 39.7 GB −7.72 GB —
8K 23.0 GB 9.05 GB 193–369 39.8 GB −7.80 GB —
16K 23.1 GB 8.87 GB 182–345 40.0 GB −7.98 GB —
32K 23.5 GB 8.53 GB 163–305 40.3 GB −8.32 GB —
64K 24.2 GB 7.84 GB 136–249 41.0 GB −9.01 GB —
128K 25.5 GB 6.47 GB 101–181 42.4 GB −10.4 GB —
256K (full) 28.3 GB 3.72 GB 67–117 45.1 GB −13.1 GB —

--n-cpu-moe for Qwen3.6 35B-A3B on the RTX 5090

The measured UD-Q4_K_M GGUF fits the RTX 5090 whole at 32K (21.7 GB), so --n-cpu-moe is not needed there. Measured UD-Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 0 21.3 GB 515 MB 149–276149–276149–276
32K 0 21.7 GB 515 MB 131–239131–239131–239
64K 0 22.4 GB 515 MB 113–203113–203113–203
128K 0 23.6 GB 515 MB 88–15688–15688–156
256K (full) 0 26.1 GB 515 MB 61–10661–10661–106

Qwen3.6 35B-A3B on more than one RTX 5090

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/sRent per hour
2× (64 GB) Q8_0 141–290 256K (full) 172–372 $1.38
4× (128 GB) BF16 151–316 256K (full) 215–502 $2.76

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Qwen3.6 35B-A3B on the RTX 5090 with llama-server

llama-server -hf ggml-org/Qwen3.6-35B-A3B-GGUF:Q4_K_M -c 262144

Qwen3.6-35B-A3B-Q4_K_M.gguf, 20.4 GB, from ggml-org/Qwen3.6-35B-A3B-GGUF (checked 2026-09-29). At -c 262144 on the RTX 5090: 28.3 GB of 32 GB, 3.72 GB free.

Questions

Can I run Qwen3.6 35B-A3B on an RTX 5090?

Yes: Qwen3.6 35B-A3B needs about 23.5 GB at Q4_K_M with 32K tokens of context, which fits the 32 GB RTX 5090 with 8.53 GB to spare. At 32K the RTX 5090 holds up to Q6_K (31.4 GB), and Q4_K_M runs up to 256K (full) tokens.

How fast is Qwen3.6 35B-A3B on an RTX 5090?

At Q4_K_M with 32K tokens of context it writes about 163–305 tokens/s for one request on an RTX 5090.

Does Qwen3.6 35B-A3B need --n-cpu-moe on an RTX 5090?

The measured UD-Q4_K_M GGUF fits the RTX 5090 whole at 32K (21.7 GB), so --n-cpu-moe is not needed there.

What does a second RTX 5090 change for Qwen3.6 35B-A3B?

Two RTX 5090 cards (64 GB in one tensor-parallel group) hold Qwen3.6 35B-A3B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 172–372 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3.6 35B-A3B VRAM requirements, what LLMs an RTX 5090 can run and every pair, or detect your own GPU. Model data checked .