Can I run Qwen3.6 35B-A3B on an RTX 5090?
Yes: Qwen3.6 35B-A3B needs about 23.5 GB at Q4_K_M with 32K tokens of context, which fits the 32 GB RTX 5090 with 8.53 GB to spare. At 32K the RTX 5090 holds up to Q6_K (31.4 GB), and Q4_K_M runs up to 256K (full) tokens.
Yes Q4_K_M with 32K tokens of context
- Q4_K_M, 32K
- 23.5 GB
- RTX 5090
- 32 GB, 1,792 GB/s
- To spare
- 8.53 GB
- Tokens/s
- 163–305 tokens/s
At Q4_K_M with 32K tokens of context it writes about 163–305 tokens/s for one request on an RTX 5090.
Best precision for Qwen3.6 35B-A3B on an RTX 5090
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Q6_K | 30.9 GB | 1.13 GB | 157–291 |
| 32K | Q6_K | 31.4 GB | 626 MB | 137–250 |
| 128K | Q5_K_M | 29.4 GB | 2.65 GB | 96–170 |
| 256K (full) | Q4_K_M | 28.3 GB | 3.72 GB | 67–117 |
Qwen3.6 35B-A3B on the RTX 5090 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 22.9 GB | 9.13 GB | 199–382 | 39.7 GB | −7.72 GB | — |
| 8K | 23.0 GB | 9.05 GB | 193–369 | 39.8 GB | −7.80 GB | — |
| 16K | 23.1 GB | 8.87 GB | 182–345 | 40.0 GB | −7.98 GB | — |
| 32K | 23.5 GB | 8.53 GB | 163–305 | 40.3 GB | −8.32 GB | — |
| 64K | 24.2 GB | 7.84 GB | 136–249 | 41.0 GB | −9.01 GB | — |
| 128K | 25.5 GB | 6.47 GB | 101–181 | 42.4 GB | −10.4 GB | — |
| 256K (full) | 28.3 GB | 3.72 GB | 67–117 | 45.1 GB | −13.1 GB | — |
--n-cpu-moe for Qwen3.6 35B-A3B on the RTX 5090
The measured UD-Q4_K_M GGUF fits the RTX 5090 whole at 32K (21.7 GB), so --n-cpu-moe is not needed there. Measured UD-Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.
Qwen3.6 35B-A3B on more than one RTX 5090
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s | Rent per hour |
|---|---|---|---|---|---|
| 2× (64 GB) | Q8_0 | 141–290 | 256K (full) | 172–372 | $1.38 |
| 4× (128 GB) | BF16 | 151–316 | 256K (full) | 215–502 | $2.76 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 19.2 GB at 32K, 12.8 GB under the RTX 5090; it fits with 0.5 GB to spare up to 256K (full) tokens.
- Renting an RTX 5090 costs about $0.69 an hour (median on getdeploying.com, 2026-09-29): $0.63–1.2 per million tokens at 163–305 tokens/s.
- Smallest setup for Q4_K_M at 32K: RTX 3090 (24 GB).
Run Qwen3.6 35B-A3B on the RTX 5090 with llama-server
llama-server -hf ggml-org/Qwen3.6-35B-A3B-GGUF:Q4_K_M -c 262144 Qwen3.6-35B-A3B-Q4_K_M.gguf, 20.4 GB, from ggml-org/
Questions
Can I run Qwen3.6 35B-A3B on an RTX 5090?
Yes: Qwen3.6 35B-A3B needs about 23.5 GB at Q4_K_M with 32K tokens of context, which fits the 32 GB RTX 5090 with 8.53 GB to spare. At 32K the RTX 5090 holds up to Q6_K (31.4 GB), and Q4_K_M runs up to 256K (full) tokens.
How fast is Qwen3.6 35B-A3B on an RTX 5090?
At Q4_K_M with 32K tokens of context it writes about 163–305 tokens/s for one request on an RTX 5090.
Does Qwen3.6 35B-A3B need --n-cpu-moe on an RTX 5090?
The measured UD-Q4_K_M GGUF fits the RTX 5090 whole at 32K (21.7 GB), so --n-cpu-moe is not needed there.
What does a second RTX 5090 change for Qwen3.6 35B-A3B?
Two RTX 5090 cards (64 GB in one tensor-parallel group) hold Qwen3.6 35B-A3B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 172–372 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3.6 35B-A3B VRAM requirements, what LLMs an RTX 5090 can run and every pair, or detect your own GPU. Model data checked .