Can I run Qwen3.6 35B-A3B on an RX 7900 XTX?

Yes: Qwen3.6 35B-A3B needs about 23.5 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RX 7900 XTX with 542 MB to spare. Nothing more precise fits at 32K with 0.5 GB to spare; Q4_K_M runs up to 33K tokens on the RX 7900 XTX.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
23.5 GB
RX 7900 XTX
24 GB, 960 GB/s
To spare
542 MB
Tokens/s
99–176 tokens/s

At Q4_K_M with 32K tokens of context it writes about 99–176 tokens/s for one request on an RX 7900 XTX.

Best precision for Qwen3.6 35B-A3B on an RX 7900 XTX

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q4_K_M 23.0 GB 1.05 GB 119–216
32K Q4_K_M 23.5 GB 542 MB 99–176
128K Q3_K_M 21.3 GB 2.75 GB 63–109
256K (full) Q2_K 21.4 GB 2.58 GB 41–70

Qwen3.6 35B-A3B on the RX 7900 XTX as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 22.9 GB 1.13 GB 124–224 39.7 GB −15.7 GB —
8K 23.0 GB 1.05 GB 119–216 39.8 GB −15.8 GB —
16K 23.1 GB 894 MB 112–201 40.0 GB −16.0 GB —
32K 23.5 GB 542 MB 99–176 40.3 GB −16.3 GB —
64K 24.2 GB −162 MB — 41.0 GB −17.0 GB —
128K 25.5 GB −1.53 GB — 42.4 GB −18.4 GB —
256K (full) 28.3 GB −4.28 GB — 45.1 GB −21.1 GB —

--n-cpu-moe for Qwen3.6 35B-A3B on the RX 7900 XTX

The measured UD-Q4_K_M GGUF fits the RX 7900 XTX whole at 32K (21.7 GB), so --n-cpu-moe is not needed there. Measured UD-Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 0 21.3 GB 515 MB 89–15889–15889–158
32K 0 21.7 GB 515 MB 77–13677–13677–136
64K 0 22.4 GB 515 MB 65–11465–11465–114
128K 0 23.6 GB 515 MB 50–8650–8650–86
256K (full) 5 23.8 GB 2.77 GB 29–5031–5332–54

Qwen3.6 35B-A3B on more than one RX 7900 XTX

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/s
2× (48 GB) Q8_0 98–188 256K (full) 128–257
4× (96 GB) BF16 108–209 256K (full) 177–385

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Qwen3.6 35B-A3B on the RX 7900 XTX with llama-server

llama-server -hf ggml-org/Qwen3.6-35B-A3B-GGUF:Q4_K_M -c 33792

Qwen3.6-35B-A3B-Q4_K_M.gguf, 20.4 GB, from ggml-org/Qwen3.6-35B-A3B-GGUF (checked 2026-09-29). At -c 33792 on the RX 7900 XTX: 23.5 GB of 24 GB, 520 MB free.

Questions

Can I run Qwen3.6 35B-A3B on an RX 7900 XTX?

Yes: Qwen3.6 35B-A3B needs about 23.5 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RX 7900 XTX with 542 MB to spare. Nothing more precise fits at 32K with 0.5 GB to spare; Q4_K_M runs up to 33K tokens on the RX 7900 XTX.

How fast is Qwen3.6 35B-A3B on an RX 7900 XTX?

At Q4_K_M with 32K tokens of context it writes about 99–176 tokens/s for one request on an RX 7900 XTX.

Does Qwen3.6 35B-A3B need --n-cpu-moe on an RX 7900 XTX?

The measured UD-Q4_K_M GGUF fits the RX 7900 XTX whole at 32K (21.7 GB), so --n-cpu-moe is not needed there.

What does a second RX 7900 XTX change for Qwen3.6 35B-A3B?

Two RX 7900 XTX cards (48 GB in one tensor-parallel group) hold Qwen3.6 35B-A3B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 128–257 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3.6 35B-A3B VRAM requirements, what LLMs an RX 7900 XTX can run and every pair, or detect your own GPU. Model data checked .