Can I run Qwen3.6 35B-A3B on an RTX 4080 Super 16GB?

With --n-cpu-moe 13: the UD-Q4_K_M GGUF keeps 15.8 GB on the RTX 4080 Super and 6.39 GB in system RAM, with the experts of 13 of its 40 layers moved there. Whole, Qwen3.6 35B-A3B needs about 23.5 GB at Q4_K_M with 32K tokens of context, 7.47 GB more than the RTX 4080 Super 16GB holds. The smallest setup here that holds Qwen3.6 35B-A3B at Q4_K_M with 32K is RTX 3090 (24 GB).

With offload the UD-Q4_K_M GGUF with --n-cpu-moe 13 and 32K tokens of context

On the card
15.8 GB
RTX 4080 Super
16 GB, 736 GB/s
In RAM
6.39 GB
Tokens/s
44–75 tokens/s

With the UD-Q4_K_M GGUF, --n-cpu-moe 13 and 32K tokens of context it writes about 44–75 tokens/s for one request on an RTX 4080 Super 16GB and dual-channel DDR5-5600.

Best precision for Qwen3.6 35B-A3B on an RTX 4080 Super 16GB

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Only IQ3_XXS, tight 15.9 GB 139 MB (tight) 127–232
32K Nothing fits — — —
128K Nothing fits — — —
256K (full) Nothing fits — — —

Qwen3.6 35B-A3B on the RTX 4080 Super as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 22.9 GB −6.87 GB — 39.7 GB −23.7 GB —
8K 23.0 GB −6.95 GB — 39.8 GB −23.8 GB —
16K 23.1 GB −7.13 GB — 40.0 GB −24.0 GB —
32K 23.5 GB −7.47 GB — 40.3 GB −24.3 GB —
64K 24.2 GB −8.16 GB — 41.0 GB −25.0 GB —
128K 25.5 GB −9.53 GB — 42.4 GB −26.4 GB —
256K (full) 28.3 GB −12.3 GB — 45.1 GB −29.1 GB —

--n-cpu-moe for Qwen3.6 35B-A3B on the RTX 4080 Super

With --n-cpu-moe 13 the UD-Q4_K_M GGUF keeps 15.8 GB on the RTX 4080 Super and 6.39 GB in RAM, about 44–75 tokens/s with DDR5-5600. Measured UD-Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 12 15.8 GB 5.94 GB 40–6850–8652–90
32K 13 15.8 GB 6.39 GB 35–6044–7546–78
64K 15 15.6 GB 7.30 GB 30–5137–6339–66
128K 17 15.9 GB 8.21 GB 24–4129–5030–52
256K (full) 23 15.7 GB 10.9 GB 17–2920–3421–36

Qwen3.6 35B-A3B on more than one RTX 4080 Super

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/s
2× (32 GB) Q5_K_M 102–196 256K (full) 110–214
4× (64 GB) Q8_0 127–255 256K (full) 158–335

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Qwen3.6 35B-A3B on the RTX 4080 Super with llama-server

llama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M -c 32768 --n-cpu-moe 13

Qwen3.6-35B-A3B-UD-Q4_K_M.gguf, 22.1 GB, from unsloth/Qwen3.6-35B-A3B-GGUF (checked 2026-09-29); --n-cpu-moe 13 with -c 32768 is the plan above.

Questions

Can I run Qwen3.6 35B-A3B on an RTX 4080 Super 16GB?

With --n-cpu-moe 13: the UD-Q4_K_M GGUF keeps 15.8 GB on the RTX 4080 Super and 6.39 GB in system RAM, with the experts of 13 of its 40 layers moved there. Whole, Qwen3.6 35B-A3B needs about 23.5 GB at Q4_K_M with 32K tokens of context, 7.47 GB more than the RTX 4080 Super 16GB holds. The smallest setup here that holds Qwen3.6 35B-A3B at Q4_K_M with 32K is RTX 3090 (24 GB).

How fast is Qwen3.6 35B-A3B on an RTX 4080 Super 16GB?

With the UD-Q4_K_M GGUF, --n-cpu-moe 13 and 32K tokens of context it writes about 44–75 tokens/s for one request on an RTX 4080 Super 16GB and dual-channel DDR5-5600.

Does Qwen3.6 35B-A3B need --n-cpu-moe on an RTX 4080 Super 16GB?

With --n-cpu-moe 13 the UD-Q4_K_M GGUF keeps 15.8 GB on the RTX 4080 Super and 6.39 GB in RAM, about 44–75 tokens/s with DDR5-5600.

What does a second RTX 4080 Super 16GB change for Qwen3.6 35B-A3B?

Two RTX 4080 Super 16GB cards (32 GB in one tensor-parallel group) hold Qwen3.6 35B-A3B at Q5_K_M with 32K, and Q4_K_M up to 256K (full) tokens, at about 110–214 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3.6 35B-A3B VRAM requirements, what LLMs an RTX 4080 Super 16GB can run and every pair, or detect your own GPU. Model data checked .