Can I run Qwen3.6 35B-A3B on an RTX 5060 Ti 16GB?
With --n-cpu-moe 13: the UD-Q4_K_M GGUF keeps 15.8 GB on the RTX 5060 Ti and 6.39 GB in system RAM, with the experts of 13 of its 40 layers moved there. Whole, Qwen3.6 35B-A3B needs about 23.5 GB at Q4_K_M with 32K tokens of context, 7.47 GB more than the RTX 5060 Ti 16GB holds. The smallest setup here that holds Qwen3.6 35B-A3B at Q4_K_M with 32K is RTX 3090 (24 GB).
With offload the UD-Q4_K_M GGUF with --n-cpu-moe 13 and 32K tokens of context
- On the card
- 15.8 GB
- RTX 5060 Ti
- 16 GB, 448 GB/s
- In RAM
- 6.39 GB
- Tokens/s
- 31–53 tokens/s
With the UD-Q4_K_M GGUF, --n-cpu-moe 13 and 32K tokens of context it writes about 31–53 tokens/s for one request on an RTX 5060 Ti 16GB and dual-channel DDR5-5600.
Best precision for Qwen3.6 35B-A3B on an RTX 5060 Ti 16GB
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Only IQ3_XXS, tight | 15.9 GB | 139 MB (tight) | 84–148 |
| 32K | Nothing fits | — | — | — |
| 128K | Nothing fits | — | — | — |
| 256K (full) | Nothing fits | — | — | — |
Qwen3.6 35B-A3B on the RTX 5060 Ti as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 22.9 GB | −6.87 GB | — | 39.7 GB | −23.7 GB | — |
| 8K | 23.0 GB | −6.95 GB | — | 39.8 GB | −23.8 GB | — |
| 16K | 23.1 GB | −7.13 GB | — | 40.0 GB | −24.0 GB | — |
| 32K | 23.5 GB | −7.47 GB | — | 40.3 GB | −24.3 GB | — |
| 64K | 24.2 GB | −8.16 GB | — | 41.0 GB | −25.0 GB | — |
| 128K | 25.5 GB | −9.53 GB | — | 42.4 GB | −26.4 GB | — |
| 256K (full) | 28.3 GB | −12.3 GB | — | 45.1 GB | −29.1 GB | — |
--n-cpu-moe for Qwen3.6 35B-A3B on the RTX 5060 Ti
With --n-cpu-moe 13 the UD-Q4_K_M GGUF keeps 15.8 GB on the RTX 5060 Ti and 6.39 GB in RAM, about 31–53 tokens/s with DDR5-5600. Measured UD-Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.
Qwen3.6 35B-A3B on more than one RTX 5060 Ti
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s |
|---|---|---|---|---|
| 2× (32 GB) | Q5_K_M | 72–133 | 256K (full) | 78–146 |
| 4× (64 GB) | Q8_0 | 94–178 | 256K (full) | 123–245 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 19.2 GB at 32K, 3.19 GB over the RTX 5060 Ti; it does not fit with 0.5 GB to spare even with 1K tokens.
- Smallest setup for Q4_K_M at 32K: RTX 3090 (24 GB).
Run Qwen3.6 35B-A3B on the RTX 5060 Ti with llama-server
llama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M -c 32768 --n-cpu-moe 13 Qwen3.6-35B-A3B-UD-Q4_K_M.gguf, 22.1 GB, from unsloth/
Questions
Can I run Qwen3.6 35B-A3B on an RTX 5060 Ti 16GB?
With --n-cpu-moe 13: the UD-Q4_K_M GGUF keeps 15.8 GB on the RTX 5060 Ti and 6.39 GB in system RAM, with the experts of 13 of its 40 layers moved there. Whole, Qwen3.6 35B-A3B needs about 23.5 GB at Q4_K_M with 32K tokens of context, 7.47 GB more than the RTX 5060 Ti 16GB holds. The smallest setup here that holds Qwen3.6 35B-A3B at Q4_K_M with 32K is RTX 3090 (24 GB).
How fast is Qwen3.6 35B-A3B on an RTX 5060 Ti 16GB?
With the UD-Q4_K_M GGUF, --n-cpu-moe 13 and 32K tokens of context it writes about 31–53 tokens/s for one request on an RTX 5060 Ti 16GB and dual-channel DDR5-5600.
Does Qwen3.6 35B-A3B need --n-cpu-moe on an RTX 5060 Ti 16GB?
With --n-cpu-moe 13 the UD-Q4_K_M GGUF keeps 15.8 GB on the RTX 5060 Ti and 6.39 GB in RAM, about 31–53 tokens/s with DDR5-5600.
What does a second RTX 5060 Ti 16GB change for Qwen3.6 35B-A3B?
Two RTX 5060 Ti 16GB cards (32 GB in one tensor-parallel group) hold Qwen3.6 35B-A3B at Q5_K_M with 32K, and Q4_K_M up to 256K (full) tokens, at about 78–146 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3.6 35B-A3B VRAM requirements, what LLMs an RTX 5060 Ti 16GB can run and every pair, or detect your own GPU. Model data checked .