Can I run Qwen3-Coder-Next on an M4 Max Mac (128 GB)?
Yes: Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, which fits the M4 Max Mac (128 GB, 96 GB usable) with 45.3 GB to spare. At 32K the M4 Max 128GB holds up to Q8_0 (88.0 GB), and Q4_K_M runs up to 256K (full) tokens.
Yes Q4_K_M with 32K tokens of context
- Q4_K_M, 32K
- 50.7 GB
- M4 Max 128GB
- 96 GB usable, 546 GB/s
- To spare
- 45.3 GB
- Tokens/s
- 57–99 tokens/s
At Q4_K_M with 32K tokens of context it writes about 57–99 tokens/s for one request on an M4 Max Mac (128 GB).
Best precision for Qwen3-Coder-Next on an M4 Max Mac (128 GB)
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Q8_0 | 87.4 GB | 8.57 GB | 45–77 |
| 32K | Q8_0 | 88.0 GB | 7.95 GB | 39–66 |
| 128K | Q8_0 | 90.5 GB | 5.48 GB | 25–42 |
| 256K (full) | Q8_0 | 93.8 GB | 2.18 GB | 17–28 |
Qwen3-Coder-Next on the M4 Max 128GB as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 50.0 GB | 46.0 GB | 76–133 | 87.3 GB | 8.67 GB | 46–80 |
| 8K | 50.1 GB | 45.9 GB | 72–127 | 87.4 GB | 8.57 GB | 45–77 |
| 16K | 50.3 GB | 45.7 GB | 66–116 | 87.6 GB | 8.36 GB | 43–73 |
| 32K | 50.7 GB | 45.3 GB | 57–99 | 88.0 GB | 7.95 GB | 39–66 |
| 64K | 51.5 GB | 44.5 GB | 45–77 | 88.9 GB | 7.13 GB | 32–55 |
| 128K | 53.2 GB | 42.8 GB | 31–53 | 90.5 GB | 5.48 GB | 25–42 |
| 256K (full) | 56.5 GB | 39.5 GB | 19–33 | 93.8 GB | 2.18 GB | 17–28 |
Other options
- The next smaller setting, Q3_K_M, takes 41.2 GB at 32K, 54.8 GB under the M4 Max 128GB; it fits with 0.5 GB to spare up to 256K (full) tokens.
- Smallest setup for Q4_K_M at 32K: 2× RTX 5090 (64 GB).
Run Qwen3-Coder-Next on the M4 Max 128GB with llama-server
llama-server -hf unsloth/Qwen3-Coder-Next-GGUF:Q4_K_M -c 262144 -ngl 99 Qwen3-Coder-Next-Q4_K_M.gguf, 48.5 GB, from unsloth/
Questions
Can I run Qwen3-Coder-Next on an M4 Max Mac (128 GB)?
Yes: Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, which fits the M4 Max Mac (128 GB, 96 GB usable) with 45.3 GB to spare. At 32K the M4 Max 128GB holds up to Q8_0 (88.0 GB), and Q4_K_M runs up to 256K (full) tokens.
How fast is Qwen3-Coder-Next on an M4 Max Mac (128 GB)?
At Q4_K_M with 32K tokens of context it writes about 57–99 tokens/s for one request on an M4 Max Mac (128 GB).
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3-Coder-Next VRAM requirements, what LLMs an M4 Max Mac (128 GB) can run and every pair, or detect your own GPU. Model data checked .