Can I run Qwen3-Coder-Next on an M4 Max Mac (128 GB)?

Yes: Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, which fits the M4 Max Mac (128 GB, 96 GB usable) with 45.3 GB to spare. At 32K the M4 Max 128GB holds up to Q8_0 (88.0 GB), and Q4_K_M runs up to 256K (full) tokens.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
50.7 GB
M4 Max 128GB
96 GB usable, 546 GB/s
To spare
45.3 GB
Tokens/s
57–99 tokens/s

At Q4_K_M with 32K tokens of context it writes about 57–99 tokens/s for one request on an M4 Max Mac (128 GB).

Best precision for Qwen3-Coder-Next on an M4 Max Mac (128 GB)

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q8_0 87.4 GB 8.57 GB 45–77
32K Q8_0 88.0 GB 7.95 GB 39–66
128K Q8_0 90.5 GB 5.48 GB 25–42
256K (full) Q8_0 93.8 GB 2.18 GB 17–28

Qwen3-Coder-Next on the M4 Max 128GB as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 50.0 GB 46.0 GB 76–133 87.3 GB 8.67 GB 46–80
8K 50.1 GB 45.9 GB 72–127 87.4 GB 8.57 GB 45–77
16K 50.3 GB 45.7 GB 66–116 87.6 GB 8.36 GB 43–73
32K 50.7 GB 45.3 GB 57–99 88.0 GB 7.95 GB 39–66
64K 51.5 GB 44.5 GB 45–77 88.9 GB 7.13 GB 32–55
128K 53.2 GB 42.8 GB 31–53 90.5 GB 5.48 GB 25–42
256K (full) 56.5 GB 39.5 GB 19–33 93.8 GB 2.18 GB 17–28

Other options

Run Qwen3-Coder-Next on the M4 Max 128GB with llama-server

llama-server -hf unsloth/Qwen3-Coder-Next-GGUF:Q4_K_M -c 262144 -ngl 99

Qwen3-Coder-Next-Q4_K_M.gguf, 48.5 GB, from unsloth/Qwen3-Coder-Next-GGUF (checked 2026-09-29). At -c 262144 on the M4 Max 128GB: 56.8 GB of 96 GB, 39.2 GB free. The file is 310 MB over the estimate above, so -c counts the file.

Questions

Can I run Qwen3-Coder-Next on an M4 Max Mac (128 GB)?

Yes: Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, which fits the M4 Max Mac (128 GB, 96 GB usable) with 45.3 GB to spare. At 32K the M4 Max 128GB holds up to Q8_0 (88.0 GB), and Q4_K_M runs up to 256K (full) tokens.

How fast is Qwen3-Coder-Next on an M4 Max Mac (128 GB)?

At Q4_K_M with 32K tokens of context it writes about 57–99 tokens/s for one request on an M4 Max Mac (128 GB).

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3-Coder-Next VRAM requirements, what LLMs an M4 Max Mac (128 GB) can run and every pair, or detect your own GPU. Model data checked .