Can I run Qwen3.8 Flash Next on an M4 Max Mac (128 GB)?
Only at Q3_K_M: Qwen3.8 Flash Next needs about 113 GB at Q4_K_M with 32K tokens of context, 16.9 GB more than the M4 Max Mac (128 GB, 96 GB usable) holds, but 91.5 GB at Q3_K_M, which fits with 4.55 GB to spare. The smallest setup here that holds Qwen3.8 Flash Next at Q4_K_M with 32K is DGX Spark (128 GB, 120 GB usable).
Partly Q3_K_M with 32K tokens of context
- Q4_K_M, 32K
- 113 GB
- M4 Max 128GB
- 96 GB usable, 546 GB/s
- Short by
- 16.9 GB
- Tokens/s
- 41–70 tokens/s
At Q3_K_M with 32K tokens of context it writes about 41–70 tokens/s for one request on an M4 Max Mac (128 GB).
Best precision for Qwen3.8 Flash Next on an M4 Max Mac (128 GB)
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Q3_K_M | 90.8 GB | 5.17 GB | 48–83 |
| 32K | Q3_K_M | 91.5 GB | 4.55 GB | 41–70 |
| 128K | Q3_K_M | 93.9 GB | 2.07 GB | 26–43 |
| 256K (full) | Q2_K | 84.3 GB | 11.7 GB | 18–30 |
Qwen3.8 Flash Next on the M4 Max 128GB as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 112 GB | −16.2 GB | — | 197 GB | −101 GB | — |
| 8K | 112 GB | −16.3 GB | — | 197 GB | −101 GB | — |
| 16K | 112 GB | −16.5 GB | — | 197 GB | −101 GB | — |
| 32K | 113 GB | −16.9 GB | — | 197 GB | −101 GB | — |
| 64K | 114 GB | −17.7 GB | — | 198 GB | −102 GB | — |
| 128K | 115 GB | −19.4 GB | — | 200 GB | −104 GB | — |
| 256K (full) | 119 GB | −22.7 GB | — | 203 GB | −107 GB | — |
Other options
- The next smaller setting, Q2_K, takes 78.5 GB at 32K, 17.5 GB under the M4 Max 128GB; it fits with 0.5 GB to spare up to 256K (full) tokens.
- Smallest setup for Q4_K_M at 32K: DGX Spark (128 GB, 120 GB usable).
Run Qwen3.8 Flash Next on the M4 Max 128GB with llama-server
llama-server -hf bartowski/Qwen3.8-Flash-Next-GGUF:Q3_K_M -c 30720 Qwen3.8-Flash-Next-Q3_
Questions
Can I run Qwen3.8 Flash Next on an M4 Max Mac (128 GB)?
Only at Q3_K_M: Qwen3.8 Flash Next needs about 113 GB at Q4_K_M with 32K tokens of context, 16.9 GB more than the M4 Max Mac (128 GB, 96 GB usable) holds, but 91.5 GB at Q3_K_M, which fits with 4.55 GB to spare. The smallest setup here that holds Qwen3.8 Flash Next at Q4_K_M with 32K is DGX Spark (128 GB, 120 GB usable).
How fast is Qwen3.8 Flash Next on an M4 Max Mac (128 GB)?
At Q3_K_M with 32K tokens of context it writes about 41–70 tokens/s for one request on an M4 Max Mac (128 GB).
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3.8 Flash Next VRAM requirements, what LLMs an M4 Max Mac (128 GB) can run and every pair, or detect your own GPU. Model data checked .