Can I run Qwen3.8 Flash Next on an M4 Max Mac (128 GB)?

Only at Q3_K_M: Qwen3.8 Flash Next needs about 113 GB at Q4_K_M with 32K tokens of context, 16.9 GB more than the M4 Max Mac (128 GB, 96 GB usable) holds, but 91.5 GB at Q3_K_M, which fits with 4.55 GB to spare. The smallest setup here that holds Qwen3.8 Flash Next at Q4_K_M with 32K is DGX Spark (128 GB, 120 GB usable).

Partly Q3_K_M with 32K tokens of context

Q4_K_M, 32K
113 GB
M4 Max 128GB
96 GB usable, 546 GB/s
Short by
16.9 GB
Tokens/s
41–70 tokens/s

At Q3_K_M with 32K tokens of context it writes about 41–70 tokens/s for one request on an M4 Max Mac (128 GB).

Best precision for Qwen3.8 Flash Next on an M4 Max Mac (128 GB)

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q3_K_M 90.8 GB 5.17 GB 48–83
32K Q3_K_M 91.5 GB 4.55 GB 41–70
128K Q3_K_M 93.9 GB 2.07 GB 26–43
256K (full) Q2_K 84.3 GB 11.7 GB 18–30

Qwen3.8 Flash Next on the M4 Max 128GB as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 112 GB −16.2 GB — 197 GB −101 GB —
8K 112 GB −16.3 GB — 197 GB −101 GB —
16K 112 GB −16.5 GB — 197 GB −101 GB —
32K 113 GB −16.9 GB — 197 GB −101 GB —
64K 114 GB −17.7 GB — 198 GB −102 GB —
128K 115 GB −19.4 GB — 200 GB −104 GB —
256K (full) 119 GB −22.7 GB — 203 GB −107 GB —

Other options

Run Qwen3.8 Flash Next on the M4 Max 128GB with llama-server

llama-server -hf bartowski/Qwen3.8-Flash-Next-GGUF:Q3_K_M -c 30720

Qwen3.8-Flash-Next-Q3_K_M/Qwen3.8-Flash-Next-Q3_K_M-00001-of-00003.gguf, 92.0 GB, from bartowski/Qwen3.8-Flash-Next-GGUF (checked 2026-09-29). At -c 30720 on the M4 Max 128GB: 95.5 GB of 96 GB, 532 MB free. The file is 3.71 GB over the estimate above, so -c counts the file.

Questions

Can I run Qwen3.8 Flash Next on an M4 Max Mac (128 GB)?

Only at Q3_K_M: Qwen3.8 Flash Next needs about 113 GB at Q4_K_M with 32K tokens of context, 16.9 GB more than the M4 Max Mac (128 GB, 96 GB usable) holds, but 91.5 GB at Q3_K_M, which fits with 4.55 GB to spare. The smallest setup here that holds Qwen3.8 Flash Next at Q4_K_M with 32K is DGX Spark (128 GB, 120 GB usable).

How fast is Qwen3.8 Flash Next on an M4 Max Mac (128 GB)?

At Q3_K_M with 32K tokens of context it writes about 41–70 tokens/s for one request on an M4 Max Mac (128 GB).

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3.8 Flash Next VRAM requirements, what LLMs an M4 Max Mac (128 GB) can run and every pair, or detect your own GPU. Model data checked .