Can I run Llama 3.1 70B on an M4 Max Mac (128 GB)?

Yes: Llama 3.1 70B needs about 55.2 GB at Q4_K_M with 32K tokens of context, which fits the M4 Max Mac (128 GB, 96 GB usable) with 40.8 GB to spare. At 32K the M4 Max 128GB holds up to Q8_0 (88.3 GB), and Q4_K_M runs up to 128K (full) tokens.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
55.2 GB
M4 Max 128GB
96 GB usable, 546 GB/s
To spare
40.8 GB
Tokens/s
5.6–7.6 tokens/s

At Q4_K_M with 32K tokens of context it writes about 5.6–7.6 tokens/s for one request on an M4 Max Mac (128 GB).

Best precision for Llama 3.1 70B on an M4 Max Mac (128 GB)

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q8_0 80.0 GB 16.0 GB 3.8–5.3
32K Q8_0 88.3 GB 7.70 GB 3.5–4.8
128K (full) Q4_K_M 88.2 GB 7.77 GB 3.5–4.8

Llama 3.1 70B on the M4 Max 128GB as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 45.6 GB 50.4 GB 6.8–9.3 78.7 GB 17.3 GB 3.9–5.4
8K 47.0 GB 49.0 GB 6.6–9.0 80.0 GB 16.0 GB 3.8–5.3
16K 49.7 GB 46.3 GB 6.2–8.5 82.8 GB 13.2 GB 3.7–5.1
32K 55.2 GB 40.8 GB 5.6–7.6 88.3 GB 7.70 GB 3.5–4.8
64K 66.2 GB 29.8 GB 4.6–6.4 99.3 GB −3.30 GB —
128K (full) 88.2 GB 7.77 GB 3.5–4.8 121 GB −25.3 GB —

Other options

Run Llama 3.1 70B on the M4 Max 128GB with llama-server

llama-server -hf bartowski/Meta-Llama-3.1-70B-Instruct-GGUF:Q4_K_M -c 131072

Meta-Llama-3.1-70B-Instruct-Q4_K_M.gguf, 42.5 GB, from bartowski/Meta-Llama-3.1-70B-Instruct-GGUF (checked 2026-09-29). At -c 131072 on the M4 Max 128GB: 88.2 GB of 96 GB, 7.77 GB free.

Questions

Can I run Llama 3.1 70B on an M4 Max Mac (128 GB)?

Yes: Llama 3.1 70B needs about 55.2 GB at Q4_K_M with 32K tokens of context, which fits the M4 Max Mac (128 GB, 96 GB usable) with 40.8 GB to spare. At 32K the M4 Max 128GB holds up to Q8_0 (88.3 GB), and Q4_K_M runs up to 128K (full) tokens.

How fast is Llama 3.1 70B on an M4 Max Mac (128 GB)?

At Q4_K_M with 32K tokens of context it writes about 5.6–7.6 tokens/s for one request on an M4 Max Mac (128 GB).

Try other settings in the VRAM calculator or the speed calculator. See also Llama 3.1 70B VRAM requirements, what LLMs an M4 Max Mac (128 GB) can run and every pair, or detect your own GPU. Model data checked .