Can I run Llama 3.1 70B on an M4 Max Mac (128 GB)?
Yes: Llama 3.1 70B needs about 55.2 GB at Q4_K_M with 32K tokens of context, which fits the M4 Max Mac (128 GB, 96 GB usable) with 40.8 GB to spare. At 32K the M4 Max 128GB holds up to Q8_0 (88.3 GB), and Q4_K_M runs up to 128K (full) tokens.
Yes Q4_K_M with 32K tokens of context
- Q4_K_M, 32K
- 55.2 GB
- M4 Max 128GB
- 96 GB usable, 546 GB/s
- To spare
- 40.8 GB
- Tokens/s
- 5.6–7.6 tokens/s
At Q4_K_M with 32K tokens of context it writes about 5.6–7.6 tokens/s for one request on an M4 Max Mac (128 GB).
Best precision for Llama 3.1 70B on an M4 Max Mac (128 GB)
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Q8_0 | 80.0 GB | 16.0 GB | 3.8–5.3 |
| 32K | Q8_0 | 88.3 GB | 7.70 GB | 3.5–4.8 |
| 128K (full) | Q4_K_M | 88.2 GB | 7.77 GB | 3.5–4.8 |
Llama 3.1 70B on the M4 Max 128GB as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 45.6 GB | 50.4 GB | 6.8–9.3 | 78.7 GB | 17.3 GB | 3.9–5.4 |
| 8K | 47.0 GB | 49.0 GB | 6.6–9.0 | 80.0 GB | 16.0 GB | 3.8–5.3 |
| 16K | 49.7 GB | 46.3 GB | 6.2–8.5 | 82.8 GB | 13.2 GB | 3.7–5.1 |
| 32K | 55.2 GB | 40.8 GB | 5.6–7.6 | 88.3 GB | 7.70 GB | 3.5–4.8 |
| 64K | 66.2 GB | 29.8 GB | 4.6–6.4 | 99.3 GB | −3.30 GB | — |
| 128K (full) | 88.2 GB | 7.77 GB | 3.5–4.8 | 121 GB | −25.3 GB | — |
Other options
- The next smaller setting, Q3_K_M, takes 46.8 GB at 32K, 49.2 GB under the M4 Max 128GB; it fits with 0.5 GB to spare up to 128K (full) tokens.
- Smallest setup for Q4_K_M at 32K: 2× RTX 5090 (64 GB).
Run Llama 3.1 70B on the M4 Max 128GB with llama-server
llama-server -hf bartowski/Meta-Llama-3.1-70B-Instruct-GGUF:Q4_K_M -c 131072 Meta-Llama-3.1-70B-Instruct-Q4_K_M.gguf, 42.5 GB, from bartowski/
Questions
Can I run Llama 3.1 70B on an M4 Max Mac (128 GB)?
Yes: Llama 3.1 70B needs about 55.2 GB at Q4_K_M with 32K tokens of context, which fits the M4 Max Mac (128 GB, 96 GB usable) with 40.8 GB to spare. At 32K the M4 Max 128GB holds up to Q8_0 (88.3 GB), and Q4_K_M runs up to 128K (full) tokens.
How fast is Llama 3.1 70B on an M4 Max Mac (128 GB)?
At Q4_K_M with 32K tokens of context it writes about 5.6–7.6 tokens/s for one request on an M4 Max Mac (128 GB).
Try other settings in the VRAM calculator or the speed calculator. See also Llama 3.1 70B VRAM requirements, what LLMs an M4 Max Mac (128 GB) can run and every pair, or detect your own GPU. Model data checked .