What LLMs can an M3 Max Mac (128 GB) run?
An M3 Max Mac (128 GB) gives a model 96 GB usable of memory (about 75% of unified memory is usable by the GPU) and 400 GB/s of bandwidth. Of the 61 open models tracked here, 42 fit at Q4_K_M with an 8,192-token context, and 2 more at a lower precision, each leaving at least 0.5 GB free on the M3 Max Mac (128 GB).
Models that fit on one M3 Max Mac (128 GB)
The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first. gpt-oss-120b (67.7 GB, 36–61 tokens/s on the M3 Max Mac (128 GB)) and gpt-oss-20b (14.8 GB, 43–74 tokens/s on the M3 Max Mac (128 GB)) are listed at MXFP4 only, as published.
| Model | Best precision | Memory | Longest context | Tokens/s |
|---|---|---|---|---|
| Step 3.7 Flash 196B-A11B 201.4B, 11B active | GGUF Q2_K | 87.4 GB | 164K | 23–38 |
| Qwen3.8 Flash Next 180.0B, 6B active | GGUF Q3_K_M | 90.8 GB | 189K | 36–62 |
| Mistral Medium 3.5 128B 127.7B | GGUF Q4_K_M | 82.7 GB | 41K | 2.7–3.7 |
| Qwen3.5 122B-A10B 125.1B, 10B active | GGUF Q5_K_M | 91.5 GB | 161K | 16–27 |
| Nemotron 3 Super 120B-A12B 123.6B, 12B active | GGUF Q5_K_M | 90.3 GB | 256K (full) | 14–23 |
| gpt-oss-120b 116.8B, 5.1B active | As published (MXFP4) | 67.7 GB | 128K (full) | 36–61 |
| AliceAI Foundation 80B-A3B 81.3B, 3B active | GGUF Q8_0 | 89.2 GB | 252K | 34–57 |
| Qwen3-Coder-Next 79.7B, 3B active | GGUF Q8_0 | 87.4 GB | 256K (full) | 34–57 |
| Llama 3.1 70B 70.6B | GGUF Q8_0 | 80.0 GB | 52K | 2.8–3.9 |
| K2-Horizon MoVA 36B-A4B 37.4B, 4B active | FP16 / BF16 | 78.9 GB | 88K | 12–21 |
| Ornith 1.5 35B-A3B 36.0B, 3B active | FP16 / BF16 | 74.3 GB | 256K (full) | 19–32 |
| Qwen3.6 35B-A3B 36.0B, 3B active | FP16 / BF16 | 74.3 GB | 256K (full) | 19–32 |
| Ornith 1.0 35B 35.1B, 3B active | FP16 / BF16 | 72.6 GB | 256K (full) | 19–32 |
| LLM-jp-4.1 32B-A3B Thinking 32.1B, 3.8B active | FP16 / BF16 | 66.9 GB | 64K (full) | 14–24 |
| Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active | FP16 / BF16 | 65.3 GB | 256K (full) | 17–28 |
| Gemma 4 31B 31.3B | FP16 / BF16 | 66.6 GB | 256K (full) | 3.4–4.6 |
| GLM-4.7 Flash 31.2B, 3B active | FP16 / BF16 | 64.9 GB | 198K (full) | 18–31 |
| Xing 4.0 29B-A4B 31.2B, 4B active | FP16 / BF16 | 64.8 GB | 256K (full) | 14–24 |
| Qwen3-Coder 30B-A3B 30.5B, 3.3B active | FP16 / BF16 | 63.9 GB | 256K (full) | 16–27 |
| Qwen3 30B-A3B 30.5B, 3.3B active | FP16 / BF16 | 63.9 GB | 40K (full) | 16–27 |
| Muse Glimmer 30B 29.8B | FP16 / BF16 | 61.7 GB | 128K (full) | 3.7–5.0 |
| Granite 4.2 30B 29.3B | FP16 / BF16 | 62.7 GB | 127K | 3.6–4.9 |
| Qwen3.6 27B 27.8B | FP16 / BF16 | 58.0 GB | 256K (full) | 3.9–5.3 |
| Qwen3.8 27B 27.8B | FP16 / BF16 | 58.0 GB | 256K (full) | 3.9–5.3 |
| Hemmingway-1 27B 27.3B | FP16 / BF16 | 57.0 GB | 256K (full) | 4.0–5.4 |
| Gemma 4 26B-A4B 25.8B, 4B active | FP16 / BF16 | 53.9 GB | 256K (full) | 14–23 |
| gpt-oss-20b 20.9B, 3.6B active | As published (MXFP4) | 14.8 GB | 128K (full) | 43–74 |
| Gemma 4 12B 12.0B | FP16 / BF16 | 25.7 GB | 256K (full) | 8.8–12 |
| ZDTaichu 5.0 9B 9.8B | FP16 / BF16 | 20.8 GB | 128K (full) | 11–15 |
| Ornith 1.5 9B 9.7B | FP16 / BF16 | 20.6 GB | 256K (full) | 11–15 |
| Qwen3.5 9B 9.7B | FP16 / BF16 | 20.6 GB | 256K (full) | 11–15 |
| Ornith 1.0 9B 9.4B | FP16 / BF16 | 20.1 GB | 256K (full) | 11–16 |
| MiMo V2.6 Distill Qwen 9B 9.4B | FP16 / BF16 | 20.1 GB | 256K (full) | 11–16 |
| Granite 4.2 8B 8.8B | FP16 / BF16 | 19.9 GB | 128K (full) | 11–16 |
| LFM2.5 8B-A1B 8.5B, 1.5B active | FP16 / BF16 | 18.0 GB | 128,000 (full) | 37–62 |
| Qwen3 8B 8.2B | FP16 / BF16 | 18.5 GB | 40K (full) | 12–17 |
| Llama 3.1 8B 8.0B | FP16 / BF16 | 18.1 GB | 128K (full) | 13–17 |
| Gemma 4 E4B 8.0B | FP16 / BF16 | 17.1 GB | 128K (full) | 13–18 |
| Ling 3.0 Tiny 7.9B-A1.3B 7.9B, 1.3B active | FP16 / BF16 | 16.7 GB | 128K (full) | 42–73 |
| Spark-X2.5 4B 4.1B | FP16 / BF16 | 9.35 GB | 1M (full) | 25–34 |
| Nemotron 3 Nano 4B 4.0B | FP16 / BF16 | 8.78 GB | 256K (full) | 26–36 |
| Granite 4.2 3B 3.7B | FP16 / BF16 | 8.69 GB | 128K (full) | 26–37 |
| MiniCPM5 2B 2.5B | FP16 / BF16 | 6.02 GB | 128K (full) | 38–54 |
| Limite 1B Violetto 1.0B | FP16 / BF16 | 2.79 GB | 128K (full) | 86–126 |
Raising the GPU memory limit on an M3 Max Mac (128 GB)
macOS lets the GPU wire about 96 GB of this Mac's 128 GB by default. Running
sudo sysctl iogpu.wired_limit_mb=122880 raises that to 120 GB, leaving
8 GB to macOS and the apps next to it, until the next restart
(macOS 14 or newer; iogpu.wired_limit_mb=0 restores the default). LM Studio and Ollama pick the new limit up after a restart of the app.
At 120 GB, 43 of the 61 models fit at Q4_K_M with 8K context, against 42 by default
: Qwen3.8 Flash Next (30–51 tokens/s) join the list. Mistral Medium 3.5 128B, the biggest model that fits already, goes from 41K to 105K of context. If macOS runs short of memory it swaps or kills apps, so close big apps first and raise the limit in steps.
Too big for M3 Max Mac (128 GB) at 8K context
How much memory each of these lacks at Q4_K_M with 8K context, one request and 0.5 GB kept free, out of the 96 GB usable here; a bigger machine or a smaller quant is the way in.
- MiniMax M2.7: 48.9 GB short
- DeepSeek V4 Flash: 86.1 GB short
- DeepSeek V4 Flash 0731: 94.3 GB short
- MiMo V2.6 Flash: 98.0 GB short
- IQuest-Q1 320B-A15B: 106 GB short
- GLM-5.3 Flash: 104 GB short
- MiniMax M3: 171 GB short
- DeepSeek V3 / R1: 330 GB short
- DeepSeek V3.2: 330 GB short
- GLM-5.3: 373 GB short
- GLM-5.2: 373 GB short
- DeepSeek V4.1 Flash: 379 GB short
- Hy4 Preview 770B-A49B: 389 GB short
- MiMo V2.6 Pro: 540 GB short
- DeepSeek V4 Pro: 897 GB short
- Qwen3.8 2.4T-A95B: 1,422 GB short
- Kimi K3: 1,628 GB short
Best models for M3 Max Mac (128 GB) with room for long context
The biggest models that still fit at Q4_K_M with 32,768 tokens of context (or their whole window, if shorter), the memory left over, the longest context the Mac takes at that precision and the writing speed with that context in the cache, from 400 GB/s of bandwidth. gpt-oss-120b is at MXFP4, as published.
| Model | Context | Memory | Left over | Longest context | Tokens/s |
|---|---|---|---|---|---|
| Mistral Medium 3.5 128B 127.7B | 32K | 91.8 GB | 4.25 GB | 41K | 2.5–3.4 |
| Qwen3.5 122B-A10B 125.1B | 32K | 78.9 GB | 17.1 GB | 256K (full) | 17–29 |
| Nemotron 3 Super 120B-A12B 123.6B | 32K | 77.4 GB | 18.6 GB | 256K (full) | 16–26 |
| gpt-oss-120b 116.8B, MXFP4 | 32K | 68.6 GB | 27.4 GB | 128K (full) | 28–48 |
| AliceAI Foundation 80B-A3B 81.3B | 32K | 51.7 GB | 44.3 GB | 256K (full) | 43–74 |
llama-server commands for an M3 Max Mac (128 GB)
GGUF repos checked 2026-09-29; -c is the longest context with 0.5 GB free on the M3 Max Mac (128 GB). No -ngl: llama.cpp’s -fit, on by default since b7440, places the layers.
Mistral Medium 3.5 128B, Q4_K_M with 41K tokens: 2.4–3.2 tokens/s on the M3 Max Mac (128 GB)
llama-server -hf unsloth/Mistral-Medium-3.5-128B-GGUF:Q4_K_M -c 41984 Q4_K_M/
Qwen3.5 122B-A10B, Q4_K_M with 256K tokens: 9.5–16 tokens/s on the M3 Max Mac (128 GB)
llama-server -hf unsloth/Qwen3.5-122B-A10B-GGUF:Q4_K_M -c 262144 Q4_K_M/
How fast the M3 Max Mac (128 GB) writes as the context fills
Tokens per second at Q4_K_M (gpt-oss-120b at MXFP4) for the biggest models that fit with 32K tokens, with 8K and 32K tokens in the cache and at the longest context the Mac holds, and at Q8_0 with 8K where that fits. Each token reads the active weights and the whole cache once, so the speed follows the 400 GB/s of bandwidth and falls as the cache grows. The ceiling is that bandwidth divided by the bytes read per token, which no runtime reaches; the reply time is for 1,000 new tokens with 32K (or the longest context) in the cache, prompt processing aside. With the same memory, M4 Max Mac (128 GB), 546 GB/s, writes 35% faster on average; M5 Max Mac (128 GB), 614 GB/s, writes 50% faster on average.
| Model | 8K | 32K | Longest | At the longest | Q8_0 at 8K | Bandwidth ceiling at 8K | 1,000-token reply at 32K |
|---|---|---|---|---|---|---|---|
| Mistral Medium 3.5 128B | 2.7–3.7 | 2.5–3.4 | 41K | 2.4–3.2 | Does not fit | 5 | 297–406 s |
| Qwen3.5 122B-A10B | 19–31 | 17–29 | 256K | 9.5–16 | Does not fit | 64 | 35–59 s |
| Nemotron 3 Super 120B-A12B | 16–27 | 16–26 | 256K | 13–21 | Does not fit | 55 | 38–64 s |
| gpt-oss-120b | 36–61 | 28–48 | 128K | 15–26 | MXFP4 only | 126 | 21–36 s |
| AliceAI Foundation 80B-A3B | 55–95 | 43–74 | 256K | 14–24 | 34–57 | 198 | 14–23 s |
| Qwen3-Coder-Next | 55–95 | 43–74 | 256K | 14–24 | 34–57 | 198 | 14–23 s |
| Llama 3.1 70B | 4.8–6.6 | 4.1–5.6 | 128K | 2.6–3.5 | 2.8–3.9 | 9 | 179–244 s |
| K2-Horizon MoVA 36B-A4B | 28–48 | 13–22 | 348K | 1.7–2.8 | 20–34 | 99 | 45–75 s |
| Ornith 1.5 35B-A3B | 55–96 | 45–77 | 256K | 16–27 | 34–58 | 202 | 13–22 s |
| Qwen3.6 35B-A3B | 55–96 | 45–77 | 256K | 16–27 | 34–58 | 202 | 13–22 s |
M3 Max Mac (128 GB) against the other Macs with 96 GB usable
The same models fit on every Mac with 96 GB usable, so speed is what separates them. By bandwidth the M3 Max Mac (128 GB) is the slowest of the 3 Macs with 96 GB usable at 400 GB/s. Over the 12 biggest models that fit at Q4_K_M with 8K context, the M4 Max Mac (128 GB), at 546 GB/s, writes 35% faster and the M5 Max Mac (128 GB), at 614 GB/s, writes 50% faster than the M3 Max Mac (128 GB). Each cell below is the other Mac's tokens per second minus this one's, midpoints of the estimated ranges.
| Model | M3 Max Mac (128 GB) tokens/s | M4 Max Mac (128 GB) | M5 Max Mac (128 GB) |
|---|---|---|---|
| Mistral Medium 3.5 128B | 2.7–3.7 | +1.2 | +1.7 |
| Qwen3.5 122B-A10B | 19–31 | +8.9 | +13 |
| Nemotron 3 Super 120B-A12B | 16–27 | +7.6 | +11 |
| gpt-oss-120b | 36–61 | +17 | +24 |
| AliceAI Foundation 80B-A3B | 55–95 | +25 | +36 |
| Qwen3-Coder-Next | 55–95 | +25 | +36 |
| Llama 3.1 70B | 4.8–6.6 | +2.1 | +3.0 |
| K2-Horizon MoVA 36B-A4B | 28–48 | +13 | +20 |
| Ornith 1.5 35B-A3B | 55–96 | +25 | +37 |
| Qwen3.6 35B-A3B | 55–96 | +25 | +37 |
| Ornith 1.0 35B | 55–96 | +25 | +37 |
| LLM-jp-4.1 32B-A3B Thinking | 40–68 | +18 | +27 |
Near misses on M3 Max Mac (128 GB)
Models that miss at Q4_K_M with 32,768 tokens of context (or their whole window), or leave under 0.5 GB free there, but fit with a shorter context or 3-bit weights, or are short by at most a quarter of the memory and fit with MoE offload. Smallest shortfall first.
| Model | Short by at Q4_K_M | Fits instead |
|---|---|---|
| Qwen3.8 Flash Next 180.0B | 16.9 GB | GGUF Q3_K_M with 32K, 91.5 GB |
| Step 3.7 Flash 196B-A11B 201.4B | 31.1 GB | GGUF IQ3_XXS with 32K, 87.4 GB |
Run locally or rent an M3 Max Mac (128 GB)?
getdeploying.com lists no on-demand rental of M3 Max Mac (128 GB) (checked ). Among the GPUs this site tracks, the cheapest to rent with at least 96 GB is the RTX PRO 6000 Blackwell at a median $2.19 an hour, which holds 42 of the 61 models at Q4_K_M with 8K context against 42 here; every $1,000 equals about 457 hours of it.
How these numbers are worked out
Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.
Nearby GPUs and setups
How many of the tracked models each one holds at Q4_K_M with 8K context, against 42 here.
| Setup | How it relates | Memory | Bandwidth | Models at Q4_K_M |
|---|---|---|---|---|
| M4 Max Mac (128 GB) | Same memory | 96 GB usable | 546 GB/s | 42 (same) |
| M5 Max Mac (128 GB) | Same memory | 96 GB usable | 614 GB/s | 42 (same) |
| M5 Max Mac (64 GB) | Next size down | 48 GB usable | 614 GB/s | 36 (−6) |
| M3 Ultra Mac Studio (512 GB) | Next size up | 384 GB usable | 819 GB/s | 51 (+9) |
Model numbers read from Hugging Face on .