What LLMs can an M5 Pro Mac (64 GB) run?
An M5 Pro Mac (64 GB) gives a model 48 GB usable of memory (about 75% of unified memory is usable by the GPU) and 307 GB/s of bandwidth. Of the 61 open models tracked here, 36 fit at Q4_K_M with an 8,192-token context, and 2 more at a lower precision, each leaving at least 0.5 GB free on the M5 Pro Mac (64 GB).
Models that fit on one M5 Pro Mac (64 GB)
The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first. gpt-oss-20b (14.8 GB, 34–58 tokens/s on the M5 Pro Mac (64 GB)) is listed at MXFP4 only, as published.
| Model | Best precision | Memory | Longest context | Tokens/s |
|---|---|---|---|---|
| AliceAI Foundation 80B-A3B 81.3B, 3B active | GGUF Q3_K_M | 41.4 GB | 244K | 51–88 |
| Qwen3-Coder-Next 79.7B, 3B active | GGUF Q3_K_M | 40.6 GB | 256K (full) | 51–88 |
| Llama 3.1 70B 70.6B | GGUF Q4_K_M | 47.0 GB | 9K | 3.7–5.1 |
| K2-Horizon MoVA 36B-A4B 37.4B, 4B active | GGUF Q8_0 | 42.9 GB | 30K | 15–26 |
| Ornith 1.5 35B-A3B 36.0B, 3B active | GGUF Q8_0 | 39.8 GB | 256K (full) | 26–45 |
| Qwen3.6 35B-A3B 36.0B, 3B active | GGUF Q8_0 | 39.8 GB | 256K (full) | 26–45 |
| Ornith 1.0 35B 35.1B, 3B active | GGUF Q8_0 | 38.9 GB | 256K (full) | 26–45 |
| LLM-jp-4.1 32B-A3B Thinking 32.1B, 3.8B active | GGUF Q8_0 | 36.0 GB | 64K (full) | 19–33 |
| Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active | GGUF Q8_0 | 34.9 GB | 256K (full) | 24–40 |
| Gemma 4 31B 31.3B | GGUF Q8_0 | 36.5 GB | 135K | 4.8–6.5 |
| GLM-4.7 Flash 31.2B, 3B active | GGUF Q8_0 | 34.9 GB | 198K (full) | 24–41 |
| Xing 4.0 29B-A4B 31.2B, 4B active | GGUF Q8_0 | 34.9 GB | 256K (full) | 19–33 |
| Qwen3-Coder 30B-A3B 30.5B, 3.3B active | GGUF Q8_0 | 34.6 GB | 133K | 21–35 |
| Qwen3 30B-A3B 30.5B, 3.3B active | GGUF Q8_0 | 34.6 GB | 40K (full) | 21–35 |
| Muse Glimmer 30B 29.8B | GGUF Q8_0 | 33.1 GB | 128K (full) | 5.3–7.2 |
| Granite 4.2 30B 29.3B | GGUF Q8_0 | 34.6 GB | 54K | 5.0–6.9 |
| Qwen3.6 27B 27.8B | GGUF Q8_0 | 31.3 GB | 243K | 5.6–7.6 |
| Qwen3.8 27B 27.8B | GGUF Q8_0 | 31.3 GB | 243K | 5.6–7.6 |
| Hemmingway-1 27B 27.3B | GGUF Q8_0 | 30.8 GB | 251K | 5.7–7.8 |
| Gemma 4 26B-A4B 25.8B, 4B active | GGUF Q8_0 | 29.1 GB | 256K (full) | 19–32 |
| gpt-oss-20b 20.9B, 3.6B active | As published (MXFP4) | 14.8 GB | 128K (full) | 34–58 |
| Gemma 4 12B 12.0B | FP16 / BF16 | 25.7 GB | 256K (full) | 6.8–9.3 |
| ZDTaichu 5.0 9B 9.8B | FP16 / BF16 | 20.8 GB | 128K (full) | 8.4–12 |
| Ornith 1.5 9B 9.7B | FP16 / BF16 | 20.6 GB | 256K (full) | 8.5–12 |
| Qwen3.5 9B 9.7B | FP16 / BF16 | 20.6 GB | 256K (full) | 8.5–12 |
| Ornith 1.0 9B 9.4B | FP16 / BF16 | 20.1 GB | 256K (full) | 8.7–12 |
| MiMo V2.6 Distill Qwen 9B 9.4B | FP16 / BF16 | 20.1 GB | 256K (full) | 8.7–12 |
| Granite 4.2 8B 8.8B | FP16 / BF16 | 19.9 GB | 128K (full) | 8.8–12 |
| LFM2.5 8B-A1B 8.5B, 1.5B active | FP16 / BF16 | 18.0 GB | 128,000 (full) | 28–48 |
| Qwen3 8B 8.2B | FP16 / BF16 | 18.5 GB | 40K (full) | 9.5–13 |
| Llama 3.1 8B 8.0B | FP16 / BF16 | 18.1 GB | 128K (full) | 9.7–13 |
| Gemma 4 E4B 8.0B | FP16 / BF16 | 17.1 GB | 128K (full) | 10–14 |
| Ling 3.0 Tiny 7.9B-A1.3B 7.9B, 1.3B active | FP16 / BF16 | 16.7 GB | 128K (full) | 33–56 |
| Spark-X2.5 4B 4.1B | FP16 / BF16 | 9.35 GB | 994K | 19–26 |
| Nemotron 3 Nano 4B 4.0B | FP16 / BF16 | 8.78 GB | 256K (full) | 20–28 |
| Granite 4.2 3B 3.7B | FP16 / BF16 | 8.69 GB | 128K (full) | 20–28 |
| MiniCPM5 2B 2.5B | FP16 / BF16 | 6.02 GB | 128K (full) | 30–42 |
| Limite 1B Violetto 1.0B | FP16 / BF16 | 2.79 GB | 128K (full) | 68–98 |
Raising the GPU memory limit on an M5 Pro Mac (64 GB)
macOS lets the GPU wire about 48 GB of this Mac's 64 GB by default. Running
sudo sysctl iogpu.wired_limit_mb=57344 raises that to 56 GB, leaving
8 GB to macOS and the apps next to it, until the next restart
(macOS 14 or newer; iogpu.wired_limit_mb=0 restores the default). LM Studio and Ollama pick the new limit up after a restart of the app.
At 56 GB, 38 of the 61 models fit at Q4_K_M with 8K context, against 36 by default
: Qwen3-Coder-Next (43–73 tokens/s) and AliceAI Foundation 80B-A3B (43–73 tokens/s) join the list. Llama 3.1 70B, the biggest model that fits already, goes from 9K to 32K of context. If macOS runs short of memory it swaps or kills apps, so close big apps first and raise the limit in steps.
Too big for M5 Pro Mac (64 GB) at 8K context
How much memory each of these lacks at Q4_K_M (gpt-oss-120b at MXFP4) with 8K context, one request and 0.5 GB kept free, out of the 48 GB usable here; a bigger machine or a smaller quant is the way in.
- gpt-oss-120b: 20.2 GB short
- Nemotron 3 Super 120B-A12B: 29.7 GB short
- Qwen3.5 122B-A10B: 30.7 GB short
- Mistral Medium 3.5 128B: 35.2 GB short
- Qwen3.8 Flash Next: 64.8 GB short
- Step 3.7 Flash 196B-A11B: 78.4 GB short
- MiniMax M2.7: 96.9 GB short
- DeepSeek V4 Flash: 134 GB short
- DeepSeek V4 Flash 0731: 142 GB short
- MiMo V2.6 Flash: 146 GB short
- IQuest-Q1 320B-A15B: 154 GB short
- GLM-5.3 Flash: 152 GB short
- MiniMax M3: 219 GB short
- DeepSeek V3 / R1: 378 GB short
- DeepSeek V3.2: 378 GB short
- GLM-5.3: 421 GB short
- GLM-5.2: 421 GB short
- DeepSeek V4.1 Flash: 427 GB short
- Hy4 Preview 770B-A49B: 437 GB short
- MiMo V2.6 Pro: 588 GB short
- DeepSeek V4 Pro: 945 GB short
- Qwen3.8 2.4T-A95B: 1,470 GB short
- Kimi K3: 1,676 GB short
Best models for M5 Pro Mac (64 GB) with room for long context
The biggest models that still fit at Q4_K_M with 32,768 tokens of context (or their whole window, if shorter), the memory left over, the longest context the Mac takes at that precision and the writing speed with that context in the cache, from 307 GB/s of bandwidth.
| Model | Context | Memory | Left over | Longest context | Tokens/s |
|---|---|---|---|---|---|
| K2-Horizon MoVA 36B-A4B 37.4B | 32K | 30.3 GB | 17.7 GB | 115K | 10–17 |
| Ornith 1.5 35B-A3B 36.0B | 32K | 23.5 GB | 24.5 GB | 256K (full) | 35–60 |
| Qwen3.6 35B-A3B 36.0B | 32K | 23.5 GB | 24.5 GB | 256K (full) | 35–60 |
| Ornith 1.0 35B 35.1B | 32K | 22.9 GB | 25.1 GB | 256K (full) | 35–60 |
| LLM-jp-4.1 32B-A3B Thinking 32.1B | 32K | 22.6 GB | 25.4 GB | 64K (full) | 20–34 |
llama-server commands for an M5 Pro Mac (64 GB)
GGUF repos checked 2026-09-29; -c is the longest context with 0.5 GB free on the M5 Pro Mac (64 GB).
K2-Horizon MoVA 36B-A4B, Q4_K_M with 115K tokens: 3.6–6.0 tokens/s on the M5 Pro Mac (64 GB)
llama-server -hf IFM/K2-Horizon-MoVA-36B-A4B-GGUF:Q4_K_M -c 117760 -ngl 99 K2-Horizon-MoVA-36B-A4B-Q4_K_M.gguf, 22.4 GB: 47.4 GB used and 587 MB free of the M5 Pro Mac (64 GB)'s 48 GB usable, 3.6–6.0 tokens/s once the 115K cache is full.
Ornith 1.5 35B-A3B, Q4_K_M with 256K tokens: 13–21 tokens/s on the M5 Pro Mac (64 GB)
llama-server -hf bartowski/Ornith-1.5-35B-A3B-GGUF:Q4_K_M -c 262144 -ngl 99 Ornith-1.5-35B-A3B-Q4_K_M.gguf, 21.9 GB: 28.4 GB used and 19.6 GB free of the M5 Pro Mac (64 GB)'s 48 GB usable, 13–21 tokens/s once the 256K cache is full.
How fast the M5 Pro Mac (64 GB) writes as the context fills
Tokens per second at Q4_K_M for the biggest models that fit with 32K tokens, with 8K and 32K tokens in the cache and at the longest context the Mac holds, and at Q8_0 with 8K where that fits. Each token reads the active weights and the whole cache once, so the speed follows the 307 GB/s of bandwidth and falls as the cache grows. The ceiling is that bandwidth divided by the bytes read per token, which no runtime reaches; the reply time is for 1,000 new tokens with 32K (or the longest context) in the cache, prompt processing aside. With the same memory, M4 Pro Mac (64 GB), 273 GB/s, writes 11% slower on average; M1–M3 Max Mac (64 GB), 400 GB/s, writes 29% faster on average; M4 Max Mac (64 GB), 546 GB/s, writes 73% faster on average; M5 Max Mac (64 GB), 614 GB/s, writes 94% faster on average.
| Model | 8K | 32K | Longest | At the longest | Q8_0 at 8K | Bandwidth ceiling at 8K | 1,000-token reply at 32K |
|---|---|---|---|---|---|---|---|
| K2-Horizon MoVA 36B-A4B | 22–37 | 10–17 | 115K | 3.6–6.0 | 15–26 | 76 | 58–98 s |
| Ornith 1.5 35B-A3B | 43–75 | 35–60 | 256K | 13–21 | 26–45 | 155 | 17–28 s |
| Qwen3.6 35B-A3B | 43–75 | 35–60 | 256K | 13–21 | 26–45 | 155 | 17–28 s |
| Ornith 1.0 35B | 43–75 | 35–60 | 256K | 13–21 | 26–45 | 155 | 17–28 s |
| LLM-jp-4.1 32B-A3B Thinking | 31–52 | 20–34 | 64K | 14–23 | 19–33 | 108 | 30–50 s |
| Nemotron 3 Nano 30B-A3B | 40–68 | 37–64 | 256K | 24–40 | 24–40 | 142 | 16–27 s |
| Gemma 4 31B | 8.0–11 | 7.3–10 | 256K | 4.0–5.5 | 4.8–6.5 | 15 | 100–137 s |
| GLM-4.7 Flash | 38–66 | 25–42 | 198K | 7.1–12 | 24–41 | 136 | 24–40 s |
| Xing 4.0 29B-A4B | 31–53 | 23–38 | 256K | 6.3–11 | 19–33 | 110 | 26–44 s |
| Qwen3-Coder 30B-A3B | 31–53 | 17–29 | 256K | 3.3–5.5 | 21–35 | 110 | 34–58 s |
M5 Pro Mac (64 GB) against the other Macs with 48 GB usable
The same models fit on every Mac with 48 GB usable, so speed is what separates them. By bandwidth the M5 Pro Mac (64 GB) ranks 4 of 5 at 307 GB/s. Over the 12 biggest models that fit at Q4_K_M with 8K context, the M4 Pro Mac (64 GB), at 273 GB/s, writes 11% slower, the M1–M3 Max Mac (64 GB), at 400 GB/s, writes 29% faster, the M4 Max Mac (64 GB), at 546 GB/s, writes 73% faster and the M5 Max Mac (64 GB), at 614 GB/s, writes 94% faster than the M5 Pro Mac (64 GB). Each cell below is the other Mac's tokens per second minus this one's, midpoints of the estimated ranges.
| Model | M5 Pro Mac (64 GB) tokens/s | M4 Pro Mac (64 GB) | M1–M3 Max Mac (64 GB) | M4 Max Mac (64 GB) | M5 Max Mac (64 GB) |
|---|---|---|---|---|---|
| Llama 3.1 70B | 3.7–5.1 | −0.5 | +1.3 | +3.4 | +4.3 |
| K2-Horizon MoVA 36B-A4B | 22–37 | −3.2 | +8.7 | +22 | +28 |
| Ornith 1.5 35B-A3B | 43–75 | −6.3 | +17 | +42 | +54 |
| Qwen3.6 35B-A3B | 43–75 | −6.3 | +17 | +42 | +54 |
| Ornith 1.0 35B | 43–75 | −6.3 | +17 | +42 | +54 |
| LLM-jp-4.1 32B-A3B Thinking | 31–52 | −4.5 | +12 | +31 | +39 |
| Nemotron 3 Nano 30B-A3B | 40–68 | −5.8 | +15 | +39 | +50 |
| Gemma 4 31B | 8.0–11 | −1.0 | +2.8 | +7.3 | +9.3 |
| GLM-4.7 Flash | 38–66 | −5.6 | +15 | +38 | +48 |
| Xing 4.0 29B-A4B | 31–53 | −4.6 | +12 | +31 | +40 |
| Qwen3-Coder 30B-A3B | 31–53 | −4.5 | +12 | +31 | +40 |
| Qwen3 30B-A3B | 31–53 | −4.5 | +12 | +31 | +40 |
Near misses on M5 Pro Mac (64 GB)
Models that miss at Q4_K_M with 32,768 tokens of context (or their whole window), or leave under 0.5 GB free there, but fit with a shorter context or 3-bit weights, or are short by at most a quarter of the memory and fit with MoE offload. Smallest shortfall first.
| Model | Short by at Q4_K_M | Fits instead |
|---|---|---|
| Qwen3-Coder-Next 79.7B | 2.71 GB | GGUF Q3_K_M with 32K, 41.2 GB |
| AliceAI Foundation 80B-A3B 81.3B | 3.71 GB | GGUF Q3_K_M with 32K, 42.0 GB |
| Llama 3.1 70B 70.6B | 7.23 GB | Q4_K_M with 9K context |
Run locally or rent an M5 Pro Mac (64 GB)?
getdeploying.com lists no on-demand rental of M5 Pro Mac (64 GB) (checked ). Among the GPUs this site tracks, the cheapest to rent with at least 48 GB is the L40S at a median $1.57 an hour, which holds 36 of the 61 models at Q4_K_M with 8K context against 36 here; every $1,000 equals about 637 hours of it.
How these numbers are worked out
Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.
Nearby GPUs and setups
How many of the tracked models each one holds at Q4_K_M with 8K context, against 36 here.
| Setup | How it relates | Memory | Bandwidth | Models at Q4_K_M |
|---|---|---|---|---|
| M4 Pro Mac (64 GB) | Same memory | 48 GB usable | 273 GB/s | 36 (same) |
| M1–M3 Max Mac (64 GB) | Same memory | 48 GB usable | 400 GB/s | 36 (same) |
| M4 Max Mac (64 GB) | Same memory | 48 GB usable | 546 GB/s | 36 (same) |
| M5 Max Mac (64 GB) | Same memory | 48 GB usable | 614 GB/s | 36 (same) |
| M5 Max Mac (48 GB) | Next size down | 36 GB usable | 614 GB/s | 35 (−1) |
| M3 Max Mac (128 GB) | Next size up | 96 GB usable | 400 GB/s | 42 (+6) |
Model numbers read from Hugging Face on .