What LLMs can an M5 Mac (32 GB) run?
An M5 Mac (32 GB) gives a model 21.3 GB usable of memory (about two thirds of unified memory is usable by the GPU) and 153 GB/s of bandwidth. Of the 61 open models tracked here, 28 fit at Q4_K_M with an 8,192-token context, and 7 more at a lower precision, each leaving at least 0.5 GB free on the M5 Mac (32 GB).
Models that fit on one M5 Mac (32 GB)
The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first. gpt-oss-20b (14.8 GB, 17–29 tokens/s on the M5 Mac (32 GB)) is listed at MXFP4 only, as published.
| Model | Best precision | Memory | Longest context | Tokens/s |
|---|---|---|---|---|
| K2-Horizon MoVA 36B-A4B 37.4B, 4B active | GGUF Q2_K | 18.2 GB | 20K | 14–23 |
| Ornith 1.5 35B-A3B 36.0B, 3B active | GGUF Q3_K_M | 18.7 GB | 106K | 27–46 |
| Qwen3.6 35B-A3B 36.0B, 3B active | GGUF Q3_K_M | 18.7 GB | 106K | 27–46 |
| Ornith 1.0 35B 35.1B, 3B active | GGUF Q3_K_M | 18.3 GB | 126K | 27–46 |
| LLM-jp-4.1 32B-A3B Thinking 32.1B, 3.8B active | GGUF Q3_K_M | 17.1 GB | 61K | 19–31 |
| Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active | GGUF Q4_K_M | 20.1 GB | 112K | 21–35 |
| Gemma 4 31B 31.3B | GGUF Q3_K_M | 18.1 GB | 38K | 4.9–6.6 |
| GLM-4.7 Flash 31.2B, 3B active | GGUF Q4_K_M | 20.3 GB | 16K | 20–33 |
| Xing 4.0 29B-A4B 31.2B, 4B active | GGUF Q4_K_M | 20.2 GB | 19K | 16–27 |
| Qwen3-Coder 30B-A3B 30.5B, 3.3B active | GGUF Q4_K_M | 20.2 GB | 13K | 16–27 |
| Qwen3 30B-A3B 30.5B, 3.3B active | GGUF Q4_K_M | 20.2 GB | 13K | 16–27 |
| Muse Glimmer 30B 29.8B | GGUF Q4_K_M | 19.2 GB | 124K | 4.6–6.3 |
| Granite 4.2 30B 29.3B | GGUF Q3_K_M | 17.4 GB | 20K | 5.1–6.9 |
| Qwen3.6 27B 27.8B | GGUF Q4_K_M | 18.3 GB | 44K | 4.8–6.6 |
| Qwen3.8 27B 27.8B | GGUF Q4_K_M | 18.3 GB | 44K | 4.8–6.6 |
| Hemmingway-1 27B 27.3B | GGUF Q4_K_M | 18.0 GB | 48K | 4.9–6.7 |
| Gemma 4 26B-A4B 25.8B, 4B active | GGUF Q5_K_M | 19.7 GB | 57K | 14–23 |
| gpt-oss-20b 20.9B, 3.6B active | As published (MXFP4) | 14.8 GB | 128K (full) | 17–29 |
| Gemma 4 12B 12.0B | GGUF Q8_0 | 14.2 GB | 256K (full) | 6.2–8.6 |
| ZDTaichu 5.0 9B 9.8B | GGUF Q8_0 | 11.4 GB | 128K (full) | 7.8–11 |
| Ornith 1.5 9B 9.7B | FP16 / BF16 | 20.6 GB | 14K | 4.3–5.8 |
| Qwen3.5 9B 9.7B | FP16 / BF16 | 20.6 GB | 14K | 4.3–5.8 |
| Ornith 1.0 9B 9.4B | FP16 / BF16 | 20.1 GB | 29K | 4.4–6.0 |
| MiMo V2.6 Distill Qwen 9B 9.4B | FP16 / BF16 | 20.1 GB | 29K | 4.4–6.0 |
| Granite 4.2 8B 8.8B | FP16 / BF16 | 19.9 GB | 13K | 4.4–6.0 |
| LFM2.5 8B-A1B 8.5B, 1.5B active | FP16 / BF16 | 18.0 GB | 128,000 (full) | 14–24 |
| Qwen3 8B 8.2B | FP16 / BF16 | 18.5 GB | 22K | 4.8–6.5 |
| Llama 3.1 8B 8.0B | FP16 / BF16 | 18.1 GB | 27K | 4.9–6.7 |
| Gemma 4 E4B 8.0B | FP16 / BF16 | 17.1 GB | 128K (full) | 5.2–7.1 |
| Ling 3.0 Tiny 7.9B-A1.3B 7.9B, 1.3B active | FP16 / BF16 | 16.7 GB | 128K (full) | 17–28 |
| Spark-X2.5 4B 4.1B | FP16 / BF16 | 9.35 GB | 303K | 9.6–13 |
| Nemotron 3 Nano 4B 4.0B | FP16 / BF16 | 8.78 GB | 256K (full) | 10–14 |
| Granite 4.2 3B 3.7B | FP16 / BF16 | 8.69 GB | 128K (full) | 10–14 |
| MiniCPM5 2B 2.5B | FP16 / BF16 | 6.02 GB | 128K (full) | 15–21 |
| Limite 1B Violetto 1.0B | FP16 / BF16 | 2.79 GB | 128K (full) | 36–50 |
Raising the GPU memory limit on an M5 Mac (32 GB)
macOS lets the GPU wire about 21.3 GB of this Mac's 32 GB by default. Running
sudo sysctl iogpu.wired_limit_mb=24576 raises that to 24 GB, leaving
8 GB to macOS and the apps next to it, until the next restart
(macOS 14 or newer; iogpu.wired_limit_mb=0 restores the default). LM Studio and Ollama pick the new limit up after a restart of the app.
At 24 GB, 34 of the 61 models fit at Q4_K_M with 8K context, against 28 by default
: Granite 4.2 30B (4.2–5.8 tokens/s), Gemma 4 31B (4.0–5.5 tokens/s), LLM-jp-4.1 32B-A3B Thinking (16–26 tokens/s), Ornith 1.0 35B (22–38 tokens/s), Ornith 1.5 35B-A3B (22–38 tokens/s) and Qwen3.6 35B-A3B (22–38 tokens/s) join the list. Nemotron 3 Nano 30B-A3B, the biggest model that fits already, goes from 112K to 256K of context. If macOS runs short of memory it swaps or kills apps, so close big apps first and raise the limit in steps.
Too big for M5 Mac (32 GB) at 8K context
How much memory each of these lacks at Q4_K_M (gpt-oss-120b at MXFP4) with 8K context, one request and 0.5 GB kept free, out of the 21.3 GB usable here; a bigger machine or a smaller quant is the way in.
- Llama 3.1 70B: 26.2 GB short
- Qwen3-Coder-Next: 29.3 GB short
- AliceAI Foundation 80B-A3B: 30.3 GB short
- gpt-oss-120b: 46.9 GB short
- Nemotron 3 Super 120B-A12B: 56.4 GB short
- Qwen3.5 122B-A10B: 57.4 GB short
- Mistral Medium 3.5 128B: 61.9 GB short
- Qwen3.8 Flash Next: 91.5 GB short
- Step 3.7 Flash 196B-A11B: 105 GB short
- MiniMax M2.7: 124 GB short
- DeepSeek V4 Flash: 161 GB short
- DeepSeek V4 Flash 0731: 169 GB short
- MiMo V2.6 Flash: 173 GB short
- IQuest-Q1 320B-A15B: 180 GB short
- GLM-5.3 Flash: 179 GB short
- MiniMax M3: 245 GB short
- DeepSeek V3 / R1: 405 GB short
- DeepSeek V3.2: 405 GB short
- GLM-5.3: 447 GB short
- GLM-5.2: 447 GB short
- DeepSeek V4.1 Flash: 453 GB short
- Hy4 Preview 770B-A49B: 464 GB short
- MiMo V2.6 Pro: 615 GB short
- DeepSeek V4 Pro: 972 GB short
- Qwen3.8 2.4T-A95B: 1,497 GB short
- Kimi K3: 1,703 GB short
Best models for M5 Mac (32 GB) with room for long context
The biggest models that still fit at Q4_K_M with 32,768 tokens of context (or their whole window, if shorter), the memory left over, the longest context the Mac takes at that precision and the writing speed with that context in the cache, from 153 GB/s of bandwidth.
| Model | Context | Memory | Left over | Longest context | Tokens/s |
|---|---|---|---|---|---|
| Nemotron 3 Nano 30B-A3B 31.6B | 32K | 20.3 GB | 1.02 GB | 112K | 19–32 |
| Muse Glimmer 30B 29.8B | 32K | 19.5 GB | 1.79 GB | 124K | 4.5–6.2 |
| Qwen3.6 27B 27.8B | 32K | 19.9 GB | 1.38 GB | 44K | 4.4–6.0 |
| Qwen3.8 27B 27.8B | 32K | 19.9 GB | 1.38 GB | 44K | 4.4–6.0 |
| Hemmingway-1 27B 27.3B | 32K | 19.6 GB | 1.67 GB | 48K | 4.5–6.1 |
llama-server commands for an M5 Mac (32 GB)
GGUF repos checked 2026-09-29; -c is the longest context with 0.5 GB free on the M5 Mac (32 GB).
Muse Glimmer 30B, Q4_K_M with 124K tokens: 4.2–5.8 tokens/s on the M5 Mac (32 GB)
llama-server -hf bartowski/Muse-Glimmer-30B-GGUF:Q4_K_M -c 126976 -ngl 99 -np 1 Muse-Glimmer-30B-Q4_K_M.gguf, 17.3 GB: 20.8 GB used and 520 MB free of the M5 Mac (32 GB)'s 21.3 GB usable, 4.2–5.8 tokens/s once the 124K cache is full. -np 1: one slot, one sliding window.
Qwen3.6 27B, Q4_K_M with 10K tokens: 4.8–6.5 tokens/s on the M5 Mac (32 GB)
llama-server -hf ggml-org/Qwen3.6-27B-GGUF:Q4_K_M -c 10240 -ngl 99 Qwen3.6-27B-Q4_K_M.gguf, 19.1 GB: 20.8 GB used and 563 MB free of the M5 Mac (32 GB)'s 21.3 GB usable, 4.8–6.5 tokens/s once the 10K cache is full.
Qwen3.8 27B, Q4_K_M with 12K tokens: 4.7–6.5 tokens/s on the M5 Mac (32 GB)
llama-server -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M -c 12288 -ngl 99 Qwen3.8-27B-Q4_K_M.gguf, 19.0 GB: 20.8 GB used and 550 MB free of the M5 Mac (32 GB)'s 21.3 GB usable, 4.7–6.5 tokens/s once the 12K cache is full.
How fast the M5 Mac (32 GB) writes as the context fills
Tokens per second at Q4_K_M (gpt-oss-20b at MXFP4) for the biggest models that fit with 32K tokens, with 8K and 32K tokens in the cache and at the longest context the Mac holds, and at Q8_0 with 8K where that fits. Each token reads the active weights and the whole cache once, so the speed follows the 153 GB/s of bandwidth and falls as the cache grows. The ceiling is that bandwidth divided by the bytes read per token, which no runtime reaches; the reply time is for 1,000 new tokens with 32K (or the longest context) in the cache, prompt processing aside. With the same memory, M4 Mac (32 GB), 120 GB/s, writes 21% slower on average; M1 Pro or M2 Pro Mac (32 GB), 200 GB/s, writes 30% faster on average; M1 Max or M2 Max Mac (32 GB), 400 GB/s, writes 156% faster on average.
| Model | 8K | 32K | Longest | At the longest | Q8_0 at 8K | Bandwidth ceiling at 8K | 1,000-token reply at 32K |
|---|---|---|---|---|---|---|---|
| Nemotron 3 Nano 30B-A3B | 21–35 | 19–32 | 112K | 16–27 | Does not fit | 71 | 31–52 s |
| Muse Glimmer 30B | 4.6–6.3 | 4.5–6.2 | 124K | 4.2–5.8 | Does not fit | 8 | 162–222 s |
| Qwen3.6 27B | 4.8–6.6 | 4.4–6.0 | 44K | 4.2–5.8 | Does not fit | 9 | 166–227 s |
| Qwen3.8 27B | 4.8–6.6 | 4.4–6.0 | 44K | 4.2–5.8 | Does not fit | 9 | 166–227 s |
| Hemmingway-1 27B | 4.9–6.7 | 4.5–6.1 | 48K | 4.2–5.8 | Does not fit | 9 | 163–223 s |
| Gemma 4 26B-A4B | 15–26 | 13–22 | 185K | 6.9–11 | Does not fit | 53 | 45–76 s |
| gpt-oss-20b | 17–29 | 14–24 | 128K | 8.1–14 | MXFP4 only | 59 | 42–71 s |
| Gemma 4 12B | 11–14 | 10–14 | 256K | 6.9–9.5 | 6.2–8.6 | 19 | 73–100 s |
| ZDTaichu 5.0 9B | 13–18 | 12–16 | 128K | 8.1–11 | 7.8–11 | 25 | 61–85 s |
| Ornith 1.5 9B | 13–19 | 12–16 | 256K | 5.8–7.9 | 7.9–11 | 25 | 61–84 s |
M5 Mac (32 GB) against the other Macs with 21.3 GB usable
The same models fit on every Mac with 21.3 GB usable, so speed is what separates them. By bandwidth the M5 Mac (32 GB) ranks 3 of 4 at 153 GB/s. Over the 12 biggest models that fit at Q4_K_M with 8K context, the M4 Mac (32 GB), at 120 GB/s, writes 21% slower, the M1 Pro or M2 Pro Mac (32 GB), at 200 GB/s, writes 30% faster and the M1 Max or M2 Max Mac (32 GB), at 400 GB/s, writes 156% faster than the M5 Mac (32 GB). Each cell below is the other Mac's tokens per second minus this one's, midpoints of the estimated ranges.
| Model | M5 Mac (32 GB) tokens/s | M4 Mac (32 GB) | M1 Pro or M2 Pro Mac (32 GB) | M1 Max or M2 Max Mac (32 GB) |
|---|---|---|---|---|
| Nemotron 3 Nano 30B-A3B | 21–35 | −5.8 | +8.2 | +42 |
| GLM-4.7 Flash | 20–33 | −5.6 | +7.9 | +40 |
| Xing 4.0 29B-A4B | 16–27 | −4.6 | +6.5 | +33 |
| Qwen3-Coder 30B-A3B | 16–27 | −4.6 | +6.4 | +33 |
| Qwen3 30B-A3B | 16–27 | −4.6 | +6.4 | +33 |
| Muse Glimmer 30B | 4.6–6.3 | −1.2 | +1.7 | +8.7 |
| Qwen3.6 27B | 4.8–6.6 | −1.2 | +1.7 | +9.1 |
| Qwen3.8 27B | 4.8–6.6 | −1.2 | +1.7 | +9.1 |
| Hemmingway-1 27B | 4.9–6.7 | −1.2 | +1.8 | +9.2 |
| Gemma 4 26B-A4B | 15–26 | −4.4 | +6.2 | +32 |
| gpt-oss-20b | 17–29 | −4.9 | +7.0 | +36 |
| Gemma 4 12B | 11–14 | −2.7 | +3.8 | +20 |
Near misses on M5 Mac (32 GB)
Models that miss at Q4_K_M with 32,768 tokens of context (or their whole window), or leave under 0.5 GB free there, but fit with a shorter context or 3-bit weights, or are short by at most a quarter of the memory and fit with MoE offload. Smallest shortfall first.
| Model | Short by at Q4_K_M | Fits instead |
|---|---|---|
| Xing 4.0 29B-A4B 31.2B | 96 MB | Q4_K_M with 19K context |
| GLM-4.7 Flash 31.2B | 377 MB | Q4_K_M with 16K context |
| LLM-jp-4.1 32B-A3B Thinking 32.1B | 1.32 GB | GGUF Q3_K_M with 32K, 18.8 GB |
| Qwen3-Coder 30B-A3B 30.5B | 1.42 GB | Q4_K_M with 13K context |
| Qwen3 30B-A3B 30.5B | 1.42 GB | Q4_K_M with 13K context |
How these numbers are worked out
Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.
Nearby GPUs and setups
How many of the tracked models each one holds at Q4_K_M with 8K context, against 28 here.
| Setup | How it relates | Memory | Bandwidth | Models at Q4_K_M |
|---|---|---|---|---|
| M4 Mac (32 GB) | Same memory | 21.3 GB usable | 120 GB/s | 28 (same) |
| M1 Pro or M2 Pro Mac (32 GB) | Same memory | 21.3 GB usable | 200 GB/s | 28 (same) |
| M1 Max or M2 Max Mac (32 GB) | Same memory | 21.3 GB usable | 400 GB/s | 28 (same) |
| M5 Pro Mac (24 GB) | Next size down | 16 GB usable | 307 GB/s | 18 (−10) |
| M3 Pro Mac (36 GB) | Next size up | 27 GB usable | 150 GB/s | 35 (+7) |
Model numbers read from Hugging Face on .