What LLMs can an M3 Pro Mac (18 GB) run?
An M3 Pro Mac (18 GB) gives a model 12 GB usable of memory (about two thirds of unified memory is usable by the GPU) and 150 GB/s of bandwidth. Of the 61 open models tracked here, 17 fit at Q4_K_M with an 8,192-token context, each leaving at least 0.5 GB free on the M3 Pro Mac (18 GB).
Models that fit on one M3 Pro Mac (18 GB)
The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first.
| Model | Best precision | Memory | Longest context | Tokens/s |
|---|---|---|---|---|
| Gemma 4 12B 12.0B | GGUF Q6_K | 11.2 GB | 25K | 7.8–11 |
| ZDTaichu 5.0 9B 9.8B | GGUF Q8_0 | 11.4 GB | 9K | 7.6–10 |
| Ornith 1.5 9B 9.7B | GGUF Q8_0 | 11.3 GB | 14K | 7.7–11 |
| Qwen3.5 9B 9.7B | GGUF Q8_0 | 11.3 GB | 14K | 7.7–11 |
| Ornith 1.0 9B 9.4B | GGUF Q8_0 | 11.0 GB | 21K | 7.9–11 |
| MiMo V2.6 Distill Qwen 9B 9.4B | GGUF Q8_0 | 11.0 GB | 21K | 7.9–11 |
| Granite 4.2 8B 8.8B | GGUF Q8_0 | 11.4 GB | 8K | 7.6–10 |
| LFM2.5 8B-A1B 8.5B, 1.5B active | GGUF Q8_0 | 9.82 GB | 128,000 (full) | 26–43 |
| Qwen3 8B 8.2B | GGUF Q8_0 | 10.7 GB | 13K | 8.2–11 |
| Llama 3.1 8B 8.0B | GGUF Q8_0 | 10.3 GB | 16K | 8.5–12 |
| Gemma 4 E4B 8.0B | GGUF Q8_0 | 9.38 GB | 128K (full) | 9.4–13 |
| Ling 3.0 Tiny 7.9B-A1.3B 7.9B, 1.3B active | GGUF Q8_0 | 9.15 GB | 128K (full) | 30–51 |
| Spark-X2.5 4B 4.1B | FP16 / BF16 | 9.35 GB | 63K | 9.4–13 |
| Nemotron 3 Nano 4B 4.0B | FP16 / BF16 | 8.78 GB | 166K | 10–14 |
| Granite 4.2 3B 3.7B | FP16 / BF16 | 8.69 GB | 40K | 10–14 |
| MiniCPM5 2B 2.5B | FP16 / BF16 | 6.02 GB | 128K (full) | 15–21 |
| Limite 1B Violetto 1.0B | FP16 / BF16 | 2.79 GB | 128K (full) | 35–49 |
Raising the GPU memory limit on an M3 Pro Mac (18 GB)
macOS lets the GPU wire about 12 GB of this Mac's 18 GB by default. Running
sudo sysctl iogpu.wired_limit_mb=14336 raises that to 14 GB, leaving
4 GB to macOS and the apps next to it, until the next restart
(macOS 14 or newer; iogpu.wired_limit_mb=0 restores the default). LM Studio and Ollama pick the new limit up after a restart of the app.
At 14 GB, 17 of the 61 models fit at Q4_K_M with 8K context, against 17 by default
. Gemma 4 12B, the biggest model that fits already, goes from 178K to 256K of context. If macOS runs short of memory it swaps or kills apps, so close big apps first and raise the limit in steps.
Too big for M3 Pro Mac (18 GB) at 8K context
How much memory each of these lacks at Q4_K_M (gpt-oss-20b and gpt-oss-120b at MXFP4) with 8K context, one request and 0.5 GB kept free, out of the 12 GB usable here; a bigger machine or a smaller quant is the way in.
- gpt-oss-20b: 3.32 GB short
- Gemma 4 26B-A4B: 5.49 GB short
- Hemmingway-1 27B: 6.48 GB short
- Qwen3.6 27B: 6.77 GB short
- Qwen3.8 27B: 6.77 GB short
- Granite 4.2 30B: 9.35 GB short
- Muse Glimmer 30B: 7.67 GB short
- Qwen3-Coder 30B-A3B: 8.75 GB short
- Qwen3 30B-A3B: 8.75 GB short
- Xing 4.0 29B-A4B: 8.73 GB short
- GLM-4.7 Flash: 8.81 GB short
- Gemma 4 31B: 10.4 GB short
- Nemotron 3 Nano 30B-A3B: 8.62 GB short
- LLM-jp-4.1 32B-A3B Thinking: 9.47 GB short
- Ornith 1.0 35B: 10.9 GB short
- Ornith 1.5 35B-A3B: 11.5 GB short
- Qwen3.6 35B-A3B: 11.5 GB short
- K2-Horizon MoVA 36B-A4B: 13.9 GB short
- Llama 3.1 70B: 35.5 GB short
- Qwen3-Coder-Next: 38.6 GB short
- AliceAI Foundation 80B-A3B: 39.6 GB short
- gpt-oss-120b: 56.2 GB short
- Nemotron 3 Super 120B-A12B: 65.7 GB short
- Qwen3.5 122B-A10B: 66.7 GB short
- Mistral Medium 3.5 128B: 71.2 GB short
- Qwen3.8 Flash Next: 101 GB short
- Step 3.7 Flash 196B-A11B: 114 GB short
- MiniMax M2.7: 133 GB short
- DeepSeek V4 Flash: 170 GB short
- DeepSeek V4 Flash 0731: 178 GB short
- MiMo V2.6 Flash: 182 GB short
- IQuest-Q1 320B-A15B: 190 GB short
- GLM-5.3 Flash: 188 GB short
- MiniMax M3: 255 GB short
- DeepSeek V3 / R1: 414 GB short
- DeepSeek V3.2: 414 GB short
- GLM-5.3: 457 GB short
- GLM-5.2: 457 GB short
- DeepSeek V4.1 Flash: 463 GB short
- Hy4 Preview 770B-A49B: 473 GB short
- MiMo V2.6 Pro: 624 GB short
- DeepSeek V4 Pro: 981 GB short
- Qwen3.8 2.4T-A95B: 1,506 GB short
- Kimi K3: 1,712 GB short
Best models for M3 Pro Mac (18 GB) with room for long context
The biggest models that still fit at Q4_K_M with 32,768 tokens of context (or their whole window, if shorter), the memory left over, the longest context the Mac takes at that precision and the writing speed with that context in the cache, from 150 GB/s of bandwidth.
| Model | Context | Memory | Left over | Longest context | Tokens/s |
|---|---|---|---|---|---|
| Gemma 4 12B 12.0B | 32K | 8.98 GB | 3.02 GB | 178K | 9.8–14 |
| ZDTaichu 5.0 9B 9.8B | 32K | 7.67 GB | 4.33 GB | 128K (full) | 12–16 |
| Ornith 1.5 9B 9.7B | 32K | 7.58 GB | 4.42 GB | 145K | 12–16 |
| Qwen3.5 9B 9.7B | 32K | 7.58 GB | 4.42 GB | 145K | 12–16 |
| Ornith 1.0 9B 9.4B | 32K | 7.43 GB | 4.57 GB | 150K | 12–16 |
llama-server commands for an M3 Pro Mac (18 GB)
GGUF repos checked 2026-09-29; -c is the longest context with 0.5 GB free on the M3 Pro Mac (18 GB). No -ngl: llama.cpp’s -fit, on by default since b7440, places the layers.
Gemma 4 12B, Q4_K_M with 178K tokens: 7.6–10 tokens/s on the M3 Pro Mac (18 GB)
llama-server -hf unsloth/gemma-4-12b-it-GGUF:Q4_K_M -c 182272 -np 1 gemma-4-12b-it-Q4_K_M.gguf, 7.1 GB: 11.5 GB used and 525 MB free of the M3 Pro Mac (18 GB)'s 12 GB usable, 7.6–10 tokens/s once the 178K cache is full. -np 1: one slot, one sliding window.
Ornith 1.5 9B, Q4_K_M with 143K tokens: 7.7–11 tokens/s on the M3 Pro Mac (18 GB)
llama-server -hf bartowski/Ornith-1.5-9B-GGUF:Q4_K_M -c 146432 Ornith-1.5-9B-Q4_K_M.gguf, 5.9 GB: 11.5 GB used and 542 MB free of the M3 Pro Mac (18 GB)'s 12 GB usable, 7.7–11 tokens/s once the 143K cache is full.
Qwen3.5 9B, Q4_K_M with 145K tokens: 7.6–10 tokens/s on the M3 Pro Mac (18 GB)
llama-server -hf unsloth/Qwen3.5-9B-GGUF:Q4_K_M -c 148480 Qwen3.5-9B-Q4_K_M.gguf, 5.7 GB: 11.5 GB used and 545 MB free of the M3 Pro Mac (18 GB)'s 12 GB usable, 7.6–10 tokens/s once the 145K cache is full.
How fast the M3 Pro Mac (18 GB) writes as the context fills
Tokens per second at Q4_K_M for the biggest models that fit with 32K tokens, with 8K and 32K tokens in the cache and at the longest context the Mac holds, and at Q8_0 with 8K where that fits. Each token reads the active weights and the whole cache once, so the speed follows the 150 GB/s of bandwidth and falls as the cache grows. The ceiling is that bandwidth divided by the bytes read per token, which no runtime reaches; the reply time is for 1,000 new tokens with 32K (or the longest context) in the cache, prompt processing aside.
| Model | 8K | 32K | Longest | At the longest | Q8_0 at 8K | Bandwidth ceiling at 8K | 1,000-token reply at 32K |
|---|---|---|---|---|---|---|---|
| Gemma 4 12B | 10–14 | 9.8–14 | 178K | 7.6–10 | Does not fit | 19 | 74–102 s |
| ZDTaichu 5.0 9B | 13–18 | 12–16 | 128K | 8.0–11 | 7.6–10 | 24 | 63–86 s |
| Ornith 1.5 9B | 13–18 | 12–16 | 145K | 7.6–10 | 7.7–11 | 25 | 62–85 s |
| Qwen3.5 9B | 13–18 | 12–16 | 145K | 7.6–10 | 7.7–11 | 25 | 62–85 s |
| Ornith 1.0 9B | 14–19 | 12–16 | 150K | 7.6–10 | 7.9–11 | 25 | 61–84 s |
| MiMo V2.6 Distill Qwen 9B | 14–19 | 12–16 | 150K | 7.6–10 | 7.9–11 | 25 | 61–84 s |
| Granite 4.2 8B | 12–17 | 7.6–10 | 32K | 7.6–10 | 7.6–10 | 23 | 96–131 s |
| LFM2.5 8B-A1B | 42–72 | 33–56 | 128,000 | 18–30 | 26–43 | 149 | 18–31 s |
| Qwen3 8B | 13–18 | 8.3–11 | 38K | 7.6–10 | 8.2–11 | 24 | 87–120 s |
| Llama 3.1 8B | 14–19 | 8.9–12 | 43K | 7.7–11 | 8.5–12 | 25 | 82–112 s |
How these numbers are worked out
Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.
Nearby GPUs and setups
How many of the tracked models each one holds at Q4_K_M with 8K context, against 17 here.
| Setup | How it relates | Memory | Bandwidth | Models at Q4_K_M |
|---|---|---|---|---|
| M1 Pro or M2 Pro Mac (16 GB) | Next size down | 10.7 GB usable | 200 GB/s | 17 (same) |
| M2 or M3 Mac (24 GB) | Next size up | 16 GB usable | 100 GB/s | 18 (+1) |
Model numbers read from Hugging Face on .