What LLMs can an M1 Pro or M2 Pro Mac (16 GB) run?

An M1 Pro or M2 Pro Mac (16 GB) gives a model 10.7 GB usable of memory (about two thirds of unified memory is usable by the GPU) and 200 GB/s of bandwidth. Of the 61 open models tracked here, 17 fit at Q4_K_M with an 8,192-token context, each leaving at least 0.5 GB free on the M1 Pro or M2 Pro Mac (16 GB).

Models that fit on one M1 Pro or M2 Pro Mac (16 GB)

The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first.

Model Best precision Memory Longest context Tokens/s
Gemma 4 12B 12.0B GGUF Q5_K_M 9.84 GB 28K 12–16
ZDTaichu 5.0 9B 9.8B GGUF Q6_K 9.00 GB 42K 13–18
Ornith 1.5 9B 9.7B GGUF Q6_K 8.88 GB 46K 13–18
Qwen3.5 9B 9.7B GGUF Q6_K 8.88 GB 46K 13–18
Ornith 1.0 9B 9.4B GGUF Q6_K 8.68 GB 52K 13–19
MiMo V2.6 Distill Qwen 9B 9.4B GGUF Q6_K 8.68 GB 52K 13–19
Granite 4.2 8B 8.8B GGUF Q6_K 9.26 GB 13K 13–17
LFM2.5 8B-A1B 8.5B, 1.5B active GGUF Q8_0 9.82 GB 37K 34–57
Qwen3 8B 8.2B FP8 / INT8 10.1 GB 8K 12–16
Llama 3.1 8B 8.0B FP8 / INT8 9.83 GB 10K 12–16
Gemma 4 E4B 8.0B GGUF Q8_0 9.38 GB 55K 12–17
Ling 3.0 Tiny 7.9B-A1.3B 7.9B, 1.3B active GGUF Q8_0 9.15 GB 128K (full) 39–67
Spark-X2.5 4B 4.1B FP16 / BF16 9.35 GB 29K 12–17
Nemotron 3 Nano 4B 4.0B FP16 / BF16 8.78 GB 90K 13–18
Granite 4.2 3B 3.7B FP16 / BF16 8.69 GB 25K 13–19
MiniCPM5 2B 2.5B FP16 / BF16 6.02 GB 100K 20–27
Limite 1B Violetto 1.0B FP16 / BF16 2.79 GB 128K (full) 46–65

Raising the GPU memory limit on an M1 Pro or M2 Pro Mac (16 GB)

macOS lets the GPU wire about 10.7 GB of this Mac's 16 GB by default. Running sudo sysctl iogpu.wired_limit_mb=12288 raises that to 12 GB, leaving 4 GB to macOS and the apps next to it, until the next restart (macOS 14 or newer; iogpu.wired_limit_mb=0 restores the default). LM Studio and Ollama pick the new limit up after a restart of the app. At 12 GB, 17 of the 61 models fit at Q4_K_M with 8K context, against 17 by default . Gemma 4 12B, the biggest model that fits already, goes from 102K to 178K of context. If macOS runs short of memory it swaps or kills apps, so close big apps first and raise the limit in steps.

Too big for M1 Pro or M2 Pro Mac (16 GB) at 8K context

How much memory each of these lacks at Q4_K_M (gpt-oss-20b and gpt-oss-120b at MXFP4) with 8K context, one request and 0.5 GB kept free, out of the 10.7 GB usable here; a bigger machine or a smaller quant is the way in.

Best models for M1 Pro or M2 Pro Mac (16 GB) with room for long context

The biggest models that still fit at Q4_K_M with 32,768 tokens of context (or their whole window, if shorter), the memory left over, the longest context the Mac takes at that precision and the writing speed with that context in the cache, from 200 GB/s of bandwidth.

ModelContextMemoryLeft overLongest contextTokens/s
Gemma 4 12B 12.0B 32K 8.98 GB 1.72 GB 102K 13–18
ZDTaichu 5.0 9B 9.8B 32K 7.67 GB 3.03 GB 105K 15–21
Ornith 1.5 9B 9.7B 32K 7.58 GB 3.12 GB 108K 16–21
Qwen3.5 9B 9.7B 32K 7.58 GB 3.12 GB 108K 16–21
Ornith 1.0 9B 9.4B 32K 7.43 GB 3.27 GB 112K 16–22

llama-server commands for an M1 Pro or M2 Pro Mac (16 GB)

GGUF repos checked 2026-09-29; -c is the longest context with 0.5 GB free on the M1 Pro or M2 Pro Mac (16 GB).

Gemma 4 12B, Q4_K_M with 102K tokens: 11–16 tokens/s on the M1 Pro or M2 Pro Mac (16 GB)

llama-server -hf unsloth/gemma-4-12b-it-GGUF:Q4_K_M -c 104448 -ngl 99 -np 1

gemma-4-12b-it-Q4_K_M.gguf, 7.1 GB: 10.2 GB used and 531 MB free of the M1 Pro or M2 Pro Mac (16 GB)'s 10.7 GB usable, 11–16 tokens/s once the 102K cache is full. -np 1: one slot, one sliding window.

Ornith 1.5 9B, Q4_K_M with 105K tokens: 12–16 tokens/s on the M1 Pro or M2 Pro Mac (16 GB)

llama-server -hf bartowski/Ornith-1.5-9B-GGUF:Q4_K_M -c 107520 -ngl 99

Ornith-1.5-9B-Q4_K_M.gguf, 5.9 GB: 10.2 GB used and 548 MB free of the M1 Pro or M2 Pro Mac (16 GB)'s 10.7 GB usable, 12–16 tokens/s once the 105K cache is full.

Qwen3.5 9B, Q4_K_M with 108K tokens: 11–16 tokens/s on the M1 Pro or M2 Pro Mac (16 GB)

llama-server -hf unsloth/Qwen3.5-9B-GGUF:Q4_K_M -c 110592 -ngl 99

Qwen3.5-9B-Q4_K_M.gguf, 5.7 GB: 10.2 GB used and 517 MB free of the M1 Pro or M2 Pro Mac (16 GB)'s 10.7 GB usable, 11–16 tokens/s once the 108K cache is full.

How fast the M1 Pro or M2 Pro Mac (16 GB) writes as the context fills

Tokens per second at Q4_K_M for the biggest models that fit with 32K tokens, with 8K and 32K tokens in the cache and at the longest context the Mac holds, and at Q8_0 with 8K where that fits. Each token reads the active weights and the whole cache once, so the speed follows the 200 GB/s of bandwidth and falls as the cache grows. The ceiling is that bandwidth divided by the bytes read per token, which no runtime reaches; the reply time is for 1,000 new tokens with 32K (or the longest context) in the cache, prompt processing aside. With the same memory, M1 Mac (16 GB), 68.25 GB/s, writes 65% slower on average; M2 or M3 Mac (16 GB), 100 GB/s, writes 49% slower on average; M4 Mac (16 GB), 120 GB/s, writes 39% slower on average; M5 Mac (16 GB), 153 GB/s, writes 23% slower on average.

Model8K32KLongestAt the longestQ8_0 at 8KBandwidth ceiling at 8K1,000-token reply at 32K
Gemma 4 12B 14–19 13–18 102K 11–16 Does not fit 25 56–77 s
ZDTaichu 5.0 9B 17–24 15–21 105K 11–16 Does not fit 32 47–65 s
Ornith 1.5 9B 18–24 16–21 108K 11–16 Does not fit 33 47–64 s
Qwen3.5 9B 18–24 16–21 108K 11–16 Does not fit 33 47–64 s
Ornith 1.0 9B 18–25 16–22 112K 11–16 Does not fit 34 46–63 s
MiMo V2.6 Distill Qwen 9B 18–25 16–22 112K 11–16 Does not fit 34 46–63 s
LFM2.5 8B-A1B 55–95 43–74 128,000 23–40 34–57 198 14–23 s
Llama 3.1 8B 18–25 12–16 34K 11–16 Does not fit 34 62–85 s
Gemma 4 E4B 21–29 20–27 128K 15–21 12–17 40 37–51 s
Ling 3.0 Tiny 7.9B-A1.3B 64–112 54–94 128K 34–57 39–67 237 11–18 s

M1 Pro or M2 Pro Mac (16 GB) against the other Macs with 10.7 GB usable

The same models fit on every Mac with 10.7 GB usable, so speed is what separates them. By bandwidth the M1 Pro or M2 Pro Mac (16 GB) is the fastest of the 5 Macs with 10.7 GB usable at 200 GB/s. Over the 12 biggest models that fit at Q4_K_M with 8K context, the M1 Mac (16 GB), at 68.25 GB/s, writes 65% slower, the M2 or M3 Mac (16 GB), at 100 GB/s, writes 49% slower, the M4 Mac (16 GB), at 120 GB/s, writes 39% slower and the M5 Mac (16 GB), at 153 GB/s, writes 23% slower than the M1 Pro or M2 Pro Mac (16 GB). Each cell below is the other Mac's tokens per second minus this one's, midpoints of the estimated ranges.

Model M1 Pro or M2 Pro Mac (16 GB) tokens/s M1 Mac (16 GB)M2 or M3 Mac (16 GB)M4 Mac (16 GB)M5 Mac (16 GB)
Gemma 4 12B 14–19 −11−8.1−6.5−3.8
ZDTaichu 5.0 9B 17–24 −13−10−8.2−4.8
Ornith 1.5 9B 18–24 −14−10−8.3−4.8
Qwen3.5 9B 18–24 −14−10−8.3−4.8
Ornith 1.0 9B 18–25 −14−11−8.5−5.0
MiMo V2.6 Distill Qwen 9B 18–25 −14−11−8.5−5.0
Granite 4.2 8B 16–22 −13−9.5−7.6−4.5
LFM2.5 8B-A1B 55–95 −48−36−29−17
Qwen3 8B 17–24 −14−10−8.2−4.8
Llama 3.1 8B 18–25 −14−11−8.5−5.0
Gemma 4 E4B 21–29 −17−13−10−5.9
Ling 3.0 Tiny 7.9B-A1.3B 64–112 −57−42−34−20

Near misses on M1 Pro or M2 Pro Mac (16 GB)

Models that miss at Q4_K_M with 32,768 tokens of context (or their whole window), or leave under 0.5 GB free there, but fit with a shorter context or 3-bit weights, or are short by at most a quarter of the memory and fit with MoE offload. Smallest shortfall first.

ModelShort by at Q4_K_MFits instead
Qwen3 8B 8.2B Tight, 178 MB free Q4_K_M with 29K context
Granite 4.2 8B 8.8B 767 MB Q4_K_M with 24K context

How these numbers are worked out

Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.

Nearby GPUs and setups

How many of the tracked models each one holds at Q4_K_M with 8K context, against 17 here.

SetupHow it relatesMemoryBandwidthModels at Q4_K_M
M1 Mac (16 GB) Same memory 10.7 GB usable 68.25 GB/s 17 (same)
M2 or M3 Mac (16 GB) Same memory 10.7 GB usable 100 GB/s 17 (same)
M4 Mac (16 GB) Same memory 10.7 GB usable 120 GB/s 17 (same)
M5 Mac (16 GB) Same memory 10.7 GB usable 153 GB/s 17 (same)
M2 or M3 Mac (8 GB) Next size down 5.3 GB usable 100 GB/s 5 (−12)
M3 Pro Mac (18 GB) Next size up 12 GB usable 150 GB/s 17 (same)

Model numbers read from Hugging Face on .