What LLMs can an M5 Pro Mac (24 GB) run?

An M5 Pro Mac (24 GB) gives a model 16 GB usable of memory (about two thirds of unified memory is usable by the GPU) and 307 GB/s of bandwidth. Of the 61 open models tracked here, 18 fit at Q4_K_M with an 8,192-token context, and 12 more at a lower precision, each leaving at least 0.5 GB free on the M5 Pro Mac (24 GB).

Models that fit on one M5 Pro Mac (24 GB)

The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first. gpt-oss-20b (14.8 GB, 34–58 tokens/s on the M5 Pro Mac (24 GB)) is listed at MXFP4 only, as published.

Model Best precision Memory Longest context Tokens/s
LLM-jp-4.1 32B-A3B Thinking 32.1B, 3.8B active GGUF Q2_K 14.8 GB 17K 40–69
Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active GGUF Q2_K 14.1 GB 225K 56–96
GLM-4.7 Flash 31.2B, 3B active GGUF Q2_K 14.3 GB 28K 50–86
Xing 4.0 29B-A4B 31.2B, 4B active GGUF Q2_K 14.3 GB 33K 42–72
Qwen3-Coder 30B-A3B 30.5B, 3.3B active GGUF Q2_K 14.4 GB 18K 40–68
Qwen3 30B-A3B 30.5B, 3.3B active GGUF Q2_K 14.4 GB 18K 40–68
Muse Glimmer 30B 29.8B GGUF Q2_K 13.5 GB 128K (full) 13–18
Granite 4.2 30B 29.3B GGUF Q2_K 15.3 GB 8K 12–16
Qwen3.6 27B 27.8B GGUF Q3_K_M 15.0 GB 15K 12–16
Qwen3.8 27B 27.8B GGUF Q3_K_M 15.0 GB 15K 12–16
Hemmingway-1 27B 27.3B GGUF Q3_K_M 14.7 GB 19K 12–16
Gemma 4 26B-A4B 25.8B, 4B active GGUF Q3_K_M 13.9 GB 81K 36–61
gpt-oss-20b 20.9B, 3.6B active As published (MXFP4) 14.8 GB 34K 34–58
Gemma 4 12B 12.0B GGUF Q8_0 14.2 GB 85K 12–17
ZDTaichu 5.0 9B 9.8B GGUF Q8_0 11.4 GB 126K 15–21
Ornith 1.5 9B 9.7B GGUF Q8_0 11.3 GB 130K 16–22
Qwen3.5 9B 9.7B GGUF Q8_0 11.3 GB 130K 16–22
Ornith 1.0 9B 9.4B GGUF Q8_0 11.0 GB 138K 16–22
MiMo V2.6 Distill Qwen 9B 9.4B GGUF Q8_0 11.0 GB 138K 16–22
Granite 4.2 8B 8.8B GGUF Q8_0 11.4 GB 31K 15–21
LFM2.5 8B-A1B 8.5B, 1.5B active GGUF Q8_0 9.82 GB 128,000 (full) 50–87
Qwen3 8B 8.2B GGUF Q8_0 10.7 GB 39K 17–23
Llama 3.1 8B 8.0B GGUF Q8_0 10.3 GB 45K 17–24
Gemma 4 E4B 8.0B GGUF Q8_0 9.38 GB 128K (full) 19–26
Ling 3.0 Tiny 7.9B-A1.3B 7.9B, 1.3B active GGUF Q8_0 9.15 GB 128K (full) 58–101
Spark-X2.5 4B 4.1B FP16 / BF16 9.35 GB 166K 19–26
Nemotron 3 Nano 4B 4.0B FP16 / BF16 8.78 GB 256K (full) 20–28
Granite 4.2 3B 3.7B FP16 / BF16 8.69 GB 87K 20–28
MiniCPM5 2B 2.5B FP16 / BF16 6.02 GB 128K (full) 30–42
Limite 1B Violetto 1.0B FP16 / BF16 2.79 GB 128K (full) 68–98

Raising the GPU memory limit on an M5 Pro Mac (24 GB)

macOS lets the GPU wire about 16 GB of this Mac's 24 GB by default. Running sudo sysctl iogpu.wired_limit_mb=20480 raises that to 20 GB, leaving 4 GB to macOS and the apps next to it, until the next restart (macOS 14 or newer; iogpu.wired_limit_mb=0 restores the default). LM Studio and Ollama pick the new limit up after a restart of the app. At 20 GB, 23 of the 61 models fit at Q4_K_M with 8K context, against 18 by default : Gemma 4 26B-A4B (30–52 tokens/s), Hemmingway-1 27B (9.7–13 tokens/s), Qwen3.6 27B (9.6–13 tokens/s), Qwen3.8 27B (9.6–13 tokens/s) and Muse Glimmer 30B (9.1–13 tokens/s) join the list. gpt-oss-20b, the biggest model that fits already, goes from 34K to 128K of context. If macOS runs short of memory it swaps or kills apps, so close big apps first and raise the limit in steps.

Too big for M5 Pro Mac (24 GB) at 8K context

How much memory each of these lacks at Q4_K_M (gpt-oss-120b at MXFP4) with 8K context, one request and 0.5 GB kept free, out of the 16 GB usable here; a bigger machine or a smaller quant is the way in.

Best models for M5 Pro Mac (24 GB) with room for long context

The biggest models that still fit at Q4_K_M with 32,768 tokens of context (or their whole window, if shorter), the memory left over, the longest context the Mac takes at that precision and the writing speed with that context in the cache, from 307 GB/s of bandwidth. gpt-oss-20b is at MXFP4, as published.

ModelContextMemoryLeft overLongest contextTokens/s
gpt-oss-20b 20.9B, MXFP4 32K 15.4 GB 571 MB 34K 28–47
Gemma 4 12B 12.0B 32K 8.98 GB 7.02 GB 256K (full) 20–27
ZDTaichu 5.0 9B 9.8B 32K 7.67 GB 8.33 GB 128K (full) 23–32
Ornith 1.5 9B 9.7B 32K 7.58 GB 8.42 GB 256K (full) 24–33
Qwen3.5 9B 9.7B 32K 7.58 GB 8.42 GB 256K (full) 24–33

llama-server commands for an M5 Pro Mac (24 GB)

GGUF repos checked 2026-09-29; -c is the longest context with 0.5 GB free on the M5 Pro Mac (24 GB). No -ngl: llama.cpp’s -fit, on by default since b7440, places the layers.

gpt-oss-20b, MXFP4 with 34K tokens: 27–46 tokens/s on the M5 Pro Mac (24 GB)

llama-server -hf ggml-org/gpt-oss-20b-GGUF:MXFP4 -c 34816 -np 1

gpt-oss-20b-MXFP4.gguf, 12.1 GB: 15.5 GB used and 518 MB free of the M5 Pro Mac (24 GB)'s 16 GB usable, 27–46 tokens/s once the 34K cache is full. -np 1: one slot, one sliding window.

Gemma 4 12B, Q4_K_M with 256K tokens: 14–19 tokens/s on the M5 Pro Mac (24 GB)

llama-server -hf unsloth/gemma-4-12b-it-GGUF:Q4_K_M -c 262144 -np 1

gemma-4-12b-it-Q4_K_M.gguf, 7.1 GB: 12.8 GB used and 3.17 GB free of the M5 Pro Mac (24 GB)'s 16 GB usable, 14–19 tokens/s once the 256K cache is full. -np 1: one slot, one sliding window.

How fast the M5 Pro Mac (24 GB) writes as the context fills

Tokens per second at Q4_K_M (gpt-oss-20b at MXFP4) for the biggest models that fit with 32K tokens, with 8K and 32K tokens in the cache and at the longest context the Mac holds, and at Q8_0 with 8K where that fits. Each token reads the active weights and the whole cache once, so the speed follows the 307 GB/s of bandwidth and falls as the cache grows. The ceiling is that bandwidth divided by the bytes read per token, which no runtime reaches; the reply time is for 1,000 new tokens with 32K (or the longest context) in the cache, prompt processing aside. With the same memory, M2 or M3 Mac (24 GB), 100 GB/s, writes 67% slower on average; M4 Mac (24 GB), 120 GB/s, writes 60% slower on average; M5 Mac (24 GB), 153 GB/s, writes 49% slower on average; M4 Pro Mac (24 GB), 273 GB/s, writes 11% slower on average.

Model8K32KLongestAt the longestQ8_0 at 8KBandwidth ceiling at 8K1,000-token reply at 32K
gpt-oss-20b 34–58 28–47 34K 27–46 MXFP4 only 119 21–36 s
Gemma 4 12B 21–29 20–27 256K 14–19 12–17 39 36–51 s
ZDTaichu 5.0 9B 26–36 23–32 128K 16–22 15–21 50 31–43 s
Ornith 1.5 9B 27–37 24–33 256K 11–16 16–22 50 31–42 s
Qwen3.5 9B 27–37 24–33 256K 11–16 16–22 50 31–42 s
Ornith 1.0 9B 27–38 24–33 256K 12–16 16–22 51 30–42 s
MiMo V2.6 Distill Qwen 9B 27–38 24–33 256K 12–16 16–22 51 30–42 s
Granite 4.2 8B 24–34 15–21 55K 11–16 15–21 46 47–65 s
LFM2.5 8B-A1B 80–141 64–111 128,000 35–60 50–87 305 9–16 s
Qwen3 8B 26–37 17–23 40K 15–21 17–23 50 43–59 s

M5 Pro Mac (24 GB) against the other Macs with 16 GB usable

The same models fit on every Mac with 16 GB usable, so speed is what separates them. By bandwidth the M5 Pro Mac (24 GB) is the fastest of the 5 Macs with 16 GB usable at 307 GB/s. Over the 12 biggest models that fit at Q4_K_M with 8K context, the M2 or M3 Mac (24 GB), at 100 GB/s, writes 67% slower, the M4 Mac (24 GB), at 120 GB/s, writes 60% slower, the M5 Mac (24 GB), at 153 GB/s, writes 49% slower and the M4 Pro Mac (24 GB), at 273 GB/s, writes 11% slower than the M5 Pro Mac (24 GB). Each cell below is the other Mac's tokens per second minus this one's, midpoints of the estimated ranges.

Model M5 Pro Mac (24 GB) tokens/s M2 or M3 Mac (24 GB)M4 Mac (24 GB)M5 Mac (24 GB)M4 Pro Mac (24 GB)
gpt-oss-20b 34–58 −30−27−22−4.9
Gemma 4 12B 21–29 −17−15−12−2.7
ZDTaichu 5.0 9B 26–36 −21−19−16−3.4
Ornith 1.5 9B 27–37 −21−19−16−3.4
Qwen3.5 9B 27–37 −21−19−16−3.4
Ornith 1.0 9B 27–38 −22−20−16−3.5
MiMo V2.6 Distill Qwen 9B 27–38 −22−20−16−3.5
Granite 4.2 8B 24–34 −20−18−14−3.2
LFM2.5 8B-A1B 80–141 −72−65−53−11
Qwen3 8B 26–37 −21−19−16−3.4
Llama 3.1 8B 27–38 −22−20−16−3.5
Gemma 4 E4B 32–45 −26−23−19−4.1

Near misses on M5 Pro Mac (24 GB)

Models that miss at Q4_K_M with 32,768 tokens of context (or their whole window), or leave under 0.5 GB free there, but fit with a shorter context or 3-bit weights, or are short by at most a quarter of the memory and fit with MoE offload. Smallest shortfall first.

ModelShort by at Q4_K_MFits instead
Gemma 4 26B-A4B 25.8B 1.50 GB GGUF Q3_K_M with 32K, 14.4 GB
Muse Glimmer 30B 29.8B 3.51 GB GGUF IQ3_XXS with 32K, 13.6 GB
Hemmingway-1 27B 27.3B 3.63 GB GGUF IQ3_XXS with 32K, 14.2 GB
Qwen3.6 27B 27.8B 3.92 GB GGUF IQ3_XXS with 32K, 14.4 GB
Qwen3.8 27B 27.8B 3.92 GB GGUF IQ3_XXS with 32K, 14.4 GB

How these numbers are worked out

Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.

Nearby GPUs and setups

How many of the tracked models each one holds at Q4_K_M with 8K context, against 18 here.

SetupHow it relatesMemoryBandwidthModels at Q4_K_M
M2 or M3 Mac (24 GB) Same memory 16 GB usable 100 GB/s 18 (same)
M4 Mac (24 GB) Same memory 16 GB usable 120 GB/s 18 (same)
M5 Mac (24 GB) Same memory 16 GB usable 153 GB/s 18 (same)
M4 Pro Mac (24 GB) Same memory 16 GB usable 273 GB/s 18 (same)
M3 Pro Mac (18 GB) Next size down 12 GB usable 150 GB/s 17 (−1)
M4 Mac (32 GB) Next size up 21.3 GB usable 120 GB/s 28 (+10)

Model numbers read from Hugging Face on .