What LLMs can an M5 Max Mac (36 GB) run?

An M5 Max Mac (36 GB) gives a model 27 GB usable of memory (about 75% of unified memory is usable by the GPU) and 460 GB/s of bandwidth. Of the 61 open models tracked here, 35 fit at Q4_K_M with an 8,192-token context, each leaving at least 0.5 GB free on the M5 Max Mac (36 GB).

Models that fit on one M5 Max Mac (36 GB)

The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first. gpt-oss-20b (14.8 GB, 49–85 tokens/s on the M5 Max Mac (36 GB)) is listed at MXFP4 only, as published.

Model Best precision Memory Longest context Tokens/s
K2-Horizon MoVA 36B-A4B 37.4B, 4B active GGUF Q4_K_M 25.4 GB 13K 33–55
Ornith 1.5 35B-A3B 36.0B, 3B active GGUF Q4_K_M 23.0 GB 172K 63–110
Qwen3.6 35B-A3B 36.0B, 3B active GGUF Q4_K_M 23.0 GB 172K 63–110
Ornith 1.0 35B 35.1B, 3B active GGUF Q5_K_M 26.2 GB 23K 55–95
LLM-jp-4.1 32B-A3B Thinking 32.1B, 3.8B active GGUF Q5_K_M 24.4 GB 38K 40–68
Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active GGUF Q5_K_M 23.5 GB 256K (full) 50–87
Gemma 4 31B 31.3B GGUF Q5_K_M 25.2 GB 23K 10–14
GLM-4.7 Flash 31.2B, 3B active GGUF Q5_K_M 23.6 GB 58K 50–86
Xing 4.0 29B-A4B 31.2B, 4B active GGUF Q5_K_M 23.6 GB 68K 40–69
Qwen3-Coder 30B-A3B 30.5B, 3.3B active GGUF Q5_K_M 23.5 GB 37K 41–71
Qwen3 30B-A3B 30.5B, 3.3B active GGUF Q5_K_M 23.5 GB 37K 41–71
Muse Glimmer 30B 29.8B GGUF Q6_K 25.7 GB 63K 10–14
Granite 4.2 30B 29.3B GGUF Q5_K_M 24.0 GB 17K 11–15
Qwen3.6 27B 27.8B GGUF Q6_K 24.4 GB 38K 11–15
Qwen3.8 27B 27.8B GGUF Q6_K 24.4 GB 38K 11–15
Hemmingway-1 27B 27.3B GGUF Q6_K 24.0 GB 44K 11–15
Gemma 4 26B-A4B 25.8B, 4B active GGUF Q6_K 22.7 GB 186K 35–59
gpt-oss-20b 20.9B, 3.6B active As published (MXFP4) 14.8 GB 128K (full) 49–85
Gemma 4 12B 12.0B FP16 / BF16 25.7 GB 56K 10–14
ZDTaichu 5.0 9B 9.8B FP16 / BF16 20.8 GB 128K (full) 13–17
Ornith 1.5 9B 9.7B FP16 / BF16 20.6 GB 180K 13–17
Qwen3.5 9B 9.7B FP16 / BF16 20.6 GB 180K 13–17
Ornith 1.0 9B 9.4B FP16 / BF16 20.1 GB 195K 13–18
MiMo V2.6 Distill Qwen 9B 9.4B FP16 / BF16 20.1 GB 195K 13–18
Granite 4.2 8B 8.8B FP16 / BF16 19.9 GB 46K 13–18
LFM2.5 8B-A1B 8.5B, 1.5B active FP16 / BF16 18.0 GB 128,000 (full) 42–72
Qwen3 8B 8.2B FP16 / BF16 18.5 GB 40K (full) 14–19
Llama 3.1 8B 8.0B FP16 / BF16 18.1 GB 69K 14–20
Gemma 4 E4B 8.0B FP16 / BF16 17.1 GB 128K (full) 15–21
Ling 3.0 Tiny 7.9B-A1.3B 7.9B, 1.3B active FP16 / BF16 16.7 GB 128K (full) 48–83
Spark-X2.5 4B 4.1B FP16 / BF16 9.35 GB 451K 28–39
Nemotron 3 Nano 4B 4.0B FP16 / BF16 8.78 GB 256K (full) 30–42
Granite 4.2 3B 3.7B FP16 / BF16 8.69 GB 128K (full) 30–42
MiniCPM5 2B 2.5B FP16 / BF16 6.02 GB 128K (full) 44–62
Limite 1B Violetto 1.0B FP16 / BF16 2.79 GB 128K (full) 97–143

Raising the GPU memory limit on an M5 Max Mac (36 GB)

macOS lets the GPU wire about 27 GB of this Mac's 36 GB by default. Running sudo sysctl iogpu.wired_limit_mb=28672 raises that to 28 GB, leaving 8 GB to macOS and the apps next to it, until the next restart (macOS 14 or newer; iogpu.wired_limit_mb=0 restores the default). LM Studio and Ollama pick the new limit up after a restart of the app. At 28 GB, 35 of the 61 models fit at Q4_K_M with 8K context, against 35 by default . K2-Horizon MoVA 36B-A4B, the biggest model that fits already, goes from 13K to 18K of context. If macOS runs short of memory it swaps or kills apps, so close big apps first and raise the limit in steps.

Too big for M5 Max Mac (36 GB) at 8K context

How much memory each of these lacks at Q4_K_M (gpt-oss-120b at MXFP4) with 8K context, one request and 0.5 GB kept free, out of the 27 GB usable here; a bigger machine or a smaller quant is the way in.

Best models for M5 Max Mac (36 GB) with room for long context

The biggest models that still fit at Q4_K_M with 32,768 tokens of context (or their whole window, if shorter), the memory left over, the longest context the Mac takes at that precision and the writing speed with that context in the cache, from 460 GB/s of bandwidth.

ModelContextMemoryLeft overLongest contextTokens/s
Ornith 1.5 35B-A3B 36.0B 32K 23.5 GB 3.53 GB 172K 51–88
Qwen3.6 35B-A3B 36.0B 32K 23.5 GB 3.53 GB 172K 51–88
Ornith 1.0 35B 35.1B 32K 22.9 GB 4.05 GB 197K 51–88
LLM-jp-4.1 32B-A3B Thinking 32.1B 32K 22.6 GB 4.38 GB 64K (full) 30–50
Nemotron 3 Nano 30B-A3B 31.6B 32K 20.3 GB 6.72 GB 256K (full) 55–95

llama-server commands for an M5 Max Mac (36 GB)

GGUF repos checked 2026-09-29; -c is the longest context with 0.5 GB free on the M5 Max Mac (36 GB).

Ornith 1.5 35B-A3B, Q4_K_M with 167K tokens: 25–42 tokens/s on the M5 Max Mac (36 GB)

llama-server -hf bartowski/Ornith-1.5-35B-A3B-GGUF:Q4_K_M -c 171008 -ngl 99

Ornith-1.5-35B-A3B-Q4_K_M.gguf, 21.9 GB: 26.5 GB used and 526 MB free of the M5 Max Mac (36 GB)'s 27 GB usable, 25–42 tokens/s once the 167K cache is full.

Qwen3.6 35B-A3B, Q4_K_M with 172K tokens: 25–42 tokens/s on the M5 Max Mac (36 GB)

llama-server -hf ggml-org/Qwen3.6-35B-A3B-GGUF:Q4_K_M -c 176128 -ngl 99

Qwen3.6-35B-A3B-Q4_K_M.gguf, 20.4 GB: 26.5 GB used and 534 MB free of the M5 Max Mac (36 GB)'s 27 GB usable, 25–42 tokens/s once the 172K cache is full.

Ornith 1.0 35B, Q4_K_M with 190K tokens: 23–39 tokens/s on the M5 Max Mac (36 GB)

llama-server -hf bartowski/deepreinforce-ai_Ornith-1.0-35B-GGUF:Q4_K_M -c 194560 -ngl 99

deepreinforce-ai_Ornith-1.0-35B-Q4_K_M.gguf, 21.4 GB: 26.5 GB used and 515 MB free of the M5 Max Mac (36 GB)'s 27 GB usable, 23–39 tokens/s once the 190K cache is full.

How fast the M5 Max Mac (36 GB) writes as the context fills

Tokens per second at Q4_K_M for the biggest models that fit with 32K tokens, with 8K and 32K tokens in the cache and at the longest context the Mac holds, and at Q8_0 with 8K where that fits. Each token reads the active weights and the whole cache once, so the speed follows the 460 GB/s of bandwidth and falls as the cache grows. The ceiling is that bandwidth divided by the bytes read per token, which no runtime reaches; the reply time is for 1,000 new tokens with 32K (or the longest context) in the cache, prompt processing aside. With the same memory, M3 Pro Mac (36 GB), 150 GB/s, writes 66% slower on average; M3 Max Mac (36 GB), 300 GB/s, writes 34% slower on average; M4 Max Mac (36 GB), 410 GB/s, writes 10% slower on average.

Model8K32KLongestAt the longestQ8_0 at 8KBandwidth ceiling at 8K1,000-token reply at 32K
Ornith 1.5 35B-A3B 63–110 51–88 172K 25–42 Does not fit 232 11–20 s
Qwen3.6 35B-A3B 63–110 51–88 172K 25–42 Does not fit 232 11–20 s
Ornith 1.0 35B 63–110 51–88 197K 22–38 Does not fit 232 11–20 s
LLM-jp-4.1 32B-A3B Thinking 45–78 30–50 64K 20–34 Does not fit 161 20–34 s
Nemotron 3 Nano 30B-A3B 58–101 55–95 256K 35–60 Does not fit 212 11–18 s
Gemma 4 31B 12–16 11–15 61K 9.9–14 Does not fit 22 67–92 s
GLM-4.7 Flash 56–97 36–62 117K 16–27 Does not fit 204 16–28 s
Xing 4.0 29B-A4B 46–79 33–57 137K 15–26 Does not fit 164 18–30 s
Qwen3-Coder 30B-A3B 46–79 25–43 68K 15–26 Does not fit 164 23–39 s
Qwen3 30B-A3B 46–79 25–43 40K 22–37 Does not fit 164 23–39 s

M5 Max Mac (36 GB) against the other Macs with 27 GB usable

The same models fit on every Mac with 27 GB usable, so speed is what separates them. By bandwidth the M5 Max Mac (36 GB) is the fastest of the 4 Macs with 27 GB usable at 460 GB/s. Over the 12 biggest models that fit at Q4_K_M with 8K context, the M3 Pro Mac (36 GB), at 150 GB/s, writes 66% slower, the M3 Max Mac (36 GB), at 300 GB/s, writes 34% slower and the M4 Max Mac (36 GB), at 410 GB/s, writes 10% slower than the M5 Max Mac (36 GB). Each cell below is the other Mac's tokens per second minus this one's, midpoints of the estimated ranges.

Model M5 Max Mac (36 GB) tokens/s M3 Pro Mac (36 GB)M3 Max Mac (36 GB)M4 Max Mac (36 GB)
K2-Horizon MoVA 36B-A4B 33–55 −29−15−4.6
Ornith 1.5 35B-A3B 63–110 −57−29−8.8
Qwen3.6 35B-A3B 63–110 −57−29−8.8
Ornith 1.0 35B 63–110 −57−29−8.8
LLM-jp-4.1 32B-A3B Thinking 45–78 −41−21−6.4
Nemotron 3 Nano 30B-A3B 58–101 −52−26−8.1
Gemma 4 31B 12–16 −9.5−4.9−1.5
GLM-4.7 Flash 56–97 −50−25−7.8
Xing 4.0 29B-A4B 46–79 −41−21−6.5
Qwen3-Coder 30B-A3B 46–79 −41−21−6.5
Qwen3 30B-A3B 46–79 −41−21−6.5
Muse Glimmer 30B 14–19 −11−5.6−1.7

Near misses on M5 Max Mac (36 GB)

Models that miss at Q4_K_M with 32,768 tokens of context (or their whole window), or leave under 0.5 GB free there, but fit with a shorter context or 3-bit weights, or are short by at most a quarter of the memory and fit with MoE offload. Smallest shortfall first.

ModelShort by at Q4_K_MFits instead
Granite 4.2 30B 29.3B 456 MB Q4_K_M with 28K context
K2-Horizon MoVA 36B-A4B 37.4B 3.31 GB Q4_K_M with 13K context

Run locally or rent an M5 Max Mac (36 GB)?

getdeploying.com lists no on-demand rental of M5 Max Mac (36 GB) (checked ). Among the GPUs this site tracks, the cheapest to rent with at least 27 GB is the RTX 5090 at a median $0.69 an hour, which holds 35 of the 61 models at Q4_K_M with 8K context against 35 here; every $1,000 equals about 1,449 hours of it.

How these numbers are worked out

Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.

Nearby GPUs and setups

How many of the tracked models each one holds at Q4_K_M with 8K context, against 35 here.

SetupHow it relatesMemoryBandwidthModels at Q4_K_M
M3 Pro Mac (36 GB) Same memory 27 GB usable 150 GB/s 35 (same)
M3 Max Mac (36 GB) Same memory 27 GB usable 300 GB/s 35 (same)
M4 Max Mac (36 GB) Same memory 27 GB usable 410 GB/s 35 (same)
M1 Max or M2 Max Mac (32 GB) Next size down 21.3 GB usable 400 GB/s 28 (−7)
M4 Pro Mac (48 GB) Next size up 36 GB usable 273 GB/s 35 (same)

Model numbers read from Hugging Face on .