What LLMs can an M1 Max or M2 Max Mac (32 GB) run?

An M1 Max or M2 Max Mac (32 GB) gives a model 21.3 GB usable of memory (about two thirds of unified memory is usable by the GPU) and 400 GB/s of bandwidth. Of the 61 open models tracked here, 28 fit at Q4_K_M with an 8,192-token context, and 7 more at a lower precision, each leaving at least 0.5 GB free on the M1 Max or M2 Max Mac (32 GB).

Models that fit on one M1 Max or M2 Max Mac (32 GB)

The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first. gpt-oss-20b (14.8 GB, 43–74 tokens/s on the M1 Max or M2 Max Mac (32 GB)) is listed at MXFP4 only, as published.

Model Best precision Memory Longest context Tokens/s
K2-Horizon MoVA 36B-A4B 37.4B, 4B active GGUF Q2_K 18.2 GB 20K 35–59
Ornith 1.5 35B-A3B 36.0B, 3B active GGUF Q3_K_M 18.7 GB 106K 66–115
Qwen3.6 35B-A3B 36.0B, 3B active GGUF Q3_K_M 18.7 GB 106K 66–115
Ornith 1.0 35B 35.1B, 3B active GGUF Q3_K_M 18.3 GB 126K 66–115
LLM-jp-4.1 32B-A3B Thinking 32.1B, 3.8B active GGUF Q3_K_M 17.1 GB 61K 46–80
Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active GGUF Q4_K_M 20.1 GB 112K 51–88
Gemma 4 31B 31.3B GGUF Q3_K_M 18.1 GB 38K 13–17
GLM-4.7 Flash 31.2B, 3B active GGUF Q4_K_M 20.3 GB 16K 49–85
Xing 4.0 29B-A4B 31.2B, 4B active GGUF Q4_K_M 20.2 GB 19K 40–69
Qwen3-Coder 30B-A3B 30.5B, 3.3B active GGUF Q4_K_M 20.2 GB 13K 40–69
Qwen3 30B-A3B 30.5B, 3.3B active GGUF Q4_K_M 20.2 GB 13K 40–69
Muse Glimmer 30B 29.8B GGUF Q4_K_M 19.2 GB 124K 12–16
Granite 4.2 30B 29.3B GGUF Q3_K_M 17.4 GB 20K 13–18
Qwen3.6 27B 27.8B GGUF Q4_K_M 18.3 GB 44K 12–17
Qwen3.8 27B 27.8B GGUF Q4_K_M 18.3 GB 44K 12–17
Hemmingway-1 27B 27.3B GGUF Q4_K_M 18.0 GB 48K 13–17
Gemma 4 26B-A4B 25.8B, 4B active GGUF Q5_K_M 19.7 GB 57K 34–59
gpt-oss-20b 20.9B, 3.6B active As published (MXFP4) 14.8 GB 128K (full) 43–74
Gemma 4 12B 12.0B GGUF Q8_0 14.2 GB 256K (full) 16–22
ZDTaichu 5.0 9B 9.8B GGUF Q8_0 11.4 GB 128K (full) 20–28
Ornith 1.5 9B 9.7B FP16 / BF16 20.6 GB 14K 11–15
Qwen3.5 9B 9.7B FP16 / BF16 20.6 GB 14K 11–15
Ornith 1.0 9B 9.4B FP16 / BF16 20.1 GB 29K 11–16
MiMo V2.6 Distill Qwen 9B 9.4B FP16 / BF16 20.1 GB 29K 11–16
Granite 4.2 8B 8.8B FP16 / BF16 19.9 GB 13K 11–16
LFM2.5 8B-A1B 8.5B, 1.5B active FP16 / BF16 18.0 GB 128,000 (full) 37–62
Qwen3 8B 8.2B FP16 / BF16 18.5 GB 22K 12–17
Llama 3.1 8B 8.0B FP16 / BF16 18.1 GB 27K 13–17
Gemma 4 E4B 8.0B FP16 / BF16 17.1 GB 128K (full) 13–18
Ling 3.0 Tiny 7.9B-A1.3B 7.9B, 1.3B active FP16 / BF16 16.7 GB 128K (full) 42–73
Spark-X2.5 4B 4.1B FP16 / BF16 9.35 GB 303K 25–34
Nemotron 3 Nano 4B 4.0B FP16 / BF16 8.78 GB 256K (full) 26–36
Granite 4.2 3B 3.7B FP16 / BF16 8.69 GB 128K (full) 26–37
MiniCPM5 2B 2.5B FP16 / BF16 6.02 GB 128K (full) 38–54
Limite 1B Violetto 1.0B FP16 / BF16 2.79 GB 128K (full) 86–126

Raising the GPU memory limit on an M1 Max or M2 Max Mac (32 GB)

macOS lets the GPU wire about 21.3 GB of this Mac's 32 GB by default. Running sudo sysctl iogpu.wired_limit_mb=24576 raises that to 24 GB, leaving 8 GB to macOS and the apps next to it, until the next restart (macOS 14 or newer; iogpu.wired_limit_mb=0 restores the default). LM Studio and Ollama pick the new limit up after a restart of the app. At 24 GB, 34 of the 61 models fit at Q4_K_M with 8K context, against 28 by default : Granite 4.2 30B (11–15 tokens/s), Gemma 4 31B (10–14 tokens/s), LLM-jp-4.1 32B-A3B Thinking (40–68 tokens/s), Ornith 1.0 35B (55–96 tokens/s), Ornith 1.5 35B-A3B (55–96 tokens/s) and Qwen3.6 35B-A3B (55–96 tokens/s) join the list. Nemotron 3 Nano 30B-A3B, the biggest model that fits already, goes from 112K to 256K of context. If macOS runs short of memory it swaps or kills apps, so close big apps first and raise the limit in steps.

Too big for M1 Max or M2 Max Mac (32 GB) at 8K context

How much memory each of these lacks at Q4_K_M (gpt-oss-120b at MXFP4) with 8K context, one request and 0.5 GB kept free, out of the 21.3 GB usable here; a bigger machine or a smaller quant is the way in.

Best models for M1 Max or M2 Max Mac (32 GB) with room for long context

The biggest models that still fit at Q4_K_M with 32,768 tokens of context (or their whole window, if shorter), the memory left over, the longest context the Mac takes at that precision and the writing speed with that context in the cache, from 400 GB/s of bandwidth.

ModelContextMemoryLeft overLongest contextTokens/s
Nemotron 3 Nano 30B-A3B 31.6B 32K 20.3 GB 1.02 GB 112K 48–83
Muse Glimmer 30B 29.8B 32K 19.5 GB 1.79 GB 124K 12–16
Qwen3.6 27B 27.8B 32K 19.9 GB 1.38 GB 44K 11–16
Qwen3.8 27B 27.8B 32K 19.9 GB 1.38 GB 44K 11–16
Hemmingway-1 27B 27.3B 32K 19.6 GB 1.67 GB 48K 12–16

llama-server commands for an M1 Max or M2 Max Mac (32 GB)

GGUF repos checked 2026-09-29; -c is the longest context with 0.5 GB free on the M1 Max or M2 Max Mac (32 GB). No -ngl: llama.cpp’s -fit, on by default since b7440, places the layers.

Muse Glimmer 30B, Q4_K_M with 124K tokens: 11–15 tokens/s on the M1 Max or M2 Max Mac (32 GB)

llama-server -hf bartowski/Muse-Glimmer-30B-GGUF:Q4_K_M -c 126976 -np 1

Muse-Glimmer-30B-Q4_K_M.gguf, 17.3 GB: 20.8 GB used and 520 MB free of the M1 Max or M2 Max Mac (32 GB)'s 21.3 GB usable, 11–15 tokens/s once the 124K cache is full. -np 1: one slot, one sliding window.

Qwen3.6 27B, Q4_K_M with 10K tokens: 12–17 tokens/s on the M1 Max or M2 Max Mac (32 GB)

llama-server -hf ggml-org/Qwen3.6-27B-GGUF:Q4_K_M -c 10240

Qwen3.6-27B-Q4_K_M.gguf, 19.1 GB: 20.8 GB used and 563 MB free of the M1 Max or M2 Max Mac (32 GB)'s 21.3 GB usable, 12–17 tokens/s once the 10K cache is full.

Qwen3.8 27B, Q4_K_M with 12K tokens: 12–17 tokens/s on the M1 Max or M2 Max Mac (32 GB)

llama-server -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M -c 12288

Qwen3.8-27B-Q4_K_M.gguf, 19.0 GB: 20.8 GB used and 550 MB free of the M1 Max or M2 Max Mac (32 GB)'s 21.3 GB usable, 12–17 tokens/s once the 12K cache is full.

How fast the M1 Max or M2 Max Mac (32 GB) writes as the context fills

Tokens per second at Q4_K_M (gpt-oss-20b at MXFP4) for the biggest models that fit with 32K tokens, with 8K and 32K tokens in the cache and at the longest context the Mac holds, and at Q8_0 with 8K where that fits. Each token reads the active weights and the whole cache once, so the speed follows the 400 GB/s of bandwidth and falls as the cache grows. The ceiling is that bandwidth divided by the bytes read per token, which no runtime reaches; the reply time is for 1,000 new tokens with 32K (or the longest context) in the cache, prompt processing aside. With the same memory, M4 Mac (32 GB), 120 GB/s, writes 69% slower on average; M5 Mac (32 GB), 153 GB/s, writes 61% slower on average; M1 Pro or M2 Pro Mac (32 GB), 200 GB/s, writes 49% slower on average.

Model8K32KLongestAt the longestQ8_0 at 8KBandwidth ceiling at 8K1,000-token reply at 32K
Nemotron 3 Nano 30B-A3B 51–88 48–83 112K 40–68 Does not fit 185 12–21 s
Muse Glimmer 30B 12–16 12–16 124K 11–15 Does not fit 22 62–86 s
Qwen3.6 27B 12–17 11–16 44K 11–15 Does not fit 23 64–88 s
Qwen3.8 27B 12–17 11–16 44K 11–15 Does not fit 23 64–88 s
Hemmingway-1 27B 13–17 12–16 48K 11–15 Does not fit 23 63–86 s
Gemma 4 26B-A4B 39–67 33–57 185K 18–30 Does not fit 138 18–30 s
gpt-oss-20b 43–74 36–61 128K 21–35 MXFP4 only 155 16–28 s
Gemma 4 12B 27–37 26–36 256K 18–25 16–22 51 28–39 s
ZDTaichu 5.0 9B 34–47 30–42 128K 21–29 20–28 65 24–33 s
Ornith 1.5 9B 34–48 30–42 256K 15–21 20–28 65 24–33 s

M1 Max or M2 Max Mac (32 GB) against the other Macs with 21.3 GB usable

The same models fit on every Mac with 21.3 GB usable, so speed is what separates them. By bandwidth the M1 Max or M2 Max Mac (32 GB) is the fastest of the 4 Macs with 21.3 GB usable at 400 GB/s. Over the 12 biggest models that fit at Q4_K_M with 8K context, the M4 Mac (32 GB), at 120 GB/s, writes 69% slower, the M5 Mac (32 GB), at 153 GB/s, writes 61% slower and the M1 Pro or M2 Pro Mac (32 GB), at 200 GB/s, writes 49% slower than the M1 Max or M2 Max Mac (32 GB). Each cell below is the other Mac's tokens per second minus this one's, midpoints of the estimated ranges.

Model M1 Max or M2 Max Mac (32 GB) tokens/s M4 Mac (32 GB)M5 Mac (32 GB)M1 Pro or M2 Pro Mac (32 GB)
Nemotron 3 Nano 30B-A3B 51–88 −48−42−34
GLM-4.7 Flash 49–85 −46−40−33
Xing 4.0 29B-A4B 40–69 −38−33−27
Qwen3-Coder 30B-A3B 40–69 −38−33−27
Qwen3 30B-A3B 40–69 −38−33−27
Muse Glimmer 30B 12–16 −9.8−8.7−7.0
Qwen3.6 27B 12–17 −10−9.1−7.4
Qwen3.8 27B 12–17 −10−9.1−7.4
Hemmingway-1 27B 13–17 −10−9.2−7.5
Gemma 4 26B-A4B 39–67 −36−32−26
gpt-oss-20b 43–74 −41−36−29
Gemma 4 12B 27–37 −22−20−16

Near misses on M1 Max or M2 Max Mac (32 GB)

Models that miss at Q4_K_M with 32,768 tokens of context (or their whole window), or leave under 0.5 GB free there, but fit with a shorter context or 3-bit weights, or are short by at most a quarter of the memory and fit with MoE offload. Smallest shortfall first.

ModelShort by at Q4_K_MFits instead
Xing 4.0 29B-A4B 31.2B 96 MB Q4_K_M with 19K context
GLM-4.7 Flash 31.2B 377 MB Q4_K_M with 16K context
LLM-jp-4.1 32B-A3B Thinking 32.1B 1.32 GB GGUF Q3_K_M with 32K, 18.8 GB
Qwen3-Coder 30B-A3B 30.5B 1.42 GB Q4_K_M with 13K context
Qwen3 30B-A3B 30.5B 1.42 GB Q4_K_M with 13K context

How these numbers are worked out

Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.

Nearby GPUs and setups

How many of the tracked models each one holds at Q4_K_M with 8K context, against 28 here.

SetupHow it relatesMemoryBandwidthModels at Q4_K_M
M4 Mac (32 GB) Same memory 21.3 GB usable 120 GB/s 28 (same)
M5 Mac (32 GB) Same memory 21.3 GB usable 153 GB/s 28 (same)
M1 Pro or M2 Pro Mac (32 GB) Same memory 21.3 GB usable 200 GB/s 28 (same)
M5 Pro Mac (24 GB) Next size down 16 GB usable 307 GB/s 18 (−10)
M3 Pro Mac (36 GB) Next size up 27 GB usable 150 GB/s 35 (+7)

Model numbers read from Hugging Face on .