What LLMs can an M4 Max Mac (48 GB) run?

An M4 Max Mac (48 GB) gives a model 36 GB usable of memory (about 75% of unified memory is usable by the GPU) and 546 GB/s of bandwidth. Of the 61 open models tracked here, 35 fit at Q4_K_M with an 8,192-token context, and 3 more at a lower precision, each leaving at least 0.5 GB free on the M4 Max Mac (48 GB).

Models that fit on one M4 Max Mac (48 GB)

The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first. gpt-oss-20b (14.8 GB, 58–100 tokens/s on the M4 Max Mac (48 GB)) is listed at MXFP4 only, as published.

Model Best precision Memory Longest context Tokens/s
AliceAI Foundation 80B-A3B 81.3B, 3B active GGUF IQ3_XXS 35.1 GB 25K 97–173
Qwen3-Coder-Next 79.7B, 3B active GGUF Q2_K 34.9 GB 31K 96–171
Llama 3.1 70B 70.6B GGUF Q2_K 33.5 GB 13K 9.2–13
K2-Horizon MoVA 36B-A4B 37.4B, 4B active GGUF Q6_K 33.6 GB 16K 32–54
Ornith 1.5 35B-A3B 36.0B, 3B active GGUF Q6_K 30.9 GB 223K 57–99
Qwen3.6 35B-A3B 36.0B, 3B active GGUF Q6_K 30.9 GB 223K 57–99
Ornith 1.0 35B 35.1B, 3B active GGUF Q6_K 30.2 GB 256K (full) 57–99
LLM-jp-4.1 32B-A3B Thinking 32.1B, 3.8B active FP8 / INT8 34.0 GB 30K 36–61
Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active GGUF Q8_0 34.9 GB 97K 41–70
Gemma 4 31B 31.3B FP8 / INT8 34.5 GB 19K 8.9–12
GLM-4.7 Flash 31.2B, 3B active GGUF Q8_0 34.9 GB 17K 42–72
Xing 4.0 29B-A4B 31.2B, 4B active GGUF Q8_0 34.9 GB 20K 34–57
Qwen3-Coder 30B-A3B 30.5B, 3.3B active GGUF Q8_0 34.6 GB 16K 36–61
Qwen3 30B-A3B 30.5B, 3.3B active GGUF Q8_0 34.6 GB 16K 36–61
Muse Glimmer 30B 29.8B GGUF Q8_0 33.1 GB 128K (full) 9.3–13
Granite 4.2 30B 29.3B GGUF Q8_0 34.6 GB 11K 8.9–12
Qwen3.6 27B 27.8B GGUF Q8_0 31.3 GB 69K 9.8–14
Qwen3.8 27B 27.8B GGUF Q8_0 31.3 GB 69K 9.8–14
Hemmingway-1 27B 27.3B GGUF Q8_0 30.8 GB 76K 10–14
Gemma 4 26B-A4B 25.8B, 4B active GGUF Q8_0 29.1 GB 256K (full) 33–56
gpt-oss-20b 20.9B, 3.6B active As published (MXFP4) 14.8 GB 128K (full) 58–100
Gemma 4 12B 12.0B FP16 / BF16 25.7 GB 256K (full) 12–17
ZDTaichu 5.0 9B 9.8B FP16 / BF16 20.8 GB 128K (full) 15–20
Ornith 1.5 9B 9.7B FP16 / BF16 20.6 GB 256K (full) 15–21
Qwen3.5 9B 9.7B FP16 / BF16 20.6 GB 256K (full) 15–21
Ornith 1.0 9B 9.4B FP16 / BF16 20.1 GB 256K (full) 15–21
MiMo V2.6 Distill Qwen 9B 9.4B FP16 / BF16 20.1 GB 256K (full) 15–21
Granite 4.2 8B 8.8B FP16 / BF16 19.9 GB 98K 15–21
LFM2.5 8B-A1B 8.5B, 1.5B active FP16 / BF16 18.0 GB 128,000 (full) 49–84
Qwen3 8B 8.2B FP16 / BF16 18.5 GB 40K (full) 17–23
Llama 3.1 8B 8.0B FP16 / BF16 18.1 GB 128K (full) 17–24
Gemma 4 E4B 8.0B FP16 / BF16 17.1 GB 128K (full) 18–25
Ling 3.0 Tiny 7.9B-A1.3B 7.9B, 1.3B active FP16 / BF16 16.7 GB 128K (full) 56–98
Spark-X2.5 4B 4.1B FP16 / BF16 9.35 GB 684K 33–46
Nemotron 3 Nano 4B 4.0B FP16 / BF16 8.78 GB 256K (full) 35–49
Granite 4.2 3B 3.7B FP16 / BF16 8.69 GB 128K (full) 36–50
MiniCPM5 2B 2.5B FP16 / BF16 6.02 GB 128K (full) 51–73
Limite 1B Violetto 1.0B FP16 / BF16 2.79 GB 128K (full) 112–168

Raising the GPU memory limit on an M4 Max Mac (48 GB)

macOS lets the GPU wire about 36 GB of this Mac's 48 GB by default. Running sudo sysctl iogpu.wired_limit_mb=40960 raises that to 40 GB, leaving 8 GB to macOS and the apps next to it, until the next restart (macOS 14 or newer; iogpu.wired_limit_mb=0 restores the default). LM Studio and Ollama pick the new limit up after a restart of the app. At 40 GB, 35 of the 61 models fit at Q4_K_M with 8K context, against 35 by default . K2-Horizon MoVA 36B-A4B, the biggest model that fits already, goes from 57K to 76K of context. If macOS runs short of memory it swaps or kills apps, so close big apps first and raise the limit in steps.

Too big for M4 Max Mac (48 GB) at 8K context

How much memory each of these lacks at Q4_K_M (gpt-oss-120b at MXFP4) with 8K context, one request and 0.5 GB kept free, out of the 36 GB usable here; a bigger machine or a smaller quant is the way in.

Best models for M4 Max Mac (48 GB) with room for long context

The biggest models that still fit at Q4_K_M with 32,768 tokens of context (or their whole window, if shorter), the memory left over, the longest context the Mac takes at that precision and the writing speed with that context in the cache, from 546 GB/s of bandwidth.

ModelContextMemoryLeft overLongest contextTokens/s
K2-Horizon MoVA 36B-A4B 37.4B 32K 30.3 GB 5.69 GB 57K 18–30
Ornith 1.5 35B-A3B 36.0B 32K 23.5 GB 12.5 GB 256K (full) 60–104
Qwen3.6 35B-A3B 36.0B 32K 23.5 GB 12.5 GB 256K (full) 60–104
Ornith 1.0 35B 35.1B 32K 22.9 GB 13.1 GB 256K (full) 60–104
LLM-jp-4.1 32B-A3B Thinking 32.1B 32K 22.6 GB 13.4 GB 64K (full) 35–59

llama-server commands for an M4 Max Mac (48 GB)

GGUF repos checked 2026-09-29; -c is the longest context with 0.5 GB free on the M4 Max Mac (48 GB). No -ngl: llama.cpp’s -fit, on by default since b7440, places the layers.

K2-Horizon MoVA 36B-A4B, Q4_K_M with 57K tokens: 12–19 tokens/s on the M4 Max Mac (48 GB)

llama-server -hf IFM/K2-Horizon-MoVA-36B-A4B-GGUF:Q4_K_M -c 58368

K2-Horizon-MoVA-36B-A4B-Q4_K_M.gguf, 22.4 GB: 35.5 GB used and 549 MB free of the M4 Max Mac (48 GB)'s 36 GB usable, 12–19 tokens/s once the 57K cache is full.

Ornith 1.5 35B-A3B, Q4_K_M with 256K tokens: 22–37 tokens/s on the M4 Max Mac (48 GB)

llama-server -hf bartowski/Ornith-1.5-35B-A3B-GGUF:Q4_K_M -c 262144

Ornith-1.5-35B-A3B-Q4_K_M.gguf, 21.9 GB: 28.4 GB used and 7.60 GB free of the M4 Max Mac (48 GB)'s 36 GB usable, 22–37 tokens/s once the 256K cache is full.

How fast the M4 Max Mac (48 GB) writes as the context fills

Tokens per second at Q4_K_M for the biggest models that fit with 32K tokens, with 8K and 32K tokens in the cache and at the longest context the Mac holds, and at Q8_0 with 8K where that fits. Each token reads the active weights and the whole cache once, so the speed follows the 546 GB/s of bandwidth and falls as the cache grows. The ceiling is that bandwidth divided by the bytes read per token, which no runtime reaches; the reply time is for 1,000 new tokens with 32K (or the longest context) in the cache, prompt processing aside. With the same memory, M4 Pro Mac (48 GB), 273 GB/s, writes 48% slower on average; M5 Pro Mac (48 GB), 307 GB/s, writes 42% slower on average; M5 Max Mac (48 GB), 614 GB/s, writes 12% faster on average.

Model8K32KLongestAt the longestQ8_0 at 8KBandwidth ceiling at 8K1,000-token reply at 32K
K2-Horizon MoVA 36B-A4B 38–66 18–30 57K 12–19 Does not fit 135 33–56 s
Ornith 1.5 35B-A3B 74–129 60–104 256K 22–37 Does not fit 275 10–17 s
Qwen3.6 35B-A3B 74–129 60–104 256K 22–37 Does not fit 275 10–17 s
Ornith 1.0 35B 74–129 60–104 256K 22–37 Does not fit 275 10–17 s
LLM-jp-4.1 32B-A3B Thinking 53–91 35–59 64K 24–40 Does not fit 191 17–29 s
Nemotron 3 Nano 30B-A3B 68–118 64–111 256K 41–71 41–70 252 9–16 s
Gemma 4 31B 14–19 13–18 166K 8.7–12 Does not fit 26 56–78 s
GLM-4.7 Flash 65–114 43–73 198K 13–21 42–72 242 14–23 s
Xing 4.0 29B-A4B 54–93 39–67 256K 11–19 34–57 195 15–25 s
Qwen3-Coder 30B-A3B 54–93 30–51 155K 9.2–15 36–61 195 20–33 s

M4 Max Mac (48 GB) against the other Macs with 36 GB usable

The same models fit on every Mac with 36 GB usable, so speed is what separates them. By bandwidth the M4 Max Mac (48 GB) ranks 2 of 4 at 546 GB/s. Over the 12 biggest models that fit at Q4_K_M with 8K context, the M4 Pro Mac (48 GB), at 273 GB/s, writes 48% slower, the M5 Pro Mac (48 GB), at 307 GB/s, writes 42% slower and the M5 Max Mac (48 GB), at 614 GB/s, writes 12% faster than the M4 Max Mac (48 GB). Each cell below is the other Mac's tokens per second minus this one's, midpoints of the estimated ranges.

Model M4 Max Mac (48 GB) tokens/s M4 Pro Mac (48 GB)M5 Pro Mac (48 GB)M5 Max Mac (48 GB)
K2-Horizon MoVA 36B-A4B 38–66 −25−22+6.2
Ornith 1.5 35B-A3B 74–129 −48−42+11
Qwen3.6 35B-A3B 74–129 −48−42+11
Ornith 1.0 35B 74–129 −48−42+11
LLM-jp-4.1 32B-A3B Thinking 53–91 −35−31+8.4
Nemotron 3 Nano 30B-A3B 68–118 −45−39+11
Gemma 4 31B 14–19 −8.3−7.3+2.1
GLM-4.7 Flash 65–114 −43−38+10
Xing 4.0 29B-A4B 54–93 −36−31+8.5
Qwen3-Coder 30B-A3B 54–93 −36−31+8.5
Qwen3 30B-A3B 54–93 −36−31+8.5
Muse Glimmer 30B 16–22 −9.5−8.3+2.3

Near misses on M4 Max Mac (48 GB)

Models that miss at Q4_K_M with 32,768 tokens of context (or their whole window), or leave under 0.5 GB free there, but fit with a shorter context or 3-bit weights, or are short by at most a quarter of the memory and fit with MoE offload. Smallest shortfall first.

ModelShort by at Q4_K_MFits instead
Qwen3-Coder-Next 79.7B 14.7 GB GGUF IQ3_XXS with 32K, 35.0 GB

Run locally or rent an M4 Max Mac (48 GB)?

getdeploying.com lists no on-demand rental of M4 Max Mac (48 GB) (checked ). Among the GPUs this site tracks, the cheapest to rent with at least 36 GB is the L40S at a median $1.57 an hour, which holds 36 of the 61 models at Q4_K_M with 8K context against 35 here; every $1,000 equals about 637 hours of it.

How these numbers are worked out

Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.

Nearby GPUs and setups

How many of the tracked models each one holds at Q4_K_M with 8K context, against 35 here.

SetupHow it relatesMemoryBandwidthModels at Q4_K_M
M4 Pro Mac (48 GB) Same memory 36 GB usable 273 GB/s 35 (same)
M5 Pro Mac (48 GB) Same memory 36 GB usable 307 GB/s 35 (same)
M5 Max Mac (48 GB) Same memory 36 GB usable 614 GB/s 35 (same)
M5 Max Mac (36 GB) Next size down 27 GB usable 460 GB/s 35 (same)
M4 Pro Mac (64 GB) Next size up 48 GB usable 273 GB/s 36 (+1)

Model numbers read from Hugging Face on .