How much of a Mac's unified memory can the GPU use for LLMs?

By default macOS lets the GPU use about two thirds of a Mac's unified memory up to 32 GB and three quarters from 36 GB up: 10.7 GB of 16 GB, 16 GB of 24 GB, 21.3 GB of 32 GB, 48 GB of 64 GB and 96 GB of 128 GB; sudo sysctl iogpu.wired_limit_mb raises that limit until the next restart.

Where the limit comes from

The limit is Metal's recommendedMaxWorkingSetSize, which Apple describes as “an approximation of how much memory, in bytes, this GPU device can allocate without affecting its runtime performance”. In its Metal Compute on MacBook Pro tech talk Apple gave two values: on an M1 Pro or M1 Max with 32 GB the GPU can access 21 GB, and on an M1 Max with 64 GB, 48 GB.

llama.cpp reads that property as the Metal device's working-set size (ggml-metal-device.m:1289-1290) and reports free memory as the limit minus what is already allocated (1501-1511). The same source notes that it is possible to allocate more than the limit, and llama.cpp warns when it does (2009-2021), so the limit is a recommendation that apps follow rather than a hard wall. MLX uses the same value as its system wired limit (mlx memory.cpp).

Apple does not publish a rule for every memory size. The site uses two thirds up to 32 GB and three quarters above, which matches Apple's two figures and the public llama.cpp logs in the table below. Every Mac page, the LLM VRAM calculator and the 2026 VRAM report use this same rule.

Default GPU limit and the largest model, by memory size

“Largest model” is the biggest of the 64 models the site tracks that fits at Q4_K_M (MXFP4 for gpt-oss, which ships in it) with 0.5 GB left free, one request and an FP16 KV cache. The last column uses a raised limit that leaves 4 GB to macOS up to 24 GB of memory and 8 GB above: that is the site's rule of thumb, not Apple's. GB means GiB throughout.

Unified memory GPU limit by default Logged by llama.cpp Largest model, Q4_K_M, 8K contextLargest model, Q4_K_M, 32K context Raised limit: largest at 32K
16 GB 10.7 GB (2/3) 10,922.67 MB = 10.67 GiB 2023-07-13 Gemma 4 12B 8.57 GB; 19 models fitGemma 4 12B 8.98 GB; 16 models fit 12 GB: Gemma 4 12B
24 GB 16 GB (2/3) 17,179.89 MB = 16.00 GiB M2 (MacBook Air), 2026-04-27 gpt-oss-20b 14.82 GB; 20 models fitgpt-oss-20b 15.44 GB; 19 models fit 20 GB: Gemma 4 26B-A4B
32 GB 21.3 GB (2/3) 22,906.50 MB = 21.33 GiB 2025-02-11 Nemotron 3 Nano 30B-A3B 20.12 GB; 31 models fitNemotron 3 Nano 30B-A3B 20.28 GB; 26 models fit 24 GB: Ornith 1.5 35B-A3B
36 GB 27 GB (3/4) 27,648.00 MiB = 27.00 GiB M3 Pro, 2023-12-14 K2-Horizon MoVA 36B-A4B 25.36 GB; 38 models fitOrnith 1.5 35B-A3B 23.47 GB; 35 models fit 28 GB: Ornith 1.5 35B-A3B
48 GB 36 GB (3/4) 38,654.71 MB = 36.00 GiB M4 Pro, 2024-12-30; memory not stated, the value is 75% of 48 GB K2-Horizon MoVA 36B-A4B 25.36 GB; 38 models fitK2-Horizon MoVA 36B-A4B 30.31 GB; 37 models fit 40 GB: K2-Horizon MoVA 36B-A4B
64 GB 48 GB (3/4) 49,152.00 MB = 48.00 GiB M2 Max, 2023-08-06 Llama 3.1 70B 46.98 GB; 39 models fitK2-Horizon MoVA 36B-A4B 30.31 GB; 37 models fit 56 GB: AliceAI Foundation 80B-A3B
96 GB 72 GB (3/4) 77,309.41 MB = 72.00 GiB M2 Max, 2025-09-10; memory not stated, the value is 75% of 96 GB gpt-oss-120b 67.68 GB; 42 models fitgpt-oss-120b 68.61 GB; 41 models fit 88 GB: Qwen3.5 122B-A10B
128 GB 96 GB (3/4) 103,079.22 MB = 96.00 GiB M1 Ultra, 2024-03-03; memory not stated, the value is 75% of 128 GB Mistral Medium 3.5 128B 82.68 GB; 45 models fitMistral Medium 3.5 128B 91.75 GB; 44 models fit 120 GB: Qwen3.8 Flash Next (180B MoE)
192 GB 144 GB (3/4) 154,618.82 MB = 144.00 GiB M2 Ultra, 2024-09-13; memory not stated, the value is 75% of 192 GB Step 3.7 Flash 196B-A11B 125.86 GB; 47 models fitStep 3.7 Flash 196B-A11B 127.1 GB; 46 models fit 184 GB: MiniMax M2.7
256 GB 192 GB (3/4) 206,158.43 MB = 192.00 GiB M3 Ultra, 2025-03-18 DeepSeek V4 Flash 0731 189.77 GB; 50 models fitDeepSeek V4 Flash 183.78 GB; 48 models fit 248 GB: GLM-5.3 Flash

Linked memory sizes open a Mac page with every model and its speed there; the GPU index lists all of them. Logs from llama.cpp builds before b2000 divided the bytes by 1024² (b1400 labelled that “MB”, b1600 “MiB”; b1400, b1600); from b2000 the value is divided by 10⁶. The limit can be higher than the rule: a newer M4 Max log (2025-11-29) shows 115,448.73 MB = 107.52 GiB, 84% of 128 GB (the only M4 Max size above 64 GB), so on such a Mac the site's numbers err low.

Raising the limit with iogpu.wired_limit_mb

sysctl iogpu.wired_limit_mb              # 0 means the macOS default
sudo sysctl iogpu.wired_limit_mb=24576   # 24 GB for the GPU on a 32 GB Mac
sudo sysctl iogpu.wired_limit_mb=0       # back to the default

For a single Mac, its page on this site shows the exact command for the raised limit and which models it adds, for example the M1 Max or M2 Max Mac (32 GB) or the M5 Max Mac (128 GB).

Reading the real limit in the llama.cpp log

llama.cpp and apps built on it print the limit when the Metal device starts (ggml-metal-device.m:1362-1366):

ggml_metal_device_init: recommendedMaxWorkingSetSize  = 22906.50 MB

To check a model against your Mac before downloading it, use the LLM VRAM calculator, which counts the Mac at the default limit, and the LLM speed calculator for tokens per second from the Mac's memory bandwidth.

Questions

How much unified memory can the GPU use on a 16 GB or 24 GB Mac?

About two thirds: 10.7 GB on a 16 GB Mac and 16 GB on a 24 GB one, which is what llama.cpp logs as recommendedMaxWorkingSetSize on those machines. At Q4_K_M with 8K context the largest model the site tracks that fits is Gemma 4 12B on 16 GB and gpt-oss-20b on 24 GB, with 0.5 GB left free.

How do I let the GPU use more of a Mac's memory?

Run sudo sysctl iogpu.wired_limit_mb=<MiB>, for example 24576 for 24 GB on a 32 GB Mac. It takes effect at once, apps such as LM Studio see it after a relaunch, and it goes back to the default at the next restart; sudo sysctl iogpu.wired_limit_mb=0 restores the default without one. On macOS 13 Ventura the setting was called debug.iogpu.wired_limit.

Is raising iogpu.wired_limit_mb safe?

It changes nothing permanent: a restart undoes it. The risk is leaving macOS and your other apps too little memory, which makes the Mac swap or become unstable; one llama.cpp user with a 24 GB M4 Pro reported kernel panics running large MoE models with the limit at 22000. Leave several GB to macOS, quit large apps and save your work before loading a model that uses the extra room.

Where do I see the GPU memory limit in llama.cpp or Ollama?

llama.cpp prints it when the Metal device starts: ggml_metal_device_init: recommendedMaxWorkingSetSize = 22906.50 MB on a 32 GB Mac. Current builds divide bytes by 1,000,000, so divide by 1,073.74 to get GiB (21.33). Ollama writes the same line to its server log in ~/.ollama/logs/.

Does MLX use the same limit?

Yes. MLX reports the system wired limit as max_recommended_working_set_size in device_info(), its set_wired_limit() cannot go above it, and both MLX and mlx-lm tell you to raise it with sudo sysctl iogpu.wired_limit_mb. Wiring model memory in MLX needs macOS 15 or newer.

Sources read on .