How much of a Mac's unified memory can the GPU use for LLMs?
By default macOS lets the GPU use about two thirds of a Mac's unified memory up to 32 GB and three quarters from 36 GB up: 10.7 GB of 16 GB, 16 GB of 24 GB, 21.3 GB of 32 GB, 48 GB of 64 GB and 96 GB of 128 GB; sudo sysctl iogpu.wired_limit_mb raises that limit until the next restart.
Where the limit comes from
The limit is Metal's recommendedMaxWorkingSetSize, which Apple
describes as “an approximation of how much memory, in bytes, this GPU device can allocate without affecting its runtime
performance”. In its Metal Compute on MacBook Pro tech talk Apple gave two
values: on an M1 Pro or M1 Max with 32 GB the GPU can access 21 GB, and on an M1 Max with 64 GB, 48 GB.
llama.cpp reads that property as the Metal device's working-set size (ggml-metal-device.m:1289-1290) and reports free memory as the limit minus what is already allocated (1501-1511). The same source notes that it is possible to allocate more than the limit, and llama.cpp warns when it does (2009-2021), so the limit is a recommendation that apps follow rather than a hard wall. MLX uses the same value as its system wired limit (mlx memory.cpp).
Apple does not publish a rule for every memory size. The site uses two thirds up to 32 GB and three quarters above, which matches Apple's two figures and the public llama.cpp logs in the table below. Every Mac page, the LLM VRAM calculator and the 2026 VRAM report use this same rule.
Default GPU limit and the largest model, by memory size
“Largest model” is the biggest of the 64 models the site tracks that fits at Q4_K_M (MXFP4 for gpt-oss, which ships in it) with 0.5 GB left free, one request and an FP16 KV cache. The last column uses a raised limit that leaves 4 GB to macOS up to 24 GB of memory and 8 GB above: that is the site's rule of thumb, not Apple's. GB means GiB throughout.
| Unified memory | GPU limit by default | Logged by llama.cpp | Largest model, Q4_K_M, 8K context | Largest model, Q4_K_M, 32K context | Raised limit: largest at 32K |
|---|---|---|---|---|---|
| 16 GB | 10.7 GB (2/3) | 10,922.67 MB = 10.67 GiB 2023-07-13 | Gemma 4 12B 8.57 GB; 19 models fit | Gemma 4 12B 8.98 GB; 16 models fit | 12 GB: Gemma 4 12B |
| 24 GB | 16 GB (2/3) | 17,179.89 MB = 16.00 GiB M2 (MacBook Air), 2026-04-27 | gpt-oss-20b 14.82 GB; 20 models fit | gpt-oss-20b 15.44 GB; 19 models fit | 20 GB: Gemma 4 26B-A4B |
| 32 GB | 21.3 GB (2/3) | 22,906.50 MB = 21.33 GiB 2025-02-11 | Nemotron 3 Nano 30B-A3B 20.12 GB; 31 models fit | Nemotron 3 Nano 30B-A3B 20.28 GB; 26 models fit | 24 GB: Ornith 1.5 35B-A3B |
| 36 GB | 27 GB (3/4) | 27,648.00 MiB = 27.00 GiB M3 Pro, 2023-12-14 | K2-Horizon MoVA 36B-A4B 25.36 GB; 38 models fit | Ornith 1.5 35B-A3B 23.47 GB; 35 models fit | 28 GB: Ornith 1.5 35B-A3B |
| 48 GB | 36 GB (3/4) | 38,654.71 MB = 36.00 GiB M4 Pro, 2024-12-30; memory not stated, the value is 75% of 48 GB | K2-Horizon MoVA 36B-A4B 25.36 GB; 38 models fit | K2-Horizon MoVA 36B-A4B 30.31 GB; 37 models fit | 40 GB: K2-Horizon MoVA 36B-A4B |
| 64 GB | 48 GB (3/4) | 49,152.00 MB = 48.00 GiB M2 Max, 2023-08-06 | Llama 3.1 70B 46.98 GB; 39 models fit | K2-Horizon MoVA 36B-A4B 30.31 GB; 37 models fit | 56 GB: AliceAI Foundation 80B-A3B |
| 96 GB | 72 GB (3/4) | 77,309.41 MB = 72.00 GiB M2 Max, 2025-09-10; memory not stated, the value is 75% of 96 GB | gpt-oss-120b 67.68 GB; 42 models fit | gpt-oss-120b 68.61 GB; 41 models fit | 88 GB: Qwen3.5 122B-A10B |
| 128 GB | 96 GB (3/4) | 103,079.22 MB = 96.00 GiB M1 Ultra, 2024-03-03; memory not stated, the value is 75% of 128 GB | Mistral Medium 3.5 128B 82.68 GB; 45 models fit | Mistral Medium 3.5 128B 91.75 GB; 44 models fit | 120 GB: Qwen3.8 Flash Next (180B MoE) |
| 192 GB | 144 GB (3/4) | 154,618.82 MB = 144.00 GiB M2 Ultra, 2024-09-13; memory not stated, the value is 75% of 192 GB | Step 3.7 Flash 196B-A11B 125.86 GB; 47 models fit | Step 3.7 Flash 196B-A11B 127.1 GB; 46 models fit | 184 GB: MiniMax M2.7 |
| 256 GB | 192 GB (3/4) | 206,158.43 MB = 192.00 GiB M3 Ultra, 2025-03-18 | DeepSeek V4 Flash 0731 189.77 GB; 50 models fit | DeepSeek V4 Flash 183.78 GB; 48 models fit | 248 GB: GLM-5.3 Flash |
Linked memory sizes open a Mac page with every model and its speed there; the GPU index lists all of them. Logs from llama.cpp builds before b2000 divided the bytes by 1024² (b1400 labelled that “MB”, b1600 “MiB”; b1400, b1600); from b2000 the value is divided by 10⁶. The limit can be higher than the rule: a newer M4 Max log (2025-11-29) shows 115,448.73 MB = 107.52 GiB, 84% of 128 GB (the only M4 Max size above 64 GB), so on such a Mac the site's numbers err low.
Raising the limit with iogpu.wired_limit_mb
sysctl iogpu.wired_limit_mb # 0 means the macOS default
sudo sysctl iogpu.wired_limit_mb=24576 # 24 GB for the GPU on a 32 GB Mac
sudo sysctl iogpu.wired_limit_mb=0 # back to the default - Units are MiB. On a 32 GB Mac, 24576 took llama.cpp's reported limit from 22,906.50 MB to 25,769.80 MB, which is 24 GiB (peddals.com, community test).
- It lasts until the next restart. The change applies at once; LM Studio shows the new value after a
relaunch. A restart returns to the default, and the same post shows how to keep it with
/etc/sysctl.conf, with a warning to do that carefully. - macOS version.
iogpu.wired_limit_mbis the name on macOS 14 Sonoma and later; on 13 Ventura it wasdebug.iogpu.wired_limit, and 0 means the default split (llama.cpp #4406, community). Besides community posts, the sources for it are the MLX project's API docs and the mlx-lm README, which tell users to set it to a value above the model's size and below the machine's memory. - Risk. Memory the GPU wires is not available to macOS and your apps. Leave several GB, quit big apps and save work first. One user with a 24 GB M4 Pro running large MoE models at a limit of 22000 reported kernel panics (llama.cpp #19825).
For a single Mac, its page on this site shows the exact command for the raised limit and which models it adds, for example the M1 Max or M2 Max Mac (32 GB) or the M5 Max Mac (128 GB).
Reading the real limit in the llama.cpp log
llama.cpp and apps built on it print the limit when the Metal device starts (ggml-metal-device.m:1362-1366):
ggml_metal_device_init: recommendedMaxWorkingSetSize = 22906.50 MB - Current builds print bytes ÷ 1,000,000. Divide by 1,073.74 for GiB: 22,906.50 MB is 21.33 GiB, two thirds of 32 GB.
- The model loader then reports that limit minus what is already allocated as “MiB free” for the Metal device.
-
When a buffer takes the total past the limit, llama.cpp prints
current allocated size is greater than the recommended max working set size; at debug log level it also logs every Metal buffer as “allocated / limit” in MiB (2009-2021). - Ollama writes the same line to its server log in
~/.ollama/logs/(peddals.com). - In MLX,
device_info()reports it asmax_recommended_working_set_size.
To check a model against your Mac before downloading it, use the LLM VRAM calculator, which counts the Mac at the default limit, and the LLM speed calculator for tokens per second from the Mac's memory bandwidth.
Questions
How much unified memory can the GPU use on a 16 GB or 24 GB Mac?
About two thirds: 10.7 GB on a 16 GB Mac and 16 GB on a 24 GB one, which is what llama.cpp logs as recommended
How do I let the GPU use more of a Mac's memory?
Run sudo sysctl iogpu.wired_limit_mb=<MiB>, for example 24576 for 24 GB on a 32 GB Mac. It takes effect at once, apps such as LM Studio see it after a relaunch, and it goes back to the default at the next restart; sudo sysctl iogpu.wired_limit_mb=0 restores the default without one. On macOS 13 Ventura the setting was called debug.iogpu.wired_limit.
Is raising iogpu.wired_limit_mb safe?
It changes nothing permanent: a restart undoes it. The risk is leaving macOS and your other apps too little memory, which makes the Mac swap or become unstable; one llama.cpp user with a 24 GB M4 Pro reported kernel panics running large MoE models with the limit at 22000. Leave several GB to macOS, quit large apps and save your work before loading a model that uses the extra room.
Where do I see the GPU memory limit in llama.cpp or Ollama?
llama.cpp prints it when the Metal device starts: ggml_metal_device_init: recommended
Does MLX use the same limit?
Yes. MLX reports the system wired limit as max_recommended_working_set_size in device_info(), its set_wired_limit() cannot go above it, and both MLX and mlx-lm tell you to raise it with sudo sysctl iogpu.wired_limit_mb. Wiring model memory in MLX needs macOS 15 or newer.
Sources read on .