Can I run gpt-oss-120b on an M4 Max Mac (128 GB)?
Yes: gpt-oss-120b needs about 68.6 GB at MXFP4 with 32K tokens of context, which fits the M4 Max Mac (128 GB, 96 GB usable) with 27.4 GB to spare. MXFP4 is the format gpt-oss-120b is published in, and on the M4 Max 128GB it runs up to 128K (full) tokens with 0.5 GB to spare.
Why 68.6 GB with vLLM, but 60.6 GB on the M4 Max 128GB with llama.cpp
The 68.6 GB counts the published checkpoint as vLLM loads it: 60.8 GB of weights, the 32K KV cache, 0.5 GB and 10% overhead. llama.cpp's measured MXFP4 GGUF is 59.0 GB, and its 587 MB of token embeddings stay in RAM. With every layer on the M4 Max 128GB it takes 60.6 GB at 32K with 1 GB of buffers, 35.4 GB under its 96 GB.
Yes MXFP4 with 32K tokens of context
- MXFP4, 32K
- 68.6 GB
- M4 Max 128GB
- 96 GB usable, 546 GB/s
- To spare
- 27.4 GB
- Tokens/s
- 38–65 tokens/s
At MXFP4 with 32K tokens of context it writes about 38–65 tokens/s for one request on an M4 Max Mac (128 GB).
Best precision for gpt-oss-120b on an M4 Max Mac (128 GB)
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | MXFP4 | 67.7 GB | 28.3 GB | 48–82 |
| 32K | MXFP4 | 68.6 GB | 27.4 GB | 38–65 |
| 128K (full) | MXFP4 | 72.3 GB | 23.7 GB | 21–35 |
gpt-oss-120b on the M4 Max 128GB as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; the MXFP4 GGUF columns are the measured llama.cpp file with every layer on the GPU and 1 GB of buffers; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | MXFP4 | Free | Tokens/s | MXFP4 GGUF | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 67.5 GB | 28.5 GB | 50–86 | 59.6 GB | 36.4 GB | 41–70 |
| 8K | 67.7 GB | 28.3 GB | 48–82 | 59.8 GB | 36.2 GB | 39–67 |
| 16K | 68.0 GB | 28.0 GB | 44–75 | 60.0 GB | 36.0 GB | 37–63 |
| 32K | 68.6 GB | 27.4 GB | 38–65 | 60.6 GB | 35.4 GB | 32–55 |
| 64K | 69.8 GB | 26.2 GB | 30–50 | 61.7 GB | 34.3 GB | 26–44 |
| 128K (full) | 72.3 GB | 23.7 GB | 21–35 | 64.0 GB | 32.0 GB | 19–32 |
Other options
- Smallest setup for MXFP4 at 32K: A100 80GB.
Run gpt-oss-120b on the M4 Max 128GB with llama-server
llama-server -hf ggml-org/gpt-oss-120b-GGUF:MXFP4 -c 131072 -np 1 gpt-oss-120b-MXFP4.gguf, 63.4 GB, from ggml-org/
Questions
Can I run gpt-oss-120b on an M4 Max Mac (128 GB)?
Yes: gpt-oss-120b needs about 68.6 GB at MXFP4 with 32K tokens of context, which fits the M4 Max Mac (128 GB, 96 GB usable) with 27.4 GB to spare. MXFP4 is the format gpt-oss-120b is published in, and on the M4 Max 128GB it runs up to 128K (full) tokens with 0.5 GB to spare.
How fast is gpt-oss-120b on an M4 Max Mac (128 GB)?
At MXFP4 with 32K tokens of context it writes about 38–65 tokens/s for one request on an M4 Max Mac (128 GB).
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also gpt-oss-120b VRAM requirements, what LLMs an M4 Max Mac (128 GB) can run and every pair, or detect your own GPU. Model data checked .