Can I run gpt-oss-120b on an M4 Max Mac (128 GB)?

Yes: gpt-oss-120b needs about 68.6 GB at MXFP4 with 32K tokens of context, which fits the M4 Max Mac (128 GB, 96 GB usable) with 27.4 GB to spare. MXFP4 is the format gpt-oss-120b is published in, and on the M4 Max 128GB it runs up to 128K (full) tokens with 0.5 GB to spare.

Why 68.6 GB with vLLM, but 60.6 GB on the M4 Max 128GB with llama.cpp

The 68.6 GB counts the published checkpoint as vLLM loads it: 60.8 GB of weights, the 32K KV cache, 0.5 GB and 10% overhead. llama.cpp's measured MXFP4 GGUF is 59.0 GB, and its 587 MB of token embeddings stay in RAM. With every layer on the M4 Max 128GB it takes 60.6 GB at 32K with 1 GB of buffers, 35.4 GB under its 96 GB.

Yes MXFP4 with 32K tokens of context

MXFP4, 32K
68.6 GB
M4 Max 128GB
96 GB usable, 546 GB/s
To spare
27.4 GB
Tokens/s
38–65 tokens/s

At MXFP4 with 32K tokens of context it writes about 38–65 tokens/s for one request on an M4 Max Mac (128 GB).

Best precision for gpt-oss-120b on an M4 Max Mac (128 GB)

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K MXFP4 67.7 GB 28.3 GB 48–82
32K MXFP4 68.6 GB 27.4 GB 38–65
128K (full) MXFP4 72.3 GB 23.7 GB 21–35

gpt-oss-120b on the M4 Max 128GB as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; the MXFP4 GGUF columns are the measured llama.cpp file with every layer on the GPU and 1 GB of buffers; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context MXFP4FreeTokens/sMXFP4 GGUFFreeTokens/s
4K 67.5 GB 28.5 GB 50–86 59.6 GB 36.4 GB 41–70
8K 67.7 GB 28.3 GB 48–82 59.8 GB 36.2 GB 39–67
16K 68.0 GB 28.0 GB 44–75 60.0 GB 36.0 GB 37–63
32K 68.6 GB 27.4 GB 38–65 60.6 GB 35.4 GB 32–55
64K 69.8 GB 26.2 GB 30–50 61.7 GB 34.3 GB 26–44
128K (full) 72.3 GB 23.7 GB 21–35 64.0 GB 32.0 GB 19–32

Other options

Run gpt-oss-120b on the M4 Max 128GB with llama-server

llama-server -hf ggml-org/gpt-oss-120b-GGUF:MXFP4 -c 131072 -np 1

gpt-oss-120b-MXFP4.gguf, 63.4 GB, from ggml-org/gpt-oss-120b-GGUF (checked 2026-09-29). At -c 131072 on the M4 Max 128GB: 72.3 GB of 96 GB, 23.7 GB free. -np 1: one slot, one sliding window.

Questions

Can I run gpt-oss-120b on an M4 Max Mac (128 GB)?

Yes: gpt-oss-120b needs about 68.6 GB at MXFP4 with 32K tokens of context, which fits the M4 Max Mac (128 GB, 96 GB usable) with 27.4 GB to spare. MXFP4 is the format gpt-oss-120b is published in, and on the M4 Max 128GB it runs up to 128K (full) tokens with 0.5 GB to spare.

How fast is gpt-oss-120b on an M4 Max Mac (128 GB)?

At MXFP4 with 32K tokens of context it writes about 38–65 tokens/s for one request on an M4 Max Mac (128 GB).

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also gpt-oss-120b VRAM requirements, what LLMs an M4 Max Mac (128 GB) can run and every pair, or detect your own GPU. Model data checked .