What LLMs can an M4 Pro Mac (64 GB) run?

An M4 Pro Mac (64 GB) gives a model 48 GB usable of memory (about 75% of unified memory is usable by the GPU) and 273 GB/s of bandwidth. Of the 40 open models tracked here, 19 fit at Q4_K_M with an 8,192-token context, and 1 more at a lower precision.

Models that fit on one M4 Pro Mac (64 GB)

The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first.

Model Best precision Memory Longest context Tokens/s
Qwen3-Coder-Next 79.7B, 3B active GGUF Q3_K_M 40.6 GB 256K (full) 46–79
Llama 3.1 70B 70.6B GGUF Q4_K_M 47.0 GB 10K 3.3–4.5
Qwen3.6 35B-A3B 36.0B, 3B active GGUF Q8_0 39.8 GB 256K (full) 24–40
Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active GGUF Q8_0 34.9 GB 256K (full) 21–36
Gemma 4 31B 31.3B GGUF Q8_0 35.7 GB 256K (full) 4.3–5.9
GLM-4.7 Flash 31.2B, 3B active GGUF Q8_0 34.9 GB 198K (full) 22–37
Xing 4.0 29B-A4B 31.2B, 4B active GGUF Q8_0 34.9 GB 256K (full) 17–29
Qwen3 30B-A3B 30.5B, 3.3B active GGUF Q8_0 34.6 GB 40K (full) 18–31
Muse Glimmer 30B 29.8B GGUF Q8_0 33.1 GB 128K (full) 4.7–6.4
Qwen3.6 27B 27.8B GGUF Q8_0 31.3 GB 251K 5.0–6.8
Qwen3.8 27B 27.8B GGUF Q8_0 31.3 GB 251K 5.0–6.8
Gemma 4 26B-A4B 25.8B, 4B active GGUF Q8_0 28.9 GB 256K (full) 18–30
gpt-oss-20b 20.9B, 3.6B active FP16 / BF16 43.6 GB 128K (full) 11–18
Gemma 4 12B 12.0B FP16 / BF16 25.4 GB 256K (full) 6.1–8.4
Qwen3.5 9B 9.7B FP16 / BF16 20.6 GB 256K (full) 7.6–10
Qwen3 8B 8.2B FP16 / BF16 18.5 GB 40K (full) 8.4–12
Llama 3.1 8B 8.0B FP16 / BF16 18.1 GB 128K (full) 8.6–12
Gemma 4 E4B 8.0B FP16 / BF16 17.0 GB 128K (full) 9.2–13
Nemotron 3 Nano 4B 4.0B FP16 / BF16 8.78 GB 256K (full) 18–25
MiniCPM5 2B 2.5B FP16 / BF16 6.02 GB 128K (full) 27–37

Too big for one card

How many M4 Pro Mac (64 GB) cards these need at Q4_K_M in one tensor-parallel group; some need more than 8.

How these numbers are worked out

Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.

Other GPUs and Macs

Model numbers read from Hugging Face on .