What LLMs can an M3 Ultra Mac Studio (512 GB) run?

An M3 Ultra Mac Studio (512 GB) gives a model 384 GB usable of memory (about 75% of unified memory is usable by the GPU) and 819 GB/s of bandwidth. Of the 40 open models tracked here, 31 fit at Q4_K_M with an 8,192-token context, and 5 more at a lower precision.

Models that fit on one M3 Ultra Mac Studio (512 GB)

The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first.

Model Best precision Memory Longest context Tokens/s
DeepSeek V4.1 Flash 763.2B, 16B active GGUF Q3_K_M 383 GB 15K 28–47
GLM-5.3 753.3B, 40B active GGUF Q3_K_M 378 GB 64K 12–20
GLM-5.2 753.3B, 40B active GGUF Q3_K_M 378 GB 64K 12–20
DeepSeek V3.2 685.4B, 37B active GGUF Q3_K_M 344 GB 160K (full) 13–22
DeepSeek V3 / R1 684.5B, 37B active GGUF Q3_K_M 344 GB 160K (full) 13–22
MiniMax M3 427.0B, 23B active GGUF Q6_K 360 GB 192K 12–20
GLM-5.3 Flash 321.3B, 18B active GGUF Q8_0 350 GB 1M (full) 13–21
MiMo V2.6 Flash 310.8B, 15B active GGUF Q8_0 339 GB 1M (full) 15–25
DeepSeek V4 Flash 0731 304.2B, 13B active GGUF Q8_0 332 GB 567K 16–28
DeepSeek V4 Flash 290.9B, 13B active GGUF Q8_0 318 GB 723K 16–28
MiniMax M2.7 228.7B, 10B active GGUF Q8_0 252 GB 200K (full) 19–32
Qwen3.8 Flash Next 180.0B, 6B active FP16 / BF16 370 GB 256K (full) 20–33
Mistral Medium 3.5 128B 127.7B FP16 / BF16 265 GB 256K (full) 1.7–2.4
Qwen3.5 122B-A10B 125.1B, 10B active FP16 / BF16 257 GB 256K (full) 12–20
Nemotron 3 Super 120B-A12B 123.6B, 12B active FP16 / BF16 254 GB 256K (full) 10–17
gpt-oss-120b 116.8B, 5.1B active FP16 / BF16 240 GB 128K (full) 23–38
Qwen3-Coder-Next 79.7B, 3B active FP16 / BF16 164 GB 256K (full) 37–64
Llama 3.1 70B 70.6B FP16 / BF16 148 GB 128K (full) 3.1–4.3
Qwen3.6 35B-A3B 36.0B, 3B active FP16 / BF16 74.3 GB 256K (full) 38–64
Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active FP16 / BF16 65.3 GB 256K (full) 33–56
Gemma 4 31B 31.3B FP16 / BF16 65.8 GB 256K (full) 7.0–9.6
GLM-4.7 Flash 31.2B, 3B active FP16 / BF16 64.9 GB 198K (full) 36–62
Xing 4.0 29B-A4B 31.2B, 4B active FP16 / BF16 64.8 GB 256K (full) 28–48
Qwen3 30B-A3B 30.5B, 3.3B active FP16 / BF16 63.9 GB 40K (full) 32–54
Muse Glimmer 30B 29.8B FP16 / BF16 61.7 GB 128K (full) 7.5–10
Qwen3.6 27B 27.8B FP16 / BF16 58.0 GB 256K (full) 7.9–11
Qwen3.8 27B 27.8B FP16 / BF16 58.0 GB 256K (full) 7.9–11
Gemma 4 26B-A4B 25.8B, 4B active FP16 / BF16 53.7 GB 256K (full) 28–48
gpt-oss-20b 20.9B, 3.6B active FP16 / BF16 43.6 GB 128K (full) 32–54
Gemma 4 12B 12.0B FP16 / BF16 25.4 GB 256K (full) 18–25
Qwen3.5 9B 9.7B FP16 / BF16 20.6 GB 256K (full) 22–31
Qwen3 8B 8.2B FP16 / BF16 18.5 GB 40K (full) 25–34
Llama 3.1 8B 8.0B FP16 / BF16 18.1 GB 128K (full) 25–35
Gemma 4 E4B 8.0B FP16 / BF16 17.0 GB 128K (full) 27–37
Nemotron 3 Nano 4B 4.0B FP16 / BF16 8.78 GB 256K (full) 51–73
MiniCPM5 2B 2.5B FP16 / BF16 6.02 GB 128K (full) 74–108

Too big for one card

How many M3 Ultra Mac Studio (512 GB) cards these need at Q4_K_M in one tensor-parallel group; some need more than 8.

How these numbers are worked out

Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.

Other GPUs and Macs

Model numbers read from Hugging Face on .