What LLMs can an M2 or M3 Mac (8 GB) run?
An M2 or M3 Mac (8 GB) gives a model 5.3 GB usable of memory (about two thirds of unified memory is usable by the GPU) and 100 GB/s of bandwidth. Of the 61 open models tracked here, 5 fit at Q4_K_M with an 8,192-token context, and 5 more at a lower precision, each leaving at least 0.5 GB free on the M2 or M3 Mac (8 GB).
Models that fit on one M2 or M3 Mac (8 GB)
The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first.
| Model | Best precision | Memory | Longest context | Tokens/s |
|---|---|---|---|---|
| Ornith 1.0 9B 9.4B | GGUF IQ3_XXS | 4.75 GB | 9K | 13–18 |
| MiMo V2.6 Distill Qwen 9B 9.4B | GGUF IQ3_XXS | 4.75 GB | 9K | 13–18 |
| LFM2.5 8B-A1B 8.5B, 1.5B active | GGUF Q2_K | 4.24 GB | 51K | 39–66 |
| Gemma 4 E4B 8.0B | GGUF Q3_K_M | 4.68 GB | 14K | 13–18 |
| Ling 3.0 Tiny 7.9B-A1.3B 7.9B, 1.3B active | GGUF Q3_K_M | 4.51 GB | 47K | 41–70 |
| Spark-X2.5 4B 4.1B | GGUF Q6_K | 4.38 GB | 18K | 14–20 |
| Nemotron 3 Nano 4B 4.0B | FP8 / INT8 | 4.71 GB | 13K | 13–18 |
| Granite 4.2 3B 3.7B | GGUF Q6_K | 4.26 GB | 14K | 15–20 |
| MiniCPM5 2B 2.5B | GGUF Q8_0 | 3.60 GB | 34K | 18–24 |
| Limite 1B Violetto 1.0B | FP16 / BF16 | 2.79 GB | 128K (full) | 24–33 |
Too big for M2 or M3 Mac (8 GB) at 8K context
How much memory each of these lacks at Q4_K_M (gpt-oss-20b and gpt-oss-120b at MXFP4) with 8K context, one request and 0.5 GB kept free, out of the 5.3 GB usable here; a bigger machine or a smaller quant is the way in.
- Llama 3.1 8B: 1.78 GB short
- Qwen3 8B: 2.01 GB short
- Granite 4.2 8B: 2.52 GB short
- Ornith 1.5 9B: 1.96 GB short
- Qwen3.5 9B: 1.96 GB short
- ZDTaichu 5.0 9B: 2.05 GB short
- Gemma 4 12B: 3.77 GB short
- gpt-oss-20b: 10.0 GB short
- Gemma 4 26B-A4B: 12.2 GB short
- Hemmingway-1 27B: 13.2 GB short
- Qwen3.6 27B: 13.5 GB short
- Qwen3.8 27B: 13.5 GB short
- Granite 4.2 30B: 16.0 GB short
- Muse Glimmer 30B: 14.4 GB short
- Qwen3-Coder 30B-A3B: 15.4 GB short
- Qwen3 30B-A3B: 15.4 GB short
- Xing 4.0 29B-A4B: 15.4 GB short
- GLM-4.7 Flash: 15.5 GB short
- Gemma 4 31B: 17.1 GB short
- Nemotron 3 Nano 30B-A3B: 15.3 GB short
- LLM-jp-4.1 32B-A3B Thinking: 16.2 GB short
- Ornith 1.0 35B: 17.6 GB short
- Ornith 1.5 35B-A3B: 18.2 GB short
- Qwen3.6 35B-A3B: 18.2 GB short
- K2-Horizon MoVA 36B-A4B: 20.6 GB short
- Llama 3.1 70B: 42.2 GB short
- Qwen3-Coder-Next: 45.3 GB short
- AliceAI Foundation 80B-A3B: 46.3 GB short
- gpt-oss-120b: 62.9 GB short
- Nemotron 3 Super 120B-A12B: 72.4 GB short
- Qwen3.5 122B-A10B: 73.4 GB short
- Mistral Medium 3.5 128B: 77.9 GB short
- Qwen3.8 Flash Next: 107 GB short
- Step 3.7 Flash 196B-A11B: 121 GB short
- MiniMax M2.7: 140 GB short
- DeepSeek V4 Flash: 177 GB short
- DeepSeek V4 Flash 0731: 185 GB short
- MiMo V2.6 Flash: 189 GB short
- IQuest-Q1 320B-A15B: 196 GB short
- GLM-5.3 Flash: 195 GB short
- MiniMax M3: 261 GB short
- DeepSeek V3 / R1: 421 GB short
- DeepSeek V3.2: 421 GB short
- GLM-5.3: 463 GB short
- GLM-5.2: 463 GB short
- DeepSeek V4.1 Flash: 469 GB short
- Hy4 Preview 770B-A49B: 480 GB short
- MiMo V2.6 Pro: 631 GB short
- DeepSeek V4 Pro: 988 GB short
- Qwen3.8 2.4T-A95B: 1,513 GB short
- Kimi K3: 1,719 GB short
Best models for M2 or M3 Mac (8 GB) with room for long context
The biggest models that still fit at Q4_K_M with 32,768 tokens of context (or their whole window, if shorter), the memory left over, the longest context the Mac takes at that precision and the writing speed with that context in the cache, from 100 GB/s of bandwidth.
| Model | Context | Memory | Left over | Longest context | Tokens/s |
|---|---|---|---|---|---|
| Spark-X2.5 4B 4.1B | 32K | 4.40 GB | 919 MB | 42K | 14–19 |
| Nemotron 3 Nano 4B 4.0B | 32K | 3.51 GB | 1.79 GB | 106K | 18–25 |
| MiniCPM5 2B 2.5B | 32K | 3.50 GB | 1.80 GB | 60K | 18–25 |
| Limite 1B Violetto 1.0B | 32K | 1.62 GB | 3.68 GB | 128K (full) | 47–66 |
llama-server commands for an M2 or M3 Mac (8 GB)
GGUF repos checked 2026-09-29; -c is the longest context with 0.5 GB free on the M2 or M3 Mac (8 GB). No -ngl: llama.cpp’s -fit, on by default since b7440, places the layers.
Spark-X2.5 4B, Q4_K_M with 39K tokens: 13–18 tokens/s on the M2 or M3 Mac (8 GB)
llama-server -hf XHToken/Spark-X2.5-4B-GGUF:Q4_K_M -c 39936 -np 1 Spark-X2.5-4B-Q4_K_M.gguf, 2.6 GB: 4.79 GB used and 524 MB free of the M2 or M3 Mac (8 GB)'s 5.3 GB usable, 13–18 tokens/s once the 39K cache is full. -np 1: one slot, one sliding window.
Nemotron 3 Nano 4B, Q4_K_M with 77K tokens: 15–20 tokens/s on the M2 or M3 Mac (8 GB)
llama-server -hf unsloth/NVIDIA-Nemotron-3-Nano-4B-GGUF:Q4_K_M -c 78848 NVIDIA-Nemotron-3-Nano-4B-Q4_K_M.gguf, 2.9 GB: 4.79 GB used and 517 MB free of the M2 or M3 Mac (8 GB)'s 5.3 GB usable, 15–20 tokens/s once the 77K cache is full.
MiniCPM5 2B, Q4_K_M with 58K tokens: 13–18 tokens/s on the M2 or M3 Mac (8 GB)
llama-server -hf bartowski/MiniCPM5-2B-GGUF:Q4_K_M -c 59392 MiniCPM5-2B-Q4_K_M.gguf, 1.6 GB: 4.77 GB used and 541 MB free of the M2 or M3 Mac (8 GB)'s 5.3 GB usable, 13–18 tokens/s once the 58K cache is full.
How fast the M2 or M3 Mac (8 GB) writes as the context fills
Tokens per second at Q4_K_M for the biggest models that fit with 32K tokens, with 8K and 32K tokens in the cache and at the longest context the Mac holds, and at Q8_0 with 8K where that fits. Each token reads the active weights and the whole cache once, so the speed follows the 100 GB/s of bandwidth and falls as the cache grows. The ceiling is that bandwidth divided by the bytes read per token, which no runtime reaches; the reply time is for 1,000 new tokens with 32K (or the longest context) in the cache, prompt processing aside. With the same memory, M1 Mac (8 GB), 68.25 GB/s, writes 31% slower on average.
| Model | 8K | 32K | Longest | At the longest | Q8_0 at 8K | Bandwidth ceiling at 8K | 1,000-token reply at 32K |
|---|---|---|---|---|---|---|---|
| Spark-X2.5 4B | 18–26 | 14–19 | 42K | 13–18 | Does not fit | 34 | 51–71 s |
| Nemotron 3 Nano 4B | 21–29 | 18–25 | 106K | 13–18 | Does not fit | 39 | 40–55 s |
| MiniCPM5 2B | 28–39 | 18–25 | 60K | 13–18 | 18–24 | 53 | 40–55 s |
| Limite 1B Violetto | 63–90 | 47–66 | 128K | 23–32 | 41–58 | 126 | 15–21 s |
M2 or M3 Mac (8 GB) against the other Macs with 5.3 GB usable
The same models fit on every Mac with 5.3 GB usable, so speed is what separates them. By bandwidth the M2 or M3 Mac (8 GB) is the fastest of the 2 Macs with 5.3 GB usable at 100 GB/s. Over the 5 biggest models that fit at Q4_K_M with 8K context, the M1 Mac (8 GB), at 68.25 GB/s, writes 31% slower than the M2 or M3 Mac (8 GB). Each cell below is the other Mac's tokens per second minus this one's, midpoints of the estimated ranges.
| Model | M2 or M3 Mac (8 GB) tokens/s | M1 Mac (8 GB) |
|---|---|---|
| Spark-X2.5 4B | 18–26 | −6.9 |
| Nemotron 3 Nano 4B | 21–29 | −7.8 |
| Granite 4.2 3B | 19–26 | −6.9 |
| MiniCPM5 2B | 28–39 | −10 |
| Limite 1B Violetto | 63–90 | −23 |
Near misses on M2 or M3 Mac (8 GB)
Models that miss at Q4_K_M with 32,768 tokens of context (or their whole window), or leave under 0.5 GB free there, but fit with a shorter context or 3-bit weights, or are short by at most a quarter of the memory and fit with MoE offload. Smallest shortfall first.
| Model | Short by at Q4_K_M | Fits instead |
|---|---|---|
| Granite 4.2 3B 3.7B | 224 MB | Q4_K_M with 23K context |
| Ling 3.0 Tiny 7.9B-A1.3B 7.9B | 332 MB | GGUF Q3_K_M with 32K, 4.68 GB |
| Gemma 4 E4B 8.0B | 767 MB | GGUF IQ3_XXS with 32K, 4.47 GB |
| LFM2.5 8B-A1B 8.5B | 881 MB | GGUF IQ3_XXS with 32K, 4.49 GB |
How these numbers are worked out
Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.
Nearby GPUs and setups
How many of the tracked models each one holds at Q4_K_M with 8K context, against 5 here.
| Setup | How it relates | Memory | Bandwidth | Models at Q4_K_M |
|---|---|---|---|---|
| M1 Mac (8 GB) | Same memory | 5.3 GB usable | 68.25 GB/s | 5 (same) |
| M1 Mac (16 GB) | Next size up | 10.7 GB usable | 68.25 GB/s | 17 (+12) |
Model numbers read from Hugging Face on .