Best local LLMs for 24 GB of VRAM
On one 24 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is Muse Glimmer 30B (29.8B, 19.5 GB, about 29–40 tokens/s on one RTX 4090), the largest MoE is a tie between Ornith 1.5 35B-A3B and Qwen3.6 35B-A3B (36.0B total, 3.0B active, 23.5 GB, about 103–184 tokens/s), among models that take 12 GB or more, the longest context goes to Nemotron 3 Nano 30B-A3B (its full 256K in 21.7 GB), and the largest at Q8_0 is Gemma 4 12B (12.0B, 14.6 GB).
Updated , the latest date the data of a model listed here was checked; 61 models in total. Ordered by memory and size, not quality. Total = weights + FP16 KV cache for one request + 0.5 GB + 10% overhead; GB = GiB. A model counts as fitting when it leaves at least 0.5 GB free, as on the Can I run it? pages. Speeds are bandwidth estimates for one request with the 32K context full, on one RTX 4090 (1008 GB/s), not measurements. Not sure how much VRAM you have? Detect your GPU.
gpt-oss-20b is listed at MXFP4, the format it is published in, in the Q4_K_M lists and not in the Q8_0 list: its GGUF quants keep the experts in MXFP4, so every type is about the same size.
Top picks for 24 GB
| Pick | Model | Parameters | Active | VRAM | Tokens/s (RTX 4090) |
|---|---|---|---|---|---|
| Largest dense (Q4_K_M) | Muse Glimmer 30B | 29.8B | dense | 19.5 GB | 29–40 |
| Largest MoE (Q4_K_M) | Ornith 1.5 35B-A3B (MoE) | 36.0B | 3.0B | 23.5 GB, tied with Qwen3.6 35B-A3B | 103–184 |
| Longest context, 12 GB+ models (Q4_K_M) | Nemotron 3 Nano 30B-A3B (MoE) | 31.6B | 3.5B | 256K (full), 21.7 GB | 109–196 |
| Largest at Q8_0 | Gemma 4 12B | 12.0B | dense | 14.6 GB | 38–54 |
What is the best local LLM for 24 GB VRAM?
By size, the largest model that fits one 24 GB GPU at Q4_K_M with 32K context is a tie between Ornith 1.5 35B-A3B and Qwen3.6 35B-A3B (MoE, 36.0B total, 3.0B active, 23.5 GB, 0.5 GB spare); 32 of the 61 models we track fit. This page ranks by memory and size, not by benchmark scores, which our data does not include.
Can 24 GB run a 70B model?
Not on one card: Llama 3.1 70B (70.6B) needs 55.2 GB at Q4_K_M with 32K context, or 4 × 24 GB cards, and still 41.3 GB at IQ3_XXS.
Every step is on the Llama 3.1 70B VRAM page.
Is Q8_0 or a bigger model better on 24 GB?
On one 24 GB GPU with 32K context, the largest model at Q8_0 is Gemma 4 12B (12.0B, 14.6 GB, about 38–54 tokens/s) and at Q4_K_M it is a tie between Ornith 1.5 35B-A3B and Qwen3.6 35B-A3B (MoE, 36.0B total, 3.0B active, 23.5 GB, about 103–184 tokens/s), 3.0× the total parameters. Q8_0 stores about 8.5 bits per weight and Q4_K_M 4.84; our data covers memory and speed, not output quality, so this page does not score the trade-off.
Bits per weight for every GGUF type are in GGUF quantization explained.
Newest models that fit 24 GB
Sorted by the date each model's data was added or checked, newest first (Q4_K_M, 32K context). New models appear here as they are added.
| Model | Added / checked | Parameters | Active | VRAM (GB) | Tokens/s |
|---|---|---|---|---|---|
| Qwen3-Coder 30B-A3B (MoE) | 30.5B | 3.3B | 22.72 | 53–92 | |
| Ornith 1.0 35B (MoE) | 35.1B | 3.0B | 22.95 | 103–184 | |
| Ornith 1.0 9B | 9.4B | dense | 7.43 | 73–106 | |
| Granite 4.2 8B | 8.8B | dense | 11.45 | 48–68 | |
| Granite 4.2 3B | 3.7B | dense | 5.52 | 97–143 | |
| Spark-X2.5 4B | 4.1B | dense | 4.40 | 119–181 | |
| Ling 3.0 Tiny 7.9B-A1.3B (MoE) | 7.9B | 1.3B | 5.62 | 206–398 | |
| LFM2.5 8B-A1B (MoE) | 8.5B | 1.5B | 6.16 | 171–323 |
All 32 models that fit 24 GB at Q4_K_M
| Model | Parameters | Active | VRAM (GB) | Spare (GB) | Tokens/s | Calculator |
|---|---|---|---|---|---|---|
| Ornith 1.5 35B-A3B (MoE) | 36.0B | 3.0B | 23.47 | 0.53 | 103–184 | Open |
| Qwen3.6 35B-A3B (MoE) | 36.0B | 3.0B | 23.47 | 0.53 | 103–184 | Open |
| Ornith 1.0 35B (MoE) | 35.1B | 3.0B | 22.95 | 1.05 | 103–184 | Open |
| LLM-jp-4.1 32B-A3B Thinking (MoE) | 32.1B | 3.8B | 22.62 | 1.38 | 62–107 | Open |
| Nemotron 3 Nano 30B-A3B (MoE) | 31.6B | 3.5B | 20.28 | 3.72 | 109–196 | Open |
| GLM-4.7 Flash | 31.2B | 3.0B | 21.67 | 2.33 | 75–131 | Open |
| Xing 4.0 29B-A4B (MoE) | 31.2B | 4.0B | 21.39 | 2.61 | 69–121 | Open |
| Qwen3 30B-A3B (MoE) | 30.5B | 3.3B | 22.72 | 1.28 | 53–92 | Open |
| Qwen3-Coder 30B-A3B (MoE) | 30.5B | 3.3B | 22.72 | 1.28 | 53–92 | Open |
| Muse Glimmer 30B | 29.8B | dense | 19.51 | 4.49 | 29–40 | Open |
| Qwen3.6 27B | 27.8B | dense | 19.92 | 4.08 | 28–39 | Open |
| Qwen3.8 27B | 27.8B | dense | 19.92 | 4.08 | 28–39 | Open |
| Hemmingway-1 27B | 27.3B | dense | 19.63 | 4.37 | 28–40 | Open |
| Gemma 4 26B-A4B (MoE) | 25.8B | 4.0B | 17.50 | 6.50 | 78–138 | Open |
| gpt-oss-20b MXFP4 | 20.9B | 3.6B | 15.44 | 8.56 | 83–146 | Open |
| Gemma 4 12B | 12.0B | dense | 8.98 | 15.02 | 61–87 | Open |
| ZDTaichu 5.0 9B | 9.8B | dense | 7.67 | 16.33 | 71–102 | Open |
| Ornith 1.5 9B | 9.7B | dense | 7.58 | 16.42 | 72–104 | Open |
| Qwen3.5 9B | 9.7B | dense | 7.58 | 16.42 | 72–104 | Open |
| MiMo V2.6 Distill Qwen 9B | 9.4B | dense | 7.43 | 16.57 | 73–106 | Open |
| Ornith 1.0 9B | 9.4B | dense | 7.43 | 16.57 | 73–106 | Open |
| Granite 4.2 8B | 8.8B | dense | 11.45 | 12.55 | 48–68 | Open |
| LFM2.5 8B-A1B (MoE) | 8.5B | 1.5B | 6.16 | 17.84 | 171–323 | Open |
| Qwen3 8B | 8.2B | dense | 10.53 | 13.47 | 52–74 | Open |
| Llama 3.1 8B | 8.0B | dense | 9.88 | 14.12 | 56–79 | Open |
| Gemma 4 E4B | 8.0B | dense | 6.05 | 17.95 | 89–130 | Open |
| Ling 3.0 Tiny 7.9B-A1.3B (MoE) | 7.9B | 1.3B | 5.62 | 18.38 | 206–398 | Open |
| Spark-X2.5 4B | 4.1B | dense | 4.40 | 19.60 | 119–181 | Open |
| Nemotron 3 Nano 4B | 4.0B | dense | 3.51 | 20.49 | 147–228 | Open |
| Granite 4.2 3B | 3.7B | dense | 5.52 | 18.48 | 97–143 | Open |
| MiniCPM5 2B | 2.5B | dense | 3.50 | 20.50 | 147–228 | Open |
| Limite 1B Violetto | 1.0B | dense | 1.62 | 22.38 | 288–513 | Open |
All 17 models that fit 24 GB at Q8_0
| Model | Parameters | Active | VRAM (GB) | Spare (GB) | Tokens/s | Calculator |
|---|---|---|---|---|---|---|
| Gemma 4 12B | 12.0B | dense | 14.58 | 9.42 | 38–54 | Open |
| ZDTaichu 5.0 9B | 9.8B | dense | 12.26 | 11.74 | 45–64 | Open |
| Ornith 1.5 9B | 9.7B | dense | 12.11 | 11.89 | 46–65 | Open |
| Qwen3.5 9B | 9.7B | dense | 12.11 | 11.89 | 46–65 | Open |
| MiMo V2.6 Distill Qwen 9B | 9.4B | dense | 11.84 | 12.16 | 47–66 | Open |
| Ornith 1.0 9B | 9.4B | dense | 11.84 | 12.16 | 47–66 | Open |
| Granite 4.2 8B | 8.8B | dense | 15.57 | 8.43 | 36–50 | Open |
| LFM2.5 8B-A1B (MoE) | 8.5B | 1.5B | 10.13 | 13.87 | 123–224 | Open |
| Qwen3 8B | 8.2B | dense | 14.37 | 9.63 | 39–54 | Open |
| Llama 3.1 8B | 8.0B | dense | 13.64 | 10.36 | 41–57 | Open |
| Gemma 4 E4B | 8.0B | dense | 9.80 | 14.20 | 56–80 | Open |
| Ling 3.0 Tiny 7.9B-A1.3B (MoE) | 7.9B | 1.3B | 9.32 | 14.68 | 147–271 | Open |
| Spark-X2.5 4B | 4.1B | dense | 6.33 | 17.67 | 85–125 | Open |
| Nemotron 3 Nano 4B | 4.0B | dense | 5.38 | 18.62 | 99–147 | Open |
| Granite 4.2 3B | 3.7B | dense | 7.23 | 16.77 | 75–109 | Open |
| MiniCPM5 2B | 2.5B | dense | 4.68 | 19.32 | 113–169 | Open |
| Limite 1B Violetto | 1.0B | dense | 2.11 | 21.89 | 231–388 | Open |
Macs and CPU-only PCs with about 24 GB to use
These machines leave a model between 24 GB and the next size up: on a Mac, the unified memory macOS lets the GPU use by default (about two thirds up to 32 GB, 75% above; sudo sysctl iogpu.wired_limit_mb raises it), on a CPU-only PC the RAM minus about 6 GB for the system. The lists above apply to them too; speed follows each one's bandwidth, and a CPU is also much slower at reading long prompts.
| Machine | Usable memory | Bandwidth | Models at Q4_K_M (8K) | Biggest at 32K | Tokens/s |
|---|---|---|---|---|---|
| M3 Pro Mac (36 GB) | 27 GB | 150 GB/s | 35 | Ornith 1.5 35B-A3B | 18–30 |
| M3 Max Mac (36 GB) | 27 GB | 300 GB/s | 35 | Ornith 1.5 35B-A3B | 34–59 |
| M4 Max Mac (36 GB) | 27 GB | 410 GB/s | 35 | Ornith 1.5 35B-A3B | 46–79 |
| M5 Max Mac (36 GB) | 27 GB | 460 GB/s | 35 | Ornith 1.5 35B-A3B | 51–88 |
| CPU only, DDR4-3200 dual channel (32 GB) | 26 GB | 51.2 GB/s | 35 | Ornith 1.5 35B-A3B | 6.1–11 |
| CPU only, DDR5-5600 dual channel (32 GB) | 26 GB | 89.6 GB/s | 35 | Ornith 1.5 35B-A3B | 11–20 |
Other VRAM sizes
All numbers, including 8K and 128K context, are in the open dataset, with summary figures on LLM VRAM statistics; any other setting can be worked out in the LLM VRAM Calculator.