Best local LLMs for 48 GB of VRAM

On one 48 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is Gemma 4 31B (31.3B, 23.9 GB, about 20–28 tokens/s on one L40S), the largest MoE is K2-Horizon MoVA 36B-A4B (37.4B total, 4.0B active, 30.3 GB, about 28–48 tokens/s), among models that take 24 GB or more, the longest context goes to K2-Horizon MoVA 36B-A4B (up to 115K tokens in 47.5 GB), and the largest at Q8_0 is a tie between Ornith 1.5 35B-A3B and Qwen3.6 35B-A3B (MoE, 36.0B total, 3.0B active, 40.3 GB).

Updated , the latest date the data of a model listed here was checked; 61 models in total. Ordered by memory and size, not quality. Total = weights + FP16 KV cache for one request + 0.5 GB + 10% overhead; GB = GiB. A model counts as fitting when it leaves at least 0.5 GB free, as on the Can I run it? pages. Speeds are bandwidth estimates for one request with the 32K context full, on one L40S (864 GB/s), not measurements. Not sure how much VRAM you have? Detect your GPU.

gpt-oss-20b is listed at MXFP4, the format it is published in, in the Q4_K_M lists and not in the Q8_0 list: its GGUF quants keep the experts in MXFP4, so every type is about the same size.

Top picks for 48 GB

Pick Model Parameters Active VRAM Tokens/s (L40S)
Largest dense (Q4_K_M) Gemma 4 31B 31.3B dense 23.9 GB 20–28
Largest MoE (Q4_K_M) K2-Horizon MoVA 36B-A4B (MoE) 37.4B 4.0B 30.3 GB 28–48
Longest context, 24 GB+ models (Q4_K_M) K2-Horizon MoVA 36B-A4B (MoE) 37.4B 4.0B 115K, 47.5 GB 28–48
Largest at Q8_0 Ornith 1.5 35B-A3B (MoE) 36.0B 3.0B 40.3 GB, tied with Qwen3.6 35B-A3B 61–106

What is the best local LLM for 48 GB VRAM?

By size, the largest model that fits one 48 GB GPU at Q4_K_M with 32K context is K2-Horizon MoVA 36B-A4B (MoE, 37.4B total, 4.0B active, 30.3 GB, 17.7 GB spare); 35 of the 61 models we track fit. This page ranks by memory and size, not by benchmark scores, which our data does not include.

Can 48 GB run a 70B model?

Only below Q4_K_M: Llama 3.1 70B (70.6B) needs 55.2 GB at Q4_K_M with 32K context, but fits one 48 GB card at Q3_K_M (46.8 GB). At Q4_K_M it fits one card only with up to 9K tokens of context.

Every step is on the Llama 3.1 70B VRAM page.

Is Q8_0 or a bigger model better on 48 GB?

On one 48 GB GPU with 32K context, the largest model at Q8_0 is a tie between Ornith 1.5 35B-A3B and Qwen3.6 35B-A3B (MoE, 36.0B total, 3.0B active, 40.3 GB, about 61–106 tokens/s) and at Q4_K_M it is K2-Horizon MoVA 36B-A4B (MoE, 37.4B total, 4.0B active, 30.3 GB, about 28–48 tokens/s), 1.0× the total parameters. Q8_0 stores about 8.5 bits per weight and Q4_K_M 4.84; our data covers memory and speed, not output quality, so this page does not score the trade-off.

Bits per weight for every GGUF type are in GGUF quantization explained.

Newest models that fit 48 GB

Sorted by the date each model's data was added or checked, newest first (Q4_K_M, 32K context). New models appear here as they are added.

Model Added / checked Parameters Active VRAM (GB) Tokens/s
Qwen3-Coder 30B-A3B (MoE) 30.5B 3.3B 22.72 46–80
Ornith 1.0 35B (MoE) 35.1B 3.0B 22.95 90–160
Ornith 1.0 9B 9.4B dense 7.43 64–91
Granite 4.2 30B 29.3B dense 27.45 18–24
Granite 4.2 8B 8.8B dense 11.45 42–59
Granite 4.2 3B 3.7B dense 5.52 85–124
Spark-X2.5 4B 4.1B dense 4.40 105–157
Ling 3.0 Tiny 7.9B-A1.3B (MoE) 7.9B 1.3B 5.62 185–352

All 35 models that fit 48 GB at Q4_K_M

Model Parameters Active VRAM (GB) Spare (GB) Tokens/s Calculator
K2-Horizon MoVA 36B-A4B (MoE) 37.4B 4.0B 30.31 17.69 28–48 Open
Ornith 1.5 35B-A3B (MoE) 36.0B 3.0B 23.47 24.53 90–160 Open
Qwen3.6 35B-A3B (MoE) 36.0B 3.0B 23.47 24.53 90–160 Open
Ornith 1.0 35B (MoE) 35.1B 3.0B 22.95 25.05 90–160 Open
LLM-jp-4.1 32B-A3B Thinking (MoE) 32.1B 3.8B 22.62 25.38 53–92 Open
Nemotron 3 Nano 30B-A3B (MoE) 31.6B 3.5B 20.28 27.72 96–170 Open
Gemma 4 31B 31.3B dense 23.92 24.08 20–28 Open
GLM-4.7 Flash 31.2B 3.0B 21.67 26.33 65–114 Open
Xing 4.0 29B-A4B (MoE) 31.2B 4.0B 21.39 26.61 60–104 Open
Qwen3 30B-A3B (MoE) 30.5B 3.3B 22.72 25.28 46–80 Open
Qwen3-Coder 30B-A3B (MoE) 30.5B 3.3B 22.72 25.28 46–80 Open
Muse Glimmer 30B 29.8B dense 19.51 28.49 25–34 Open
Granite 4.2 30B 29.3B dense 27.45 20.55 18–24 Open
Qwen3.6 27B 27.8B dense 19.92 28.08 24–34 Open
Qwen3.8 27B 27.8B dense 19.92 28.08 24–34 Open
Hemmingway-1 27B 27.3B dense 19.63 28.37 25–34 Open
Gemma 4 26B-A4B (MoE) 25.8B 4.0B 17.50 30.50 68–119 Open
gpt-oss-20b MXFP4 20.9B 3.6B 15.44 32.56 72–127 Open
Gemma 4 12B 12.0B dense 8.98 39.02 53–75 Open
ZDTaichu 5.0 9B 9.8B dense 7.67 40.33 62–88 Open
Ornith 1.5 9B 9.7B dense 7.58 40.42 62–90 Open
Qwen3.5 9B 9.7B dense 7.58 40.42 62–90 Open
MiMo V2.6 Distill Qwen 9B 9.4B dense 7.43 40.57 64–91 Open
Ornith 1.0 9B 9.4B dense 7.43 40.57 64–91 Open
Granite 4.2 8B 8.8B dense 11.45 36.55 42–59 Open
LFM2.5 8B-A1B (MoE) 8.5B 1.5B 6.16 41.84 153–283 Open
Qwen3 8B 8.2B dense 10.53 37.47 45–64 Open
Llama 3.1 8B 8.0B dense 9.88 38.12 48–68 Open
Gemma 4 E4B 8.0B dense 6.05 41.95 78–113 Open
Ling 3.0 Tiny 7.9B-A1.3B (MoE) 7.9B 1.3B 5.62 42.38 185–352 Open
Spark-X2.5 4B 4.1B dense 4.40 43.60 105–157 Open
Nemotron 3 Nano 4B 4.0B dense 3.51 44.49 130–198 Open
Granite 4.2 3B 3.7B dense 5.52 42.48 85–124 Open
MiniCPM5 2B 2.5B dense 3.50 44.50 130–199 Open
Limite 1B Violetto 1.0B dense 1.62 46.38 263–457 Open

All 33 models that fit 48 GB at Q8_0

Model Parameters Active VRAM (GB) Spare (GB) Tokens/s Calculator
Ornith 1.5 35B-A3B (MoE) 36.0B 3.0B 40.32 7.68 61–106 Open
Qwen3.6 35B-A3B (MoE) 36.0B 3.0B 40.32 7.68 61–106 Open
Ornith 1.0 35B (MoE) 35.1B 3.0B 39.40 8.60 61–106 Open
LLM-jp-4.1 32B-A3B Thinking (MoE) 32.1B 3.8B 37.68 10.32 39–67 Open
Nemotron 3 Nano 30B-A3B (MoE) 31.6B 3.5B 35.08 12.92 60–104 Open
Gemma 4 31B 31.3B dense 38.58 9.42 13–17 Open
GLM-4.7 Flash 31.2B 3.0B 36.30 11.70 48–83 Open
Xing 4.0 29B-A4B (MoE) 31.2B 4.0B 36.02 11.98 42–72 Open
Qwen3 30B-A3B (MoE) 30.5B 3.3B 37.03 10.97 36–62 Open
Qwen3-Coder 30B-A3B (MoE) 30.5B 3.3B 37.03 10.97 36–62 Open
Muse Glimmer 30B 29.8B dense 33.46 14.54 14–20 Open
Granite 4.2 30B 29.3B dense 41.17 6.83 12–16 Open
Qwen3.6 27B 27.8B dense 32.94 15.06 15–20 Open
Qwen3.8 27B 27.8B dense 32.94 15.06 15–20 Open
Hemmingway-1 27B 27.3B dense 32.44 15.56 15–21 Open
Gemma 4 26B-A4B (MoE) 25.8B 4.0B 29.60 18.40 46–79 Open
Gemma 4 12B 12.0B dense 14.58 33.42 33–46 Open
ZDTaichu 5.0 9B 9.8B dense 12.26 35.74 39–55 Open
Ornith 1.5 9B 9.7B dense 12.11 35.89 39–56 Open
Qwen3.5 9B 9.7B dense 12.11 35.89 39–56 Open
MiMo V2.6 Distill Qwen 9B 9.4B dense 11.84 36.16 40–57 Open
Ornith 1.0 9B 9.4B dense 11.84 36.16 40–57 Open
Granite 4.2 8B 8.8B dense 15.57 32.43 31–43 Open
LFM2.5 8B-A1B (MoE) 8.5B 1.5B 10.13 37.87 109–195 Open
Qwen3 8B 8.2B dense 14.37 33.63 33–47 Open
Llama 3.1 8B 8.0B dense 13.64 34.36 35–49 Open
Gemma 4 E4B 8.0B dense 9.80 38.20 49–69 Open
Ling 3.0 Tiny 7.9B-A1.3B (MoE) 7.9B 1.3B 9.32 38.68 130–237 Open
Spark-X2.5 4B 4.1B dense 6.33 41.67 74–108 Open
Nemotron 3 Nano 4B 4.0B dense 5.38 42.62 87–127 Open
Granite 4.2 3B 3.7B dense 7.23 40.77 65–94 Open
MiniCPM5 2B 2.5B dense 4.68 43.32 99–147 Open
Limite 1B Violetto 1.0B dense 2.11 45.89 208–342 Open

Macs and CPU-only PCs with about 48 GB to use

These machines leave a model between 48 GB and the next size up: on a Mac, the unified memory macOS lets the GPU use by default (about two thirds up to 32 GB, 75% above; sudo sysctl iogpu.wired_limit_mb raises it), on a CPU-only PC the RAM minus about 6 GB for the system. The lists above apply to them too; speed follows each one's bandwidth, and a CPU is also much slower at reading long prompts.

Machine Usable memory Bandwidth Models at Q4_K_M (8K) Biggest at 32K Tokens/s
M4 Pro Mac (64 GB) 48 GB 273 GB/s 36 K2-Horizon MoVA 36B-A4B 9.1–15
M5 Pro Mac (64 GB) 48 GB 307 GB/s 36 K2-Horizon MoVA 36B-A4B 10–17
M1–M3 Max Mac (64 GB) 48 GB 400 GB/s 36 K2-Horizon MoVA 36B-A4B 13–22
M4 Max Mac (64 GB) 48 GB 546 GB/s 36 K2-Horizon MoVA 36B-A4B 18–30
M5 Max Mac (64 GB) 48 GB 614 GB/s 36 K2-Horizon MoVA 36B-A4B 20–34
CPU only, DDR4-3200 dual channel (64 GB) 58 GB 51.2 GB/s 38 AliceAI Foundation 80B-A3B 5.8–11
CPU only, DDR5-5600 dual channel (64 GB) 58 GB 89.6 GB/s 38 AliceAI Foundation 80B-A3B 10–19

Other VRAM sizes

All numbers, including 8K and 128K context, are in the open dataset, with summary figures on LLM VRAM statistics; any other setting can be worked out in the LLM VRAM Calculator.