Best local LLMs for 96 GB of VRAM

On one 96 GB GPU at Q4_K_M with 32K context, the largest dense model that fits is Mistral Medium 3.5 128B (127.7B, 91.8 GB, about 11–15 tokens/s on one RTX PRO 6000 Blackwell), the largest MoE is Qwen3.5 122B-A10B (125.1B total, 10.0B active, 78.8 GB, about 70–123 tokens/s), among models that take 48 GB or more, the longest context goes to Qwen3.5 122B-A10B (its full 256K in 84.6 GB), and the largest at Q8_0 is AliceAI Foundation 80B-A3B (MoE, 81.3B total, 3.0B active, 89.8 GB).

Updated , the latest date the data of a model listed here was checked; 55 models in total. Ordered by memory and size, not quality. Total = weights + FP16 KV cache for one request + 0.5 GB + 10% overhead; GB = GiB. A model counts as fitting when it leaves at least 0.5 GB free, as on the Can I run it? pages. Speeds are bandwidth estimates for one request with the 32K context full, on one RTX PRO 6000 Blackwell (1792 GB/s), not measurements. Not sure how much VRAM you have? Detect your GPU.

gpt-oss-120b and gpt-oss-20b are listed at MXFP4, the format they are published in, in the Q4_K_M lists and not in the Q8_0 list: their GGUF quants keep the experts in MXFP4, so every type is about the same size.

Top picks for 96 GB

Pick Model Parameters Active VRAM Tokens/s (RTX PRO 6000 Blackwell)
Largest dense (Q4_K_M) Mistral Medium 3.5 128B 127.7B dense 91.8 GB 11–15
Largest MoE (Q4_K_M) Qwen3.5 122B-A10B (MoE) 125.1B 10.0B 78.8 GB 70–123
Longest context, 48 GB+ models (Q4_K_M) Qwen3.5 122B-A10B (MoE) 125.1B 10.0B 256K (full), 84.6 GB 70–123
Largest at Q8_0 AliceAI Foundation 80B-A3B (MoE) 81.3B 3.0B 89.8 GB 112–202

What is the best local LLM for 96 GB VRAM?

By size, the largest model that fits one 96 GB GPU at Q4_K_M with 32K context is Mistral Medium 3.5 128B (127.7B, 91.8 GB, 4.3 GB spare); 36 of the 55 models we track fit. This page ranks by memory and size, not by benchmark scores, which our data does not include.

Can 96 GB run a 70B model?

Yes: Llama 3.1 70B (70.6B) needs 55.2 GB at Q4_K_M with 32K context, leaving 40.8 GB on one 96 GB card, and it fits up to Q8_0 (88.3 GB). Among MoE models of 65B+ total parameters, Qwen3.5 122B-A10B (78.8 GB), Nemotron 3 Super 120B-A12B (77.4 GB), gpt-oss-120b (MXFP4, 68.6 GB) fit at Q4_K_M with 32K context.

Every step is on the Llama 3.1 70B VRAM page.

Is Q8_0 or a bigger model better on 96 GB?

On one 96 GB GPU with 32K context, the largest model at Q8_0 is AliceAI Foundation 80B-A3B (MoE, 81.3B total, 3.0B active, 89.8 GB, about 112–202 tokens/s) and at Q4_K_M it is Mistral Medium 3.5 128B (127.7B, 91.8 GB, about 11–15 tokens/s), 1.6× the total parameters. Q8_0 stores about 8.5 bits per weight and Q4_K_M 4.84; our data covers memory and speed, not output quality, so this page does not score the trade-off.

Bits per weight for every GGUF type are in GGUF quantization explained.

Newest models that fit 96 GB

Sorted by the date each model's data was added or checked, newest first (Q4_K_M, 32K context). New models appear here as they are added.

Model Added / checked Parameters Active VRAM (GB) Tokens/s
Spark-X2.5 4B 4.1B dense 4.40 186–300
Ling 3.0 Tiny 7.9B-A1.3B (MoE) 7.9B 1.3B 5.62 295–613
LFM2.5 8B-A1B (MoE) 8.5B 1.5B 6.16 254–510
Limite 1B Violetto 1.0B dense 1.62 383–761
ZDTaichu 5.0 9B 9.8B dense 7.67 116–175
Hemmingway-1 27B 27.3B dense 19.63 49–69
MiMo V2.6 Distill Qwen 9B 9.4B dense 7.43 120–181
K2-Horizon MoVA 36B-A4B (MoE) 37.4B 4.0B 30.31 56–96

All 36 models that fit 96 GB at Q4_K_M

Model Parameters Active VRAM (GB) Spare (GB) Tokens/s Calculator
Mistral Medium 3.5 128B 127.7B dense 91.75 4.25 11–15 Open
Qwen3.5 122B-A10B (MoE) 125.1B 10.0B 78.85 17.15 70–123 Open
Nemotron 3 Super 120B-A12B (MoE) 123.6B 12.0B 77.39 18.61 65–112 Open
gpt-oss-120b MXFP4 116.8B 5.1B 68.61 27.39 110–198 Open
AliceAI Foundation 80B-A3B (MoE) 81.3B 3.0B 51.71 44.29 157–292 Open
Qwen3-Coder-Next (80B MoE) 79.7B 3.0B 50.71 45.29 157–292 Open
Llama 3.1 70B 70.6B dense 55.23 40.77 18–25 Open
K2-Horizon MoVA 36B-A4B (MoE) 37.4B 4.0B 30.31 65.69 56–96 Open
Ornith 1.5 35B-A3B (MoE) 36.0B 3.0B 23.47 72.53 163–305 Open
Qwen3.6 35B-A3B (MoE) 36.0B 3.0B 23.47 72.53 163–305 Open
LLM-jp-4.1 32B-A3B Thinking (MoE) 32.1B 3.8B 22.62 73.38 102–182 Open
Nemotron 3 Nano 30B-A3B (MoE) 31.6B 3.5B 20.28 75.72 172–324 Open
Gemma 4 31B 31.3B dense 23.92 72.08 40–57 Open
GLM-4.7 Flash 31.2B 3.0B 21.67 74.33 122–222 Open
Xing 4.0 29B-A4B (MoE) 31.2B 4.0B 21.39 74.61 114–205 Open
Qwen3 30B-A3B (MoE) 30.5B 3.3B 22.72 73.28 89–158 Open
Muse Glimmer 30B 29.8B dense 19.51 76.49 49–70 Open
Qwen3.6 27B 27.8B dense 19.92 76.08 48–68 Open
Qwen3.8 27B 27.8B dense 19.92 76.08 48–68 Open
Hemmingway-1 27B 27.3B dense 19.63 76.37 49–69 Open
Gemma 4 26B-A4B (MoE) 25.8B 4.0B 17.50 78.50 128–233 Open
gpt-oss-20b MXFP4 20.9B 3.6B 15.44 80.56 134–246 Open
Gemma 4 12B 12.0B dense 8.98 87.02 101–150 Open
ZDTaichu 5.0 9B 9.8B dense 7.67 88.33 116–175 Open
Ornith 1.5 9B 9.7B dense 7.58 88.42 117–177 Open
Qwen3.5 9B 9.7B dense 7.58 88.42 117–177 Open
MiMo V2.6 Distill Qwen 9B 9.4B dense 7.43 88.57 120–181 Open
LFM2.5 8B-A1B (MoE) 8.5B 1.5B 6.16 89.84 254–510 Open
Qwen3 8B 8.2B dense 10.53 85.47 87–128 Open
Llama 3.1 8B 8.0B dense 9.88 86.12 93–137 Open
Gemma 4 E4B 8.0B dense 6.05 89.95 143–221 Open
Ling 3.0 Tiny 7.9B-A1.3B (MoE) 7.9B 1.3B 5.62 90.38 295–613 Open
Spark-X2.5 4B 4.1B dense 4.40 91.60 186–300 Open
Nemotron 3 Nano 4B 4.0B dense 3.51 92.49 223–372 Open
MiniCPM5 2B 2.5B dense 3.50 92.50 223–373 Open
Limite 1B Violetto 1.0B dense 1.62 94.38 383–761 Open

All 31 models that fit 96 GB at Q8_0

Model Parameters Active VRAM (GB) Spare (GB) Tokens/s Calculator
AliceAI Foundation 80B-A3B (MoE) 81.3B 3.0B 89.80 6.20 112–202 Open
Qwen3-Coder-Next (80B MoE) 79.7B 3.0B 88.05 7.95 112–202 Open
Llama 3.1 70B 70.6B dense 88.30 7.70 11–16 Open
K2-Horizon MoVA 36B-A4B (MoE) 37.4B 4.0B 47.86 48.14 47–80 Open
Ornith 1.5 35B-A3B (MoE) 36.0B 3.0B 40.32 55.68 115–208 Open
Qwen3.6 35B-A3B (MoE) 36.0B 3.0B 40.32 55.68 115–208 Open
LLM-jp-4.1 32B-A3B Thinking (MoE) 32.1B 3.8B 37.68 58.32 77–134 Open
Nemotron 3 Nano 30B-A3B (MoE) 31.6B 3.5B 35.08 60.92 114–205 Open
Gemma 4 31B 31.3B dense 38.58 57.42 26–36 Open
GLM-4.7 Flash 31.2B 3.0B 36.30 59.70 93–166 Open
Xing 4.0 29B-A4B (MoE) 31.2B 4.0B 36.02 59.98 82–144 Open
Qwen3 30B-A3B (MoE) 30.5B 3.3B 37.03 58.97 71–125 Open
Muse Glimmer 30B 29.8B dense 33.46 62.54 29–41 Open
Qwen3.6 27B 27.8B dense 32.94 63.06 30–42 Open
Qwen3.8 27B 27.8B dense 32.94 63.06 30–42 Open
Hemmingway-1 27B 27.3B dense 32.44 63.56 30–42 Open
Gemma 4 26B-A4B (MoE) 25.8B 4.0B 29.60 66.40 89–158 Open
Gemma 4 12B 12.0B dense 14.58 81.42 65–93 Open
ZDTaichu 5.0 9B 9.8B dense 12.26 83.74 76–111 Open
Ornith 1.5 9B 9.7B dense 12.11 83.89 77–112 Open
Qwen3.5 9B 9.7B dense 12.11 83.89 77–112 Open
MiMo V2.6 Distill Qwen 9B 9.4B dense 11.84 84.16 79–114 Open
LFM2.5 8B-A1B (MoE) 8.5B 1.5B 10.13 85.87 192–367 Open
Qwen3 8B 8.2B dense 14.37 81.63 66–95 Open
Llama 3.1 8B 8.0B dense 13.64 82.36 69–100 Open
Gemma 4 E4B 8.0B dense 9.80 86.20 93–138 Open
Ling 3.0 Tiny 7.9B-A1.3B (MoE) 7.9B 1.3B 9.32 86.68 223–436 Open
Spark-X2.5 4B 4.1B dense 6.33 89.67 137–211 Open
Nemotron 3 Nano 4B 4.0B dense 5.38 90.62 158–247 Open
MiniCPM5 2B 2.5B dense 4.68 91.32 177–283 Open
Limite 1B Violetto 1.0B dense 2.11 93.89 323–600 Open

Macs and CPU-only PCs with about 96 GB to use

These machines leave a model between 96 GB and the next size up: on a Mac, the unified memory macOS lets the GPU use by default (about two thirds up to 32 GB, 75% above; sudo sysctl iogpu.wired_limit_mb raises it), on a CPU-only PC the RAM minus about 6 GB for the system. The lists above apply to them too; speed follows each one's bandwidth, and a CPU is also much slower at reading long prompts.

Machine Usable memory Bandwidth Models at Q4_K_M (8K) Biggest at 32K Tokens/s
Ryzen AI Max+ 395 (128 GB) 96 GB 256 GB/s 36 Mistral Medium 3.5 128B 1.6–2.2
M3 Max Mac (128 GB) 96 GB 400 GB/s 36 Mistral Medium 3.5 128B 2.5–3.4
M4 Max Mac (128 GB) 96 GB 546 GB/s 36 Mistral Medium 3.5 128B 3.4–4.6
M5 Max Mac (128 GB) 96 GB 614 GB/s 36 Mistral Medium 3.5 128B 3.8–5.2
DGX Spark (128 GB) 120 GB 273 GB/s 37 Qwen3.8 Flash Next (180B MoE) 18–30
M3 Ultra Mac Studio (512 GB) 384 GB 819 GB/s 45 MiniMax M3 13–23
CPU only, DDR5-5600 dual channel (128 GB) 122 GB 89.6 GB/s 37 Qwen3.8 Flash Next (180B MoE) 6.0–11
CPU only, DDR4-3200 quad channel (128 GB) 122 GB 102.4 GB/s 37 Qwen3.8 Flash Next (180B MoE) 6.9–13
CPU only, DDR4-3200 quad channel (256 GB) 250 GB 102.4 GB/s 44 GLM-5.3 Flash 2.7–5.0

Other VRAM sizes

All numbers, including 8K and 128K context, are in the open dataset, with summary figures on LLM VRAM statistics; any other setting can be worked out in the LLM VRAM Calculator.