What LLMs can an H100 SXM run?

An H100 SXM gives a model 80 GB of memory and 3,350 GB/s of bandwidth. Of the 40 open models tracked here, 23 fit at Q4_K_M with an 8,192-token context, and 2 more at a lower precision.

Models that fit on one H100 SXM

The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first.

Model Best precision Memory Longest context Tokens/s
Qwen3.8 Flash Next 180.0B, 6B active GGUF Q2_K 77.9 GB 88K 238–472
Mistral Medium 3.5 128B 127.7B GGUF Q3_K_M 67.5 GB 41K 27–38
Qwen3.5 122B-A10B 125.1B, 10B active GGUF Q4_K_M 78.2 GB 76K 130–236
Nemotron 3 Super 120B-A12B 123.6B, 12B active GGUF Q4_K_M 77.2 GB 256K (full) 114–205
gpt-oss-120b 116.8B, 5.1B active GGUF Q4_K_M 73.2 GB 128K (full) 205–396
Qwen3-Coder-Next 79.7B, 3B active GGUF Q6_K 67.6 GB 256K (full) 241–479
Llama 3.1 70B 70.6B FP8 / INT8 75.5 GB 20K 24–34
Qwen3.6 35B-A3B 36.0B, 3B active FP16 / BF16 74.3 GB 256K (full) 131–239
Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active FP16 / BF16 65.3 GB 256K (full) 117–212
Gemma 4 31B 31.3B FP16 / BF16 65.8 GB 256K (full) 28–39
GLM-4.7 Flash 31.2B, 3B active FP16 / BF16 64.9 GB 198K (full) 126–230
Xing 4.0 29B-A4B 31.2B, 4B active FP16 / BF16 64.8 GB 256K (full) 102–182
Qwen3 30B-A3B 30.5B, 3.3B active FP16 / BF16 63.9 GB 40K (full) 113–203
Muse Glimmer 30B 29.8B FP16 / BF16 61.7 GB 128K (full) 29–41
Qwen3.6 27B 27.8B FP16 / BF16 58.0 GB 256K (full) 31–44
Qwen3.8 27B 27.8B FP16 / BF16 58.0 GB 256K (full) 31–44
Gemma 4 26B-A4B 25.8B, 4B active FP16 / BF16 53.7 GB 256K (full) 103–183
gpt-oss-20b 20.9B, 3.6B active FP16 / BF16 43.6 GB 128K (full) 113–203
Gemma 4 12B 12.0B FP16 / BF16 25.4 GB 256K (full) 68–98
Qwen3.5 9B 9.7B FP16 / BF16 20.6 GB 256K (full) 82–121
Qwen3 8B 8.2B FP16 / BF16 18.5 GB 40K (full) 91–133
Llama 3.1 8B 8.0B FP16 / BF16 18.1 GB 128K (full) 93–137
Gemma 4 E4B 8.0B FP16 / BF16 17.0 GB 128K (full) 97–144
Nemotron 3 Nano 4B 4.0B FP16 / BF16 8.78 GB 256K (full) 170–269
MiniCPM5 2B 2.5B FP16 / BF16 6.02 GB 128K (full) 226–378

Too big for one card

How many H100 SXM cards these need at Q4_K_M in one tensor-parallel group; some need more than 8.

How these numbers are worked out

Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.

Other GPUs and Macs

Model numbers read from Hugging Face on .