What LLMs can an H100 SXM run?
An H100 SXM gives a model 80 GB of memory and 3,350 GB/s of bandwidth. Of the 40 open models tracked here, 23 fit at Q4_K_M with an 8,192-token context, and 2 more at a lower precision.
Models that fit on one H100 SXM
The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first.
| Model | Best precision | Memory | Longest context | Tokens/s |
|---|---|---|---|---|
| Qwen3.8 Flash Next 180.0B, 6B active | GGUF Q2_K | 77.9 GB | 88K | 238–472 |
| Mistral Medium 3.5 128B 127.7B | GGUF Q3_K_M | 67.5 GB | 41K | 27–38 |
| Qwen3.5 122B-A10B 125.1B, 10B active | GGUF Q4_K_M | 78.2 GB | 76K | 130–236 |
| Nemotron 3 Super 120B-A12B 123.6B, 12B active | GGUF Q4_K_M | 77.2 GB | 256K (full) | 114–205 |
| gpt-oss-120b 116.8B, 5.1B active | GGUF Q4_K_M | 73.2 GB | 128K (full) | 205–396 |
| Qwen3-Coder-Next 79.7B, 3B active | GGUF Q6_K | 67.6 GB | 256K (full) | 241–479 |
| Llama 3.1 70B 70.6B | FP8 / INT8 | 75.5 GB | 20K | 24–34 |
| Qwen3.6 35B-A3B 36.0B, 3B active | FP16 / BF16 | 74.3 GB | 256K (full) | 131–239 |
| Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active | FP16 / BF16 | 65.3 GB | 256K (full) | 117–212 |
| Gemma 4 31B 31.3B | FP16 / BF16 | 65.8 GB | 256K (full) | 28–39 |
| GLM-4.7 Flash 31.2B, 3B active | FP16 / BF16 | 64.9 GB | 198K (full) | 126–230 |
| Xing 4.0 29B-A4B 31.2B, 4B active | FP16 / BF16 | 64.8 GB | 256K (full) | 102–182 |
| Qwen3 30B-A3B 30.5B, 3.3B active | FP16 / BF16 | 63.9 GB | 40K (full) | 113–203 |
| Muse Glimmer 30B 29.8B | FP16 / BF16 | 61.7 GB | 128K (full) | 29–41 |
| Qwen3.6 27B 27.8B | FP16 / BF16 | 58.0 GB | 256K (full) | 31–44 |
| Qwen3.8 27B 27.8B | FP16 / BF16 | 58.0 GB | 256K (full) | 31–44 |
| Gemma 4 26B-A4B 25.8B, 4B active | FP16 / BF16 | 53.7 GB | 256K (full) | 103–183 |
| gpt-oss-20b 20.9B, 3.6B active | FP16 / BF16 | 43.6 GB | 128K (full) | 113–203 |
| Gemma 4 12B 12.0B | FP16 / BF16 | 25.4 GB | 256K (full) | 68–98 |
| Qwen3.5 9B 9.7B | FP16 / BF16 | 20.6 GB | 256K (full) | 82–121 |
| Qwen3 8B 8.2B | FP16 / BF16 | 18.5 GB | 40K (full) | 91–133 |
| Llama 3.1 8B 8.0B | FP16 / BF16 | 18.1 GB | 128K (full) | 93–137 |
| Gemma 4 E4B 8.0B | FP16 / BF16 | 17.0 GB | 128K (full) | 97–144 |
| Nemotron 3 Nano 4B 4.0B | FP16 / BF16 | 8.78 GB | 256K (full) | 170–269 |
| MiniCPM5 2B 2.5B | FP16 / BF16 | 6.02 GB | 128K (full) | 226–378 |
Too big for one card
How many H100 SXM cards these need at Q4_K_M in one tensor-parallel group; some need more than 8.
- MiniMax M2.7: 2 cards
- DeepSeek V4 Flash: 4 cards
- DeepSeek V4 Flash 0731: 4 cards
- MiMo V2.6 Flash: 4 cards
- GLM-5.3 Flash: 4 cards
- MiniMax M3: 4 cards
- DeepSeek V3 / R1: 8 cards
- DeepSeek V3.2: 8 cards
- GLM-5.3: 8 cards
- GLM-5.2: 8 cards
- DeepSeek V4.1 Flash: 8 cards
- MiMo V2.6 Pro: 8 cards
- DeepSeek V4 Pro: more than 8 cards
- Qwen3.8 2.4T-A95B: more than 8 cards
- Kimi K3: more than 8 cards
How these numbers are worked out
Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.
Other GPUs and Macs
- RTX 4060 8GB
- RTX 3060 12GB
- Arc B580 12GB
- RTX 4070 12GB
- RTX 5070 12GB
- RTX 4060 Ti 16GB
- RTX 5060 Ti 16GB
- RX 9070 XT 16GB
- RTX 4070 Ti Super 16GB
- RTX 4080 Super 16GB
- RTX 5070 Ti 16GB
- RTX 5080 16GB
- RTX 3090
- RTX 4090
- RX 7900 XTX
- RTX 5090
- M4 Pro Mac (64 GB)
- M4 Max Mac (128 GB)
- M3 Ultra Mac Studio (512 GB)
- Ryzen AI Max+ 395 (128 GB)
- DGX Spark (128 GB)
- L40S
- RTX PRO 6000 Blackwell
- A100 80GB
- H200
- B200
- 2× RTX 3060 12GB
- 2× RTX 3090
- 4× RTX 3090
- 2× RTX 4090
- 2× RTX 5090
- 8× H100 SXM
- 8× H200
Model numbers read from Hugging Face on .