What LLMs can eight H200 GPUs run?

Eight H200 GPUs give a model 1128 GB (8 × 141 GB) of memory and 38,400 GB/s of bandwidth in total. Of the 40 open models tracked here, 38 fit at Q4_K_M with an 8,192-token context, and 1 more at a lower precision. The speeds assume tensor parallelism, as in vLLM or SGLang; llama.cpp splits layers across cards by default and then writes at about the speed of one card.

Models that fit on 8× H200

The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first.

Model Best precision Memory Longest context Tokens/s
Qwen3.8 2.4T-A95B 2.45T, 95B active GGUF Q2_K 1,051 GB 256K (full) 142–293
DeepSeek V4 Pro 1.60T, 49B active GGUF Q4_K_M 993 GB 1015K 162–345
MiMo V2.6 Pro 1.02T, 42B active GGUF Q8_0 1,116 GB 169K 135–274
DeepSeek V4.1 Flash 763.2B, 16B active GGUF Q8_0 832 GB 1M (full) 199–450
GLM-5.3 753.3B, 40B active GGUF Q8_0 821 GB 1M (full) 138–281
GLM-5.2 753.3B, 40B active GGUF Q8_0 821 GB 1M (full) 138–281
DeepSeek V3.2 685.4B, 37B active GGUF Q8_0 747 GB 160K (full) 144–296
DeepSeek V3 / R1 684.5B, 37B active GGUF Q8_0 746 GB 160K (full) 144–296
MiniMax M3 427.0B, 23B active FP16 / BF16 876 GB 1M (full) 132–267
GLM-5.3 Flash 321.3B, 18B active FP16 / BF16 659 GB 1M (full) 151–314
MiMo V2.6 Flash 310.8B, 15B active FP16 / BF16 637 GB 1M (full) 163–348
DeepSeek V4 Flash 0731 304.2B, 13B active FP16 / BF16 624 GB 1M (full) 172–372
DeepSeek V4 Flash 290.9B, 13B active FP16 / BF16 597 GB 1M (full) 172–372
MiniMax M2.7 228.7B, 10B active FP16 / BF16 471 GB 200K (full) 185–408
Qwen3.8 Flash Next 180.0B, 6B active FP16 / BF16 370 GB 256K (full) 219–517
Mistral Medium 3.5 128B 127.7B FP16 / BF16 265 GB 256K (full) 64–97
Qwen3.5 122B-A10B 125.1B, 10B active FP16 / BF16 257 GB 256K (full) 190–425
Nemotron 3 Super 120B-A12B 123.6B, 12B active FP16 / BF16 254 GB 256K (full) 179–392
gpt-oss-120b 116.8B, 5.1B active FP16 / BF16 240 GB 128K (full) 227–541
Qwen3-Coder-Next 79.7B, 3B active FP16 / BF16 164 GB 256K (full) 248–616
Llama 3.1 70B 70.6B FP16 / BF16 148 GB 128K (full) 97–159
Qwen3.6 35B-A3B 36.0B, 3B active FP16 / BF16 74.3 GB 256K (full) 248–617
Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active FP16 / BF16 65.3 GB 256K (full) 243–600
Gemma 4 31B 31.3B FP16 / BF16 65.8 GB 256K (full) 153–285
GLM-4.7 Flash 31.2B, 3B active FP16 / BF16 64.9 GB 198K (full) 246–611
Xing 4.0 29B-A4B 31.2B, 4B active FP16 / BF16 64.8 GB 256K (full) 237–576
Qwen3 30B-A3B 30.5B, 3.3B active FP16 / BF16 63.9 GB 40K (full) 241–593
Muse Glimmer 30B 29.8B FP16 / BF16 61.7 GB 128K (full) 158–296
Qwen3.6 27B 27.8B FP16 / BF16 58.0 GB 256K (full) 162–308
Qwen3.8 27B 27.8B FP16 / BF16 58.0 GB 256K (full) 162–308
Gemma 4 26B-A4B 25.8B, 4B active FP16 / BF16 53.7 GB 256K (full) 237–577
gpt-oss-20b 20.9B, 3.6B active FP16 / BF16 43.6 GB 128K (full) 241–593
Gemma 4 12B 12.0B FP16 / BF16 25.4 GB 256K (full) 215–466
Qwen3.5 9B 9.7B FP16 / BF16 20.6 GB 256K (full) 226–505
Qwen3 8B 8.2B FP16 / BF16 18.5 GB 40K (full) 231–523
Llama 3.1 8B 8.0B FP16 / BF16 18.1 GB 128K (full) 232–528
Gemma 4 E4B 8.0B FP16 / BF16 17.0 GB 128K (full) 234–537
Nemotron 3 Nano 4B 4.0B FP16 / BF16 8.78 GB 256K (full) 258–633
MiniCPM5 2B 2.5B FP16 / BF16 6.02 GB 128K (full) 266–672

Too big for 8× H200

How many H200 cards these need at Q4_K_M in one tensor-parallel group; some need more than 8.

How these numbers are worked out

Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.

Other GPUs and Macs

Model numbers read from Hugging Face on .