Can I run Gemma 4 31B on an H100 SXM?

Yes: Gemma 4 31B needs about 23.9 GB at Q4_K_M with 32K tokens of context, which fits the 80 GB H100 SXM with 56.1 GB to spare. At 32K the H100 holds up to BF16 (68.6 GB), and Q4_K_M runs up to 256K (full) tokens.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
23.9 GB
H100
80 GB, 3,350 GB/s
To spare
56.1 GB
Tokens/s
72–104 tokens/s

At Q4_K_M with 32K tokens of context it writes about 72–104 tokens/s for one request on an H100 SXM.

Best precision for Gemma 4 31B on an H100 SXM

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K BF16 66.6 GB 13.4 GB 27–38
32K BF16 68.6 GB 11.4 GB 27–37
128K BF16 76.9 GB 3.14 GB 24–33
256K (full) Q8_0 57.8 GB 22.2 GB 31–44

Gemma 4 31B on the H100 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 21.5 GB 58.5 GB 79–115 36.2 GB 43.8 GB 49–70
8K 21.9 GB 58.1 GB 78–114 36.5 GB 43.5 GB 49–69
16K 22.5 GB 57.5 GB 76–110 37.2 GB 42.8 GB 48–68
32K 23.9 GB 56.1 GB 72–104 38.6 GB 41.4 GB 46–65
64K 26.7 GB 53.3 GB 65–94 41.3 GB 38.7 GB 43–61
128K 32.2 GB 47.8 GB 55–78 46.8 GB 33.2 GB 38–54
256K (full) 43.2 GB 36.8 GB 41–59 57.8 GB 22.2 GB 31–44

Gemma 4 31B on more than one H100

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/sRent per hour
2× (160 GB) BF16 46–69 256K (full) 103–171 $6.78
4× (320 GB) BF16 80–126 256K (full) 151–280 $13.56
8× (640 GB) BF16 125–217 256K (full) 198–410 $27.12

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Gemma 4 31B on the H100 with llama-server

llama-server -hf unsloth/gemma-4-31B-it-GGUF:Q4_K_M -c 262144 -np 1

gemma-4-31B-it-Q4_K_M.gguf, 18.3 GB, from unsloth/gemma-4-31B-it-GGUF (checked 2026-09-29). At -c 262144 on the H100: 43.2 GB of 80 GB, 36.8 GB free. -np 1: one slot, one sliding window.

Questions

Can I run Gemma 4 31B on an H100 SXM?

Yes: Gemma 4 31B needs about 23.9 GB at Q4_K_M with 32K tokens of context, which fits the 80 GB H100 SXM with 56.1 GB to spare. At 32K the H100 holds up to BF16 (68.6 GB), and Q4_K_M runs up to 256K (full) tokens.

How fast is Gemma 4 31B on an H100 SXM?

At Q4_K_M with 32K tokens of context it writes about 72–104 tokens/s for one request on an H100 SXM.

What does a second H100 SXM change for Gemma 4 31B?

Two H100 SXM cards (160 GB in one tensor-parallel group) hold Gemma 4 31B at BF16 with 32K, and Q4_K_M up to 256K (full) tokens, at about 103–171 tokens/s.

Try other settings in the VRAM calculator or the speed calculator. See also Gemma 4 31B VRAM requirements, what LLMs an H100 SXM can run and every pair, or detect your own GPU. Model data checked .