Can I run Gemma 4 31B on an H100 SXM?
Yes: Gemma 4 31B needs about 23.9 GB at Q4_K_M with 32K tokens of context, which fits the 80 GB H100 SXM with 56.1 GB to spare. At 32K the H100 holds up to BF16 (68.6 GB), and Q4_K_M runs up to 256K (full) tokens.
Yes Q4_K_M with 32K tokens of context
- Q4_K_M, 32K
- 23.9 GB
- H100
- 80 GB, 3,350 GB/s
- To spare
- 56.1 GB
- Tokens/s
- 72–104 tokens/s
At Q4_K_M with 32K tokens of context it writes about 72–104 tokens/s for one request on an H100 SXM.
Best precision for Gemma 4 31B on an H100 SXM
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | BF16 | 66.6 GB | 13.4 GB | 27–38 |
| 32K | BF16 | 68.6 GB | 11.4 GB | 27–37 |
| 128K | BF16 | 76.9 GB | 3.14 GB | 24–33 |
| 256K (full) | Q8_0 | 57.8 GB | 22.2 GB | 31–44 |
Gemma 4 31B on the H100 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 21.5 GB | 58.5 GB | 79–115 | 36.2 GB | 43.8 GB | 49–70 |
| 8K | 21.9 GB | 58.1 GB | 78–114 | 36.5 GB | 43.5 GB | 49–69 |
| 16K | 22.5 GB | 57.5 GB | 76–110 | 37.2 GB | 42.8 GB | 48–68 |
| 32K | 23.9 GB | 56.1 GB | 72–104 | 38.6 GB | 41.4 GB | 46–65 |
| 64K | 26.7 GB | 53.3 GB | 65–94 | 41.3 GB | 38.7 GB | 43–61 |
| 128K | 32.2 GB | 47.8 GB | 55–78 | 46.8 GB | 33.2 GB | 38–54 |
| 256K (full) | 43.2 GB | 36.8 GB | 41–59 | 57.8 GB | 22.2 GB | 31–44 |
Gemma 4 31B on more than one H100
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s | Rent per hour |
|---|---|---|---|---|---|
| 2× (160 GB) | BF16 | 46–69 | 256K (full) | 103–171 | $6.78 |
| 4× (320 GB) | BF16 | 80–126 | 256K (full) | 151–280 | $13.56 |
| 8× (640 GB) | BF16 | 125–217 | 256K (full) | 198–410 | $27.12 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 20.2 GB at 32K, 59.8 GB under the H100; it fits with 0.5 GB to spare up to 256K (full) tokens.
- Renting an H100 SXM costs about $3.39 an hour (median on getdeploying.com, 2026-09-29): $9.0–13 per million tokens at 72–104 tokens/s.
- Smallest setup for Q4_K_M at 32K: RTX 3090 (24 GB).
Run Gemma 4 31B on the H100 with llama-server
llama-server -hf unsloth/gemma-4-31B-it-GGUF:Q4_K_M -c 262144 -np 1 gemma-4-31B-it-Q4_K_M.gguf, 18.3 GB, from unsloth/
Questions
Can I run Gemma 4 31B on an H100 SXM?
Yes: Gemma 4 31B needs about 23.9 GB at Q4_K_M with 32K tokens of context, which fits the 80 GB H100 SXM with 56.1 GB to spare. At 32K the H100 holds up to BF16 (68.6 GB), and Q4_K_M runs up to 256K (full) tokens.
How fast is Gemma 4 31B on an H100 SXM?
At Q4_K_M with 32K tokens of context it writes about 72–104 tokens/s for one request on an H100 SXM.
What does a second H100 SXM change for Gemma 4 31B?
Two H100 SXM cards (160 GB in one tensor-parallel group) hold Gemma 4 31B at BF16 with 32K, and Q4_K_M up to 256K (full) tokens, at about 103–171 tokens/s.
Try other settings in the VRAM calculator or the speed calculator. See also Gemma 4 31B VRAM requirements, what LLMs an H100 SXM can run and every pair, or detect your own GPU. Model data checked .