Can I run Llama 3.1 70B on an H100 SXM?
Yes: Llama 3.1 70B needs about 55.2 GB at Q4_K_M with 32K tokens of context, which fits the 80 GB H100 SXM with 24.8 GB to spare. At 32K the H100 holds up to Q6_K (70.8 GB), and Q4_K_M runs up to 102K tokens.
Yes Q4_K_M with 32K tokens of context
- Q4_K_M, 32K
- 55.2 GB
- H100
- 80 GB, 3,350 GB/s
- To spare
- 24.8 GB
- Tokens/s
- 33–46 tokens/s
At Q4_K_M with 32K tokens of context it writes about 33–46 tokens/s for one request on an H100 SXM.
Best precision for Llama 3.1 70B on an H100 SXM
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | FP8 | 75.5 GB | 4.47 GB | 24–34 |
| 32K | Q6_K | 70.8 GB | 9.23 GB | 26–36 |
| 128K (full) | Q2_K | 74.8 GB | 5.23 GB | 24–34 |
Llama 3.1 70B on the H100 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 45.6 GB | 34.4 GB | 39–55 | 78.7 GB | 1.33 GB | 23–32 |
| 8K | 47.0 GB | 33.0 GB | 38–54 | 80.0 GB | −48 MB | — |
| 16K | 49.7 GB | 30.3 GB | 36–51 | 82.8 GB | −2.80 GB | — |
| 32K | 55.2 GB | 24.8 GB | 33–46 | 88.3 GB | −8.30 GB | — |
| 64K | 66.2 GB | 13.8 GB | 28–38 | 99.3 GB | −19.3 GB | — |
| 128K (full) | 88.2 GB | −8.23 GB | — | 121 GB | −41.3 GB | — |
Llama 3.1 70B on more than one H100
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s | Rent per hour |
|---|---|---|---|---|---|
| 2× (160 GB) | BF16 | 22–32 | 128K (full) | 56–84 | $6.78 |
| 4× (320 GB) | BF16 | 41–61 | 128K (full) | 93–151 | $13.56 |
| 8× (640 GB) | BF16 | 72–113 | 128K (full) | 140–253 | $27.12 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 46.8 GB at 32K, 33.2 GB under the H100; it fits with 0.5 GB to spare up to 127K tokens.
- Renting an H100 SXM costs about $3.39 an hour (median on getdeploying.com, 2026-09-29): $20–29 per million tokens at 33–46 tokens/s.
- Smallest setup for Q4_K_M at 32K: 2× RTX 5090 (64 GB).
Run Llama 3.1 70B on the H100 with llama-server
llama-server -hf bartowski/Meta-Llama-3.1-70B-Instruct-GGUF:Q4_K_M -c 104448 Meta-Llama-3.1-70B-Instruct-Q4_K_M.gguf, 42.5 GB, from bartowski/
Questions
Can I run Llama 3.1 70B on an H100 SXM?
Yes: Llama 3.1 70B needs about 55.2 GB at Q4_K_M with 32K tokens of context, which fits the 80 GB H100 SXM with 24.8 GB to spare. At 32K the H100 holds up to Q6_K (70.8 GB), and Q4_K_M runs up to 102K tokens.
How fast is Llama 3.1 70B on an H100 SXM?
At Q4_K_M with 32K tokens of context it writes about 33–46 tokens/s for one request on an H100 SXM.
What does a second H100 SXM change for Llama 3.1 70B?
Two H100 SXM cards (160 GB in one tensor-parallel group) hold Llama 3.1 70B at BF16 with 32K, and Q4_K_M up to 128K (full) tokens, at about 56–84 tokens/s.
Try other settings in the VRAM calculator or the speed calculator. See also Llama 3.1 70B VRAM requirements, what LLMs an H100 SXM can run and every pair, or detect your own GPU. Model data checked .