Can I run Llama 3.1 70B on an H100 SXM?

Yes: Llama 3.1 70B needs about 55.2 GB at Q4_K_M with 32K tokens of context, which fits the 80 GB H100 SXM with 24.8 GB to spare. At 32K the H100 holds up to Q6_K (70.8 GB), and Q4_K_M runs up to 102K tokens.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
55.2 GB
H100
80 GB, 3,350 GB/s
To spare
24.8 GB
Tokens/s
33–46 tokens/s

At Q4_K_M with 32K tokens of context it writes about 33–46 tokens/s for one request on an H100 SXM.

Best precision for Llama 3.1 70B on an H100 SXM

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K FP8 75.5 GB 4.47 GB 24–34
32K Q6_K 70.8 GB 9.23 GB 26–36
128K (full) Q2_K 74.8 GB 5.23 GB 24–34

Llama 3.1 70B on the H100 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 45.6 GB 34.4 GB 39–55 78.7 GB 1.33 GB 23–32
8K 47.0 GB 33.0 GB 38–54 80.0 GB −48 MB —
16K 49.7 GB 30.3 GB 36–51 82.8 GB −2.80 GB —
32K 55.2 GB 24.8 GB 33–46 88.3 GB −8.30 GB —
64K 66.2 GB 13.8 GB 28–38 99.3 GB −19.3 GB —
128K (full) 88.2 GB −8.23 GB — 121 GB −41.3 GB —

Llama 3.1 70B on more than one H100

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/sRent per hour
2× (160 GB) BF16 22–32 128K (full) 56–84 $6.78
4× (320 GB) BF16 41–61 128K (full) 93–151 $13.56
8× (640 GB) BF16 72–113 128K (full) 140–253 $27.12

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Llama 3.1 70B on the H100 with llama-server

llama-server -hf bartowski/Meta-Llama-3.1-70B-Instruct-GGUF:Q4_K_M -c 104448

Meta-Llama-3.1-70B-Instruct-Q4_K_M.gguf, 42.5 GB, from bartowski/Meta-Llama-3.1-70B-Instruct-GGUF (checked 2026-09-29). At -c 104448 on the H100: 79.3 GB of 80 GB, 726 MB free.

Questions

Can I run Llama 3.1 70B on an H100 SXM?

Yes: Llama 3.1 70B needs about 55.2 GB at Q4_K_M with 32K tokens of context, which fits the 80 GB H100 SXM with 24.8 GB to spare. At 32K the H100 holds up to Q6_K (70.8 GB), and Q4_K_M runs up to 102K tokens.

How fast is Llama 3.1 70B on an H100 SXM?

At Q4_K_M with 32K tokens of context it writes about 33–46 tokens/s for one request on an H100 SXM.

What does a second H100 SXM change for Llama 3.1 70B?

Two H100 SXM cards (160 GB in one tensor-parallel group) hold Llama 3.1 70B at BF16 with 32K, and Q4_K_M up to 128K (full) tokens, at about 56–84 tokens/s.

Try other settings in the VRAM calculator or the speed calculator. See also Llama 3.1 70B VRAM requirements, what LLMs an H100 SXM can run and every pair, or detect your own GPU. Model data checked .