Can I run gpt-oss-120b on an H100 SXM?
Yes: gpt-oss-120b needs about 68.6 GB at MXFP4 with 32K tokens of context, which fits the 80 GB H100 SXM with 11.4 GB to spare. MXFP4 is the format gpt-oss-120b is published in, and on the H100 it runs up to 128K (full) tokens with 0.5 GB to spare.
Why 68.6 GB with vLLM, but 60.6 GB on the H100 with llama.cpp
The 68.6 GB counts the published checkpoint as vLLM loads it: 60.8 GB of weights, the 32K KV cache, 0.5 GB and 10% overhead. llama.cpp's measured MXFP4 GGUF is 59.0 GB, and its 587 MB of token embeddings stay in RAM. With every layer on the H100 it takes 60.6 GB at 32K with 1 GB of buffers, 19.4 GB under its 80 GB.
Yes MXFP4 with 32K tokens of context
- MXFP4, 32K
- 68.6 GB
- H100
- 80 GB, 3,350 GB/s
- To spare
- 11.4 GB
- Tokens/s
- 180–340 tokens/s
At MXFP4 with 32K tokens of context it writes about 180–340 tokens/s for one request on an H100 SXM.
Best precision for gpt-oss-120b on an H100 SXM
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | MXFP4 | 67.7 GB | 12.3 GB | 214–417 |
| 32K | MXFP4 | 68.6 GB | 11.4 GB | 180–340 |
| 128K (full) | MXFP4 | 72.3 GB | 7.68 GB | 109–196 |
gpt-oss-120b on the H100 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; the MXFP4 GGUF columns are the measured llama.cpp file with every layer on the GPU and 1 GB of buffers; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | MXFP4 | Free | Tokens/s | MXFP4 GGUF | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 67.5 GB | 12.5 GB | 222–433 | 59.6 GB | 20.4 GB | 190–363 |
| 8K | 67.7 GB | 12.3 GB | 214–417 | 59.8 GB | 20.2 GB | 185–352 |
| 16K | 68.0 GB | 12.0 GB | 201–388 | 60.0 GB | 20.0 GB | 175–331 |
| 32K | 68.6 GB | 11.4 GB | 180–340 | 60.6 GB | 19.4 GB | 159–296 |
| 64K | 69.8 GB | 10.2 GB | 148–273 | 61.7 GB | 18.3 GB | 133–244 |
| 128K (full) | 72.3 GB | 7.68 GB | 109–196 | 64.0 GB | 16.0 GB | 101–180 |
gpt-oss-120b on more than one H100
| Cards | Best at 32K | Tokens/s | MXFP4 longest | MXFP4 tokens/s | Rent per hour |
|---|---|---|---|---|---|
| 2× (160 GB) | MXFP4 | 181–397 | 128K (full) | 181–397 | $6.78 |
| 4× (320 GB) | MXFP4 | 221–524 | 128K (full) | 221–524 | $13.56 |
| 8× (640 GB) | MXFP4 | 249–623 | 128K (full) | 249–623 | $27.12 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- Renting an H100 SXM costs about $3.39 an hour (median on getdeploying.com, 2026-09-29): $2.8–5.2 per million tokens at 180–340 tokens/s.
- Smallest setup for MXFP4 at 32K: A100 80GB.
Run gpt-oss-120b on the H100 with llama-server
llama-server -hf ggml-org/gpt-oss-120b-GGUF:MXFP4 -c 131072 -np 1 gpt-oss-120b-MXFP4.gguf, 63.4 GB, from ggml-org/
Questions
Can I run gpt-oss-120b on an H100 SXM?
Yes: gpt-oss-120b needs about 68.6 GB at MXFP4 with 32K tokens of context, which fits the 80 GB H100 SXM with 11.4 GB to spare. MXFP4 is the format gpt-oss-120b is published in, and on the H100 it runs up to 128K (full) tokens with 0.5 GB to spare.
How fast is gpt-oss-120b on an H100 SXM?
At MXFP4 with 32K tokens of context it writes about 180–340 tokens/s for one request on an H100 SXM.
What does a second H100 SXM change for gpt-oss-120b?
Two H100 SXM cards (160 GB in one tensor-parallel group) hold gpt-oss-120b at MXFP4 with 32K, and MXFP4 up to 128K (full) tokens, at about 181–397 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also gpt-oss-120b VRAM requirements, what LLMs an H100 SXM can run and every pair, or detect your own GPU. Model data checked .