Can I run Qwen3-Coder-Next on an H100 SXM?
Yes: Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, which fits the 80 GB H100 SXM with 29.3 GB to spare. At 32K the H100 holds up to Q6_K (68.3 GB), and Q4_K_M runs up to 256K (full) tokens.
Yes Q4_K_M with 32K tokens of context
- Q4_K_M, 32K
- 50.7 GB
- H100
- 80 GB, 3,350 GB/s
- To spare
- 29.3 GB
- Tokens/s
- 243–484 tokens/s
At Q4_K_M with 32K tokens of context it writes about 243–484 tokens/s for one request on an H100 SXM.
Best precision for Qwen3-Coder-Next on an H100 SXM
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Q6_K | 67.6 GB | 12.4 GB | 241–479 |
| 32K | Q6_K | 68.3 GB | 11.7 GB | 211–408 |
| 128K | Q6_K | 70.7 GB | 9.27 GB | 140–257 |
| 256K (full) | Q6_K | 74.0 GB | 5.97 GB | 97–172 |
Qwen3-Coder-Next on the H100 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 50.0 GB | 30.0 GB | 294–608 | 87.3 GB | −7.33 GB | — |
| 8K | 50.1 GB | 29.9 GB | 285–587 | 87.4 GB | −7.43 GB | — |
| 16K | 50.3 GB | 29.7 GB | 270–548 | 87.6 GB | −7.64 GB | — |
| 32K | 50.7 GB | 29.3 GB | 243–484 | 88.0 GB | −8.05 GB | — |
| 64K | 51.5 GB | 28.5 GB | 204–393 | 88.9 GB | −8.87 GB | — |
| 128K | 53.2 GB | 26.8 GB | 154–285 | 90.5 GB | −10.5 GB | — |
| 256K (full) | 56.5 GB | 23.5 GB | 103–184 | 93.8 GB | −13.8 GB | — |
Qwen3-Coder-Next on more than one H100
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s | Rent per hour |
|---|---|---|---|---|---|
| 2× (160 GB) | Q8_0 | 182–401 | 256K (full) | 208–480 | $6.78 |
| 4× (320 GB) | BF16 | 193–432 | 256K (full) | 241–591 | $13.56 |
| 8× (640 GB) | BF16 | 230–553 | 256K (full) | 261–669 | $27.12 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 41.2 GB at 32K, 38.8 GB under the H100; it fits with 0.5 GB to spare up to 256K (full) tokens.
- Renting an H100 SXM costs about $3.39 an hour (median on getdeploying.com, 2026-09-29): $1.9–3.9 per million tokens at 243–484 tokens/s.
- Smallest setup for Q4_K_M at 32K: 2× RTX 5090 (64 GB).
Run Qwen3-Coder-Next on the H100 with llama-server
llama-server -hf unsloth/Qwen3-Coder-Next-GGUF:Q4_K_M -c 262144 Qwen3-Coder-Next-Q4_K_M.gguf, 48.5 GB, from unsloth/
Questions
Can I run Qwen3-Coder-Next on an H100 SXM?
Yes: Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, which fits the 80 GB H100 SXM with 29.3 GB to spare. At 32K the H100 holds up to Q6_K (68.3 GB), and Q4_K_M runs up to 256K (full) tokens.
How fast is Qwen3-Coder-Next on an H100 SXM?
At Q4_K_M with 32K tokens of context it writes about 243–484 tokens/s for one request on an H100 SXM.
What does a second H100 SXM change for Qwen3-Coder-Next?
Two H100 SXM cards (160 GB in one tensor-parallel group) hold Qwen3-Coder-Next at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 208–480 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3-Coder-Next VRAM requirements, what LLMs an H100 SXM can run and every pair, or detect your own GPU. Model data checked .