Can I run Qwen3-Coder-Next on an H100 SXM?

Yes: Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, which fits the 80 GB H100 SXM with 29.3 GB to spare. At 32K the H100 holds up to Q6_K (68.3 GB), and Q4_K_M runs up to 256K (full) tokens.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
50.7 GB
H100
80 GB, 3,350 GB/s
To spare
29.3 GB
Tokens/s
243–484 tokens/s

At Q4_K_M with 32K tokens of context it writes about 243–484 tokens/s for one request on an H100 SXM.

Best precision for Qwen3-Coder-Next on an H100 SXM

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q6_K 67.6 GB 12.4 GB 241–479
32K Q6_K 68.3 GB 11.7 GB 211–408
128K Q6_K 70.7 GB 9.27 GB 140–257
256K (full) Q6_K 74.0 GB 5.97 GB 97–172

Qwen3-Coder-Next on the H100 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 50.0 GB 30.0 GB 294–608 87.3 GB −7.33 GB —
8K 50.1 GB 29.9 GB 285–587 87.4 GB −7.43 GB —
16K 50.3 GB 29.7 GB 270–548 87.6 GB −7.64 GB —
32K 50.7 GB 29.3 GB 243–484 88.0 GB −8.05 GB —
64K 51.5 GB 28.5 GB 204–393 88.9 GB −8.87 GB —
128K 53.2 GB 26.8 GB 154–285 90.5 GB −10.5 GB —
256K (full) 56.5 GB 23.5 GB 103–184 93.8 GB −13.8 GB —

Qwen3-Coder-Next on more than one H100

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/sRent per hour
2× (160 GB) Q8_0 182–401 256K (full) 208–480 $6.78
4× (320 GB) BF16 193–432 256K (full) 241–591 $13.56
8× (640 GB) BF16 230–553 256K (full) 261–669 $27.12

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Qwen3-Coder-Next on the H100 with llama-server

llama-server -hf unsloth/Qwen3-Coder-Next-GGUF:Q4_K_M -c 262144

Qwen3-Coder-Next-Q4_K_M.gguf, 48.5 GB, from unsloth/Qwen3-Coder-Next-GGUF (checked 2026-09-29). At -c 262144 on the H100: 56.8 GB of 80 GB, 23.2 GB free. The file is 310 MB over the estimate above, so -c counts the file.

Questions

Can I run Qwen3-Coder-Next on an H100 SXM?

Yes: Qwen3-Coder-Next needs about 50.7 GB at Q4_K_M with 32K tokens of context, which fits the 80 GB H100 SXM with 29.3 GB to spare. At 32K the H100 holds up to Q6_K (68.3 GB), and Q4_K_M runs up to 256K (full) tokens.

How fast is Qwen3-Coder-Next on an H100 SXM?

At Q4_K_M with 32K tokens of context it writes about 243–484 tokens/s for one request on an H100 SXM.

What does a second H100 SXM change for Qwen3-Coder-Next?

Two H100 SXM cards (160 GB in one tensor-parallel group) hold Qwen3-Coder-Next at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 208–480 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3-Coder-Next VRAM requirements, what LLMs an H100 SXM can run and every pair, or detect your own GPU. Model data checked .