Can I run Qwen3.8 Flash Next on an H100 SXM?
Only at Q2_K: Qwen3.8 Flash Next needs about 113 GB at Q4_K_M with 32K tokens of context, 32.9 GB more than the 80 GB H100 SXM holds, but 78.5 GB at Q2_K, which fits with 1.46 GB to spare. The smallest setup here that holds Qwen3.8 Flash Next at Q4_K_M with 32K is DGX Spark (128 GB, 120 GB usable).
Partly Q2_K with 32K tokens of context
- Q4_K_M, 32K
- 113 GB
- H100
- 80 GB, 3,350 GB/s
- Short by
- 32.9 GB
- Tokens/s
- 208–403 tokens/s
At Q2_K with 32K tokens of context it writes about 208–403 tokens/s for one request on an H100 SXM.
Best precision for Qwen3.8 Flash Next on an H100 SXM
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Q2_K | 77.9 GB | 2.08 GB | 238–472 |
| 32K | Q2_K | 78.5 GB | 1.46 GB | 208–403 |
| 128K | Only IQ3_XXS, tight | 79.9 GB | 137 MB (tight) | 140–256 |
| 256K (full) | Nothing fits | — | — | — |
Qwen3.8 Flash Next on the H100 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 112 GB | −32.2 GB | — | 197 GB | −117 GB | — |
| 8K | 112 GB | −32.3 GB | — | 197 GB | −117 GB | — |
| 16K | 112 GB | −32.5 GB | — | 197 GB | −117 GB | — |
| 32K | 113 GB | −32.9 GB | — | 197 GB | −117 GB | — |
| 64K | 114 GB | −33.7 GB | — | 198 GB | −118 GB | — |
| 128K | 115 GB | −35.4 GB | — | 200 GB | −120 GB | — |
| 256K (full) | 119 GB | −38.7 GB | — | 203 GB | −123 GB | — |
Qwen3.8 Flash Next on more than one H100
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s | Rent per hour |
|---|---|---|---|---|---|
| 2× (160 GB) | Q6_K | 158–332 | 256K (full) | 175–381 | $6.78 |
| 4× (320 GB) | Q8_0 | 189–422 | 256K (full) | 217–510 | $13.56 |
| 8× (640 GB) | BF16 | 196–443 | 256K (full) | 247–613 | $27.12 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, IQ3_XXS, takes 77.4 GB at 32K, 2.61 GB under the H100; it fits with 0.5 GB to spare up to 113K tokens.
- Renting an H100 SXM costs about $3.39 an hour (median on getdeploying.com, 2026-09-29): $2.3–4.5 per million tokens at 208–403 tokens/s.
- Smallest setup for Q4_K_M at 32K: DGX Spark (128 GB, 120 GB usable).
llama-server command
llama-server -m Qwen3.8-Flash-Next-Q2_K.gguf -ngl 99 -c 32768 Questions
Can I run Qwen3.8 Flash Next on an H100 SXM?
Only at Q2_K: Qwen3.8 Flash Next needs about 113 GB at Q4_K_M with 32K tokens of context, 32.9 GB more than the 80 GB H100 SXM holds, but 78.5 GB at Q2_K, which fits with 1.46 GB to spare. The smallest setup here that holds Qwen3.8 Flash Next at Q4_K_M with 32K is DGX Spark (128 GB, 120 GB usable).
How fast is Qwen3.8 Flash Next on an H100 SXM?
At Q2_K with 32K tokens of context it writes about 208–403 tokens/s for one request on an H100 SXM.
What does a second H100 SXM change for Qwen3.8 Flash Next?
Two H100 SXM cards (160 GB in one tensor-parallel group) hold Qwen3.8 Flash Next at Q6_K with 32K, and Q4_K_M up to 256K (full) tokens, at about 175–381 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3.8 Flash Next VRAM requirements, what LLMs an H100 SXM can run and every pair, or detect your own GPU. Model data checked .