Can I run Qwen3.8 Flash Next on an H100 SXM?

Only at Q2_K: Qwen3.8 Flash Next needs about 113 GB at Q4_K_M with 32K tokens of context, 32.9 GB more than the 80 GB H100 SXM holds, but 78.5 GB at Q2_K, which fits with 1.46 GB to spare. The smallest setup here that holds Qwen3.8 Flash Next at Q4_K_M with 32K is DGX Spark (128 GB, 120 GB usable).

Partly Q2_K with 32K tokens of context

Q4_K_M, 32K
113 GB
H100
80 GB, 3,350 GB/s
Short by
32.9 GB
Tokens/s
208–403 tokens/s

At Q2_K with 32K tokens of context it writes about 208–403 tokens/s for one request on an H100 SXM.

Best precision for Qwen3.8 Flash Next on an H100 SXM

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q2_K 77.9 GB 2.08 GB 238–472
32K Q2_K 78.5 GB 1.46 GB 208–403
128K Only IQ3_XXS, tight 79.9 GB 137 MB (tight) 140–256
256K (full) Nothing fits — — —

Qwen3.8 Flash Next on the H100 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 112 GB −32.2 GB — 197 GB −117 GB —
8K 112 GB −32.3 GB — 197 GB −117 GB —
16K 112 GB −32.5 GB — 197 GB −117 GB —
32K 113 GB −32.9 GB — 197 GB −117 GB —
64K 114 GB −33.7 GB — 198 GB −118 GB —
128K 115 GB −35.4 GB — 200 GB −120 GB —
256K (full) 119 GB −38.7 GB — 203 GB −123 GB —

Qwen3.8 Flash Next on more than one H100

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/sRent per hour
2× (160 GB) Q6_K 158–332 256K (full) 175–381 $6.78
4× (320 GB) Q8_0 189–422 256K (full) 217–510 $13.56
8× (640 GB) BF16 196–443 256K (full) 247–613 $27.12

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

llama-server command

llama-server -m Qwen3.8-Flash-Next-Q2_K.gguf -ngl 99 -c 32768

Questions

Can I run Qwen3.8 Flash Next on an H100 SXM?

Only at Q2_K: Qwen3.8 Flash Next needs about 113 GB at Q4_K_M with 32K tokens of context, 32.9 GB more than the 80 GB H100 SXM holds, but 78.5 GB at Q2_K, which fits with 1.46 GB to spare. The smallest setup here that holds Qwen3.8 Flash Next at Q4_K_M with 32K is DGX Spark (128 GB, 120 GB usable).

How fast is Qwen3.8 Flash Next on an H100 SXM?

At Q2_K with 32K tokens of context it writes about 208–403 tokens/s for one request on an H100 SXM.

What does a second H100 SXM change for Qwen3.8 Flash Next?

Two H100 SXM cards (160 GB in one tensor-parallel group) hold Qwen3.8 Flash Next at Q6_K with 32K, and Q4_K_M up to 256K (full) tokens, at about 175–381 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Qwen3.8 Flash Next VRAM requirements, what LLMs an H100 SXM can run and every pair, or detect your own GPU. Model data checked .