Can I run GLM-4.7 Flash on an RTX 5080 16GB?

With --n-cpu-moe 12: the Q4_K_M GGUF keeps 15.8 GB on the RTX 5080 and 3.94 GB in system RAM, with the experts of 12 of its 47 layers moved there. Whole, GLM-4.7 Flash needs about 21.7 GB at Q4_K_M with 32K tokens of context, 5.67 GB more than the RTX 5080 16GB holds. The smallest setup here that holds GLM-4.7 Flash at Q4_K_M with 32K is RTX 3090 (24 GB).

With offload the Q4_K_M GGUF with --n-cpu-moe 12 and 32K tokens of context

On the card
15.8 GB
RTX 5080
16 GB, 960 GB/s
In RAM
3.94 GB
Tokens/s
41–70 tokens/s

With the Q4_K_M GGUF, --n-cpu-moe 12 and 32K tokens of context it writes about 41–70 tokens/s for one request on an RTX 5080 16GB and dual-channel DDR5-5600.

Best precision for GLM-4.7 Flash on an RTX 5080 16GB

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q2_K 14.3 GB 1.65 GB 135–247
32K Only Q2_K, tight 15.7 GB 296 MB (tight) 83–147
128K Nothing fits — — —
198K (full) Nothing fits — — —

GLM-4.7 Flash on the RTX 5080 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 20.1 GB −4.08 GB — 34.7 GB −18.7 GB —
8K 20.3 GB −4.31 GB — 34.9 GB −18.9 GB —
16K 20.8 GB −4.76 GB — 35.4 GB −19.4 GB —
32K 21.7 GB −5.67 GB — 36.3 GB −20.3 GB —
64K 23.5 GB −7.49 GB — 38.1 GB −22.1 GB —
128K 27.1 GB −11.1 GB — 41.8 GB −25.8 GB —
198K (full) 31.1 GB −15.1 GB — 45.7 GB −29.7 GB —

--n-cpu-moe for GLM-4.7 Flash on the RTX 5080

With --n-cpu-moe 12 the Q4_K_M GGUF keeps 15.8 GB on the RTX 5080 and 3.94 GB in RAM, about 41–70 tokens/s with DDR5-5600. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 8 15.8 GB 2.62 GB 46–8059–10262–107
32K 12 15.8 GB 3.94 GB 32–5441–7043–73
64K 17 15.7 GB 5.62 GB 22–3829–4930–52
128K 27 15.7 GB 8.92 GB 14–2418–3119–33
198K (full) 38 15.7 GB 12.6 GB 10–1713–2214–23

GLM-4.7 Flash on more than one RTX 5080

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/s
2× (32 GB) Q6_K 92–175 196K 103–198
4× (64 GB) Q8_0 128–257 198K (full) 151–316

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run GLM-4.7 Flash on the RTX 5080 with llama-server

llama-server -hf unsloth/GLM-4.7-Flash-GGUF:Q4_K_M -c 32768 --n-cpu-moe 12

GLM-4.7-Flash-Q4_K_M.gguf, 18.3 GB, from unsloth/GLM-4.7-Flash-GGUF (checked 2026-09-29); --n-cpu-moe 12 with -c 32768 is the plan above.

Questions

Can I run GLM-4.7 Flash on an RTX 5080 16GB?

With --n-cpu-moe 12: the Q4_K_M GGUF keeps 15.8 GB on the RTX 5080 and 3.94 GB in system RAM, with the experts of 12 of its 47 layers moved there. Whole, GLM-4.7 Flash needs about 21.7 GB at Q4_K_M with 32K tokens of context, 5.67 GB more than the RTX 5080 16GB holds. The smallest setup here that holds GLM-4.7 Flash at Q4_K_M with 32K is RTX 3090 (24 GB).

How fast is GLM-4.7 Flash on an RTX 5080 16GB?

With the Q4_K_M GGUF, --n-cpu-moe 12 and 32K tokens of context it writes about 41–70 tokens/s for one request on an RTX 5080 16GB and dual-channel DDR5-5600.

Does GLM-4.7 Flash need --n-cpu-moe on an RTX 5080 16GB?

With --n-cpu-moe 12 the Q4_K_M GGUF keeps 15.8 GB on the RTX 5080 and 3.94 GB in RAM, about 41–70 tokens/s with DDR5-5600.

What does a second RTX 5080 16GB change for GLM-4.7 Flash?

Two RTX 5080 16GB cards (32 GB in one tensor-parallel group) hold GLM-4.7 Flash at Q6_K with 32K, and Q4_K_M up to 196K tokens, at about 103–198 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also GLM-4.7 Flash VRAM requirements, what LLMs an RTX 5080 16GB can run and every pair, or detect your own GPU. Model data checked .