Can I run GLM-4.7 Flash on an RTX 5080 16GB?
With --n-cpu-moe 12: the Q4_K_M GGUF keeps 15.8 GB on the RTX 5080 and 3.94 GB in system RAM, with the experts of 12 of its 47 layers moved there. Whole, GLM-4.7 Flash needs about 21.7 GB at Q4_K_M with 32K tokens of context, 5.67 GB more than the RTX 5080 16GB holds. The smallest setup here that holds GLM-4.7 Flash at Q4_K_M with 32K is RTX 3090 (24 GB).
With offload the Q4_K_M GGUF with --n-cpu-moe 12 and 32K tokens of context
- On the card
- 15.8 GB
- RTX 5080
- 16 GB, 960 GB/s
- In RAM
- 3.94 GB
- Tokens/s
- 41–70 tokens/s
With the Q4_K_M GGUF, --n-cpu-moe 12 and 32K tokens of context it writes about 41–70 tokens/s for one request on an RTX 5080 16GB and dual-channel DDR5-5600.
Best precision for GLM-4.7 Flash on an RTX 5080 16GB
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Q2_K | 14.3 GB | 1.65 GB | 135–247 |
| 32K | Only Q2_K, tight | 15.7 GB | 296 MB (tight) | 83–147 |
| 128K | Nothing fits | — | — | — |
| 198K (full) | Nothing fits | — | — | — |
GLM-4.7 Flash on the RTX 5080 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 20.1 GB | −4.08 GB | — | 34.7 GB | −18.7 GB | — |
| 8K | 20.3 GB | −4.31 GB | — | 34.9 GB | −18.9 GB | — |
| 16K | 20.8 GB | −4.76 GB | — | 35.4 GB | −19.4 GB | — |
| 32K | 21.7 GB | −5.67 GB | — | 36.3 GB | −20.3 GB | — |
| 64K | 23.5 GB | −7.49 GB | — | 38.1 GB | −22.1 GB | — |
| 128K | 27.1 GB | −11.1 GB | — | 41.8 GB | −25.8 GB | — |
| 198K (full) | 31.1 GB | −15.1 GB | — | 45.7 GB | −29.7 GB | — |
--n-cpu-moe for GLM-4.7 Flash on the RTX 5080
With --n-cpu-moe 12 the Q4_K_M GGUF keeps 15.8 GB on the RTX 5080 and 3.94 GB in RAM, about 41–70 tokens/s with DDR5-5600. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.
GLM-4.7 Flash on more than one RTX 5080
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s |
|---|---|---|---|---|
| 2× (32 GB) | Q6_K | 92–175 | 196K | 103–198 |
| 4× (64 GB) | Q8_0 | 128–257 | 198K (full) | 151–316 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 18.0 GB at 32K, 1.95 GB over the RTX 5080; it does not fit with 0.5 GB to spare even with 1K tokens.
- Smallest setup for Q4_K_M at 32K: RTX 3090 (24 GB).
Run GLM-4.7 Flash on the RTX 5080 with llama-server
llama-server -hf unsloth/GLM-4.7-Flash-GGUF:Q4_K_M -c 32768 --n-cpu-moe 12 GLM-4.7-Flash-Q4_K_M.gguf, 18.3 GB, from unsloth/
Questions
Can I run GLM-4.7 Flash on an RTX 5080 16GB?
With --n-cpu-moe 12: the Q4_K_M GGUF keeps 15.8 GB on the RTX 5080 and 3.94 GB in system RAM, with the experts of 12 of its 47 layers moved there. Whole, GLM-4.7 Flash needs about 21.7 GB at Q4_K_M with 32K tokens of context, 5.67 GB more than the RTX 5080 16GB holds. The smallest setup here that holds GLM-4.7 Flash at Q4_K_M with 32K is RTX 3090 (24 GB).
How fast is GLM-4.7 Flash on an RTX 5080 16GB?
With the Q4_K_M GGUF, --n-cpu-moe 12 and 32K tokens of context it writes about 41–70 tokens/s for one request on an RTX 5080 16GB and dual-channel DDR5-5600.
Does GLM-4.7 Flash need --n-cpu-moe on an RTX 5080 16GB?
With --n-cpu-moe 12 the Q4_K_M GGUF keeps 15.8 GB on the RTX 5080 and 3.94 GB in RAM, about 41–70 tokens/s with DDR5-5600.
What does a second RTX 5080 16GB change for GLM-4.7 Flash?
Two RTX 5080 16GB cards (32 GB in one tensor-parallel group) hold GLM-4.7 Flash at Q6_K with 32K, and Q4_K_M up to 196K tokens, at about 103–198 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also GLM-4.7 Flash VRAM requirements, what LLMs an RTX 5080 16GB can run and every pair, or detect your own GPU. Model data checked .