Can I run GLM-4.7 Flash on an RTX 5090?

Yes: GLM-4.7 Flash needs about 21.7 GB at Q4_K_M with 32K tokens of context, which fits the 32 GB RTX 5090 with 10.3 GB to spare. At 32K the RTX 5090 holds up to Q6_K (28.5 GB), and Q4_K_M runs up to 198K (full) tokens.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
21.7 GB
RTX 5090
32 GB, 1,792 GB/s
To spare
10.3 GB
Tokens/s
122–222 tokens/s

At Q4_K_M with 32K tokens of context it writes about 122–222 tokens/s for one request on an RTX 5090.

Best precision for GLM-4.7 Flash on an RTX 5090

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q6_K 27.2 GB 4.82 GB 145–267
32K Q6_K 28.5 GB 3.45 GB 107–191
128K Q5_K_M 30.4 GB 1.56 GB 54–93
198K (full) Q4_K_M 31.1 GB 924 MB 40–68

GLM-4.7 Flash on the RTX 5090 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 20.1 GB 11.9 GB 189–361 34.7 GB −2.71 GB —
8K 20.3 GB 11.7 GB 175–331 34.9 GB −2.94 GB —
16K 20.8 GB 11.2 GB 153–284 35.4 GB −3.39 GB —
32K 21.7 GB 10.3 GB 122–222 36.3 GB −4.30 GB —
64K 23.5 GB 8.51 GB 87–154 38.1 GB −6.12 GB —
128K 27.1 GB 4.88 GB 55–96 41.8 GB −9.75 GB —
198K (full) 31.1 GB 924 MB 40–68 45.7 GB −13.7 GB —

--n-cpu-moe for GLM-4.7 Flash on the RTX 5090

The measured Q4_K_M GGUF fits the RTX 5090 whole at 32K (19.5 GB), so --n-cpu-moe is not needed there. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 0 18.3 GB 170 MB 147–272147–272147–272
32K 0 19.5 GB 170 MB 108–194108–194108–194
64K 0 21.2 GB 170 MB 80–14080–14080–140
128K 0 24.5 GB 170 MB 52–9052–9052–90
198K (full) 0 28.1 GB 170 MB 38–6538–6538–65

GLM-4.7 Flash on more than one RTX 5090

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/sRent per hour
2× (64 GB) Q8_0 123–246 198K (full) 146–303 $1.38
4× (128 GB) BF16 141–288 198K (full) 193–435 $2.76

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run GLM-4.7 Flash on the RTX 5090 with llama-server

llama-server -hf unsloth/GLM-4.7-Flash-GGUF:Q4_K_M -c 202752

GLM-4.7-Flash-Q4_K_M.gguf, 18.3 GB, from unsloth/GLM-4.7-Flash-GGUF (checked 2026-09-29). At -c 202752 on the RTX 5090: 31.1 GB of 32 GB, 924 MB free.

Questions

Can I run GLM-4.7 Flash on an RTX 5090?

Yes: GLM-4.7 Flash needs about 21.7 GB at Q4_K_M with 32K tokens of context, which fits the 32 GB RTX 5090 with 10.3 GB to spare. At 32K the RTX 5090 holds up to Q6_K (28.5 GB), and Q4_K_M runs up to 198K (full) tokens.

How fast is GLM-4.7 Flash on an RTX 5090?

At Q4_K_M with 32K tokens of context it writes about 122–222 tokens/s for one request on an RTX 5090.

Does GLM-4.7 Flash need --n-cpu-moe on an RTX 5090?

The measured Q4_K_M GGUF fits the RTX 5090 whole at 32K (19.5 GB), so --n-cpu-moe is not needed there.

What does a second RTX 5090 change for GLM-4.7 Flash?

Two RTX 5090 cards (64 GB in one tensor-parallel group) hold GLM-4.7 Flash at Q8_0 with 32K, and Q4_K_M up to 198K (full) tokens, at about 146–303 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also GLM-4.7 Flash VRAM requirements, what LLMs an RTX 5090 can run and every pair, or detect your own GPU. Model data checked .