Can I run GLM-4.7 Flash on an RX 7900 XTX?

Yes: GLM-4.7 Flash needs about 21.7 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RX 7900 XTX with 2.33 GB to spare. Nothing more precise fits at 32K with 0.5 GB to spare; Q4_K_M runs up to 64K tokens on the RX 7900 XTX.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
21.7 GB
RX 7900 XTX
24 GB, 960 GB/s
To spare
2.33 GB
Tokens/s
72–125 tokens/s

At Q4_K_M with 32K tokens of context it writes about 72–125 tokens/s for one request on an RX 7900 XTX.

Best precision for GLM-4.7 Flash on an RX 7900 XTX

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q4_K_M 20.3 GB 3.69 GB 107–192
32K Q4_K_M 21.7 GB 2.33 GB 72–125
128K Q3_K_M 23.4 GB 611 MB 32–55
198K (full) Nothing fits — — —

GLM-4.7 Flash on the RX 7900 XTX as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 20.1 GB 3.92 GB 117–211 34.7 GB −10.7 GB —
8K 20.3 GB 3.69 GB 107–192 34.9 GB −10.9 GB —
16K 20.8 GB 3.24 GB 92–163 35.4 GB −11.4 GB —
32K 21.7 GB 2.33 GB 72–125 36.3 GB −12.3 GB —
64K 23.5 GB 526 MB 50–86 38.1 GB −14.1 GB —
128K 27.1 GB −3.12 GB — 41.8 GB −17.8 GB —
198K (full) 31.1 GB −7.10 GB — 45.7 GB −21.7 GB —

--n-cpu-moe for GLM-4.7 Flash on the RX 7900 XTX

The measured Q4_K_M GGUF fits the RX 7900 XTX whole at 32K (19.5 GB), so --n-cpu-moe is not needed there. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 0 18.3 GB 170 MB 88–15688–15688–156
32K 0 19.5 GB 170 MB 62–10962–10962–109
64K 0 21.2 GB 170 MB 45–7845–7845–78
128K 3 23.8 GB 917 MB 27–4528–4728–47
198K (full) 14 23.7 GB 4.62 GB 15–2517–2918–30

GLM-4.7 Flash on more than one RX 7900 XTX

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/s
2× (48 GB) Q8_0 83–155 198K (full) 103–198
4× (96 GB) BF16 98–187 198K (full) 151–316

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run GLM-4.7 Flash on the RX 7900 XTX with llama-server

llama-server -hf unsloth/GLM-4.7-Flash-GGUF:Q4_K_M -c 65536

GLM-4.7-Flash-Q4_K_M.gguf, 18.3 GB, from unsloth/GLM-4.7-Flash-GGUF (checked 2026-09-29). At -c 65536 on the RX 7900 XTX: 23.5 GB of 24 GB, 526 MB free.

Questions

Can I run GLM-4.7 Flash on an RX 7900 XTX?

Yes: GLM-4.7 Flash needs about 21.7 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RX 7900 XTX with 2.33 GB to spare. Nothing more precise fits at 32K with 0.5 GB to spare; Q4_K_M runs up to 64K tokens on the RX 7900 XTX.

How fast is GLM-4.7 Flash on an RX 7900 XTX?

At Q4_K_M with 32K tokens of context it writes about 72–125 tokens/s for one request on an RX 7900 XTX.

Does GLM-4.7 Flash need --n-cpu-moe on an RX 7900 XTX?

The measured Q4_K_M GGUF fits the RX 7900 XTX whole at 32K (19.5 GB), so --n-cpu-moe is not needed there.

What does a second RX 7900 XTX change for GLM-4.7 Flash?

Two RX 7900 XTX cards (48 GB in one tensor-parallel group) hold GLM-4.7 Flash at Q8_0 with 32K, and Q4_K_M up to 198K (full) tokens, at about 103–198 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also GLM-4.7 Flash VRAM requirements, what LLMs an RX 7900 XTX can run and every pair, or detect your own GPU. Model data checked .