Can I run GLM-4.7 Flash on an RX 7900 XTX?
Yes: GLM-4.7 Flash needs about 21.7 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RX 7900 XTX with 2.33 GB to spare. Nothing more precise fits at 32K with 0.5 GB to spare; Q4_K_M runs up to 64K tokens on the RX 7900 XTX.
Yes Q4_K_M with 32K tokens of context
- Q4_K_M, 32K
- 21.7 GB
- RX 7900 XTX
- 24 GB, 960 GB/s
- To spare
- 2.33 GB
- Tokens/s
- 72–125 tokens/s
At Q4_K_M with 32K tokens of context it writes about 72–125 tokens/s for one request on an RX 7900 XTX.
Best precision for GLM-4.7 Flash on an RX 7900 XTX
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Q4_K_M | 20.3 GB | 3.69 GB | 107–192 |
| 32K | Q4_K_M | 21.7 GB | 2.33 GB | 72–125 |
| 128K | Q3_K_M | 23.4 GB | 611 MB | 32–55 |
| 198K (full) | Nothing fits | — | — | — |
GLM-4.7 Flash on the RX 7900 XTX as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 20.1 GB | 3.92 GB | 117–211 | 34.7 GB | −10.7 GB | — |
| 8K | 20.3 GB | 3.69 GB | 107–192 | 34.9 GB | −10.9 GB | — |
| 16K | 20.8 GB | 3.24 GB | 92–163 | 35.4 GB | −11.4 GB | — |
| 32K | 21.7 GB | 2.33 GB | 72–125 | 36.3 GB | −12.3 GB | — |
| 64K | 23.5 GB | 526 MB | 50–86 | 38.1 GB | −14.1 GB | — |
| 128K | 27.1 GB | −3.12 GB | — | 41.8 GB | −17.8 GB | — |
| 198K (full) | 31.1 GB | −7.10 GB | — | 45.7 GB | −21.7 GB | — |
--n-cpu-moe for GLM-4.7 Flash on the RX 7900 XTX
The measured Q4_K_M GGUF fits the RX 7900 XTX whole at 32K (19.5 GB), so --n-cpu-moe is not needed there. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.
GLM-4.7 Flash on more than one RX 7900 XTX
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s |
|---|---|---|---|---|
| 2× (48 GB) | Q8_0 | 83–155 | 198K (full) | 103–198 |
| 4× (96 GB) | BF16 | 98–187 | 198K (full) | 151–316 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 18.0 GB at 32K, 6.05 GB under the RX 7900 XTX; it fits with 0.5 GB to spare up to 129K tokens.
- Smallest setup for Q4_K_M at 32K: RTX 3090 (24 GB).
Run GLM-4.7 Flash on the RX 7900 XTX with llama-server
llama-server -hf unsloth/GLM-4.7-Flash-GGUF:Q4_K_M -c 65536 GLM-4.7-Flash-Q4_K_M.gguf, 18.3 GB, from unsloth/
Questions
Can I run GLM-4.7 Flash on an RX 7900 XTX?
Yes: GLM-4.7 Flash needs about 21.7 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RX 7900 XTX with 2.33 GB to spare. Nothing more precise fits at 32K with 0.5 GB to spare; Q4_K_M runs up to 64K tokens on the RX 7900 XTX.
How fast is GLM-4.7 Flash on an RX 7900 XTX?
At Q4_K_M with 32K tokens of context it writes about 72–125 tokens/s for one request on an RX 7900 XTX.
Does GLM-4.7 Flash need --n-cpu-moe on an RX 7900 XTX?
The measured Q4_K_M GGUF fits the RX 7900 XTX whole at 32K (19.5 GB), so --n-cpu-moe is not needed there.
What does a second RX 7900 XTX change for GLM-4.7 Flash?
Two RX 7900 XTX cards (48 GB in one tensor-parallel group) hold GLM-4.7 Flash at Q8_0 with 32K, and Q4_K_M up to 198K (full) tokens, at about 103–198 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also GLM-4.7 Flash VRAM requirements, what LLMs an RX 7900 XTX can run and every pair, or detect your own GPU. Model data checked .