Can I run GLM-4.7 Flash on an RTX 4060 Ti 16GB?

With --n-cpu-moe 12: the Q4_K_M GGUF keeps 15.8 GB on the RTX 4060 Ti and 3.94 GB in system RAM, with the experts of 12 of its 47 layers moved there. Whole, GLM-4.7 Flash needs about 21.7 GB at Q4_K_M with 32K tokens of context, 5.67 GB more than the RTX 4060 Ti 16GB holds. The smallest setup here that holds GLM-4.7 Flash at Q4_K_M with 32K is RTX 3090 (24 GB).

With offload the Q4_K_M GGUF with --n-cpu-moe 12 and 32K tokens of context

On the card
15.8 GB
RTX 4060 Ti
16 GB, 288 GB/s
In RAM
3.94 GB
Tokens/s
18–30 tokens/s

With the Q4_K_M GGUF, --n-cpu-moe 12 and 32K tokens of context it writes about 18–30 tokens/s for one request on an RTX 4060 Ti 16GB and dual-channel DDR5-5600.

Best precision for GLM-4.7 Flash on an RTX 4060 Ti 16GB

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q2_K 14.3 GB 1.65 GB 47–81
32K Only Q2_K, tight 15.7 GB 296 MB (tight) 27–46
128K Nothing fits — — —
198K (full) Nothing fits — — —

GLM-4.7 Flash on the RTX 4060 Ti as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 20.1 GB −4.08 GB — 34.7 GB −18.7 GB —
8K 20.3 GB −4.31 GB — 34.9 GB −18.9 GB —
16K 20.8 GB −4.76 GB — 35.4 GB −19.4 GB —
32K 21.7 GB −5.67 GB — 36.3 GB −20.3 GB —
64K 23.5 GB −7.49 GB — 38.1 GB −22.1 GB —
128K 27.1 GB −11.1 GB — 41.8 GB −25.8 GB —
198K (full) 31.1 GB −15.1 GB — 45.7 GB −29.7 GB —

--n-cpu-moe for GLM-4.7 Flash on the RTX 4060 Ti

With --n-cpu-moe 12 the Q4_K_M GGUF keeps 15.8 GB on the RTX 4060 Ti and 3.94 GB in RAM, about 18–30 tokens/s with DDR5-5600. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 8 15.8 GB 2.62 GB 23–3926–4426–45
32K 12 15.8 GB 3.94 GB 16–2718–3018–31
64K 17 15.7 GB 5.62 GB 11–1913–2113–22
128K 27 15.7 GB 8.92 GB 7.0–127.9–138.1–14
198K (full) 38 15.7 GB 12.6 GB 5.0–8.35.6–9.45.8–9.6

GLM-4.7 Flash on more than one RTX 4060 Ti

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/s
2× (32 GB) Q6_K 36–62 196K 41–73
4× (64 GB) Q8_0 56–101 198K (full) 72–133

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run GLM-4.7 Flash on the RTX 4060 Ti with llama-server

llama-server -hf unsloth/GLM-4.7-Flash-GGUF:Q4_K_M -c 32768 --n-cpu-moe 12

GLM-4.7-Flash-Q4_K_M.gguf, 18.3 GB, from unsloth/GLM-4.7-Flash-GGUF (checked 2026-09-29); --n-cpu-moe 12 with -c 32768 is the plan above.

Questions

Can I run GLM-4.7 Flash on an RTX 4060 Ti 16GB?

With --n-cpu-moe 12: the Q4_K_M GGUF keeps 15.8 GB on the RTX 4060 Ti and 3.94 GB in system RAM, with the experts of 12 of its 47 layers moved there. Whole, GLM-4.7 Flash needs about 21.7 GB at Q4_K_M with 32K tokens of context, 5.67 GB more than the RTX 4060 Ti 16GB holds. The smallest setup here that holds GLM-4.7 Flash at Q4_K_M with 32K is RTX 3090 (24 GB).

How fast is GLM-4.7 Flash on an RTX 4060 Ti 16GB?

With the Q4_K_M GGUF, --n-cpu-moe 12 and 32K tokens of context it writes about 18–30 tokens/s for one request on an RTX 4060 Ti 16GB and dual-channel DDR5-5600.

Does GLM-4.7 Flash need --n-cpu-moe on an RTX 4060 Ti 16GB?

With --n-cpu-moe 12 the Q4_K_M GGUF keeps 15.8 GB on the RTX 4060 Ti and 3.94 GB in RAM, about 18–30 tokens/s with DDR5-5600.

What does a second RTX 4060 Ti 16GB change for GLM-4.7 Flash?

Two RTX 4060 Ti 16GB cards (32 GB in one tensor-parallel group) hold GLM-4.7 Flash at Q6_K with 32K, and Q4_K_M up to 196K tokens, at about 41–73 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also GLM-4.7 Flash VRAM requirements, what LLMs an RTX 4060 Ti 16GB can run and every pair, or detect your own GPU. Model data checked .