Can I run Gemma 4 26B-A4B on an RTX 4070 12GB?

With --n-cpu-moe 12: the UD-Q4_K_M GGUF keeps 11.6 GB on the RTX 4070 and 6.05 GB in system RAM, with the experts of 12 of its 30 layers moved there. Whole, Gemma 4 26B-A4B needs about 17.5 GB at Q4_K_M with 32K tokens of context, 5.50 GB more than the RTX 4070 12GB holds. The smallest setup here that holds Gemma 4 26B-A4B at Q4_K_M with 32K is RTX 3090 (24 GB).

With offload the UD-Q4_K_M GGUF with --n-cpu-moe 12 and 32K tokens of context

On the card
11.6 GB
RTX 4070
12 GB, 504 GB/s
In RAM
6.05 GB
Tokens/s
27–46 tokens/s

With the UD-Q4_K_M GGUF, --n-cpu-moe 12 and 32K tokens of context it writes about 27–46 tokens/s for one request on an RTX 4070 12GB and dual-channel DDR5-5600.

Best precision for Gemma 4 26B-A4B on an RTX 4070 12GB

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Only IQ3_XXS, tight 11.9 GB 103 MB (tight) 64–112
32K Nothing fits — — —
128K Nothing fits — — —
256K (full) Nothing fits — — —

Gemma 4 26B-A4B on the RTX 4070 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 16.9 GB −4.90 GB — 29.0 GB −17.0 GB —
8K 17.0 GB −4.99 GB — 29.1 GB −17.1 GB —
16K 17.2 GB −5.16 GB — 29.3 GB −17.3 GB —
32K 17.5 GB −5.50 GB — 29.6 GB −17.6 GB —
64K 18.2 GB −6.19 GB — 30.3 GB −18.3 GB —
128K 19.6 GB −7.57 GB — 31.7 GB −19.7 GB —
256K (full) 22.3 GB −10.3 GB — 34.4 GB −22.4 GB —

--n-cpu-moe for Gemma 4 26B-A4B on the RTX 4070

With --n-cpu-moe 12 the UD-Q4_K_M GGUF keeps 11.6 GB on the RTX 4070 and 6.05 GB in RAM, about 27–46 tokens/s with DDR5-5600. Measured UD-Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 11 11.6 GB 5.60 GB 24–4131–5232–55
32K 12 11.6 GB 6.05 GB 21–3627–4629–48
64K 13 11.8 GB 6.49 GB 19–3224–4025–42
128K 16 11.7 GB 7.82 GB 15–2519–3119–33
256K (full) 22 11.6 GB 10.5 GB 11–1813–2214–23

Gemma 4 26B-A4B on more than one RTX 4070

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/s
2× (24 GB) Q5_K_M 62–113 256K (full) 68–124
4× (48 GB) Q8_0 82–154 256K (full) 110–214

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Gemma 4 26B-A4B on the RTX 4070 with llama-server

llama-server -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_M -c 32768 --n-cpu-moe 12 -np 1

gemma-4-26B-A4B-it-UD-Q4_K_M.gguf, 16.9 GB, from unsloth/gemma-4-26B-A4B-it-GGUF (checked 2026-09-29); --n-cpu-moe 12 with -c 32768 is the plan above. -np 1: one slot, one sliding window.

Questions

Can I run Gemma 4 26B-A4B on an RTX 4070 12GB?

With --n-cpu-moe 12: the UD-Q4_K_M GGUF keeps 11.6 GB on the RTX 4070 and 6.05 GB in system RAM, with the experts of 12 of its 30 layers moved there. Whole, Gemma 4 26B-A4B needs about 17.5 GB at Q4_K_M with 32K tokens of context, 5.50 GB more than the RTX 4070 12GB holds. The smallest setup here that holds Gemma 4 26B-A4B at Q4_K_M with 32K is RTX 3090 (24 GB).

How fast is Gemma 4 26B-A4B on an RTX 4070 12GB?

With the UD-Q4_K_M GGUF, --n-cpu-moe 12 and 32K tokens of context it writes about 27–46 tokens/s for one request on an RTX 4070 12GB and dual-channel DDR5-5600.

Does Gemma 4 26B-A4B need --n-cpu-moe on an RTX 4070 12GB?

With --n-cpu-moe 12 the UD-Q4_K_M GGUF keeps 11.6 GB on the RTX 4070 and 6.05 GB in RAM, about 27–46 tokens/s with DDR5-5600.

What does a second RTX 4070 12GB change for Gemma 4 26B-A4B?

Two RTX 4070 12GB cards (24 GB in one tensor-parallel group) hold Gemma 4 26B-A4B at Q5_K_M with 32K, and Q4_K_M up to 256K (full) tokens, at about 68–124 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Gemma 4 26B-A4B VRAM requirements, what LLMs an RTX 4070 12GB can run and every pair, or detect your own GPU. Model data checked .