Can I run Gemma 4 26B-A4B on an RTX 4090?

Yes: Gemma 4 26B-A4B needs about 17.5 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RTX 4090 with 6.50 GB to spare. At 32K the RTX 4090 holds up to Q6_K (23.2 GB), and Q4_K_M runs up to 256K (full) tokens.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
17.5 GB
RTX 4090
24 GB, 1,008 GB/s
To spare
6.50 GB
Tokens/s
78–138 tokens/s

At Q4_K_M with 32K tokens of context it writes about 78–138 tokens/s for one request on an RTX 4090.

Best precision for Gemma 4 26B-A4B on an RTX 4090

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q6_K 22.7 GB 1.33 GB 72–126
32K Q6_K 23.2 GB 831 MB 64–112
128K Q5_K_M 22.3 GB 1.69 GB 48–83
256K (full) Q4_K_M 22.3 GB 1.68 GB 35–60

Gemma 4 26B-A4B on the RTX 4090 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 16.9 GB 7.10 GB 92–164 29.0 GB −5.00 GB —
8K 17.0 GB 7.01 GB 90–160 29.1 GB −5.08 GB —
16K 17.2 GB 6.84 GB 86–152 29.3 GB −5.26 GB —
32K 17.5 GB 6.50 GB 78–138 29.6 GB −5.60 GB —
64K 18.2 GB 5.81 GB 67–116 30.3 GB −6.29 GB —
128K 19.6 GB 4.43 GB 51–89 31.7 GB −7.66 GB —
256K (full) 22.3 GB 1.68 GB 35–60 34.4 GB −10.4 GB —

--n-cpu-moe for Gemma 4 26B-A4B on the RTX 4090

The measured UD-Q4_K_M GGUF fits the RTX 4090 whole at 32K (17.0 GB), so --n-cpu-moe is not needed there. Measured UD-Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 0 16.5 GB 748 MB 83–14783–14783–147
32K 0 17.0 GB 748 MB 73–12873–12873–128
64K 0 17.6 GB 748 MB 63–11063–11063–110
128K 0 18.8 GB 748 MB 49–8549–8549–85
256K (full) 0 21.3 GB 748 MB 34–5834–5834–58

Gemma 4 26B-A4B on more than one RTX 4090

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/sRent per hour
2× (48 GB) Q8_0 82–154 256K (full) 110–214 $1.06
4× (96 GB) BF16 92–174 256K (full) 158–335 $2.12

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Gemma 4 26B-A4B on the RTX 4090 with llama-server

llama-server -hf unsloth/gemma-4-26B-A4B-it-GGUF:Q4_K_M -c 252928 -np 1

gemma-4-26B-A4B-it-UD-Q4_K_M.gguf, 16.9 GB, from unsloth/gemma-4-26B-A4B-it-GGUF (checked 2026-09-29). At -c 252928 on the RTX 4090: 23.5 GB of 24 GB, 521 MB free. The file is 1.24 GB over the estimate above, so -c counts the file. -np 1: one slot, one sliding window.

Questions

Can I run Gemma 4 26B-A4B on an RTX 4090?

Yes: Gemma 4 26B-A4B needs about 17.5 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RTX 4090 with 6.50 GB to spare. At 32K the RTX 4090 holds up to Q6_K (23.2 GB), and Q4_K_M runs up to 256K (full) tokens.

How fast is Gemma 4 26B-A4B on an RTX 4090?

At Q4_K_M with 32K tokens of context it writes about 78–138 tokens/s for one request on an RTX 4090.

Does Gemma 4 26B-A4B need --n-cpu-moe on an RTX 4090?

The measured UD-Q4_K_M GGUF fits the RTX 4090 whole at 32K (17.0 GB), so --n-cpu-moe is not needed there.

What does a second RTX 4090 change for Gemma 4 26B-A4B?

Two RTX 4090 cards (48 GB in one tensor-parallel group) hold Gemma 4 26B-A4B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 110–214 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Gemma 4 26B-A4B VRAM requirements, what LLMs an RTX 4090 can run and every pair, or detect your own GPU. Model data checked .