Can I run Gemma 4 26B-A4B on an RTX 5090?

Yes: Gemma 4 26B-A4B needs about 17.5 GB at Q4_K_M with 32K tokens of context, which fits the 32 GB RTX 5090 with 14.5 GB to spare. At 32K the RTX 5090 holds up to Q8_0 (29.6 GB), and Q4_K_M runs up to 256K (full) tokens.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
17.5 GB
RTX 5090
32 GB, 1,792 GB/s
To spare
14.5 GB
Tokens/s
128–233 tokens/s

At Q4_K_M with 32K tokens of context it writes about 128–233 tokens/s for one request on an RTX 5090.

Best precision for Gemma 4 26B-A4B on an RTX 5090

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q8_0 29.1 GB 2.92 GB 97–173
32K Q8_0 29.6 GB 2.40 GB 89–158
128K FP8 30.0 GB 1.99 GB 69–120
256K (full) Q6_K 28.0 GB 4.00 GB 55–95

Gemma 4 26B-A4B on the RTX 5090 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 16.9 GB 15.1 GB 148–274 29.0 GB 3.00 GB 99–176
8K 17.0 GB 15.0 GB 145–267 29.1 GB 2.92 GB 97–173
16K 17.2 GB 14.8 GB 139–255 29.3 GB 2.74 GB 94–168
32K 17.5 GB 14.5 GB 128–233 29.6 GB 2.40 GB 89–158
64K 18.2 GB 13.8 GB 110–198 30.3 GB 1.71 GB 80–141
128K 19.6 GB 12.4 GB 86–153 31.7 GB 347 MB (tight) 67–116
256K (full) 22.3 GB 9.68 GB 60–105 34.4 GB −2.41 GB —

--n-cpu-moe for Gemma 4 26B-A4B on the RTX 5090

The measured UD-Q4_K_M GGUF fits the RTX 5090 whole at 32K (17.0 GB), so --n-cpu-moe is not needed there. Measured UD-Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 0 16.5 GB 748 MB 135–247135–247135–247
32K 0 17.0 GB 748 MB 120–217120–217120–217
64K 0 17.6 GB 748 MB 104–187104–187104–187
128K 0 18.8 GB 748 MB 83–14683–14683–146
256K (full) 0 21.3 GB 748 MB 59–10259–10259–102

Gemma 4 26B-A4B on more than one RTX 5090

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/sRent per hour
2× (64 GB) BF16 84–158 256K (full) 150–312 $1.38
4× (128 GB) BF16 130–263 256K (full) 197–444 $2.76

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Gemma 4 26B-A4B on the RTX 5090 with llama-server

llama-server -hf unsloth/gemma-4-26B-A4B-it-GGUF:Q4_K_M -c 262144 -np 1

gemma-4-26B-A4B-it-UD-Q4_K_M.gguf, 16.9 GB, from unsloth/gemma-4-26B-A4B-it-GGUF (checked 2026-09-29). At -c 262144 on the RTX 5090: 23.7 GB of 32 GB, 8.32 GB free. The file is 1.24 GB over the estimate above, so -c counts the file. -np 1: one slot, one sliding window.

Questions

Can I run Gemma 4 26B-A4B on an RTX 5090?

Yes: Gemma 4 26B-A4B needs about 17.5 GB at Q4_K_M with 32K tokens of context, which fits the 32 GB RTX 5090 with 14.5 GB to spare. At 32K the RTX 5090 holds up to Q8_0 (29.6 GB), and Q4_K_M runs up to 256K (full) tokens.

How fast is Gemma 4 26B-A4B on an RTX 5090?

At Q4_K_M with 32K tokens of context it writes about 128–233 tokens/s for one request on an RTX 5090.

Does Gemma 4 26B-A4B need --n-cpu-moe on an RTX 5090?

The measured UD-Q4_K_M GGUF fits the RTX 5090 whole at 32K (17.0 GB), so --n-cpu-moe is not needed there.

What does a second RTX 5090 change for Gemma 4 26B-A4B?

Two RTX 5090 cards (64 GB in one tensor-parallel group) hold Gemma 4 26B-A4B at BF16 with 32K, and Q4_K_M up to 256K (full) tokens, at about 150–312 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Gemma 4 26B-A4B VRAM requirements, what LLMs an RTX 5090 can run and every pair, or detect your own GPU. Model data checked .