Can I run Gemma 4 26B-A4B on an RTX 5060 Ti 16GB?

Only at Q3_K_M: Gemma 4 26B-A4B needs about 17.5 GB at Q4_K_M with 32K tokens of context, 1.50 GB more than the RTX 5060 Ti 16GB holds, but 14.4 GB at Q3_K_M, which fits with 1.57 GB to spare. The smallest setup here that holds Gemma 4 26B-A4B at Q4_K_M with 32K is RTX 3090 (24 GB).

Partly Q3_K_M with 32K tokens of context

Q4_K_M, 32K
17.5 GB
RTX 5060 Ti
16 GB, 448 GB/s
Short by
1.50 GB
Tokens/s
43–73 tokens/s

At Q3_K_M with 32K tokens of context it writes about 43–73 tokens/s for one request on an RTX 5060 Ti 16GB.

Best precision for Gemma 4 26B-A4B on an RTX 5060 Ti 16GB

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q3_K_M 13.9 GB 2.08 GB 51–88
32K Q3_K_M 14.4 GB 1.57 GB 43–73
128K Q2_K 14.6 GB 1.36 GB 28–47
256K (full) Nothing fits — — —

Gemma 4 26B-A4B on the RTX 5060 Ti as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 16.9 GB −924 MB — 29.0 GB −13.0 GB —
8K 17.0 GB −1,012 MB — 29.1 GB −13.1 GB —
16K 17.2 GB −1.16 GB — 29.3 GB −13.3 GB —
32K 17.5 GB −1.50 GB — 29.6 GB −13.6 GB —
64K 18.2 GB −2.19 GB — 30.3 GB −14.3 GB —
128K 19.6 GB −3.57 GB — 31.7 GB −15.7 GB —
256K (full) 22.3 GB −6.32 GB — 34.4 GB −18.4 GB —

--n-cpu-moe for Gemma 4 26B-A4B on the RTX 5060 Ti

With --n-cpu-moe 3 the UD-Q4_K_M GGUF keeps 15.6 GB on the RTX 5060 Ti and 2.06 GB in RAM, about 32–54 tokens/s with DDR5-5600. Measured UD-Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 2 15.6 GB 1.62 GB 35–6037–6438–64
32K 3 15.6 GB 2.06 GB 29–5032–5432–55
64K 4 15.8 GB 2.50 GB 25–4227–4527–46
128K 7 15.7 GB 3.83 GB 18–3020–3420–34
256K (full) 13 15.6 GB 6.49 GB 12–2013–2214–23

Gemma 4 26B-A4B on more than one RTX 5060 Ti

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/s
2× (32 GB) Q8_0 44–77 256K (full) 62–112
4× (64 GB) BF16 49–88 256K (full) 102–196

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Gemma 4 26B-A4B on the RTX 5060 Ti with llama-server

llama-server -hf unsloth/gemma-4-26B-A4B-it-GGUF:Q3_K_M -c 77824 -np 1

gemma-4-26B-A4B-it-UD-Q3_K_M.gguf, 12.7 GB, from unsloth/gemma-4-26B-A4B-it-GGUF (checked 2026-09-29). At -c 77824 on the RTX 5060 Ti: 15.5 GB of 16 GB, 517 MB free. -np 1: one slot, one sliding window.

Questions

Can I run Gemma 4 26B-A4B on an RTX 5060 Ti 16GB?

Only at Q3_K_M: Gemma 4 26B-A4B needs about 17.5 GB at Q4_K_M with 32K tokens of context, 1.50 GB more than the RTX 5060 Ti 16GB holds, but 14.4 GB at Q3_K_M, which fits with 1.57 GB to spare. The smallest setup here that holds Gemma 4 26B-A4B at Q4_K_M with 32K is RTX 3090 (24 GB).

How fast is Gemma 4 26B-A4B on an RTX 5060 Ti 16GB?

At Q3_K_M with 32K tokens of context it writes about 43–73 tokens/s for one request on an RTX 5060 Ti 16GB.

Does Gemma 4 26B-A4B need --n-cpu-moe on an RTX 5060 Ti 16GB?

With --n-cpu-moe 3 the UD-Q4_K_M GGUF keeps 15.6 GB on the RTX 5060 Ti and 2.06 GB in RAM, about 32–54 tokens/s with DDR5-5600.

What does a second RTX 5060 Ti 16GB change for Gemma 4 26B-A4B?

Two RTX 5060 Ti 16GB cards (32 GB in one tensor-parallel group) hold Gemma 4 26B-A4B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 62–112 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Gemma 4 26B-A4B VRAM requirements, what LLMs an RTX 5060 Ti 16GB can run and every pair, or detect your own GPU. Model data checked .