Can I run Gemma 4 26B-A4B on an RTX 5080 16GB?
Only at Q3_K_M: Gemma 4 26B-A4B needs about 17.5 GB at Q4_K_M with 32K tokens of context, 1.50 GB more than the RTX 5080 16GB holds, but 14.4 GB at Q3_K_M, which fits with 1.57 GB to spare. The smallest setup here that holds Gemma 4 26B-A4B at Q4_K_M with 32K is RTX 3090 (24 GB).
Partly Q3_K_M with 32K tokens of context
- Q4_K_M, 32K
- 17.5 GB
- RTX 5080
- 16 GB, 960 GB/s
- Short by
- 1.50 GB
- Tokens/s
- 85–151 tokens/s
At Q3_K_M with 32K tokens of context it writes about 85–151 tokens/s for one request on an RTX 5080 16GB.
Best precision for Gemma 4 26B-A4B on an RTX 5080 16GB
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Q3_K_M | 13.9 GB | 2.08 GB | 100–179 |
| 32K | Q3_K_M | 14.4 GB | 1.57 GB | 85–151 |
| 128K | Q2_K | 14.6 GB | 1.36 GB | 56–98 |
| 256K (full) | Nothing fits | — | — | — |
Gemma 4 26B-A4B on the RTX 5080 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 16.9 GB | −924 MB | — | 29.0 GB | −13.0 GB | — |
| 8K | 17.0 GB | −1,012 MB | — | 29.1 GB | −13.1 GB | — |
| 16K | 17.2 GB | −1.16 GB | — | 29.3 GB | −13.3 GB | — |
| 32K | 17.5 GB | −1.50 GB | — | 29.6 GB | −13.6 GB | — |
| 64K | 18.2 GB | −2.19 GB | — | 30.3 GB | −14.3 GB | — |
| 128K | 19.6 GB | −3.57 GB | — | 31.7 GB | −15.7 GB | — |
| 256K (full) | 22.3 GB | −6.32 GB | — | 34.4 GB | −18.4 GB | — |
--n-cpu-moe for Gemma 4 26B-A4B on the RTX 5080
With --n-cpu-moe 3 the UD-Q4_K_M GGUF keeps 15.6 GB on the RTX 5080 and 2.06 GB in RAM, about 58–100 tokens/s with DDR5-5600. Measured UD-Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.
Gemma 4 26B-A4B on more than one RTX 5080
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s |
|---|---|---|---|---|
| 2× (32 GB) | Q8_0 | 79–148 | 256K (full) | 106–206 |
| 4× (64 GB) | BF16 | 88–167 | 256K (full) | 155–325 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q2_K, takes 12.6 GB at 32K, 3.42 GB under the RTX 5080; it fits with 0.5 GB to spare up to 167K tokens.
- Smallest setup for Q4_K_M at 32K: RTX 3090 (24 GB).
Run Gemma 4 26B-A4B on the RTX 5080 with llama-server
llama-server -hf unsloth/gemma-4-26B-A4B-it-GGUF:Q3_K_M -c 77824 -np 1 gemma-4-26B-A4B-it-UD-Q3_K_M.gguf, 12.7 GB, from unsloth/
Questions
Can I run Gemma 4 26B-A4B on an RTX 5080 16GB?
Only at Q3_K_M: Gemma 4 26B-A4B needs about 17.5 GB at Q4_K_M with 32K tokens of context, 1.50 GB more than the RTX 5080 16GB holds, but 14.4 GB at Q3_K_M, which fits with 1.57 GB to spare. The smallest setup here that holds Gemma 4 26B-A4B at Q4_K_M with 32K is RTX 3090 (24 GB).
How fast is Gemma 4 26B-A4B on an RTX 5080 16GB?
At Q3_K_M with 32K tokens of context it writes about 85–151 tokens/s for one request on an RTX 5080 16GB.
Does Gemma 4 26B-A4B need --n-cpu-moe on an RTX 5080 16GB?
With --n-cpu-moe 3 the UD-Q4_K_M GGUF keeps 15.6 GB on the RTX 5080 and 2.06 GB in RAM, about 58–100 tokens/s with DDR5-5600.
What does a second RTX 5080 16GB change for Gemma 4 26B-A4B?
Two RTX 5080 16GB cards (32 GB in one tensor-parallel group) hold Gemma 4 26B-A4B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 106–206 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Gemma 4 26B-A4B VRAM requirements, what LLMs an RTX 5080 16GB can run and every pair, or detect your own GPU. Model data checked .