Can I run Gemma 4 26B-A4B on an RTX 3090?
Yes: Gemma 4 26B-A4B needs about 17.5 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RTX 3090 with 6.50 GB to spare. At 32K the RTX 3090 holds up to Q6_K (23.2 GB), and Q4_K_M runs up to 256K (full) tokens.
Yes Q4_K_M with 32K tokens of context
- Q4_K_M, 32K
- 17.5 GB
- RTX 3090
- 24 GB, 936 GB/s
- To spare
- 6.50 GB
- Tokens/s
- 73–129 tokens/s
At Q4_K_M with 32K tokens of context it writes about 73–129 tokens/s for one request on an RTX 3090.
Best precision for Gemma 4 26B-A4B on an RTX 3090
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Q6_K | 22.7 GB | 1.33 GB | 67–117 |
| 32K | Q6_K | 23.2 GB | 831 MB | 60–104 |
| 128K | Q5_K_M | 22.3 GB | 1.69 GB | 45–77 |
| 256K (full) | Q4_K_M | 22.3 GB | 1.68 GB | 33–56 |
Gemma 4 26B-A4B on the RTX 3090 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 16.9 GB | 7.10 GB | 87–153 | 29.0 GB | −5.00 GB | — |
| 8K | 17.0 GB | 7.01 GB | 84–149 | 29.1 GB | −5.08 GB | — |
| 16K | 17.2 GB | 6.84 GB | 80–142 | 29.3 GB | −5.26 GB | — |
| 32K | 17.5 GB | 6.50 GB | 73–129 | 29.6 GB | −5.60 GB | — |
| 64K | 18.2 GB | 5.81 GB | 62–109 | 30.3 GB | −6.29 GB | — |
| 128K | 19.6 GB | 4.43 GB | 48–83 | 31.7 GB | −7.66 GB | — |
| 256K (full) | 22.3 GB | 1.68 GB | 33–56 | 34.4 GB | −10.4 GB | — |
--n-cpu-moe for Gemma 4 26B-A4B on the RTX 3090
The measured UD-Q4_K_M GGUF fits the RTX 3090 whole at 32K (17.0 GB), so --n-cpu-moe is not needed there. Measured UD-Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.
Gemma 4 26B-A4B on more than one RTX 3090
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s | Rent per hour |
|---|---|---|---|---|---|
| 2× (48 GB) | Q8_0 | 78–145 | 256K (full) | 105–202 | $0.54 |
| 4× (96 GB) | BF16 | 87–164 | 256K (full) | 153–321 | $1.08 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 14.4 GB at 32K, 9.57 GB under the RTX 3090; it fits with 0.5 GB to spare up to 256K (full) tokens.
- Renting an RTX 3090 costs about $0.27 an hour (median on getdeploying.com, 2026-09-29): $0.58–1.0 per million tokens at 73–129 tokens/s.
- Smallest setup for Q4_K_M at 32K: RTX 3090 (24 GB).
Run Gemma 4 26B-A4B on the RTX 3090 with llama-server
llama-server -hf unsloth/gemma-4-26B-A4B-it-GGUF:Q4_K_M -c 252928 -np 1 gemma-4-26B-A4B-it-UD-Q4_K_M.gguf, 16.9 GB, from unsloth/
Questions
Can I run Gemma 4 26B-A4B on an RTX 3090?
Yes: Gemma 4 26B-A4B needs about 17.5 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RTX 3090 with 6.50 GB to spare. At 32K the RTX 3090 holds up to Q6_K (23.2 GB), and Q4_K_M runs up to 256K (full) tokens.
How fast is Gemma 4 26B-A4B on an RTX 3090?
At Q4_K_M with 32K tokens of context it writes about 73–129 tokens/s for one request on an RTX 3090.
Does Gemma 4 26B-A4B need --n-cpu-moe on an RTX 3090?
The measured UD-Q4_K_M GGUF fits the RTX 3090 whole at 32K (17.0 GB), so --n-cpu-moe is not needed there.
What does a second RTX 3090 change for Gemma 4 26B-A4B?
Two RTX 3090 cards (48 GB in one tensor-parallel group) hold Gemma 4 26B-A4B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 105–202 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Gemma 4 26B-A4B VRAM requirements, what LLMs an RTX 3090 can run and every pair, or detect your own GPU. Model data checked .