Gemma 4 26B-A4B VRAM requirements
Gemma 4 26B-A4B has 25.8B parameters, of which about 4B are used per token; all 128 experts still have to be in memory. With an 8K-token context and one request it needs about 16.8 GB of GPU memory at Q4_K_M, 27.2 GB at FP8 and 53.7 GB at FP16/BF16. The published weights take 48.1 GB (BF16). On one 24 GB RTX 3090 / 4090 it runs at Q4_K_M with its full 256K-token context.
Open Gemma 4 26B-A4B in the calculator
VRAM by quantization
Weights plus the KV cache for 8,192 tokens in FP16 and the runtime overhead (0.5 GB plus 10%). Each row opens the calculator with that setting. What the GGUF names mean.
| Precision | Weights | Total | Smallest setup |
|---|---|---|---|
| As published (BF16) | 48.1 GB | 53.7 GB | A100 / H100 80GB |
| FP16 / BF16 | 48.1 GB | 53.7 GB | A100 / H100 80GB |
| FP8 / INT8 | 24.0 GB | 27.2 GB | RTX 5090 |
| INT4 (AWQ / GPTQ) | 12.8 GB | 14.8 GB | RTX 4060 Ti 16GB |
| GGUF Q8_0 | 25.5 GB | 28.9 GB | RTX 5090 |
| GGUF Q6_K | 19.7 GB | 22.5 GB | RTX 3090 / 4090 |
| GGUF Q5_K_M | 17.0 GB | 19.5 GB | RTX 3090 / 4090 |
| GGUF Q4_K_M | 14.5 GB | 16.8 GB | RTX 3090 / 4090 |
| GGUF Q3_K_M | 11.7 GB | 13.7 GB | RTX 4060 Ti 16GB |
| GGUF Q2_K | 10.1 GB | 11.9 GB | RTX 3060 |
KV cache at long context
5 of its 30 layers use full attention and 25 keep a sliding window of 1,024 tokens. Each extra token of context adds 10 KB of FP16 cache per request once the sliding windows are full. How the KV cache works.
| Context | KV cache, FP16 | KV cache, FP8 | Total at Q4_K_M |
|---|---|---|---|
| 4K tokens | 240 MB | 120 MB | 16.8 GB |
| 32K tokens | 520 MB | 260 MB | 17.1 GB |
| 128K tokens | 1.45 GB | 740 MB | 18.1 GB |
| 256K tokens | 2.70 GB | 1.35 GB | 19.5 GB |
Which GPUs can run Gemma 4 26B-A4B
With 8,192 tokens of context. Several GPUs means one tensor-parallel group of 2, 4 or 8 cards.
| GPU | Memory | Q4_K_M | FP8 |
|---|---|---|---|
| RTX 3060 | 12 GB | Needs 2 | Needs 4 |
| RTX 4060 Ti 16GB | 16 GB | Needs 2 | Needs 2 |
| RTX 3090 / 4090 | 24 GB | Fits on one | Needs 2 |
| RTX 5090 | 32 GB | Fits on one | Fits on one |
| A100 40GB | 40 GB | Fits on one | Fits on one |
| Mac, 64 GB unified memory about 75% of it is usable by the GPU by default | 48 GB | Fits on one | Fits on one |
| L40S / RTX 6000 Ada | 48 GB | Fits on one | Fits on one |
| A100 / H100 80GB | 80 GB | Fits on one | Fits on one |
| Mac, 128 GB unified memory about 75% of it is usable by the GPU by default | 96 GB | Fits on one | Fits on one |
| H200 | 141 GB | Fits on one | Fits on one |
| B200 | 180 GB | Fits on one | Fits on one |
How fast Gemma 4 26B-A4B writes
Tokens per second for one request with 8,192 tokens of context, estimated from memory bandwidth and the parameters read per token. A dash means it does not fit on one card. Try other settings in the speed calculator.
| Hardware | Bandwidth | Q4_K_M | FP8 |
|---|---|---|---|
| RTX 3060 12GB | 360 GB/s | — | — |
| RTX 4090 | 1,008 GB/s | 95–170 | — |
| RTX 5090 | 1,792 GB/s | 153–283 | 105–189 |
| M4 Max Mac (128 GB) | 546 GB/s | 55–96 | 36–62 |
| M3 Ultra Mac Studio (512 GB) | 819 GB/s | 80–140 | 53–91 |
| H100 SXM | 3,350 GB/s | 238–472 | 173–326 |
| H200 | 4,800 GB/s | 295–613 | 223–437 |
Longest context on one GPU
How many tokens of context fit on a single card with one request and an FP16 KV cache. "Full" means the model's whole context window fits.
| GPU | Memory | Q4_K_M | Q8_0 | FP8 |
|---|---|---|---|---|
| RTX 3060 | 12 GB | No | No | No |
| RTX 4060 Ti 16GB | 16 GB | No | No | No |
| RTX 3090 / 4090 | 24 GB | 256K (full) | No | No |
| RTX 5090 | 32 GB | 256K (full) | 256K (full) | 256K (full) |
| A100 40GB | 40 GB | 256K (full) | 256K (full) | 256K (full) |
| Mac, 64 GB unified memory | 48 GB | 256K (full) | 256K (full) | 256K (full) |
| L40S / RTX 6000 Ada | 48 GB | 256K (full) | 256K (full) | 256K (full) |
| A100 / H100 80GB | 80 GB | 256K (full) | 256K (full) | 256K (full) |
| Mac, 128 GB unified memory | 96 GB | 256K (full) | 256K (full) | 256K (full) |
| H200 | 141 GB | 256K (full) | 256K (full) | 256K (full) |
| B200 | 180 GB | 256K (full) | 256K (full) | 256K (full) |
Model details
- Parameters
- 25.8B (25,805,936,206)
- Experts
- 128 routed experts, all loaded
- Active per token
- 4B
- Layers
- 5 of its 30 layers use full attention and 25 keep a sliding window of 1,024 tokens
- Attention cache
- 2 KV heads × 512 (keys double as values)
- Sliding-window layers
- 8 KV heads × 256
- Context length
- 262,144 tokens
- Published weights
- 48.1 GB (BF16)
- On Hugging Face
- google/gemma-4-26B-A4B-it
Other models
- DeepSeek V4 Flash 0731
- Qwen3.6 35B-A3B
- Qwen3.6 27B
- Qwen3.8 27B
- Qwen3.8 Flash Next
- Qwen3.8 2.4T-A95B
- Qwen3.5 9B
- Qwen3.5 122B-A10B
- Qwen3-Coder-Next
- DeepSeek V4.1 Flash
- DeepSeek V4 Flash
- DeepSeek V4 Pro
- DeepSeek V3.2
- GLM-5.3
- GLM-5.3 Flash
- GLM-5.2
- GLM-4.7 Flash
- Gemma 4 31B
- Gemma 4 12B
- Gemma 4 E4B
- Kimi K3
- MiniMax M3
- MiniMax M2.7
- MiMo V2.6 Flash
- MiMo V2.6 Pro
- Mistral Medium 3.5 128B
- Nemotron 3 Nano 4B
- Nemotron 3 Nano 30B-A3B
- Nemotron 3 Super 120B-A12B
- Xing 4.0 29B-A4B
- Muse Glimmer 30B
- MiniCPM5 2B
- Llama 3.1 8B
- Llama 3.1 70B
- Qwen3 8B
- Qwen3 30B-A3B
- gpt-oss-20b
- gpt-oss-120b
- DeepSeek V3 / R1
Numbers read from the model files on Hugging Face on .