Can I run Gemma 4 12B on an RTX 3060 12GB?
Yes: Gemma 4 12B needs about 8.98 GB at Q4_K_M with 32K tokens of context, which fits the RTX 3060 12GB with 3.02 GB to spare. At 32K the RTX 3060 holds up to Q5_K_M (10.2 GB), and Q4_K_M runs up to 178K tokens.
Yes Q4_K_M with 32K tokens of context
- Q4_K_M, 32K
- 8.98 GB
- RTX 3060
- 12 GB, 360 GB/s
- To spare
- 3.02 GB
- Tokens/s
- 23–32 tokens/s
At Q4_K_M with 32K tokens of context it writes about 23–32 tokens/s for one request on an RTX 3060 12GB.
Best precision for Gemma 4 12B on an RTX 3060 12GB
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Q6_K | 11.2 GB | 819 MB | 18–26 |
| 32K | Q5_K_M | 10.2 GB | 1.75 GB | 20–28 |
| 128K | Q4_K_M | 10.6 GB | 1.37 GB | 19–27 |
| 256K (full) | Q3_K_M | 11.4 GB | 610 MB | 18–25 |
Gemma 4 12B on the RTX 3060 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 8.50 GB | 3.50 GB | 24–34 | 14.1 GB | −2.10 GB | — |
| 8K | 8.57 GB | 3.43 GB | 24–34 | 14.2 GB | −2.17 GB | — |
| 16K | 8.70 GB | 3.30 GB | 24–33 | 14.3 GB | −2.31 GB | — |
| 32K | 8.98 GB | 3.02 GB | 23–32 | 14.6 GB | −2.58 GB | — |
| 64K | 9.53 GB | 2.47 GB | 22–30 | 15.1 GB | −3.13 GB | — |
| 128K | 10.6 GB | 1.37 GB | 19–27 | 16.2 GB | −4.23 GB | — |
| 256K (full) | 12.8 GB | −848 MB | — | 18.4 GB | −6.43 GB | — |
Gemma 4 12B on more than one RTX 3060
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s | Rent per hour |
|---|---|---|---|---|---|
| 2× (24 GB) | Q8_0 | 26–37 | 256K (full) | 41–60 | $0.16 |
| 4× (48 GB) | BF16 | 29–41 | 256K (full) | 72–112 | $0.32 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 7.55 GB at 32K, 4.45 GB under the RTX 3060; it fits with 0.5 GB to spare up to 256K (full) tokens.
- Renting an RTX 3060 12GB costs about $0.08 an hour (median on getdeploying.com, 2026-09-29): $0.69–0.96 per million tokens at 23–32 tokens/s.
- Smallest setup for Q4_K_M at 32K: RTX 3060 12GB.
Run Gemma 4 12B on the RTX 3060 with llama-server
llama-server -hf unsloth/gemma-4-12b-it-GGUF:Q4_K_M -c 182272 -np 1 gemma-4-12b-it-Q4_K_M.gguf, 7.1 GB, from unsloth/
Questions
Can I run Gemma 4 12B on an RTX 3060 12GB?
Yes: Gemma 4 12B needs about 8.98 GB at Q4_K_M with 32K tokens of context, which fits the RTX 3060 12GB with 3.02 GB to spare. At 32K the RTX 3060 holds up to Q5_K_M (10.2 GB), and Q4_K_M runs up to 178K tokens.
How fast is Gemma 4 12B on an RTX 3060 12GB?
At Q4_K_M with 32K tokens of context it writes about 23–32 tokens/s for one request on an RTX 3060 12GB.
What does a second RTX 3060 12GB change for Gemma 4 12B?
Two RTX 3060 12GB cards (24 GB in one tensor-parallel group) hold Gemma 4 12B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 41–60 tokens/s.
Try other settings in the VRAM calculator or the speed calculator. See also Gemma 4 12B VRAM requirements, what LLMs an RTX 3060 12GB can run and every pair, or detect your own GPU. Model data checked .