Can I run Gemma 4 12B on an RTX 3060 12GB?

Yes: Gemma 4 12B needs about 8.98 GB at Q4_K_M with 32K tokens of context, which fits the RTX 3060 12GB with 3.02 GB to spare. At 32K the RTX 3060 holds up to Q5_K_M (10.2 GB), and Q4_K_M runs up to 178K tokens.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
8.98 GB
RTX 3060
12 GB, 360 GB/s
To spare
3.02 GB
Tokens/s
23–32 tokens/s

At Q4_K_M with 32K tokens of context it writes about 23–32 tokens/s for one request on an RTX 3060 12GB.

Best precision for Gemma 4 12B on an RTX 3060 12GB

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q6_K 11.2 GB 819 MB 18–26
32K Q5_K_M 10.2 GB 1.75 GB 20–28
128K Q4_K_M 10.6 GB 1.37 GB 19–27
256K (full) Q3_K_M 11.4 GB 610 MB 18–25

Gemma 4 12B on the RTX 3060 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 8.50 GB 3.50 GB 24–34 14.1 GB −2.10 GB —
8K 8.57 GB 3.43 GB 24–34 14.2 GB −2.17 GB —
16K 8.70 GB 3.30 GB 24–33 14.3 GB −2.31 GB —
32K 8.98 GB 3.02 GB 23–32 14.6 GB −2.58 GB —
64K 9.53 GB 2.47 GB 22–30 15.1 GB −3.13 GB —
128K 10.6 GB 1.37 GB 19–27 16.2 GB −4.23 GB —
256K (full) 12.8 GB −848 MB — 18.4 GB −6.43 GB —

Gemma 4 12B on more than one RTX 3060

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/sRent per hour
2× (24 GB) Q8_0 26–37 256K (full) 41–60 $0.16
4× (48 GB) BF16 29–41 256K (full) 72–112 $0.32

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Gemma 4 12B on the RTX 3060 with llama-server

llama-server -hf unsloth/gemma-4-12b-it-GGUF:Q4_K_M -c 182272 -np 1

gemma-4-12b-it-Q4_K_M.gguf, 7.1 GB, from unsloth/gemma-4-12b-it-GGUF (checked 2026-09-29). At -c 182272 on the RTX 3060: 11.5 GB of 12 GB, 525 MB free. -np 1: one slot, one sliding window.

Questions

Can I run Gemma 4 12B on an RTX 3060 12GB?

Yes: Gemma 4 12B needs about 8.98 GB at Q4_K_M with 32K tokens of context, which fits the RTX 3060 12GB with 3.02 GB to spare. At 32K the RTX 3060 holds up to Q5_K_M (10.2 GB), and Q4_K_M runs up to 178K tokens.

How fast is Gemma 4 12B on an RTX 3060 12GB?

At Q4_K_M with 32K tokens of context it writes about 23–32 tokens/s for one request on an RTX 3060 12GB.

What does a second RTX 3060 12GB change for Gemma 4 12B?

Two RTX 3060 12GB cards (24 GB in one tensor-parallel group) hold Gemma 4 12B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 41–60 tokens/s.

Try other settings in the VRAM calculator or the speed calculator. See also Gemma 4 12B VRAM requirements, what LLMs an RTX 3060 12GB can run and every pair, or detect your own GPU. Model data checked .