Can I run Gemma 4 31B on an RTX 4060 Ti 16GB?
No: Gemma 4 31B needs about 23.9 GB at Q4_K_M with 32K tokens of context, 7.92 GB more than the RTX 4060 Ti 16GB holds, so it takes 2 of them (32 GB) in one tensor-parallel group. The smallest setup here that holds Gemma 4 31B at Q4_K_M with 32K is RTX 3090 (24 GB).
No Q4_K_M with 32K tokens of context on 2 cards
- Q4_K_M, 32K
- 23.9 GB
- RTX 4060 Ti
- 16 GB, 288 GB/s
- Short by
- 7.92 GB
- Tokens/s
- 13–18 tokens/s
At Q4_K_M with 32K tokens of context on 2 cards it writes about 13–18 tokens/s for one request on 2 RTX 4060 Ti 16GB cards.
Best precision for Gemma 4 31B on an RTX 4060 Ti 16GB
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Only Q2_K, tight | 15.9 GB | 110 MB (tight) | 10–14 |
| 32K | Nothing fits | — | — | — |
| 128K | Nothing fits | — | — | — |
| 256K (full) | Nothing fits | — | — | — |
Gemma 4 31B on the RTX 4060 Ti as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 21.5 GB | −5.52 GB | — | 36.2 GB | −20.2 GB | — |
| 8K | 21.9 GB | −5.86 GB | — | 36.5 GB | −20.5 GB | — |
| 16K | 22.5 GB | −6.55 GB | — | 37.2 GB | −21.2 GB | — |
| 32K | 23.9 GB | −7.92 GB | — | 38.6 GB | −22.6 GB | — |
| 64K | 26.7 GB | −10.7 GB | — | 41.3 GB | −25.3 GB | — |
| 128K | 32.2 GB | −16.2 GB | — | 46.8 GB | −30.8 GB | — |
| 256K (full) | 43.2 GB | −27.2 GB | — | 57.8 GB | −41.8 GB | — |
Gemma 4 31B on more than one RTX 4060 Ti
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s |
|---|---|---|---|---|
| 2× (32 GB) | Q6_K | 10–14 | 114K | 13–18 |
| 4× (64 GB) | Q8_0 | 16–23 | 256K (full) | 25–36 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 20.2 GB at 32K, 4.20 GB over the RTX 4060 Ti; it does not fit with 0.5 GB to spare even with 1K tokens.
- Smallest setup for Q4_K_M at 32K: RTX 3090 (24 GB).
Run Gemma 4 31B on the RTX 4060 Ti with llama-server
llama-server -hf unsloth/gemma-4-31B-it-GGUF:Q4_K_M -c 116736 -np 1 gemma-4-31B-it-Q4_K_M.gguf, 18.3 GB, from unsloth/
Questions
Can I run Gemma 4 31B on an RTX 4060 Ti 16GB?
No: Gemma 4 31B needs about 23.9 GB at Q4_K_M with 32K tokens of context, 7.92 GB more than the RTX 4060 Ti 16GB holds, so it takes 2 of them (32 GB) in one tensor-parallel group. The smallest setup here that holds Gemma 4 31B at Q4_K_M with 32K is RTX 3090 (24 GB).
How fast is Gemma 4 31B on an RTX 4060 Ti 16GB?
At Q4_K_M with 32K tokens of context on 2 cards it writes about 13–18 tokens/s for one request on 2 RTX 4060 Ti 16GB cards.
What does a second RTX 4060 Ti 16GB change for Gemma 4 31B?
Two RTX 4060 Ti 16GB cards (32 GB in one tensor-parallel group) hold Gemma 4 31B at Q6_K with 32K, and Q4_K_M up to 114K tokens, at about 13–18 tokens/s.
Try other settings in the VRAM calculator or the speed calculator. See also Gemma 4 31B VRAM requirements, what LLMs an RTX 4060 Ti 16GB can run and every pair, or detect your own GPU. Model data checked .