Can I run Gemma 4 31B on an RTX 5060 Ti 16GB?

No: Gemma 4 31B needs about 23.9 GB at Q4_K_M with 32K tokens of context, 7.92 GB more than the RTX 5060 Ti 16GB holds, so it takes 2 of them (32 GB) in one tensor-parallel group. The smallest setup here that holds Gemma 4 31B at Q4_K_M with 32K is RTX 3090 (24 GB).

No Q4_K_M with 32K tokens of context on 2 cards

Q4_K_M, 32K
23.9 GB
RTX 5060 Ti
16 GB, 448 GB/s
Short by
7.92 GB
Tokens/s
20–28 tokens/s

At Q4_K_M with 32K tokens of context on 2 cards it writes about 20–28 tokens/s for one request on 2 RTX 5060 Ti 16GB cards.

Best precision for Gemma 4 31B on an RTX 5060 Ti 16GB

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Only Q2_K, tight 15.9 GB 110 MB (tight) 16–22
32K Nothing fits — — —
128K Nothing fits — — —
256K (full) Nothing fits — — —

Gemma 4 31B on the RTX 5060 Ti as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 21.5 GB −5.52 GB — 36.2 GB −20.2 GB —
8K 21.9 GB −5.86 GB — 36.5 GB −20.5 GB —
16K 22.5 GB −6.55 GB — 37.2 GB −21.2 GB —
32K 23.9 GB −7.92 GB — 38.6 GB −22.6 GB —
64K 26.7 GB −10.7 GB — 41.3 GB −25.3 GB —
128K 32.2 GB −16.2 GB — 46.8 GB −30.8 GB —
256K (full) 43.2 GB −27.2 GB — 57.8 GB −41.8 GB —

Gemma 4 31B on more than one RTX 5060 Ti

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/s
2× (32 GB) Q6_K 16–22 114K 20–28
4× (64 GB) Q8_0 24–35 256K (full) 37–55

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Gemma 4 31B on the RTX 5060 Ti with llama-server

llama-server -hf unsloth/gemma-4-31B-it-GGUF:Q4_K_M -c 116736 -np 1

gemma-4-31B-it-Q4_K_M.gguf, 18.3 GB, from unsloth/gemma-4-31B-it-GGUF (checked 2026-09-29). At -c 116736 on the RTX 5060 Ti: 31.5 GB of 32 GB, 544 MB free, layers split over 2 cards. -np 1: one slot, one sliding window.

Questions

Can I run Gemma 4 31B on an RTX 5060 Ti 16GB?

No: Gemma 4 31B needs about 23.9 GB at Q4_K_M with 32K tokens of context, 7.92 GB more than the RTX 5060 Ti 16GB holds, so it takes 2 of them (32 GB) in one tensor-parallel group. The smallest setup here that holds Gemma 4 31B at Q4_K_M with 32K is RTX 3090 (24 GB).

How fast is Gemma 4 31B on an RTX 5060 Ti 16GB?

At Q4_K_M with 32K tokens of context on 2 cards it writes about 20–28 tokens/s for one request on 2 RTX 5060 Ti 16GB cards.

What does a second RTX 5060 Ti 16GB change for Gemma 4 31B?

Two RTX 5060 Ti 16GB cards (32 GB in one tensor-parallel group) hold Gemma 4 31B at Q6_K with 32K, and Q4_K_M up to 114K tokens, at about 20–28 tokens/s.

Try other settings in the VRAM calculator or the speed calculator. See also Gemma 4 31B VRAM requirements, what LLMs an RTX 5060 Ti 16GB can run and every pair, or detect your own GPU. Model data checked .