Can I run Gemma 4 31B on an RTX 5090?

Yes: Gemma 4 31B needs about 23.9 GB at Q4_K_M with 32K tokens of context, which fits the 32 GB RTX 5090 with 8.08 GB to spare. At 32K the RTX 5090 holds up to Q6_K (30.8 GB), and Q4_K_M runs up to 120K tokens.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
23.9 GB
RTX 5090
32 GB, 1,792 GB/s
To spare
8.08 GB
Tokens/s
40–57 tokens/s

At Q4_K_M with 32K tokens of context it writes about 40–57 tokens/s for one request on an RTX 5090.

Best precision for Gemma 4 31B on an RTX 5090

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q6_K 28.7 GB 3.25 GB 34–48
32K Q6_K 30.8 GB 1.19 GB 32–44
128K Q3_K_M 28.4 GB 3.55 GB 34–48
256K (full) Nothing fits — — —

Gemma 4 31B on the RTX 5090 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 21.5 GB 10.5 GB 45–63 36.2 GB −4.17 GB —
8K 21.9 GB 10.1 GB 44–62 36.5 GB −4.52 GB —
16K 22.5 GB 9.45 GB 43–61 37.2 GB −5.20 GB —
32K 23.9 GB 8.08 GB 40–57 38.6 GB −6.58 GB —
64K 26.7 GB 5.33 GB 36–51 41.3 GB −9.33 GB —
128K 32.2 GB −176 MB — 46.8 GB −14.8 GB —
256K (full) 43.2 GB −11.2 GB — 57.8 GB −25.8 GB —

Gemma 4 31B on more than one RTX 5090

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/sRent per hour
2× (64 GB) Q8_0 45–66 256K (full) 66–102 $1.38
4× (128 GB) BF16 49–73 256K (full) 108–180 $2.76

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Gemma 4 31B on the RTX 5090 with llama-server

llama-server -hf unsloth/gemma-4-31B-it-GGUF:Q4_K_M -c 122880 -ngl 99 -np 1

gemma-4-31B-it-Q4_K_M.gguf, 18.3 GB, from unsloth/gemma-4-31B-it-GGUF (checked 2026-09-29). At -c 122880 on the RTX 5090: 31.5 GB of 32 GB, 528 MB free. -np 1: one slot, one sliding window.

Questions

Can I run Gemma 4 31B on an RTX 5090?

Yes: Gemma 4 31B needs about 23.9 GB at Q4_K_M with 32K tokens of context, which fits the 32 GB RTX 5090 with 8.08 GB to spare. At 32K the RTX 5090 holds up to Q6_K (30.8 GB), and Q4_K_M runs up to 120K tokens.

How fast is Gemma 4 31B on an RTX 5090?

At Q4_K_M with 32K tokens of context it writes about 40–57 tokens/s for one request on an RTX 5090.

What does a second RTX 5090 change for Gemma 4 31B?

Two RTX 5090 cards (64 GB in one tensor-parallel group) hold Gemma 4 31B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 66–102 tokens/s.

Try other settings in the VRAM calculator or the speed calculator. See also Gemma 4 31B VRAM requirements, what LLMs an RTX 5090 can run and every pair, or detect your own GPU. Model data checked .