Can I run Gemma 4 31B on an RX 7900 XTX?

Only with a shorter context: Gemma 4 31B needs about 23.9 GB at Q4_K_M with 32K tokens of context, which leaves only 80 MB free on the 24 GB RX 7900 XTX, too tight to count as a fit, but Q4_K_M fits with up to 26K tokens. The smallest setup here that holds Gemma 4 31B at Q4_K_M with 32K is RTX 3090 (24 GB).

Partly Q4_K_M with 26K tokens of context

Q4_K_M, 32K
23.9 GB
RX 7900 XTX
24 GB, 960 GB/s
Free (tight)
80 MB
Tokens/s
23–32 tokens/s

At Q4_K_M with 26K tokens of context it writes about 23–32 tokens/s for one request on an RX 7900 XTX.

Best precision for Gemma 4 31B on an RX 7900 XTX

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q4_K_M 21.9 GB 2.14 GB 24–34
32K Q3_K_M 20.2 GB 3.80 GB 26–37
128K Nothing fits — — —
256K (full) Nothing fits — — —

Gemma 4 31B on the RX 7900 XTX as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 21.5 GB 2.48 GB 25–34 36.2 GB −12.2 GB —
8K 21.9 GB 2.14 GB 24–34 36.5 GB −12.5 GB —
16K 22.5 GB 1.45 GB 24–33 37.2 GB −13.2 GB —
32K 23.9 GB 80 MB (tight) 22–31 38.6 GB −14.6 GB —
64K 26.7 GB −2.67 GB — 41.3 GB −17.3 GB —
128K 32.2 GB −8.17 GB — 46.8 GB −22.8 GB —
256K (full) 43.2 GB −19.2 GB — 57.8 GB −33.8 GB —

Gemma 4 31B on more than one RX 7900 XTX

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/s
2× (48 GB) Q8_0 26–37 256K (full) 40–58
4× (96 GB) BF16 29–41 256K (full) 70–108

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Gemma 4 31B on the RX 7900 XTX with llama-server

llama-server -hf unsloth/gemma-4-31B-it-GGUF:Q4_K_M -c 26624 -np 1

gemma-4-31B-it-Q4_K_M.gguf, 18.3 GB, from unsloth/gemma-4-31B-it-GGUF (checked 2026-09-29). At -c 26624 on the RX 7900 XTX: 23.4 GB of 24 GB, 608 MB free. -np 1: one slot, one sliding window.

Questions

Can I run Gemma 4 31B on an RX 7900 XTX?

Only with a shorter context: Gemma 4 31B needs about 23.9 GB at Q4_K_M with 32K tokens of context, which leaves only 80 MB free on the 24 GB RX 7900 XTX, too tight to count as a fit, but Q4_K_M fits with up to 26K tokens. The smallest setup here that holds Gemma 4 31B at Q4_K_M with 32K is RTX 3090 (24 GB).

How fast is Gemma 4 31B on an RX 7900 XTX?

At Q4_K_M with 26K tokens of context it writes about 23–32 tokens/s for one request on an RX 7900 XTX.

What does a second RX 7900 XTX change for Gemma 4 31B?

Two RX 7900 XTX cards (48 GB in one tensor-parallel group) hold Gemma 4 31B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 40–58 tokens/s.

Try other settings in the VRAM calculator or the speed calculator. See also Gemma 4 31B VRAM requirements, what LLMs an RX 7900 XTX can run and every pair, or detect your own GPU. Model data checked .