Can I run Llama 3.1 8B on an RX 7900 XTX?
Yes: Llama 3.1 8B needs about 9.88 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RX 7900 XTX with 14.1 GB to spare. At 32K the RX 7900 XTX holds up to BF16 (21.4 GB), and Q4_K_M runs up to 128K (full) tokens.
Yes Q4_K_M with 32K tokens of context
- Q4_K_M, 32K
- 9.88 GB
- RX 7900 XTX
- 24 GB, 960 GB/s
- To spare
- 14.1 GB
- Tokens/s
- 53–76 tokens/s
At Q4_K_M with 32K tokens of context it writes about 53–76 tokens/s for one request on an RX 7900 XTX.
Best precision for Llama 3.1 8B on an RX 7900 XTX
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | BF16 | 18.1 GB | 5.95 GB | 29–41 |
| 32K | BF16 | 21.4 GB | 2.65 GB | 25–35 |
| 128K (full) | Q4_K_M | 23.1 GB | 945 MB | 23–32 |
Llama 3.1 8B on the RX 7900 XTX as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 6.03 GB | 18.0 GB | 85–125 | 9.79 GB | 14.2 GB | 54–76 |
| 8K | 6.58 GB | 17.4 GB | 79–114 | 10.3 GB | 13.7 GB | 51–72 |
| 16K | 7.68 GB | 16.3 GB | 68–98 | 11.4 GB | 12.6 GB | 46–65 |
| 32K | 9.88 GB | 14.1 GB | 53–76 | 13.6 GB | 10.4 GB | 39–55 |
| 64K | 14.3 GB | 9.72 GB | 37–52 | 18.0 GB | 5.96 GB | 29–41 |
| 128K (full) | 23.1 GB | 945 MB | 23–32 | 26.8 GB | −2.84 GB | — |
Llama 3.1 8B on more than one RX 7900 XTX
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s |
|---|---|---|---|---|
| 2× (48 GB) | BF16 | 44–65 | 128K (full) | 82–131 |
| 4× (96 GB) | BF16 | 76–120 | 128K (full) | 128–223 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 8.92 GB at 32K, 15.1 GB under the RX 7900 XTX; it fits with 0.5 GB to spare up to 128K (full) tokens.
- Smallest setup for Q4_K_M at 32K: RTX 3060 12GB.
Run Llama 3.1 8B on the RX 7900 XTX with llama-server
llama-server -hf bartowski/Meta-Llama-3.1-8B-Instruct-GGUF:Q4_K_M -c 131072 Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf, 4.9 GB, from bartowski/
Questions
Can I run Llama 3.1 8B on an RX 7900 XTX?
Yes: Llama 3.1 8B needs about 9.88 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RX 7900 XTX with 14.1 GB to spare. At 32K the RX 7900 XTX holds up to BF16 (21.4 GB), and Q4_K_M runs up to 128K (full) tokens.
How fast is Llama 3.1 8B on an RX 7900 XTX?
At Q4_K_M with 32K tokens of context it writes about 53–76 tokens/s for one request on an RX 7900 XTX.
What does a second RX 7900 XTX change for Llama 3.1 8B?
Two RX 7900 XTX cards (48 GB in one tensor-parallel group) hold Llama 3.1 8B at BF16 with 32K, and Q4_K_M up to 128K (full) tokens, at about 82–131 tokens/s.
Try other settings in the VRAM calculator or the speed calculator. See also Llama 3.1 8B VRAM requirements, what LLMs an RX 7900 XTX can run and every pair, or detect your own GPU. Model data checked .