What LLMs can an RTX 4060 Ti 16GB run?
An RTX 4060 Ti 16GB gives a model 16 GB of memory and 288 GB/s of bandwidth. Of the 40 open models tracked here, 8 fit at Q4_K_M with an 8,192-token context, and 9 more at a lower precision.
Models that fit on one RTX 4060 Ti 16GB
The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first.
| Model | Best precision | Memory | Longest context | Tokens/s |
|---|---|---|---|---|
| Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active | GGUF Q2_K | 14.1 GB | 256K (full) | 53–91 |
| Gemma 4 31B 31.3B | GGUF Q2_K | 15.1 GB | 28K | 11–15 |
| GLM-4.7 Flash 31.2B, 3B active | GGUF Q2_K | 14.3 GB | 36K | 47–81 |
| Xing 4.0 29B-A4B 31.2B, 4B active | GGUF Q2_K | 14.3 GB | 43K | 40–68 |
| Qwen3 30B-A3B 30.5B, 3.3B active | GGUF Q2_K | 14.4 GB | 23K | 37–64 |
| Muse Glimmer 30B 29.8B | GGUF Q3_K_M | 15.6 GB | 36K | 11–15 |
| Qwen3.6 27B 27.8B | GGUF Q3_K_M | 15.0 GB | 22K | 11–15 |
| Qwen3.8 27B 27.8B | GGUF Q3_K_M | 15.0 GB | 22K | 11–15 |
| Gemma 4 26B-A4B 25.8B, 4B active | GGUF Q3_K_M | 13.7 GB | 219K | 36–62 |
| gpt-oss-20b 20.9B, 3.6B active | GGUF Q5_K_M | 15.9 GB | 11K | 30–51 |
| Gemma 4 12B 12.0B | GGUF Q8_0 | 13.9 GB | 248K | 12–16 |
| Qwen3.5 9B 9.7B | GGUF Q8_0 | 11.3 GB | 145K | 15–20 |
| Qwen3 8B 8.2B | GGUF Q8_0 | 10.7 GB | 40K (full) | 16–22 |
| Llama 3.1 8B 8.0B | GGUF Q8_0 | 10.3 GB | 49K | 16–22 |
| Gemma 4 E4B 8.0B | GGUF Q8_0 | 9.36 GB | 128K (full) | 18–25 |
| Nemotron 3 Nano 4B 4.0B | FP16 / BF16 | 8.78 GB | 256K (full) | 19–26 |
| MiniCPM5 2B 2.5B | FP16 / BF16 | 6.02 GB | 128K (full) | 28–39 |
Too big for one card
How many RTX 4060 Ti 16GB cards these need at Q4_K_M in one tensor-parallel group; some need more than 8.
- Qwen3.6 35B-A3B: 2 cards
- Llama 3.1 70B: 4 cards
- Qwen3-Coder-Next: 4 cards
- gpt-oss-120b: 8 cards
- Nemotron 3 Super 120B-A12B: 8 cards
- Qwen3.5 122B-A10B: 8 cards
- Mistral Medium 3.5 128B: 8 cards
- Qwen3.8 Flash Next: 8 cards
- MiniMax M2.7: more than 8 cards
- DeepSeek V4 Flash: more than 8 cards
- DeepSeek V4 Flash 0731: more than 8 cards
- MiMo V2.6 Flash: more than 8 cards
- GLM-5.3 Flash: more than 8 cards
- MiniMax M3: more than 8 cards
- DeepSeek V3 / R1: more than 8 cards
- DeepSeek V3.2: more than 8 cards
- GLM-5.3: more than 8 cards
- GLM-5.2: more than 8 cards
- DeepSeek V4.1 Flash: more than 8 cards
- MiMo V2.6 Pro: more than 8 cards
- DeepSeek V4 Pro: more than 8 cards
- Qwen3.8 2.4T-A95B: more than 8 cards
- Kimi K3: more than 8 cards
How these numbers are worked out
Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.
Other GPUs and Macs
- RTX 4060 8GB
- RTX 3060 12GB
- Arc B580 12GB
- RTX 4070 12GB
- RTX 5070 12GB
- RTX 5060 Ti 16GB
- RX 9070 XT 16GB
- RTX 4070 Ti Super 16GB
- RTX 4080 Super 16GB
- RTX 5070 Ti 16GB
- RTX 5080 16GB
- RTX 3090
- RTX 4090
- RX 7900 XTX
- RTX 5090
- M4 Pro Mac (64 GB)
- M4 Max Mac (128 GB)
- M3 Ultra Mac Studio (512 GB)
- Ryzen AI Max+ 395 (128 GB)
- DGX Spark (128 GB)
- L40S
- RTX PRO 6000 Blackwell
- A100 80GB
- H100 SXM
- H200
- B200
- 2× RTX 3060 12GB
- 2× RTX 3090
- 4× RTX 3090
- 2× RTX 4090
- 2× RTX 5090
- 8× H100 SXM
- 8× H200
Model numbers read from Hugging Face on .