What LLMs can an RTX 5070 Ti 16GB run?

An RTX 5070 Ti 16GB gives a model 16 GB of memory and 896 GB/s of bandwidth. Of the 40 open models tracked here, 8 fit at Q4_K_M with an 8,192-token context, and 9 more at a lower precision.

Models that fit on one RTX 5070 Ti 16GB

The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first.

Model Best precision Memory Longest context Tokens/s
Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active GGUF Q2_K 14.1 GB 256K (full) 140–257
Gemma 4 31B 31.3B GGUF Q2_K 15.1 GB 28K 33–46
GLM-4.7 Flash 31.2B, 3B active GGUF Q2_K 14.3 GB 36K 128–233
Xing 4.0 29B-A4B 31.2B, 4B active GGUF Q2_K 14.3 GB 43K 109–197
Qwen3 30B-A3B 30.5B, 3.3B active GGUF Q2_K 14.4 GB 23K 104–186
Muse Glimmer 30B 29.8B GGUF Q3_K_M 15.6 GB 36K 32–45
Qwen3.6 27B 27.8B GGUF Q3_K_M 15.0 GB 22K 33–47
Qwen3.8 27B 27.8B GGUF Q3_K_M 15.0 GB 22K 33–47
Gemma 4 26B-A4B 25.8B, 4B active GGUF Q3_K_M 13.7 GB 219K 101–181
gpt-oss-20b 20.9B, 3.6B active GGUF Q5_K_M 15.9 GB 11K 85–150
Gemma 4 12B 12.0B GGUF Q8_0 13.9 GB 248K 36–50
Qwen3.5 9B 9.7B GGUF Q8_0 11.3 GB 145K 44–62
Qwen3 8B 8.2B GGUF Q8_0 10.7 GB 40K (full) 46–66
Llama 3.1 8B 8.0B GGUF Q8_0 10.3 GB 49K 48–68
Gemma 4 E4B 8.0B GGUF Q8_0 9.36 GB 128K (full) 52–75
Nemotron 3 Nano 4B 4.0B FP16 / BF16 8.78 GB 256K (full) 56–80
MiniCPM5 2B 2.5B FP16 / BF16 6.02 GB 128K (full) 80–117

Too big for one card

How many RTX 5070 Ti 16GB cards these need at Q4_K_M in one tensor-parallel group; some need more than 8.

How these numbers are worked out

Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.

Other GPUs and Macs

Model numbers read from Hugging Face on .