What LLMs can two RTX 5090 GPUs run?

Two RTX 5090 GPUs give a model 64 GB (2 × 32 GB) of memory and 3,584 GB/s of bandwidth in total. Of the 40 open models tracked here, 20 fit at Q4_K_M with an 8,192-token context, and 4 more at a lower precision. The speeds assume tensor parallelism, as in vLLM or SGLang; llama.cpp splits layers across cards by default and then writes at about the speed of one card.

Models that fit on 2× RTX 5090

The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first.

Model Best precision Memory Longest context Tokens/s
Mistral Medium 3.5 128B 127.7B GGUF Q2_K 58.3 GB 21K 31–45
Qwen3.5 122B-A10B 125.1B, 10B active GGUF Q3_K_M 63.3 GB 14K 121–242
Nemotron 3 Super 120B-A12B 123.6B, 12B active GGUF Q3_K_M 62.5 GB 128K 111–217
gpt-oss-120b 116.8B, 5.1B active GGUF Q3_K_M 59.3 GB 116K 164–349
Qwen3-Coder-Next 79.7B, 3B active GGUF Q5_K_M 58.6 GB 199K 177–385
Llama 3.1 70B 70.6B GGUF Q6_K 62.5 GB 10K 29–42
Qwen3.6 35B-A3B 36.0B, 3B active GGUF Q8_0 39.8 GB 256K (full) 151–315
Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active GGUF Q8_0 34.9 GB 256K (full) 143–294
Gemma 4 31B 31.3B GGUF Q8_0 35.7 GB 256K (full) 48–71
GLM-4.7 Flash 31.2B, 3B active GGUF Q8_0 34.9 GB 198K (full) 145–301
Xing 4.0 29B-A4B 31.2B, 4B active GGUF Q8_0 34.9 GB 256K (full) 128–258
Qwen3 30B-A3B 30.5B, 3.3B active GGUF Q8_0 34.6 GB 40K (full) 133–270
Muse Glimmer 30B 29.8B FP16 / BF16 61.7 GB 128K (full) 30–43
Qwen3.6 27B 27.8B FP16 / BF16 58.0 GB 88K 31–45
Qwen3.8 27B 27.8B FP16 / BF16 58.0 GB 88K 31–45
Gemma 4 26B-A4B 25.8B, 4B active FP16 / BF16 53.7 GB 256K (full) 89–169
gpt-oss-20b 20.9B, 3.6B active FP16 / BF16 43.6 GB 128K (full) 96–184
Gemma 4 12B 12.0B FP16 / BF16 25.4 GB 256K (full) 63–97
Qwen3.5 9B 9.7B FP16 / BF16 20.6 GB 256K (full) 74–117
Qwen3 8B 8.2B FP16 / BF16 18.5 GB 40K (full) 80–127
Llama 3.1 8B 8.0B FP16 / BF16 18.1 GB 128K (full) 82–130
Gemma 4 E4B 8.0B FP16 / BF16 17.0 GB 128K (full) 86–137
Nemotron 3 Nano 4B 4.0B FP16 / BF16 8.78 GB 256K (full) 132–232
MiniCPM5 2B 2.5B FP16 / BF16 6.02 GB 128K (full) 160–303

Too big for 2× RTX 5090

How many RTX 5090 cards these need at Q4_K_M in one tensor-parallel group; some need more than 8.

How these numbers are worked out

Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.

Other GPUs and Macs

Model numbers read from Hugging Face on .