What LLMs can two RTX 3060 12GB GPUs run?

Two RTX 3060 12GB GPUs give a model 24 GB (2 × 12 GB) of memory and 720 GB/s of bandwidth in total. Of the 40 open models tracked here, 18 fit at Q4_K_M with an 8,192-token context. The speeds assume tensor parallelism, as in vLLM or SGLang; llama.cpp splits layers across cards by default and then writes at about the speed of one card.

Models that fit on 2× RTX 3060 12GB

The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first.

Model Best precision Memory Longest context Tokens/s
Qwen3.6 35B-A3B 36.0B, 3B active GGUF Q4_K_M 23.0 GB 33K 79–147
Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active GGUF Q5_K_M 23.5 GB 10K 66–120
Gemma 4 31B 31.3B GGUF Q4_K_M 21.1 GB 64K 18–26
GLM-4.7 Flash 31.2B, 3B active GGUF Q4_K_M 20.3 GB 64K 72–132
Xing 4.0 29B-A4B 31.2B, 4B active GGUF Q4_K_M 20.2 GB 75K 61–110
Qwen3 30B-A3B 30.5B, 3.3B active GGUF Q5_K_M 23.5 GB 8K 55–100
Muse Glimmer 30B 29.8B GGUF Q5_K_M 22.3 GB 92K 17–25
Qwen3.6 27B 27.8B GGUF Q5_K_M 21.2 GB 41K 18–26
Qwen3.8 27B 27.8B GGUF Q5_K_M 21.2 GB 41K 18–26
Gemma 4 26B-A4B 25.8B, 4B active GGUF Q6_K 22.5 GB 102K 50–89
gpt-oss-20b 20.9B, 3.6B active GGUF Q8_0 23.5 GB 8K 45–80
Gemma 4 12B 12.0B GGUF Q8_0 13.9 GB 256K (full) 27–39
Qwen3.5 9B 9.7B FP16 / BF16 20.6 GB 93K 19–27
Qwen3 8B 8.2B FP16 / BF16 18.5 GB 40K (full) 21–30
Llama 3.1 8B 8.0B FP16 / BF16 18.1 GB 47K 21–30
Gemma 4 E4B 8.0B FP16 / BF16 17.0 GB 128K (full) 23–32
Nemotron 3 Nano 4B 4.0B FP16 / BF16 8.78 GB 256K (full) 42–61
MiniCPM5 2B 2.5B FP16 / BF16 6.02 GB 128K (full) 58–89

Too big for 2× RTX 3060 12GB

How many RTX 3060 12GB cards these need at Q4_K_M in one tensor-parallel group; some need more than 8.

How these numbers are worked out

Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.

Other GPUs and Macs

Model numbers read from Hugging Face on .