Can I run Nemotron 3 Nano 30B-A3B on an RTX 5060 Ti 16GB?

Only at Q2_K: Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, 4.28 GB more than the RTX 5060 Ti 16GB holds, but 14.3 GB at Q2_K, which fits with 1.75 GB to spare. The smallest setup here that holds Nemotron 3 Nano 30B-A3B at Q4_K_M with 32K is RTX 3090 (24 GB).

Partly Q2_K with 32K tokens of context

Q4_K_M, 32K
20.3 GB
RTX 5060 Ti
16 GB, 448 GB/s
Short by
4.28 GB
Tokens/s
72–126 tokens/s

At Q2_K with 32K tokens of context it writes about 72–126 tokens/s for one request on an RTX 5060 Ti 16GB.

Best precision for Nemotron 3 Nano 30B-A3B on an RTX 5060 Ti 16GB

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q2_K 14.1 GB 1.90 GB 78–138
32K Q2_K 14.3 GB 1.75 GB 72–126
128K Q2_K 14.9 GB 1.13 GB 54–94
256K (full) IQ3_XXS 15.5 GB 518 MB 41–71

Nemotron 3 Nano 30B-A3B on the RTX 5060 Ti as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 20.1 GB −4.10 GB — 34.9 GB −18.9 GB —
8K 20.1 GB −4.12 GB — 34.9 GB −18.9 GB —
16K 20.2 GB −4.17 GB — 35.0 GB −19.0 GB —
32K 20.3 GB −4.28 GB — 35.1 GB −19.1 GB —
64K 20.5 GB −4.48 GB — 35.3 GB −19.3 GB —
128K 20.9 GB −4.90 GB — 35.7 GB −19.7 GB —
256K (full) 21.7 GB −5.72 GB — 36.5 GB −20.5 GB —

--n-cpu-moe for Nemotron 3 Nano 30B-A3B on the RTX 5060 Ti

With --n-cpu-moe 21 the Q4_K_M GGUF keeps 15.4 GB on the RTX 5060 Ti and 8.70 GB in RAM, about 29–49 tokens/s with DDR5-5600. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 21 15.2 GB 8.70 GB 22–3730–5132–54
32K 21 15.4 GB 8.70 GB 21–3629–4930–52
64K 21 15.6 GB 8.70 GB 21–3528–4729–50
128K 21 15.9 GB 8.70 GB 20–3325–4327–46
256K (full) 23 15.9 GB 9.52 GB 17–2822–3623–38

Nemotron 3 Nano 30B-A3B on more than one RTX 5060 Ti

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/s
2× (32 GB) Q6_K 67–123 256K (full) 82–154
4× (64 GB) Q8_0 93–176 256K (full) 128–257

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

llama-server command

llama-server -m NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-Q2_K.gguf -ngl 99 -c 32768

Questions

Can I run Nemotron 3 Nano 30B-A3B on an RTX 5060 Ti 16GB?

Only at Q2_K: Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, 4.28 GB more than the RTX 5060 Ti 16GB holds, but 14.3 GB at Q2_K, which fits with 1.75 GB to spare. The smallest setup here that holds Nemotron 3 Nano 30B-A3B at Q4_K_M with 32K is RTX 3090 (24 GB).

How fast is Nemotron 3 Nano 30B-A3B on an RTX 5060 Ti 16GB?

At Q2_K with 32K tokens of context it writes about 72–126 tokens/s for one request on an RTX 5060 Ti 16GB.

Does Nemotron 3 Nano 30B-A3B need --n-cpu-moe on an RTX 5060 Ti 16GB?

With --n-cpu-moe 21 the Q4_K_M GGUF keeps 15.4 GB on the RTX 5060 Ti and 8.70 GB in RAM, about 29–49 tokens/s with DDR5-5600.

What does a second RTX 5060 Ti 16GB change for Nemotron 3 Nano 30B-A3B?

Two RTX 5060 Ti 16GB cards (32 GB in one tensor-parallel group) hold Nemotron 3 Nano 30B-A3B at Q6_K with 32K, and Q4_K_M up to 256K (full) tokens, at about 82–154 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Nemotron 3 Nano 30B-A3B VRAM requirements, what LLMs an RTX 5060 Ti 16GB can run and every pair, or detect your own GPU. Model data checked .