Can I run Nemotron 3 Nano 30B-A3B on an RTX 4060 Ti 16GB?
Only at Q2_K: Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, 4.28 GB more than the RTX 4060 Ti 16GB holds, but 14.3 GB at Q2_K, which fits with 1.75 GB to spare. The smallest setup here that holds Nemotron 3 Nano 30B-A3B at Q4_K_M with 32K is RTX 3090 (24 GB).
Partly Q2_K with 32K tokens of context
- Q4_K_M, 32K
- 20.3 GB
- RTX 4060 Ti
- 16 GB, 288 GB/s
- Short by
- 4.28 GB
- Tokens/s
- 48–83 tokens/s
At Q2_K with 32K tokens of context it writes about 48–83 tokens/s for one request on an RTX 4060 Ti 16GB.
Best precision for Nemotron 3 Nano 30B-A3B on an RTX 4060 Ti 16GB
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Q2_K | 14.1 GB | 1.90 GB | 53–91 |
| 32K | Q2_K | 14.3 GB | 1.75 GB | 48–83 |
| 128K | Q2_K | 14.9 GB | 1.13 GB | 36–61 |
| 256K (full) | IQ3_XXS | 15.5 GB | 518 MB | 27–46 |
Nemotron 3 Nano 30B-A3B on the RTX 4060 Ti as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 20.1 GB | −4.10 GB | — | 34.9 GB | −18.9 GB | — |
| 8K | 20.1 GB | −4.12 GB | — | 34.9 GB | −18.9 GB | — |
| 16K | 20.2 GB | −4.17 GB | — | 35.0 GB | −19.0 GB | — |
| 32K | 20.3 GB | −4.28 GB | — | 35.1 GB | −19.1 GB | — |
| 64K | 20.5 GB | −4.48 GB | — | 35.3 GB | −19.3 GB | — |
| 128K | 20.9 GB | −4.90 GB | — | 35.7 GB | −19.7 GB | — |
| 256K (full) | 21.7 GB | −5.72 GB | — | 36.5 GB | −20.5 GB | — |
--n-cpu-moe for Nemotron 3 Nano 30B-A3B on the RTX 4060 Ti
With --n-cpu-moe 21 the Q4_K_M GGUF keeps 15.4 GB on the RTX 4060 Ti and 8.70 GB in RAM, about 22–38 tokens/s with DDR5-5600. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.
Nemotron 3 Nano 30B-A3B on more than one RTX 4060 Ti
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s |
|---|---|---|---|---|
| 2× (32 GB) | Q6_K | 47–84 | 256K (full) | 59–107 |
| 4× (64 GB) | Q8_0 | 67–123 | 256K (full) | 98–188 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, IQ3_XXS, takes 14.1 GB at 32K, 1.95 GB under the RTX 4060 Ti; it fits with 0.5 GB to spare up to 256K (full) tokens.
- Smallest setup for Q4_K_M at 32K: RTX 3090 (24 GB).
llama-server command
llama-server -m NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-Q2_K.gguf -ngl 99 -c 32768 Questions
Can I run Nemotron 3 Nano 30B-A3B on an RTX 4060 Ti 16GB?
Only at Q2_K: Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, 4.28 GB more than the RTX 4060 Ti 16GB holds, but 14.3 GB at Q2_K, which fits with 1.75 GB to spare. The smallest setup here that holds Nemotron 3 Nano 30B-A3B at Q4_K_M with 32K is RTX 3090 (24 GB).
How fast is Nemotron 3 Nano 30B-A3B on an RTX 4060 Ti 16GB?
At Q2_K with 32K tokens of context it writes about 48–83 tokens/s for one request on an RTX 4060 Ti 16GB.
Does Nemotron 3 Nano 30B-A3B need --n-cpu-moe on an RTX 4060 Ti 16GB?
With --n-cpu-moe 21 the Q4_K_M GGUF keeps 15.4 GB on the RTX 4060 Ti and 8.70 GB in RAM, about 22–38 tokens/s with DDR5-5600.
What does a second RTX 4060 Ti 16GB change for Nemotron 3 Nano 30B-A3B?
Two RTX 4060 Ti 16GB cards (32 GB in one tensor-parallel group) hold Nemotron 3 Nano 30B-A3B at Q6_K with 32K, and Q4_K_M up to 256K (full) tokens, at about 59–107 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Nemotron 3 Nano 30B-A3B VRAM requirements, what LLMs an RTX 4060 Ti 16GB can run and every pair, or detect your own GPU. Model data checked .