Can I run Nemotron 3 Nano 30B-A3B on an RTX 3060 12GB?
With --n-cpu-moe 30: the Q4_K_M GGUF keeps 11.9 GB on the RTX 3060 and 12.2 GB in system RAM, with the experts of 30 of its 52 layers moved there. Whole, Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, 8.28 GB more than the RTX 3060 12GB holds. The smallest setup here that holds Nemotron 3 Nano 30B-A3B at Q4_K_M with 32K is RTX 3090 (24 GB).
With offload the Q4_K_M GGUF with --n-cpu-moe 30 and 32K tokens of context
- On the card
- 11.9 GB
- RTX 3060
- 12 GB, 360 GB/s
- In RAM
- 12.2 GB
- Tokens/s
- 23–39 tokens/s
With the Q4_K_M GGUF, --n-cpu-moe 30 and 32K tokens of context it writes about 23–39 tokens/s for one request on an RTX 3060 12GB and dual-channel DDR5-5600.
Nemotron 3 Nano 30B-A3B on the RTX 3060 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 20.1 GB | −8.10 GB | — | 34.9 GB | −22.9 GB | — |
| 8K | 20.1 GB | −8.12 GB | — | 34.9 GB | −22.9 GB | — |
| 16K | 20.2 GB | −8.17 GB | — | 35.0 GB | −23.0 GB | — |
| 32K | 20.3 GB | −8.28 GB | — | 35.1 GB | −23.1 GB | — |
| 64K | 20.5 GB | −8.48 GB | — | 35.3 GB | −23.3 GB | — |
| 128K | 20.9 GB | −8.90 GB | — | 35.7 GB | −23.7 GB | — |
| 256K (full) | 21.7 GB | −9.72 GB | — | 36.5 GB | −24.5 GB | — |
--n-cpu-moe for Nemotron 3 Nano 30B-A3B on the RTX 3060
With --n-cpu-moe 30 the Q4_K_M GGUF keeps 11.9 GB on the RTX 3060 and 12.2 GB in RAM, about 23–39 tokens/s with DDR5-5600. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.
Nemotron 3 Nano 30B-A3B on more than one RTX 3060
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s | Rent per hour |
|---|---|---|---|---|---|
| 2× (24 GB) | Q4_K_M | 70–129 | 256K (full) | 70–129 | $0.16 |
| 4× (48 GB) | Q8_0 | 80–148 | 256K (full) | 113–221 | $0.32 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 16.5 GB at 32K, 4.52 GB over the RTX 3060; it does not fit with 0.5 GB to spare even with 1K tokens.
- Smallest setup for Q4_K_M at 32K: RTX 3090 (24 GB).
Run Nemotron 3 Nano 30B-A3B on the RTX 3060 with llama-server
llama-server -hf unsloth/Nemotron-3-Nano-30B-A3B-GGUF:Q4_K_M -c 32768 --n-cpu-moe 30 Nemotron-3-Nano-30B-A3B-Q4_K_M.gguf, 24.6 GB, from unsloth/
Questions
Can I run Nemotron 3 Nano 30B-A3B on an RTX 3060 12GB?
With --n-cpu-moe 30: the Q4_K_M GGUF keeps 11.9 GB on the RTX 3060 and 12.2 GB in system RAM, with the experts of 30 of its 52 layers moved there. Whole, Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, 8.28 GB more than the RTX 3060 12GB holds. The smallest setup here that holds Nemotron 3 Nano 30B-A3B at Q4_K_M with 32K is RTX 3090 (24 GB).
How fast is Nemotron 3 Nano 30B-A3B on an RTX 3060 12GB?
With the Q4_K_M GGUF, --n-cpu-moe 30 and 32K tokens of context it writes about 23–39 tokens/s for one request on an RTX 3060 12GB and dual-channel DDR5-5600.
Does Nemotron 3 Nano 30B-A3B need --n-cpu-moe on an RTX 3060 12GB?
With --n-cpu-moe 30 the Q4_K_M GGUF keeps 11.9 GB on the RTX 3060 and 12.2 GB in RAM, about 23–39 tokens/s with DDR5-5600.
What does a second RTX 3060 12GB change for Nemotron 3 Nano 30B-A3B?
Two RTX 3060 12GB cards (24 GB in one tensor-parallel group) hold Nemotron 3 Nano 30B-A3B at Q4_K_M with 32K, and Q4_K_M up to 256K (full) tokens, at about 70–129 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Nemotron 3 Nano 30B-A3B VRAM requirements, what LLMs an RTX 3060 12GB can run and every pair, or detect your own GPU. Model data checked .