Can I run Nemotron 3 Nano 30B-A3B on an RTX 4090?
Yes: Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RTX 4090 with 3.72 GB to spare. Nothing more precise fits at 32K with 0.5 GB to spare; Q4_K_M runs up to 256K (full) tokens on the RTX 4090.
Yes Q4_K_M with 32K tokens of context
- Q4_K_M, 32K
- 20.3 GB
- RTX 4090
- 24 GB, 1,008 GB/s
- To spare
- 3.72 GB
- Tokens/s
- 109–196 tokens/s
At Q4_K_M with 32K tokens of context it writes about 109–196 tokens/s for one request on an RTX 4090.
Best precision for Nemotron 3 Nano 30B-A3B on an RTX 4090
The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.
| Context | Best fit | Memory | Free | Tokens/s |
|---|---|---|---|---|
| 8K | Q5_K_M | 23.5 GB | 533 MB | 101–181 |
| 32K | Q4_K_M | 20.3 GB | 3.72 GB | 109–196 |
| 128K | Q4_K_M | 20.9 GB | 3.10 GB | 90–159 |
| 256K (full) | Q4_K_M | 21.7 GB | 2.28 GB | 72–127 |
Nemotron 3 Nano 30B-A3B on the RTX 4090 as the context fills
One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.
| Context | Q4_K_M | Free | Tokens/s | Q8_0 | Free | Tokens/s |
|---|---|---|---|---|---|---|
| 4K | 20.1 GB | 3.90 GB | 116–210 | 34.9 GB | −10.9 GB | — |
| 8K | 20.1 GB | 3.88 GB | 115–208 | 34.9 GB | −10.9 GB | — |
| 16K | 20.2 GB | 3.83 GB | 113–204 | 35.0 GB | −11.0 GB | — |
| 32K | 20.3 GB | 3.72 GB | 109–196 | 35.1 GB | −11.1 GB | — |
| 64K | 20.5 GB | 3.52 GB | 102–182 | 35.3 GB | −11.3 GB | — |
| 128K | 20.9 GB | 3.10 GB | 90–159 | 35.7 GB | −11.7 GB | — |
| 256K (full) | 21.7 GB | 2.28 GB | 72–127 | 36.5 GB | −12.5 GB | — |
--n-cpu-moe for Nemotron 3 Nano 30B-A3B on the RTX 4090
The measured Q4_K_M GGUF fits the RTX 4090 whole at 32K (23.8 GB), so --n-cpu-moe is not needed there. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.
Nemotron 3 Nano 30B-A3B on more than one RTX 4090
| Cards | Best at 32K | Tokens/s | Q4_K_M longest | Q4_K_M tokens/s | Rent per hour |
|---|---|---|---|---|---|
| 2× (48 GB) | Q8_0 | 100–193 | 256K (full) | 136–278 | $1.06 |
| 4× (96 GB) | BF16 | 106–205 | 256K (full) | 185–408 | $2.12 |
Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.
Other options
- The next smaller setting, Q3_K_M, takes 16.5 GB at 32K, 7.48 GB under the RTX 4090; it fits with 0.5 GB to spare up to 256K (full) tokens.
- Renting an RTX 4090 costs about $0.53 an hour (median on getdeploying.com, 2026-09-29): $0.75–1.3 per million tokens at 109–196 tokens/s.
- Smallest setup for Q4_K_M at 32K: RTX 3090 (24 GB).
Run Nemotron 3 Nano 30B-A3B on the RTX 4090 with llama-server
llama-server -hf ggml-org/NVIDIA-Nemotron-3-Nano-30B-A3B-GGUF:Q4_K_M -c 4096 NVIDIA-Nemotron-3-Nano-30B-A3B-Q4_K_M.gguf, 22.4 GB, from ggml-org/
Questions
Can I run Nemotron 3 Nano 30B-A3B on an RTX 4090?
Yes: Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, which fits the 24 GB RTX 4090 with 3.72 GB to spare. Nothing more precise fits at 32K with 0.5 GB to spare; Q4_K_M runs up to 256K (full) tokens on the RTX 4090.
How fast is Nemotron 3 Nano 30B-A3B on an RTX 4090?
At Q4_K_M with 32K tokens of context it writes about 109–196 tokens/s for one request on an RTX 4090.
Does Nemotron 3 Nano 30B-A3B need --n-cpu-moe on an RTX 4090?
The measured Q4_K_M GGUF fits the RTX 4090 whole at 32K (23.8 GB), so --n-cpu-moe is not needed there.
What does a second RTX 4090 change for Nemotron 3 Nano 30B-A3B?
Two RTX 4090 cards (48 GB in one tensor-parallel group) hold Nemotron 3 Nano 30B-A3B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 136–278 tokens/s.
Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Nemotron 3 Nano 30B-A3B VRAM requirements, what LLMs an RTX 4090 can run and every pair, or detect your own GPU. Model data checked .