Can I run Nemotron 3 Nano 30B-A3B on an RTX 5090?

Yes: Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, which fits the 32 GB RTX 5090 with 11.7 GB to spare. At 32K the RTX 5090 holds up to Q6_K (27.2 GB), and Q4_K_M runs up to 256K (full) tokens.

Yes Q4_K_M with 32K tokens of context

Q4_K_M, 32K
20.3 GB
RTX 5090
32 GB, 1,792 GB/s
To spare
11.7 GB
Tokens/s
172–324 tokens/s

At Q4_K_M with 32K tokens of context it writes about 172–324 tokens/s for one request on an RTX 5090.

Best precision for Nemotron 3 Nano 30B-A3B on an RTX 5090

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K Q6_K 27.1 GB 4.92 GB 144–266
32K Q6_K 27.2 GB 4.77 GB 139–255
128K Q6_K 27.9 GB 4.15 GB 120–217
256K (full) Q6_K 28.7 GB 3.32 GB 102–182

Nemotron 3 Nano 30B-A3B on the RTX 5090 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 20.1 GB 11.9 GB 182–346 34.9 GB −2.90 GB —
8K 20.1 GB 11.9 GB 181–343 34.9 GB −2.92 GB —
16K 20.2 GB 11.8 GB 178–336 35.0 GB −2.98 GB —
32K 20.3 GB 11.7 GB 172–324 35.1 GB −3.08 GB —
64K 20.5 GB 11.5 GB 162–302 35.3 GB −3.28 GB —
128K 20.9 GB 11.1 GB 144–266 35.7 GB −3.70 GB —
256K (full) 21.7 GB 10.3 GB 119–215 36.5 GB −4.52 GB —

--n-cpu-moe for Nemotron 3 Nano 30B-A3B on the RTX 5090

The measured Q4_K_M GGUF fits the RTX 5090 whole at 32K (23.8 GB), so --n-cpu-moe is not needed there. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 0 23.7 GB 231 MB 157–292157–292157–292
32K 0 23.8 GB 231 MB 150–279150–279150–279
64K 0 24.0 GB 231 MB 142–262142–262142–262
128K 0 24.4 GB 231 MB 129–235129–235129–235
256K (full) 0 25.2 GB 231 MB 108–194108–194108–194

Nemotron 3 Nano 30B-A3B on more than one RTX 5090

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/sRent per hour
2× (64 GB) Q8_0 140–287 256K (full) 177–386 $1.38
4× (128 GB) BF16 146–302 256K (full) 218–514 $2.76

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Nemotron 3 Nano 30B-A3B on the RTX 5090 with llama-server

llama-server -hf ggml-org/NVIDIA-Nemotron-3-Nano-30B-A3B-GGUF:Q4_K_M -c 262144

NVIDIA-Nemotron-3-Nano-30B-A3B-Q4_K_M.gguf, 22.4 GB, from ggml-org/NVIDIA-Nemotron-3-Nano-30B-A3B-GGUF (checked 2026-09-29). At -c 262144 on the RTX 5090: 25.1 GB of 32 GB, 6.88 GB free. The file is 3.09 GB over the estimate above, so -c counts the file.

Questions

Can I run Nemotron 3 Nano 30B-A3B on an RTX 5090?

Yes: Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, which fits the 32 GB RTX 5090 with 11.7 GB to spare. At 32K the RTX 5090 holds up to Q6_K (27.2 GB), and Q4_K_M runs up to 256K (full) tokens.

How fast is Nemotron 3 Nano 30B-A3B on an RTX 5090?

At Q4_K_M with 32K tokens of context it writes about 172–324 tokens/s for one request on an RTX 5090.

Does Nemotron 3 Nano 30B-A3B need --n-cpu-moe on an RTX 5090?

The measured Q4_K_M GGUF fits the RTX 5090 whole at 32K (23.8 GB), so --n-cpu-moe is not needed there.

What does a second RTX 5090 change for Nemotron 3 Nano 30B-A3B?

Two RTX 5090 cards (64 GB in one tensor-parallel group) hold Nemotron 3 Nano 30B-A3B at Q8_0 with 32K, and Q4_K_M up to 256K (full) tokens, at about 177–386 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Nemotron 3 Nano 30B-A3B VRAM requirements, what LLMs an RTX 5090 can run and every pair, or detect your own GPU. Model data checked .