Can I run Nemotron 3 Nano 30B-A3B on an RTX 4070 12GB?

With --n-cpu-moe 30: the Q4_K_M GGUF keeps 11.9 GB on the RTX 4070 and 12.2 GB in system RAM, with the experts of 30 of its 52 layers moved there. Whole, Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, 8.28 GB more than the RTX 4070 12GB holds. The smallest setup here that holds Nemotron 3 Nano 30B-A3B at Q4_K_M with 32K is RTX 3090 (24 GB).

With offload the Q4_K_M GGUF with --n-cpu-moe 30 and 32K tokens of context

On the card
11.9 GB
RTX 4070
12 GB, 504 GB/s
In RAM
12.2 GB
Tokens/s
26–44 tokens/s

With the Q4_K_M GGUF, --n-cpu-moe 30 and 32K tokens of context it writes about 26–44 tokens/s for one request on an RTX 4070 12GB and dual-channel DDR5-5600.

Nemotron 3 Nano 30B-A3B on the RTX 4070 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context Q4_K_MFreeTokens/sQ8_0FreeTokens/s
4K 20.1 GB −8.10 GB — 34.9 GB −22.9 GB —
8K 20.1 GB −8.12 GB — 34.9 GB −22.9 GB —
16K 20.2 GB −8.17 GB — 35.0 GB −23.0 GB —
32K 20.3 GB −8.28 GB — 35.1 GB −23.1 GB —
64K 20.5 GB −8.48 GB — 35.3 GB −23.3 GB —
128K 20.9 GB −8.90 GB — 35.7 GB −23.7 GB —
256K (full) 21.7 GB −9.72 GB — 36.5 GB −24.5 GB —

--n-cpu-moe for Nemotron 3 Nano 30B-A3B on the RTX 4070

With --n-cpu-moe 30 the Q4_K_M GGUF keeps 11.9 GB on the RTX 4070 and 12.2 GB in RAM, about 26–44 tokens/s with DDR5-5600. Measured Q4_K_M file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 30 11.7 GB 12.2 GB 19–3127–4629–49
32K 30 11.9 GB 12.2 GB 18–3126–4428–48
64K 32 11.2 GB 13.0 GB 17–2925–4226–45
128K 32 11.6 GB 13.0 GB 16–2723–3925–42
256K (full) 35 11.5 GB 13.8 GB 14–2420–3421–36

Nemotron 3 Nano 30B-A3B on more than one RTX 4070

CardsBest at 32KTokens/sQ4_K_M longestQ4_K_M tokens/s
2× (24 GB) Q4_K_M 90–169 256K (full) 90–169
4× (48 GB) Q8_0 100–193 256K (full) 136–278

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run Nemotron 3 Nano 30B-A3B on the RTX 4070 with llama-server

llama-server -hf unsloth/Nemotron-3-Nano-30B-A3B-GGUF:Q4_K_M -c 32768 --n-cpu-moe 30

Nemotron-3-Nano-30B-A3B-Q4_K_M.gguf, 24.6 GB, from unsloth/Nemotron-3-Nano-30B-A3B-GGUF (checked 2026-09-29); --n-cpu-moe 30 with -c 32768 is the plan above.

Questions

Can I run Nemotron 3 Nano 30B-A3B on an RTX 4070 12GB?

With --n-cpu-moe 30: the Q4_K_M GGUF keeps 11.9 GB on the RTX 4070 and 12.2 GB in system RAM, with the experts of 30 of its 52 layers moved there. Whole, Nemotron 3 Nano 30B-A3B needs about 20.3 GB at Q4_K_M with 32K tokens of context, 8.28 GB more than the RTX 4070 12GB holds. The smallest setup here that holds Nemotron 3 Nano 30B-A3B at Q4_K_M with 32K is RTX 3090 (24 GB).

How fast is Nemotron 3 Nano 30B-A3B on an RTX 4070 12GB?

With the Q4_K_M GGUF, --n-cpu-moe 30 and 32K tokens of context it writes about 26–44 tokens/s for one request on an RTX 4070 12GB and dual-channel DDR5-5600.

Does Nemotron 3 Nano 30B-A3B need --n-cpu-moe on an RTX 4070 12GB?

With --n-cpu-moe 30 the Q4_K_M GGUF keeps 11.9 GB on the RTX 4070 and 12.2 GB in RAM, about 26–44 tokens/s with DDR5-5600.

What does a second RTX 4070 12GB change for Nemotron 3 Nano 30B-A3B?

Two RTX 4070 12GB cards (24 GB in one tensor-parallel group) hold Nemotron 3 Nano 30B-A3B at Q4_K_M with 32K, and Q4_K_M up to 256K (full) tokens, at about 90–169 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also Nemotron 3 Nano 30B-A3B VRAM requirements, what LLMs an RTX 4070 12GB can run and every pair, or detect your own GPU. Model data checked .