Can I run gpt-oss-120b on an RTX 3060 12GB?

With --n-cpu-moe 31: the MXFP4 GGUF keeps 11.6 GB on the RTX 3060 and 49.6 GB in system RAM, with the experts of 31 of its 36 layers moved there. Whole, gpt-oss-120b needs about 68.6 GB at MXFP4 with 32K tokens of context, 56.6 GB more than the RTX 3060 12GB holds. The smallest setup here that holds gpt-oss-120b at MXFP4 with 32K is A100 80GB.

Why 68.6 GB with vLLM, but 11.6 GB on the RTX 3060 with llama.cpp

The 68.6 GB counts the published checkpoint as vLLM loads it: 60.8 GB of weights, the 32K KV cache, 0.5 GB and 10% overhead. llama.cpp's measured MXFP4 GGUF is 59.0 GB, and its 587 MB of token embeddings stay in RAM. With every layer on the RTX 3060 it takes 60.6 GB at 32K with 1 GB of buffers, 48.6 GB more than it holds, so the experts of 31 layers (49.0 GB) move to system RAM. That leaves 388 MB of the 12 GB free, and the PC needs 49.6 GB of free RAM for those experts and the embeddings.

With offload the MXFP4 GGUF with --n-cpu-moe 31 and 32K tokens of context

On the card
11.6 GB
RTX 3060
12 GB, 360 GB/s
In RAM
49.6 GB
Tokens/s
11–18 tokens/s

With the MXFP4 GGUF, --n-cpu-moe 31 and 32K tokens of context it writes about 11–18 tokens/s for one request on an RTX 3060 12GB and dual-channel DDR5-5600.

gpt-oss-120b on the RTX 3060 as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; the MXFP4 GGUF columns are the measured llama.cpp file with every layer on the GPU and 1 GB of buffers; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context MXFP4FreeTokens/sMXFP4 GGUFFreeTokens/s
4K 67.5 GB −55.5 GB — 59.6 GB −47.6 GB —
8K 67.7 GB −55.7 GB — 59.8 GB −47.8 GB —
16K 68.0 GB −56.0 GB — 60.0 GB −48.0 GB —
32K 68.6 GB −56.6 GB — 60.6 GB −48.6 GB —
64K 69.8 GB −57.8 GB — 61.7 GB −49.7 GB —
128K (full) 72.3 GB −60.3 GB — 64.0 GB −52.0 GB —

--n-cpu-moe for gpt-oss-120b on the RTX 3060

With --n-cpu-moe 31 the MXFP4 GGUF keeps 11.6 GB on the RTX 3060 and 49.6 GB in RAM, about 11–18 tokens/s with DDR5-5600. Measured MXFP4 file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 31 10.8 GB 49.6 GB 7.7–1312–2013–22
32K 31 11.6 GB 49.6 GB 7.2–1211–1812–20
64K 32 11.2 GB 51.1 GB 6.6–119.5–1610–17
128K (full) 33 11.8 GB 52.7 GB 5.6–9.47.8–138.3–14

gpt-oss-120b on more than one RTX 3060

CardsBest at 32KTokens/sMXFP4 longestMXFP4 tokens/sRent per hour
2× (24 GB) Nothing fits — Does not fit — $0.16
4× (48 GB) Nothing fits — Does not fit — $0.32

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run gpt-oss-120b on the RTX 3060 with llama-server

llama-server -hf ggml-org/gpt-oss-120b-GGUF:MXFP4 -c 32768 --n-cpu-moe 31 -np 1

gpt-oss-120b-MXFP4.gguf, 63.4 GB, from ggml-org/gpt-oss-120b-GGUF (checked 2026-09-29); --n-cpu-moe 31 with -c 32768 is the plan above. -np 1: one slot, one sliding window.

Questions

Can I run gpt-oss-120b on an RTX 3060 12GB?

With --n-cpu-moe 31: the MXFP4 GGUF keeps 11.6 GB on the RTX 3060 and 49.6 GB in system RAM, with the experts of 31 of its 36 layers moved there. Whole, gpt-oss-120b needs about 68.6 GB at MXFP4 with 32K tokens of context, 56.6 GB more than the RTX 3060 12GB holds. The smallest setup here that holds gpt-oss-120b at MXFP4 with 32K is A100 80GB.

How fast is gpt-oss-120b on an RTX 3060 12GB?

With the MXFP4 GGUF, --n-cpu-moe 31 and 32K tokens of context it writes about 11–18 tokens/s for one request on an RTX 3060 12GB and dual-channel DDR5-5600.

Does gpt-oss-120b need --n-cpu-moe on an RTX 3060 12GB?

With --n-cpu-moe 31 the MXFP4 GGUF keeps 11.6 GB on the RTX 3060 and 49.6 GB in RAM, about 11–18 tokens/s with DDR5-5600.

What does a second RTX 3060 12GB change for gpt-oss-120b?

Even two RTX 3060 12GB cards (24 GB) do not hold gpt-oss-120b at 32K.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also gpt-oss-120b VRAM requirements, what LLMs an RTX 3060 12GB can run and every pair, or detect your own GPU. Model data checked .