Can I run gpt-oss-20b on an RTX 4060 Ti 16GB?

Yes: gpt-oss-20b needs about 15.4 GB at MXFP4 with 32K tokens of context, which fits the RTX 4060 Ti 16GB with 571 MB to spare. MXFP4 is the format gpt-oss-20b is published in, and on the RTX 4060 Ti it runs up to 34K tokens with 0.5 GB to spare.

Why 15.4 GB with vLLM, but 12.5 GB on the RTX 4060 Ti with llama.cpp

The 15.4 GB counts the published checkpoint as vLLM loads it: 12.8 GB of weights, the 32K KV cache, 0.5 GB and 10% overhead. llama.cpp's measured MXFP4 GGUF is 11.3 GB, and its 587 MB of token embeddings stay in RAM. With every layer on the RTX 4060 Ti it takes 12.5 GB at 32K with 1 GB of buffers, 3.54 GB under its 16 GB.

Yes MXFP4 with 32K tokens of context

MXFP4, 32K
15.4 GB
RTX 4060 Ti
16 GB, 288 GB/s
To spare
571 MB
Tokens/s
26–44 tokens/s

At MXFP4 with 32K tokens of context it writes about 26–44 tokens/s for one request on an RTX 4060 Ti 16GB.

Best precision for gpt-oss-20b on an RTX 4060 Ti 16GB

The most precise setting that leaves at least 0.5 GB free; one that fits with less is marked tight.

ContextBest fitMemoryFreeTokens/s
8K MXFP4 14.8 GB 1.18 GB 32–54
32K MXFP4 15.4 GB 571 MB 26–44
128K (full) Nothing fits — — —

gpt-oss-20b on the RTX 4060 Ti as the context fills

One request, FP16 KV cache, 0.5 GB plus 10% overhead; the MXFP4 GGUF columns are the measured llama.cpp file with every layer on the GPU and 1 GB of buffers; a minus sign is memory missing, and tight is less than 0.5 GB free.

Context MXFP4FreeTokens/sMXFP4 GGUFFreeTokens/s
4K 14.7 GB 1.28 GB 33–56 11.8 GB 4.20 GB 31–52
8K 14.8 GB 1.18 GB 32–54 11.9 GB 4.10 GB 30–50
16K 15.0 GB 994 MB 30–50 12.1 GB 3.91 GB 28–47
32K 15.4 GB 571 MB 26–44 12.5 GB 3.54 GB 24–41
64K 16.3 GB −274 MB — 13.2 GB 2.79 GB 20–34
128K (full) 17.9 GB −1.92 GB — 14.7 GB 1.29 GB 15–24

--n-cpu-moe for gpt-oss-20b on the RTX 4060 Ti

The measured MXFP4 GGUF fits the RTX 4060 Ti whole at 32K (12.5 GB), so --n-cpu-moe is not needed there. Measured MXFP4 file, 1 GB of buffers; tokens/s by system RAM speed.

Context--n-cpu-moeOn the cardIn RAM DDR4-3200DDR5-5600DDR5-6400
8K 0 11.9 GB 587 MB 30–5030–5030–50
32K 0 12.5 GB 587 MB 24–4124–4124–41
64K 0 13.2 GB 587 MB 20–3420–3420–34
128K (full) 0 14.7 GB 587 MB 15–2415–2415–24

gpt-oss-20b on more than one RTX 4060 Ti

CardsBest at 32KTokens/sMXFP4 longestMXFP4 tokens/s
2× (32 GB) MXFP4 46–81 128K (full) 46–81
4× (64 GB) MXFP4 79–146 128K (full) 79–146

Tensor parallel with the combined bandwidth, as in vLLM; llama.cpp splits layers by default and runs at about one card's speed.

Other options

Run gpt-oss-20b on the RTX 4060 Ti with llama-server

llama-server -hf ggml-org/gpt-oss-20b-GGUF:MXFP4 -c 34816 -np 1

gpt-oss-20b-MXFP4.gguf, 12.1 GB, from ggml-org/gpt-oss-20b-GGUF (checked 2026-09-29). At -c 34816 on the RTX 4060 Ti: 15.5 GB of 16 GB, 518 MB free. -np 1: one slot, one sliding window.

Questions

Can I run gpt-oss-20b on an RTX 4060 Ti 16GB?

Yes: gpt-oss-20b needs about 15.4 GB at MXFP4 with 32K tokens of context, which fits the RTX 4060 Ti 16GB with 571 MB to spare. MXFP4 is the format gpt-oss-20b is published in, and on the RTX 4060 Ti it runs up to 34K tokens with 0.5 GB to spare.

How fast is gpt-oss-20b on an RTX 4060 Ti 16GB?

At MXFP4 with 32K tokens of context it writes about 26–44 tokens/s for one request on an RTX 4060 Ti 16GB.

Does gpt-oss-20b need --n-cpu-moe on an RTX 4060 Ti 16GB?

The measured MXFP4 GGUF fits the RTX 4060 Ti whole at 32K (12.5 GB), so --n-cpu-moe is not needed there.

What does a second RTX 4060 Ti 16GB change for gpt-oss-20b?

Two RTX 4060 Ti 16GB cards (32 GB in one tensor-parallel group) hold gpt-oss-20b at MXFP4 with 32K, and MXFP4 up to 128K (full) tokens, at about 46–81 tokens/s.

Try other settings in the VRAM calculator, the speed calculator or the MoE offload planner. See also gpt-oss-20b VRAM requirements, what LLMs an RTX 4060 Ti 16GB can run and every pair, or detect your own GPU. Model data checked .