LLM VRAM Calculator
Pick a model or load any one from Hugging Face, choose a quantization, context length and number of parallel requests, and see how much VRAM inference needs and which GPUs it fits on. Everything runs in your browser.
Advanced: the model numbers and overhead used
Editing any of these switches the model to Custom.
– estimated GPU memory
- Weights
- –
- KV cache
- –
- Runtime overhead
- –
- KV cache per token
- –
Which GPUs can run it
| GPU | Memory | Result |
|---|---|---|
| RTX 3060 | 12 GB | – |
| RTX 4060 Ti 16GB | 16 GB | – |
| RTX 3090 / 4090 | 24 GB | – |
| RTX 5090 | 32 GB | – |
| A100 40GB | 40 GB | – |
| Mac, 64 GB unified memory about 75% of it is usable by the GPU by default | 48 GB | – |
| L40S / RTX 6000 Ada | 48 GB | – |
| A100 / H100 80GB | 80 GB | – |
| Mac, 128 GB unified memory about 75% of it is usable by the GPU by default | 96 GB | – |
| H200 | 141 GB | – |
| B200 | 180 GB | – |
Fits? Then see how many tokens per second it will write.
VRAM for popular models
Total GPU memory with 8,192 tokens of context, one request and an FP16 KV cache. Each model name leads to its own page, with every quantization, the cache at long context and the GPUs that fit.
| Model | Parameters | As published | Q8_0 | Q4_K_M | Q4_K_M fits on |
|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 304.2B | 172 GB (4.4 bits/weight) | 332 GB | 190 GB | 2 × H200 |
| Qwen3.6 35B-A3B (MoE) | 36.0B | 74.3 GB (BF16) | 39.8 GB | 23.0 GB | RTX 3090 / 4090 |
| Qwen3.6 27B | 27.8B | 58.0 GB (BF16) | 31.3 GB | 18.3 GB | RTX 3090 / 4090 |
| Qwen3.8 27B | 27.8B | 58.0 GB (BF16) | 31.3 GB | 18.3 GB | RTX 3090 / 4090 |
| Qwen3.8 Flash Next (180B MoE) | 180.0B | 370 GB (BF16) | 197 GB | 112 GB | H200 |
| Qwen3.8 2.4T-A95B (MoE) | 2.45T | 5,013 GB (BF16) | 2,664 GB | 1,517 GB | More than 8 GPUs |
| Qwen3.5 9B | 9.7B | 20.6 GB (BF16) | 11.3 GB | 6.76 GB | RTX 3060 |
| Qwen3.5 122B-A10B (MoE) | 125.1B | 257 GB (BF16) | 137 GB | 78.2 GB | A100 / H100 80GB |
| Qwen3-Coder-Next (80B MoE) | 79.7B | 164 GB (BF16) | 87.4 GB | 50.1 GB | A100 / H100 80GB |
| DeepSeek V4.1 Flash | 763.2B | 524 GB (5.3 bits/weight) | 832 GB | 474 GB | 4 × H200 |
| DeepSeek V4 Flash | 290.9B | 165 GB (4.4 bits/weight) | 318 GB | 182 GB | 2 × H200 |
| DeepSeek V4 Pro | 1.60T | 887 GB (4.3 bits/weight) | 1,742 GB | 993 GB | 8 × H200 |
| DeepSeek V3.2 | 685.4B | 707 GB (FP8) | 747 GB | 426 GB | 4 × H200 |
| GLM-5.3 | 753.3B | 775 GB (FP8) | 821 GB | 468 GB | 4 × H200 |
| GLM-5.3 Flash | 321.3B | 337 GB (FP8) | 350 GB | 200 GB | 2 × H200 |
| GLM-5.2 | 753.3B | 1,545 GB (BF16) | 821 GB | 468 GB | 4 × H200 |
| GLM-4.7 Flash | 31.2B | 64.9 GB (BF16) | 34.9 GB | 20.3 GB | RTX 3090 / 4090 |
| Gemma 4 31B | 31.3B | 65.8 GB (BF16) | 35.7 GB | 21.1 GB | RTX 3090 / 4090 |
| Gemma 4 26B-A4B (MoE) | 25.8B | 53.7 GB (BF16) | 28.9 GB | 16.8 GB | RTX 3090 / 4090 |
| Gemma 4 12B | 12.0B | 25.4 GB (BF16) | 13.9 GB | 8.33 GB | RTX 3060 |
| Gemma 4 E4B | 8.0B | 17.0 GB (BF16) | 9.36 GB | 5.61 GB | RTX 3060 |
| Kimi K3 | 2.78T | 1,600 GB (4.5 bits/weight) | 3,027 GB | 1,724 GB | More than 8 GPUs |
| MiniMax M3 | 427.0B | 877 GB (BF16) | 466 GB | 266 GB | 2 × H200 |
| MiniMax M2.7 | 228.7B | 238 GB (FP8) | 252 GB | 144 GB | B200 |
| MiMo V2.6 Flash | 310.8B | 178 GB (4.5 bits/weight) | 339 GB | 193 GB | 2 × H200 |
| MiMo V2.6 Pro | 1.02T | 581 GB (4.4 bits/weight) | 1,116 GB | 636 GB | 4 × B200 |
| Mistral Medium 3.5 128B | 127.7B | 140 GB (FP8) | 143 GB | 82.7 GB | H200 |
| Nemotron 3 Nano 4B | 4.0B | 8.78 GB (BF16) | 4.96 GB | 3.10 GB | RTX 3060 |
| Nemotron 3 Nano 30B-A3B (MoE) | 31.6B | 65.3 GB (BF16) | 34.9 GB | 20.1 GB | RTX 3090 / 4090 |
| Nemotron 3 Super 120B-A12B (MoE) | 123.6B | 254 GB (BF16) | 135 GB | 77.2 GB | A100 / H100 80GB |
| Xing 4.0 29B-A4B (MoE) | 31.2B | 64.8 GB (BF16) | 34.9 GB | 20.2 GB | RTX 3090 / 4090 |
| Muse Glimmer 30B | 29.8B | 61.7 GB (BF16) | 33.1 GB | 19.2 GB | RTX 3090 / 4090 |
| MiniCPM5 2B | 2.5B | 6.02 GB (BF16) | 3.60 GB | 2.42 GB | RTX 3060 |
| Llama 3.1 8B | 8.0B | 18.1 GB (BF16) | 10.3 GB | 6.58 GB | RTX 3060 |
| Llama 3.1 70B | 70.6B | 148 GB (BF16) | 80.0 GB | 47.0 GB | L40S / RTX 6000 Ada |
| Qwen3 8B | 8.2B | 18.5 GB (BF16) | 10.7 GB | 6.81 GB | RTX 3060 |
| Qwen3 30B-A3B (MoE) | 30.5B | 63.9 GB (BF16) | 34.6 GB | 20.2 GB | RTX 3090 / 4090 |
| gpt-oss-20b | 20.9B | 14.8 GB (MXFP4) | 23.5 GB | 13.7 GB | RTX 4060 Ti 16GB |
| gpt-oss-120b | 116.8B | 67.7 GB (MXFP4) | 128 GB | 73.2 GB | A100 / H100 80GB |
| DeepSeek V3 / R1 (671B) | 684.5B | 707 GB (FP8) | 746 GB | 425 GB | 4 × H200 |
What can your GPU run?
Every model above, checked against one card or Mac: the best precision that fits, the longest context and the speed.
- RTX 4060 8GB
- RTX 3060 12GB
- Arc B580 12GB
- RTX 4070 12GB
- RTX 5070 12GB
- RTX 4060 Ti 16GB
- RTX 5060 Ti 16GB
- RX 9070 XT 16GB
- RTX 4070 Ti Super 16GB
- RTX 4080 Super 16GB
- RTX 5070 Ti 16GB
- RTX 5080 16GB
- RTX 3090
- RTX 4090
- RX 7900 XTX
- RTX 5090
- M4 Pro Mac (64 GB)
- M4 Max Mac (128 GB)
- M3 Ultra Mac Studio (512 GB)
- Ryzen AI Max+ 395 (128 GB)
- DGX Spark (128 GB)
- L40S
- RTX PRO 6000 Blackwell
- A100 80GB
- H100 SXM
- H200
- B200
- 2× RTX 3060 12GB
- 2× RTX 3090
- 4× RTX 3090
- 2× RTX 4090
- 2× RTX 5090
- 8× H100 SXM
- 8× H200
Every precision for these models is also in a CSV file on GitHub, free to reuse. Not sure which GGUF type to pick? See GGUF quantization explained.
How to use
- Choose a model, or type a Hugging Face id such as Qwen/Qwen3-8B and press Load.
- Pick the weight precision you plan to run, for example Q4_K_M for llama.cpp or FP8 for vLLM.
- Set the context length and how many requests run at the same time.
- Read the total and the GPU table. Open Advanced to see or change every number behind the estimate.
FAQ
How is the estimate calculated?
Weights = parameters × bits per weight ÷ 8. KV cache = 2 × layers × KV heads × head dimension × bytes per value × tokens × requests. Runtime overhead = 0.5 GB for the CUDA context plus 10% of weights and cache for activations and fragmentation; you can change the 10% under Advanced.
Why does context length matter so much?
The KV cache grows with every token of every request. Llama 3.1 8B needs 128 KB of FP16 cache per token, so one 128K-token conversation adds 16 GB on top of the weights. An FP8 cache halves that.
Do MoE models like Qwen3-30B-A3B need less memory?
They compute with a few experts per token, but every expert has to be loaded, so weight memory follows the total parameter count: 30.5 billion for Qwen3-30B-A3B, not 3 billion.
Which GGUF quantization should I pick?
Q4_K_M is the usual default at about a third of the BF16 size; Q5_K_M and Q6_K keep more quality for coding and small models; Q3_K_M and Q2_K make a model fit at a clear cost in quality. The GGUF quantization guide on this site lists the bits per weight of each type and the size of popular models.
Does quantizing the weights shrink the KV cache?
No. Weight precision and KV cache precision are separate settings, so Q4 weights with an FP16 cache still need the full-size cache.
How are newer architectures handled?
Everything comes from the model’s config.json. Multi-head latent attention (DeepSeek V3, GLM-5.3, Kimi K3) caches 576 values per token per layer instead of full keys and values; DeepSeek V3.2 and GLM-5 add a 128-value FP8 key per token for their sparse-attention indexer. Sliding-window layers keep only their window: 128 tokens in half of gpt-oss, 1,024 tokens in 50 of Gemma 4 31B’s 60 layers, whose other 10 layers cache 4 heads of 512 values that serve as both keys and values. Linear-attention layers keep a fixed-size state instead of a growing cache: 48 of Qwen3.8 27B’s 64 layers and 69 of Kimi K3’s 93. DeepSeek V4 also shares and compresses its cache across layers, which the calculator does not model, so its KV figure is an upper bound.
What does "As published" mean?
It uses the size of the model’s weight files on Hugging Face, so it matches the download even for checkpoints that mix formats, such as MXFP4 experts in gpt-oss or the 4-bit experts of DeepSeek V4 and MiMo V2.6.
How accurate is it?
It is an estimate. Inference engines add their own buffers: vLLM reserves a fixed share of GPU memory up front, and llama.cpp allocates the cache for the full context when it starts. Leave some headroom, especially on a GPU that also drives your display.
Is anything I enter sent anywhere?
The calculation runs in your browser. The only network requests go to Hugging Face, when you load a model.
More tools
Updated