Which GPUs can run popular LLMs?

A reproducible comparison of 41 public models at Q4_K_M, FP8 and BF16. Every row links to its model-specific calculations. Download the raw CSV.

What fits on one GPU?

Each count uses 8,192 context tokens, one request, an FP16 KV cache, and the runtime overhead described below. These are memory-capacity comparisons, not measured deployment results.

GPU memoryQ4_K_MFP8BF16
24 GiB18 / 418 / 416 / 41
48 GiB19 / 4118 / 418 / 41
80 GiB24 / 4119 / 4118 / 41

For a specific RTX 4090, L40S, H100 or multi-GPU setup, open a model row or use the interactive calculator.

All model estimates

Totals are in GiB and include weights, KV cache and runtime overhead. “Yes” means the Q4_K_M estimate fits on one GPU of the listed capacity.

ModelAttention layoutQ4_K_MFP8BF1624 GiB48 GiB80 GiB
AliceAI Foundation 80B-A3B (MoE)The 36 KDA layers keep a fixed recurrent state, which this estimate does not include; only the 12 gated-attention layers add a context-growing KV cache. Real serving memory may be higher. 12 of its 48 layers use full attention and 36 are KDA recurrent layers with no growing cache51.0983.98167.25 NoNoYes
DeepSeek V4 Flash 0731This model shares and compresses its KV cache across layers, which the calculator does not model, so the KV cache figure is an upper bound. All 43 layers use full attention189.77312.86624.48 NoNoNo
Qwen3.6 35B-A3B (MoE) 10 of its 40 layers use full attention and 30 are linear-attention layers with no growing cache22.9537.574.33 YesYesYes
Qwen3.6 27B 16 of its 64 layers use full attention and 48 are linear-attention layers with no growing cache18.2729.5157.97 YesYesYes
Qwen3.8 27B 16 of its 64 layers use full attention and 48 are linear-attention layers with no growing cache18.2729.5157.97 YesYesYes
Qwen3.8 Flash Next (180B MoE) 12 of its 48 layers use full attention and 36 are linear-attention layers with no growing cache112.27185.11369.51 NoNoNo
Qwen3.8 2.4T-A95B (MoE) 23 of its 92 layers use full attention and 69 are linear-attention layers with no growing cache1517.422507.295013.3 NoNoNo
Qwen3.5 9B 8 of its 32 layers use full attention and 24 are linear-attention layers with no growing cache6.7610.6620.55 YesYesYes
Qwen3.5 122B-A10B (MoE) 12 of its 48 layers use full attention and 36 are linear-attention layers with no growing cache78.23128.85257 NoNoYes
Qwen3-Coder-Next (80B MoE) 12 of its 48 layers use full attention and 36 are linear-attention layers with no growing cache50.0982.33163.95 NoNoYes
DeepSeek V4.1 FlashThis model shares and compresses its KV cache across layers, which the calculator does not model, so the KV cache figure is an upper bound. All 40 layers use full attention474.22783.061564.93 NoNoNo
DeepSeek V4 FlashThis model shares and compresses its KV cache across layers, which the calculator does not model, so the KV cache figure is an upper bound. All 43 layers use full attention181.57299.3597.36 NoNoNo
DeepSeek V4 ProThis model shares and compresses its KV cache across layers, which the calculator does not model, so the KV cache figure is an upper bound. All 61 layers use full attention992.51639.493277.43 NoNoNo
DeepSeek V3.2 All 61 layers use multi-head latent attention425.94703.271405.39 NoNoNo
GLM-5.3 All 78 layers use multi-head latent attention468.19773.031544.78 NoNoNo
GLM-5.3 Flash 11 of its 45 layers use multi-head latent attention and 34 are linear-attention layers with no growing cache199.76329.79658.97 NoNoNo
GLM-5.2 All 78 layers use multi-head latent attention468.19773.031544.78 NoNoNo
GLM-4.7 Flash All 47 layers use multi-head latent attention20.3132.9464.92 YesYesYes
Gemma 4 31B 10 of its 60 layers use full attention and 50 keep a sliding window of 1,024 tokens21.0933.7465.78 YesYesYes
Gemma 4 26B-A4B (MoE) 5 of its 30 layers use full attention and 25 keep a sliding window of 1,024 tokens16.827.2453.67 YesYesYes
Gemma 4 12B 8 of its 48 layers use full attention and 40 keep a sliding window of 1,024 tokens8.3313.1625.42 YesYesYes
Gemma 4 E4B 4 of its 42 layers use full attention, 20 keep a sliding window of 512 tokens and 18 reuse the cache of earlier layers5.618.8517.04 YesYesYes
Kimi K3 24 of its 93 layers use multi-head latent attention and 69 are linear-attention layers with no growing cache1723.722848.655696.56 NoNoNo
MiniMax M3 All 60 layers use full attention266.21439.01876.5 NoNoNo
MiniMax M2.7 All 62 layers use full attention144.37236.91471.2 NoNoNo
MiMo V2.6 Flash 9 of its 48 layers use full attention and 39 keep a sliding window of 128 tokens193.32319.08637.43 NoNoNo
MiMo V2.6 Pro 10 of its 70 layers use full attention and 60 keep a sliding window of 128 tokens635.771050.232099.5 NoNoNo
Mistral Medium 3.5 128B All 88 layers use full attention82.68134.35265.18 NoNoNo
Nemotron 3 Nano 4B 4 of its 42 layers use full attention and 38 are Mamba or feed-forward layers with no growing cache3.14.718.78 YesYesYes
Nemotron 3 Nano 30B-A3B (MoE) 6 of its 52 layers use full attention and 46 are Mamba or feed-forward layers with no growing cache20.1232.965.25 YesYesYes
Nemotron 3 Super 120B-A12B (MoE) 8 of its 88 layers use full attention and 80 are Mamba or feed-forward layers with no growing cache77.18127.2253.84 NoNoYes
Xing 4.0 29B-A4B (MoE) All 40 layers use multi-head latent attention20.2332.8764.84 YesYesYes
Muse Glimmer 30B 13 of its 52 layers use full attention and 39 keep a sliding window of 2,048 tokens19.1531.261.71 YesYesYes
MiniCPM5 2B All 42 layers use full attention2.423.446.02 YesYesYes
Llama 3.1 8B All 32 layers use full attention6.589.8318.05 YesYesYes
Llama 3.1 70B All 80 layers use full attention46.9875.53147.81 NoYesYes
Qwen3 8B All 36 layers use full attention6.8110.1318.52 YesYesYes
Qwen3 30B-A3B (MoE) All 48 layers use full attention20.2532.663.88 YesYesYes
gpt-oss-20b 12 of its 24 layers use full attention and 12 keep a sliding window of 128 tokens13.6722.1443.56 YesYesYes
gpt-oss-120b 18 of its 36 layers use full attention and 18 keep a sliding window of 128 tokens73.22120.5240.19 NoNoYes
DeepSeek V3 / R1 (671B) All 61 layers use multi-head latent attention425.36702.361403.63 NoNoNo

How to reproduce the estimates

Q4_K_M uses an average 4.84 bits per parameter, including GGUF quantization scales. FP8 uses 8 bits and BF16 uses 16 bits. Weight memory is total parameters × bits per parameter ÷ 8; for mixture-of-experts models, all experts count because they must be loaded even though only some run for each token.

The FP16 KV cache is calculated from each model's attention layout: full attention grows with context; sliding-window layers keep only their window; MLA stores a compressed latent; linear or state-space layers do not add a context-growing KV cache. The total is weights + cache + 0.5 GiB + 10% of weights and cache. See the KV-cache explanation and calculation source and tests.

The CSV records the weight, cache and overhead components, the three total precisions, model ID, source-check date, architecture summary and per-model URL. Model files were checked on the dates in the CSV; this report was updated .

Limits

These are estimates, not measured peak GPU usage. vLLM may reserve more memory and llama.cpp can allocate the full configured cache at startup. CUDA graphs, activation buffers, vision encoders and host-side offload can change the real requirement. A “Yes” near a GPU's limit is not a deployment guarantee. Models with a known unmodeled cache optimization carry a note in the table and on their own page.

Reuse the CSV with attribution to this report or the MIT-licensed source. There is no requirement to link back.