Image & Video Model VRAM Calculator
Image and video models in ComfyUI are three files loaded one after another: the diffusion model, a text encoder that is often as big, and a VAE. Pick the model and the file you run for each, where the text encoder lives and the job size, and see the peak VRAM range, which GPUs hold it, and how much system RAM the rest needs.
FLUX.2 VAE (FP32 file, loaded as BF16)
20.6 GB–22.2 GB estimated peak VRAM with every needed file on the GPU
Weights on the GPU at once: 18.7 GB of 35.7 GB in all files (exact sizes).
Wan 2.2 A14B is two experts, one for the high-noise steps and one for the low-noise steps. ComfyUI holds one on the GPU at a time and keeps the other in system RAM, so “all files” counts both.
This checkpoint already holds the video VAE, the audio VAE and the text connectors, so the VAE choice adds nothing.
| GPU memory | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB |
|---|---|---|---|---|---|---|
| Result | Streams | Streams | Streams | Fits | Fits | Fits |
| System RAM | 38 GB | 34 GB | 30 GB | 23 GB | 23 GB | 23 GB |
System RAM: the weights not on the GPU plus about 6 GB for ComfyUI (measured with MiniMax H3). To keep a copy of every file in RAM, which ComfyUI does by default for fast reruns, plan for 42 GB.
“Fits”: every needed file and the top of the working-memory range fit. “Tight”: only the bottom does. “Streams”: the working memory fits but not every weight, so ComfyUI’s dynamic VRAM keeps the rest in system RAM and copies it every step: it works, slower. “Risky”: even the working memory may not fit. Leave about 0.5–1 GB for the desktop if the GPU drives your screen.
GB = GiB (1024³ bytes), the unit of nvidia-smi and of ComfyUI’s “MB” (MiB); the runs keep the units their authors used.
Common setups, encoder swapped out after the prompt
Each model at its first job size: 1024×1024 for images, 832×480 for Wan 2.2 A14B, 1280×704 for Wan 2.2 5B and 1280×720 for LTX. Other sizes are in the calculator above.
| Setup | Job | Weights on GPU | All files | Peak range | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB |
|---|---|---|---|---|---|---|---|---|---|---|
| FLUX.2 [dev] (32B) FP8 mixed (ComfyUI) + FP8 (ComfyUI) | About 1 MP (1024×1024) | 33.2 GB | 50.1 GB | 35.1 GB–36.7 GB | Streams | Streams | Streams | Streams | Streams | Fits |
| FLUX.2 [dev] (32B) GGUF Q4_K_M + FP8 (ComfyUI) | About 1 MP (1024×1024) | 18.7 GB | 35.7 GB | 20.6 GB–22.2 GB | Streams | Streams | Streams | Fits | Fits | Fits |
| FLUX.2 [dev] (32B) GGUF Q2_K + GGUF Q4_K_M | About 1 MP (1024×1024) | 13.5 GB | 25.6 GB | 15.4 GB–17.0 GB | Streams | Streams | Tight | Fits | Fits | Fits |
| FLUX.2 [klein] 9B BF16 (full precision) + BF16 (full precision) | About 1 MP (1024×1024) | 17.1 GB | 32.5 GB | 18.3 GB–20.1 GB | Streams | Streams | Streams | Fits | Fits | Fits |
| FLUX.2 [klein] 9B FP8 + FP8 mixed (ComfyUI) | About 1 MP (1024×1024) | 8.9 GB | 17.2 GB | 10.1 GB–11.9 GB | Streams | Fits | Fits | Fits | Fits | Fits |
| FLUX.2 [klein] 9B GGUF Q4_K_M + GGUF Q4_K_M | About 1 MP (1024×1024) | 5.7 GB | 10.5 GB | 6.9 GB–8.7 GB | Tight | Fits | Fits | Fits | Fits | Fits |
| FLUX.2 [klein] 4B BF16 (full precision) + BF16 (full precision) | About 1 MP (1024×1024) | 7.6 GB | 15.0 GB | 8.6 GB–10.1 GB | Streams | Fits | Fits | Fits | Fits | Fits |
| FLUX.2 [klein] 4B FP8 + FP4 (ComfyUI) | About 1 MP (1024×1024) | 3.9 GB | 7.7 GB | 4.9 GB–6.4 GB | Fits | Fits | Fits | Fits | Fits | Fits |
| Qwen-Image 2.1 (7B) INT8 ConvRot (ComfyUI) + INT8 ConvRot (ComfyUI) | About 1 MP (1024×1024) | 9.3 GB | 16.1 GB | 11.4 GB–14.0 GB | Streams | Tight | Fits | Fits | Fits | Fits |
| Qwen-Image 2.1 (7B) GGUF Q4_K_M + GGUF Q4_K_M | About 1 MP (1024×1024) | 5.3 GB | 9.2 GB | 7.4 GB–10.0 GB | Tight | Fits | Fits | Fits | Fits | Fits |
| Wan 2.2 A14B (T2V / I2V) FP16 + FP16 (full precision) | 832×480, 81 frames (5 s) | 26.9 GB | 64.1 GB | 27.9 GB–29.9 GB | Streams | Streams | Streams | Streams | Fits | Fits |
| Wan 2.2 A14B (T2V / I2V) FP8 scaled + FP8 scaled (ComfyUI) | 832×480, 81 frames (5 s) | 13.5 GB | 33.1 GB | 14.5 GB–16.5 GB | Streams | Streams | Tight | Fits | Fits | Fits |
| Wan 2.2 A14B (T2V / I2V) GGUF Q4_K_M + GGUF Q5_K_M | 832×480, 81 frames (5 s) | 9.2 GB | 22.1 GB | 10.2 GB–12.2 GB | Streams | Tight | Fits | Fits | Fits | Fits |
| Wan 2.2 TI2V 5B FP16 (full precision) + FP8 scaled (ComfyUI) | 1280×704, 121 frames (5 s) | 10.6 GB | 16.9 GB | 12.0 GB–15.6 GB | Streams | Streams | Fits | Fits | Fits | Fits |
| Wan 2.2 TI2V 5B GGUF Q4_K_M + GGUF Q4_K_M | 1280×704, 121 frames (5 s) | 4.7 GB | 7.9 GB | 6.1 GB–9.7 GB | Tight | Fits | Fits | Fits | Fits | Fits |
| LTX-2 (19B) FP8 checkpoint + FP8 scaled (ComfyUI) | 1280×720, 121 frames (5 s) | 25.2 GB | 37.5 GB | 27.2 GB–30.2 GB | Streams | Streams | Streams | Streams | Fits | Fits |
| LTX-2 (19B) NVFP4 checkpoint + FP4 mixed (ComfyUI) | 1280×720, 121 frames (5 s) | 18.6 GB | 27.4 GB | 20.6 GB–23.6 GB | Streams | Streams | Streams | Fits | Fits | Fits |
| LTX-2.3 (22B) FP8 checkpoint + FP8 scaled (ComfyUI) | 1280×720, 121 frames (5 s) | 27.1 GB | 39.4 GB | 29.1 GB–32.1 GB | Streams | Streams | Streams | Streams | Tight | Fits |
| LTX-2.3 (22B) GGUF Q4_K_M + GGUF Q4_K_M | 1280×720, 121 frames (5 s) | 17.2 GB | 24.0 GB | 19.2 GB–22.2 GB | Streams | Streams | Streams | Fits | Fits | Fits |
| LTX-2.5 (22B) INT8 ConvRot (ComfyUI) + INT8 ConvRot (ComfyUI) | 1280×720, 121 frames (5 s) | 21.7 GB | 36.1 GB | 23.7 GB–26.7 GB | Streams | Streams | Streams | Tight | Fits | Fits |
| LTX-2.5 (22B) GGUF Q4_K_M (distilled) + INT8 ConvRot (ComfyUI) | 1280×720, 121 frames (5 s) | 16.3 GB | 30.6 GB | 18.3 GB–21.3 GB | Streams | Streams | Streams | Fits | Fits | Fits |
Checked against public ComfyUI runs
Two checks. First, ComfyUI’s own load log (“loaded”, “staged”, in MiB) against the file’s bytes: this is what makes the weights exact. Second, the working memory a run implies, its peak (or what was allocated when it ran out of memory, or the card it finished on) minus the weights that were on the GPU, against the range this page uses for that job.
| GPU | Run | Reported | This page | Agrees | Source |
|---|---|---|---|---|---|
| RTX 3090 24GB | FLUX.2 [dev] text encoder, Mistral FP8 | 17,180.59 MB loaded | file: 17,199.2 MiB (−0.11%) | Yes | #10891 2025-11-25 |
| RTX 3090 24GB | FLUX.2 [dev] DiT, FP8 mixed | 20,308.52 MB loaded + 13,504.50 MB offloaded | file: 33,813.1 MiB (<0.01%) | Yes | #10891 2025-11-25 |
| not stated | FLUX.2 [klein] 9B DiT, FP8 KV | 9,364 MB staged | file: 9,364.1 MiB (<0.01%) | Yes | #12906 2026-03-12 |
| not stated | FLUX.2 VAE, FP32 file loaded as BF16 | 160 MB staged | file: 160.3 MiB (−0.20%) | Yes | #12906 2026-03-12 |
| RTX 5070 Ti 16GB | Wan 2.2 I2V A14B NVFP4, high-noise expert | ~9.4 GB staged ("≈ file size") | file: 9,450.3 MiB (−0.53%) | Yes | #11864 2026-07-24 |
| RX 9060 XT 16GB | Wan 2.1 VAE (also Qwen-Image’s first VAE) | 242.03 MB loaded | file: 242.1 MiB (−0.01%) | Yes | #15220 2026-08-02 |
| RTX 3090 24GB (24,121 MB) | FLUX.2 [dev] FP8 mixed, 21,897 MB of the DiT on the GPU, 1024×1024 | “peaking close to full 24 GiB” | implies 2.1 GB of working memory; range 1.9–3.5 GB | Yes | #10891 2025-11-27 |
| RTX 5090 32GB | FLUX.2 [klein] 9B FP8 KV + FLUX.2 VAE, reference-image workflow without the KV cache node | “VRAM usage is 14 GB” | implies 4.7 GB of working memory; range 2.0–5.0 GB | Yes | #12906 2026-03-12 |
| 16 GB card | Wan 2.2 FP8 scaled, high-noise expert + Wan 2.1 VAE, 720p | “around 15GB” | implies 1.5 GB of working memory; range 1.4–5.5 GB | Yes | #9293 2025-08-13 |
| 16 GB card | same, low-noise expert | “exceeded 16GB” (spilled to shared memory) | implies at least 2.5 GB; range up to 5.5 GB | Yes | #9293 2025-08-13 |
| RTX 5070 Ti 16GB | Wan 2.2 I2V A14B NVFP4, low-noise expert (9.7 GB) + Wan 2.1 VAE, 640×640, 81 frames | finished | implies at most 6.3 GB; range from 1.0 GB | Yes | #11864 2026-07-24 |
| RTX 3090 24GB (22.06 GiB usable) | LTX-2 19B FP8 checkpoint, template workflow | OOM at 21.90 GB with --normalvram; passed with --novram (6.74 GB peak, weights streamed) | weights alone: 25.2 GB, more than the card; the page says “Streams” | Yes | #12047 2026-01-23 |
| RTX 5090 32GB (31.36 GiB usable) | LTX-2 19B FP8 checkpoint (~27 GB), 1080p, 81 frames | “can easily run” | implies at most 6.1 GB; range from 3.5 GB | Yes | #11864 2026-01-14 |
| RTX 5090 32GB (31.36 GiB usable) | LTX-2.3, 27,580 MB of the model on the GPU (+1,209 MB offloaded), 120+ frames, size not stated | OOM with 29.60 GiB allocated | implies at least 2.7 GB; range up to 5.0 GB | Yes | #14683 2026-06-30 |
How far off it can be
The weights are exact: the load logs above are within 0.2% of the file sizes (0.5% for one report rounded to 0.1 GB). The working memory is the uncertain part. The runs above fall inside the ranges, but each range rests on one to three runs, several of them from Task Manager or a rounded “around 15GB”, so the peak can be off by 1–2 GB for images and 2–4 GB for video. FLUX.2 [klein] 4B, Wan 2.2 5B and LTX-2.5 have no public run with sizes yet. If you have one, a peak from nvidia-smi or ComfyUI’s log together with the file names makes the range narrower: send it here.
Why ComfyUI runs models that do not fit
ComfyUI does not need the whole model on the GPU. It loads what fits and copies the rest from system RAM as each block runs (partial loading, and since early 2026 dynamic VRAM, the default), so the FLUX.2 [dev] FP8 DiT (33.0 GB) finishes on a 24 GB card: in issue #10891 an RTX 3090 held 19.8 GB of it and streamed 13.2 GB. What still has to fit is the working memory, which is why the video rows need more headroom than the weights suggest. Streaming needs the RAM in the table above, and it runs out when RAM does: a 16 GB card with Wan 2.2 only stopped crashing with a 100 GB page file.
File sizes read from the Hugging Face API on : black-forest-labs/
How to use
- Pick the model: FLUX.2 [dev] or [klein], Qwen-Image 2.1, Wan 2.2 A14B or 5B, LTX-2, LTX-2.3 or LTX-2.5.
- Pick the diffusion model file (BF16, FP8, INT8, NVFP4 or a GGUF quant), the text encoder file and the VAE.
- Choose where the text encoder lives while generating. ComfyUI swaps it out after the prompt when memory is short.
- Choose the image or video size, then read the peak range, the GPU table and the RAM each card needs. "Streams" means ComfyUI keeps part of the weights in system RAM: it runs, slower.
Frequently asked questions
How much VRAM does FLUX.2 [dev] need?
Everything in BF16 on the GPU takes 93.3 GB of weights (32B DiT, Mistral Small 3.2 24B encoder, VAE). ComfyUI's FP8 set needs 33.2 GB on the card at once when the encoder is swapped out after the prompt, 35.1 GB–36.7 GB at 1 MP, so a 24 GB card streams part of the DiT from system RAM. The Q4_K_M GGUF brings it to 20.6 GB–22.2 GB. The distilled FLUX.2 [klein] 9B at Q4_K_M needs 6.9 GB–8.7 GB.
Can Wan 2.2 14B run on a 12 GB or 16 GB GPU?
Yes, with quantized experts. Wan 2.2 A14B is two 14B experts, and ComfyUI holds one on the GPU at a time. At 720p, 81 frames, one FP8 expert (13.3 GB) peaks at 14.9 GB–19.0 GB: tight on 16 GB. A Q4_K_M expert peaks at 10.6 GB–14.7 GB. The other expert waits in system RAM, so the files add up to 22.1 GB for that set.
How much VRAM does LTX-2.5 or LTX-2.3 need?
LTX-2.5 with Lightricks' INT8 ConvRot transformer and Gemma 4 encoder needs 23.7 GB–26.7 GB for 1280×720, 121 frames with the encoder swapped out. LTX-2.3 as the unsloth Q4_K_M GGUF with a Q4_K_M Gemma 3 encoder needs 19.2 GB–22.2 GB. The FP8 checkpoints of LTX-2 and LTX-2.3 are 25–27 GB on their own, more than a 24 GB card holds.
Why does ComfyUI still run a model that does not fit?
ComfyUI's dynamic VRAM keeps the part of the weights that does not fit in system RAM and copies it to the GPU every step. The job finishes, slower, and needs that RAM: the page marks these cards "Streams" and shows the RAM per card. It only fails when the working memory itself does not fit.
How accurate is this?
The weights are the files' byte sizes; ComfyUI's own load logs match them within 0.2% (0.5% for one report rounded to 0.1 GB). The working memory on top is a range. 8 of 8 public ComfyUI runs fall inside it, but there are few of them, so a video estimate can be off by 2–4 GB. FLUX.2 [klein] 4B, Wan 2.2 5B and LTX-2.5 have no public run with sizes yet; their ranges are borrowed from a sibling model, and the page says so.
Is GB here GB or GiB?
GiB (1024³ bytes), the unit nvidia-smi and ComfyUI's "MB" use, as everywhere on this site. The runs keep the units their authors gave.
More calculators
- Fine-Tuning VRAM Calculator GPU memory for full fine-tuning, LoRA and QLoRA in Transformers or Unsloth, checked against published runs. Open →
- Jev Alternatives You Can Run Locally Open models that replace the API-only Jev on your own machine: Laya, zero-shot encoders and small LLMs, with sizes and VRAM. Open →
- Laya VRAM Requirements Real file sizes and run-time memory of the three Laya decision-encoder checkpoints, and three ways to run them. Open →
- Bonsai 2 27B VRAM Exact file sizes of Ternary Bonsai 2 27B, the longest context on 8–24 GB GPUs, and why sites say 5.95 GB or 8.60 GB. Open →
- vLLM KV Cache & Concurrency Calculator The KV cache pool vLLM allocates, in tokens, and how many requests fit at once, with the vllm serve command. Open →
- What Can My PC Run? Detects your GPU in the browser and lists the local LLMs it runs, with the best quantization and speed. Open →
Updated