Image & Video Model VRAM Calculator

Image and video models in ComfyUI are three files loaded one after another: the diffusion model, a text encoder that is often as big, and a VAE. Pick the model and the file you run for each, where the text encoder lives and the job size, and see the peak VRAM range, which GPUs hold it, and how much system RAM the rest needs.

FLUX.2 VAE (FP32 file, loaded as BF16)

20.6 GB–22.2 GB estimated peak VRAM with every needed file on the GPU

Weights on the GPU at once: 18.7 GB of 35.7 GB in all files (exact sizes).

GPU memory8 GB12 GB16 GB24 GB32 GB48 GB
ResultStreamsStreamsStreamsFitsFitsFits
System RAM38 GB34 GB30 GB23 GB23 GB23 GB

System RAM: the weights not on the GPU plus about 6 GB for ComfyUI (measured with MiniMax H3). To keep a copy of every file in RAM, which ComfyUI does by default for fast reruns, plan for 42 GB.

“Fits”: every needed file and the top of the working-memory range fit. “Tight”: only the bottom does. “Streams”: the working memory fits but not every weight, so ComfyUI’s dynamic VRAM keeps the rest in system RAM and copies it every step: it works, slower. “Risky”: even the working memory may not fit. Leave about 0.5–1 GB for the desktop if the GPU drives your screen.

GB = GiB (1024³ bytes), the unit of nvidia-smi and of ComfyUI’s “MB” (MiB); the runs keep the units their authors used.

Common setups, encoder swapped out after the prompt

Each model at its first job size: 1024×1024 for images, 832×480 for Wan 2.2 A14B, 1280×704 for Wan 2.2 5B and 1280×720 for LTX. Other sizes are in the calculator above.

SetupJobWeights on GPUAll filesPeak range8 GB12 GB16 GB24 GB32 GB48 GB
FLUX.2 [dev] (32B)
FP8 mixed (ComfyUI) + FP8 (ComfyUI)
About 1 MP (1024×1024) 33.2 GB 50.1 GB 35.1 GB–36.7 GB StreamsStreamsStreamsStreamsStreamsFits
FLUX.2 [dev] (32B)
GGUF Q4_K_M + FP8 (ComfyUI)
About 1 MP (1024×1024) 18.7 GB 35.7 GB 20.6 GB–22.2 GB StreamsStreamsStreamsFitsFitsFits
FLUX.2 [dev] (32B)
GGUF Q2_K + GGUF Q4_K_M
About 1 MP (1024×1024) 13.5 GB 25.6 GB 15.4 GB–17.0 GB StreamsStreamsTightFitsFitsFits
FLUX.2 [klein] 9B
BF16 (full precision) + BF16 (full precision)
About 1 MP (1024×1024) 17.1 GB 32.5 GB 18.3 GB–20.1 GB StreamsStreamsStreamsFitsFitsFits
FLUX.2 [klein] 9B
FP8 + FP8 mixed (ComfyUI)
About 1 MP (1024×1024) 8.9 GB 17.2 GB 10.1 GB–11.9 GB StreamsFitsFitsFitsFitsFits
FLUX.2 [klein] 9B
GGUF Q4_K_M + GGUF Q4_K_M
About 1 MP (1024×1024) 5.7 GB 10.5 GB 6.9 GB–8.7 GB TightFitsFitsFitsFitsFits
FLUX.2 [klein] 4B
BF16 (full precision) + BF16 (full precision)
About 1 MP (1024×1024) 7.6 GB 15.0 GB 8.6 GB–10.1 GB StreamsFitsFitsFitsFitsFits
FLUX.2 [klein] 4B
FP8 + FP4 (ComfyUI)
About 1 MP (1024×1024) 3.9 GB 7.7 GB 4.9 GB–6.4 GB FitsFitsFitsFitsFitsFits
Qwen-Image 2.1 (7B)
INT8 ConvRot (ComfyUI) + INT8 ConvRot (ComfyUI)
About 1 MP (1024×1024) 9.3 GB 16.1 GB 11.4 GB–14.0 GB StreamsTightFitsFitsFitsFits
Qwen-Image 2.1 (7B)
GGUF Q4_K_M + GGUF Q4_K_M
About 1 MP (1024×1024) 5.3 GB 9.2 GB 7.4 GB–10.0 GB TightFitsFitsFitsFitsFits
Wan 2.2 A14B (T2V / I2V)
FP16 + FP16 (full precision)
832×480, 81 frames (5 s) 26.9 GB 64.1 GB 27.9 GB–29.9 GB StreamsStreamsStreamsStreamsFitsFits
Wan 2.2 A14B (T2V / I2V)
FP8 scaled + FP8 scaled (ComfyUI)
832×480, 81 frames (5 s) 13.5 GB 33.1 GB 14.5 GB–16.5 GB StreamsStreamsTightFitsFitsFits
Wan 2.2 A14B (T2V / I2V)
GGUF Q4_K_M + GGUF Q5_K_M
832×480, 81 frames (5 s) 9.2 GB 22.1 GB 10.2 GB–12.2 GB StreamsTightFitsFitsFitsFits
Wan 2.2 TI2V 5B
FP16 (full precision) + FP8 scaled (ComfyUI)
1280×704, 121 frames (5 s) 10.6 GB 16.9 GB 12.0 GB–15.6 GB StreamsStreamsFitsFitsFitsFits
Wan 2.2 TI2V 5B
GGUF Q4_K_M + GGUF Q4_K_M
1280×704, 121 frames (5 s) 4.7 GB 7.9 GB 6.1 GB–9.7 GB TightFitsFitsFitsFitsFits
LTX-2 (19B)
FP8 checkpoint + FP8 scaled (ComfyUI)
1280×720, 121 frames (5 s) 25.2 GB 37.5 GB 27.2 GB–30.2 GB StreamsStreamsStreamsStreamsFitsFits
LTX-2 (19B)
NVFP4 checkpoint + FP4 mixed (ComfyUI)
1280×720, 121 frames (5 s) 18.6 GB 27.4 GB 20.6 GB–23.6 GB StreamsStreamsStreamsFitsFitsFits
LTX-2.3 (22B)
FP8 checkpoint + FP8 scaled (ComfyUI)
1280×720, 121 frames (5 s) 27.1 GB 39.4 GB 29.1 GB–32.1 GB StreamsStreamsStreamsStreamsTightFits
LTX-2.3 (22B)
GGUF Q4_K_M + GGUF Q4_K_M
1280×720, 121 frames (5 s) 17.2 GB 24.0 GB 19.2 GB–22.2 GB StreamsStreamsStreamsFitsFitsFits
LTX-2.5 (22B)
INT8 ConvRot (ComfyUI) + INT8 ConvRot (ComfyUI)
1280×720, 121 frames (5 s) 21.7 GB 36.1 GB 23.7 GB–26.7 GB StreamsStreamsStreamsTightFitsFits
LTX-2.5 (22B)
GGUF Q4_K_M (distilled) + INT8 ConvRot (ComfyUI)
1280×720, 121 frames (5 s) 16.3 GB 30.6 GB 18.3 GB–21.3 GB StreamsStreamsStreamsFitsFitsFits

Checked against public ComfyUI runs

Two checks. First, ComfyUI’s own load log (“loaded”, “staged”, in MiB) against the file’s bytes: this is what makes the weights exact. Second, the working memory a run implies, its peak (or what was allocated when it ran out of memory, or the card it finished on) minus the weights that were on the GPU, against the range this page uses for that job.

GPURunReportedThis pageAgreesSource
RTX 3090 24GB FLUX.2 [dev] text encoder, Mistral FP8 17,180.59 MB loaded file: 17,199.2 MiB (−0.11%) Yes #10891 2025-11-25
RTX 3090 24GB FLUX.2 [dev] DiT, FP8 mixed 20,308.52 MB loaded + 13,504.50 MB offloaded file: 33,813.1 MiB (<0.01%) Yes #10891 2025-11-25
not stated FLUX.2 [klein] 9B DiT, FP8 KV 9,364 MB staged file: 9,364.1 MiB (<0.01%) Yes #12906 2026-03-12
not stated FLUX.2 VAE, FP32 file loaded as BF16 160 MB staged file: 160.3 MiB (−0.20%) Yes #12906 2026-03-12
RTX 5070 Ti 16GB Wan 2.2 I2V A14B NVFP4, high-noise expert ~9.4 GB staged ("≈ file size") file: 9,450.3 MiB (−0.53%) Yes #11864 2026-07-24
RX 9060 XT 16GB Wan 2.1 VAE (also Qwen-Image’s first VAE) 242.03 MB loaded file: 242.1 MiB (−0.01%) Yes #15220 2026-08-02
RTX 3090 24GB (24,121 MB) FLUX.2 [dev] FP8 mixed, 21,897 MB of the DiT on the GPU, 1024×1024 “peaking close to full 24 GiB” implies 2.1 GB of working memory; range 1.9–3.5 GB Yes #10891 2025-11-27
RTX 5090 32GB FLUX.2 [klein] 9B FP8 KV + FLUX.2 VAE, reference-image workflow without the KV cache node “VRAM usage is 14 GB” implies 4.7 GB of working memory; range 2.0–5.0 GB Yes #12906 2026-03-12
16 GB card Wan 2.2 FP8 scaled, high-noise expert + Wan 2.1 VAE, 720p “around 15GB” implies 1.5 GB of working memory; range 1.4–5.5 GB Yes #9293 2025-08-13
16 GB card same, low-noise expert “exceeded 16GB” (spilled to shared memory) implies at least 2.5 GB; range up to 5.5 GB Yes #9293 2025-08-13
RTX 5070 Ti 16GB Wan 2.2 I2V A14B NVFP4, low-noise expert (9.7 GB) + Wan 2.1 VAE, 640×640, 81 frames finished implies at most 6.3 GB; range from 1.0 GB Yes #11864 2026-07-24
RTX 3090 24GB (22.06 GiB usable) LTX-2 19B FP8 checkpoint, template workflow OOM at 21.90 GB with --normalvram; passed with --novram (6.74 GB peak, weights streamed) weights alone: 25.2 GB, more than the card; the page says “Streams” Yes #12047 2026-01-23
RTX 5090 32GB (31.36 GiB usable) LTX-2 19B FP8 checkpoint (~27 GB), 1080p, 81 frames “can easily run” implies at most 6.1 GB; range from 3.5 GB Yes #11864 2026-01-14
RTX 5090 32GB (31.36 GiB usable) LTX-2.3, 27,580 MB of the model on the GPU (+1,209 MB offloaded), 120+ frames, size not stated OOM with 29.60 GiB allocated implies at least 2.7 GB; range up to 5.0 GB Yes #14683 2026-06-30

How far off it can be

The weights are exact: the load logs above are within 0.2% of the file sizes (0.5% for one report rounded to 0.1 GB). The working memory is the uncertain part. The runs above fall inside the ranges, but each range rests on one to three runs, several of them from Task Manager or a rounded “around 15GB”, so the peak can be off by 1–2 GB for images and 2–4 GB for video. FLUX.2 [klein] 4B, Wan 2.2 5B and LTX-2.5 have no public run with sizes yet. If you have one, a peak from nvidia-smi or ComfyUI’s log together with the file names makes the range narrower: send it here.

Why ComfyUI runs models that do not fit

ComfyUI does not need the whole model on the GPU. It loads what fits and copies the rest from system RAM as each block runs (partial loading, and since early 2026 dynamic VRAM, the default), so the FLUX.2 [dev] FP8 DiT (33.0 GB) finishes on a 24 GB card: in issue #10891 an RTX 3090 held 19.8 GB of it and streamed 13.2 GB. What still has to fit is the working memory, which is why the video rows need more headroom than the weights suggest. Streaming needs the RAM in the table above, and it runs out when RAM does: a 16 GB card with Wan 2.2 only stopped crashing with a 100 GB page file.

File sizes read from the Hugging Face API on : black-forest-labs/FLUX.2-dev, Comfy-Org/flux2-dev, unsloth/FLUX.2-dev-GGUF, unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF, black-forest-labs/FLUX.2-klein-9B, unsloth/FLUX.2-klein-9B-GGUF, black-forest-labs/FLUX.2-klein-9b-kv-fp8, black-forest-labs/FLUX.2-klein-9b-fp8, Comfy-Org/vae-text-encorder-for-flux-klein-9b, unsloth/Qwen3-8B-GGUF, black-forest-labs/FLUX.2-klein-4B, unsloth/FLUX.2-klein-4B-GGUF, black-forest-labs/FLUX.2-klein-4b-fp8, Comfy-Org/vae-text-encorder-for-flux-klein-4b, unsloth/Qwen3-4B-GGUF, Comfy-Org/Qwen-Image-2.1, unsloth/Qwen-Image-2.1-GGUF, unsloth/Qwen-Image-2.1-FP8, gguf-org/qwen-image-2.1-gguf, unsloth/Qwen3-VL-8B-Instruct-GGUF, Qwen/Qwen-Image-2.1, Comfy-Org/Wan_2.2_ComfyUI_Repackaged, QuantStack/Wan2.2-T2V-A14B-GGUF, city96/umt5-xxl-encoder-gguf, QuantStack/Wan2.2-TI2V-5B-GGUF, Lightricks/LTX-2, Comfy-Org/ltx-2, unsloth/gemma-3-12b-it-GGUF, Lightricks/LTX-2.3, Lightricks/LTX-2.3-fp8, unsloth/LTX-2.3-GGUF, Lightricks/LTX-2.5, Abiray/LTX-2.5-Distilled-GGUF. Qwen-Image 2.1 in depth, with every GGUF file and low-VRAM setups for 8–16 GB: the Qwen-Image 2.1 VRAM calculator. MiniMax H3 video: the MiniMax H3 calculator. Language models: the LLM VRAM calculator.

How to use

  1. Pick the model: FLUX.2 [dev] or [klein], Qwen-Image 2.1, Wan 2.2 A14B or 5B, LTX-2, LTX-2.3 or LTX-2.5.
  2. Pick the diffusion model file (BF16, FP8, INT8, NVFP4 or a GGUF quant), the text encoder file and the VAE.
  3. Choose where the text encoder lives while generating. ComfyUI swaps it out after the prompt when memory is short.
  4. Choose the image or video size, then read the peak range, the GPU table and the RAM each card needs. "Streams" means ComfyUI keeps part of the weights in system RAM: it runs, slower.

Frequently asked questions

How much VRAM does FLUX.2 [dev] need?

Everything in BF16 on the GPU takes 93.3 GB of weights (32B DiT, Mistral Small 3.2 24B encoder, VAE). ComfyUI's FP8 set needs 33.2 GB on the card at once when the encoder is swapped out after the prompt, 35.1 GB–36.7 GB at 1 MP, so a 24 GB card streams part of the DiT from system RAM. The Q4_K_M GGUF brings it to 20.6 GB–22.2 GB. The distilled FLUX.2 [klein] 9B at Q4_K_M needs 6.9 GB–8.7 GB.

Can Wan 2.2 14B run on a 12 GB or 16 GB GPU?

Yes, with quantized experts. Wan 2.2 A14B is two 14B experts, and ComfyUI holds one on the GPU at a time. At 720p, 81 frames, one FP8 expert (13.3 GB) peaks at 14.9 GB–19.0 GB: tight on 16 GB. A Q4_K_M expert peaks at 10.6 GB–14.7 GB. The other expert waits in system RAM, so the files add up to 22.1 GB for that set.

How much VRAM does LTX-2.5 or LTX-2.3 need?

LTX-2.5 with Lightricks' INT8 ConvRot transformer and Gemma 4 encoder needs 23.7 GB–26.7 GB for 1280×720, 121 frames with the encoder swapped out. LTX-2.3 as the unsloth Q4_K_M GGUF with a Q4_K_M Gemma 3 encoder needs 19.2 GB–22.2 GB. The FP8 checkpoints of LTX-2 and LTX-2.3 are 25–27 GB on their own, more than a 24 GB card holds.

Why does ComfyUI still run a model that does not fit?

ComfyUI's dynamic VRAM keeps the part of the weights that does not fit in system RAM and copies it to the GPU every step. The job finishes, slower, and needs that RAM: the page marks these cards "Streams" and shows the RAM per card. It only fails when the working memory itself does not fit.

How accurate is this?

The weights are the files' byte sizes; ComfyUI's own load logs match them within 0.2% (0.5% for one report rounded to 0.1 GB). The working memory on top is a range. 8 of 8 public ComfyUI runs fall inside it, but there are few of them, so a video estimate can be off by 2–4 GB. FLUX.2 [klein] 4B, Wan 2.2 5B and LTX-2.5 have no public run with sizes yet; their ranges are borrowed from a sibling model, and the page says so.

Is GB here GB or GiB?

GiB (1024³ bytes), the unit nvidia-smi and ComfyUI's "MB" use, as everywhere on this site. The runs keep the units their authors gave.

More calculators

Updated