Qwen-Image-2.1 VRAM Calculator

Qwen-Image-2.1 is three models: a 7B diffusion transformer, the Qwen3-VL-8B text encoder and a VAE. Pick the file you run for each, where the text encoder lives, and the image size, and see the peak VRAM range and which GPUs fit.

– estimated peak VRAM

Weights on the GPU at once: – (exact file sizes). Working memory on top is the range measured on real GPUs below.

GPU memory8 GB12 GB16 GB24 GB32 GB
Result–––––

“Tight”: fits only at the low end of the range, as in ComfyUI, which can also move part of the model to system RAM (slower). Leave about 0.5–1 GB for the desktop if the GPU drives your screen.

Common setups at 1024×1024

SetupWeights on GPUPeak range8 GB12 GB16 GB24 GB32 GB
Full precision (diffusers default) 30.8 GB 32.9 GB–35.5 GB NoNoNoNoNo
ComfyUI INT8 ConvRot set 16.1 GB 18.2 GB–20.8 GB NoNoNoFitsFits
ComfyUI INT8 set, encoder unloaded 9.3 GB 11.4 GB–14.0 GB NoTightFitsFitsFits
FP8 DiT + FP8 encoder (unsloth) 16.0 GB 18.1 GB–20.7 GB NoNoNoFitsFits
GGUF Q8_0 + encoder Q8_0 15.9 GB 18.0 GB–20.6 GB NoNoNoFitsFits
GGUF Q4_K_M + encoder INT8 13.2 GB 15.3 GB–17.9 GB NoNoTightFitsFits
GGUF Q4_K_M + encoder Q4_K_M 9.2 GB 11.3 GB–13.9 GB NoTightFitsFitsFits
GGUF Q4_K_M, encoder on CPU 4.5 GB 6.1 GB–6.9 GB FitsFitsFitsFitsFits

Low-VRAM setups for 8, 12 and 16 GB

Each row is a mix of the files above at 1024×1024. The three phase columns are the weights on the GPU while the prompt is encoded, while the image is denoised and while the VAE decodes it: exact file sizes. The peak adds the working memory measured in real runs, 1.6–2.4 GB with the text encoder on the CPU.

CardDiT + text encoderText encoder PromptDenoiseVAE decode PeakResultTry it
8 GB GGUF Q4_K_M + GGUF Q4_K_M
+ tiled VAE decode
On the CPU 0.6 GB4.5 GB4.5 GB 6.1 GB–6.9 GB Fits Try it
8 GB GGUF Q5_K_M + GGUF Q4_K_M
+ tiled VAE decode
On the CPU 0.6 GB5.6 GB5.6 GB 7.2 GB–8.0 GB Tight Try it
8 GB GGUF Q3_K_M + GGUF Q3_K_M
An RTX 3070 8GB ran out of memory in VAE decode with this pair (both on the GPU).
Everything on the GPU 7.4 GB7.4 GB7.4 GB 9.5 GB–12.1 GB No Try it
12 GB GGUF Q8_0 + GGUF Q4_K_M
+ tiled VAE decode
On the CPU 0.6 GB7.7 GB7.7 GB 9.3 GB–10.1 GB Fits Try it
12 GB GGUF Q4_K_M + GGUF Q4_K_M Unloaded after the prompt 5.3 GB4.5 GB4.5 GB 7.4 GB–10.0 GB Fits Try it
12 GB INT8 ConvRot (ComfyUI) + FP8 Unloaded after the prompt 9.4 GB7.4 GB7.4 GB 11.5 GB–14.1 GB Tight Try it
16 GB INT8 ConvRot (ComfyUI) + INT8 ConvRot (ComfyUI) Unloaded after the prompt 9.3 GB7.4 GB7.4 GB 11.4 GB–14.0 GB Fits Try it
16 GB GGUF Q6_K + GGUF Q4_K_M Everything on the GPU 11.2 GB11.2 GB11.2 GB 13.3 GB–15.9 GB Fits Try it
16 GB GGUF Q4_K_M + INT8 ConvRot (ComfyUI) Everything on the GPU 13.2 GB13.2 GB13.2 GB 15.3 GB–17.9 GB Tight Try it

On Windows, “runs on 8 GB” can mean 18 GB

NVIDIA’s Windows driver moves what does not fit into shared system memory instead of failing (System Memory Fallback, since driver 536.40), so a job finishes but runs from system RAM. One RTX 3070 8GB owner saw the INT8 set use 7.3 GB of dedicated VRAM plus 11.2 GB of shared memory. On the same card, the Q4_K_M DiT on the GPU with the Q4_K_M encoder on the CPU peaked at 6.9 GB, and Q3_K_M files for both, together 6.8 GB, ran out of memory while decoding. If Task Manager shows “Shared GPU memory” climbing during a run, it does not fit; setting “CUDA - Sysmem Fallback Policy” to “Prefer No Sysmem Fallback” in the NVIDIA Control Panel makes that an out-of-memory error instead.

On a Mac (unified memory)

A public report from an M5 Max 48 GB shows Unsloth Desktop on macOS using about 34 GB after loading and about 42 GB while generating with the Q4_K_M DiT (image size not stated), far more than the CUDA figures on this page. The load figure is close to the files with the DiT held at BF16 (32.4 GB with a BF16 encoder and VAE, against 22.4 GB with the Q4_K_M DiT kept quantized, in decimal GB), so that path likely dequantizes it; this has not been confirmed. The ranges here come from CUDA runs (ComfyUI and diffusers) and may underestimate Mac paths such as MPS or MLX. macOS also lets the GPU use only part of the RAM by default, logged at about two thirds to three quarters, so leave headroom on a 48 GB Mac; in the same thread, the user’s 8-bit MLX build in MLX Serve peaked at about 20 GB. stable-diffusion.cpp runs GGUF files on Metal, but we have no Mac measurement of it.

Settings that do this

  • ComfyUI: --disable-dynamic-vram --lowvram runs the text encoders on the CPU (with dynamic VRAM on, --lowvram does nothing); the “VAE Decode (Tiled)” node decodes in tiles; --reserve-vram 1 keeps 1 GB free for the desktop; --novram is the last resort when --lowvram is not enough.
  • stable-diffusion.cpp: --backend te=cpu runs the text encoder on the CPU (it replaces the deprecated --clip-on-cpu); --vae-tiling decodes in tiles; --diffusion-fa uses flash attention in the DiT; --offload-to-cpu keeps the weights in system RAM and copies them to the GPU when needed; --max-vram cuda0=6 caps what it plans to use.

Flags checked on 2026-09-29 in ComfyUI’s cli_args.py and nodes.py, and stable-diffusion.cpp’s performance, backend and Qwen Image 2.1 docs.

Peak VRAM people measured

GPUAppFilesImagePeakSource
RTX 3090 24GBComfyUIDiT GGUF Q4_K_M + encoder INT8 ConvRot + VAE BF161024×102415.39 GiB github.com 2026-09-21
RTX 3090 24GBComfyUIsame2048×115216.83 GiB github.com 2026-09-21
RTX 3090 24GBComfyUIsame, image editing1024×102417.15 GiB github.com 2026-09-21
RTX 3090 24GBComfyUI --lowvramsame, encoder on CPU, tiled VAE1024×10246.14 GiB github.com 2026-09-21
RTX 4060 Ti 16GBComfyUIDiT INT8 ConvRot + encoder INT8 + VAE BF16not stated~15 GB huggingface.co 2026-09-22
RTX 4070 12GBComfyUIDiT INT8 ConvRot + encoder FP8 + VAE BF16, editing2048×204811.28 GiB note.com 2026-09-24
RTX 5090 32GBdiffusersDiT + encoder FP8 (torchao), all resident1024×102421.3 GB huggingface.co 2026-09-22
RTX 5090 32GBdiffuserssame2048×204822.6 GB huggingface.co 2026-09-22
RTX 5090 32GBdiffusersDiT + encoder NF4 (bitsandbytes)1024×102415.2 GB huggingface.co 2026-09-22
RTX 3070 8GBnot statedDiT GGUF Q4_K_M on the GPU, encoder Q4_K_M on the CPU1024×1024, 40 steps6.9 GB x.com 2026-09-26
RTX 3070 8GBnot statedINT8 set, spilling into shared system memorynot stated7.3 GB dedicated + 11.2 GB shared x.com 2026-09-26
RTX 3070 8GBnot statedDiT GGUF Q3_K_M + encoder Q3_K_M, both on the GPUnot statedout of memory in VAE decode x.com 2026-09-26
Apple M5 Max 48GBUnsloth (macOS)DiT GGUF Q4_K_M; encoder and VAE not stated; the DiT appears to be held at BF16not stated~34 GB loaded, ~42 GB generating github.com 2026-09-29

File sizes read from Hugging Face on : Comfy-Org/Qwen-Image-2.1, unsloth/Qwen-Image-2.1-GGUF, unsloth/Qwen-Image-2.1-FP8, gguf-org/qwen-image-2.1-gguf, unsloth/Qwen3-VL-8B-Instruct-GGUF, Qwen/Qwen-Image-2.1. For language models, see the LLM VRAM calculator. How the site’s estimates compare with real runs: predicted vs measured.

How to use

  1. Pick the diffusion model (DiT) file: BF16, FP8, INT8 or a GGUF quant.
  2. Pick the text encoder file and where it runs: on the GPU the whole time, unloaded after the prompt (ComfyUI does this when memory is short), or on the CPU.
  3. Choose the image size, or editing with reference images.
  4. Read the peak range and the GPU table. "Tight" means it fits only at the low end of the measured range.

Frequently asked questions

How much VRAM does Qwen-Image-2.1 need?

The full-precision files take 30.8 GB (BF16 DiT 13.3 GB, BF16 Qwen3-VL-8B encoder 16.3 GB, FP32 VAE 1.3 GB). ComfyUI's INT8 ConvRot set takes 16.1 GB, and a GGUF Q4_K_M DiT with the INT8 encoder and BF16 VAE 13.2 GB. Generating adds roughly 2–6 GB, depending on the image size and the app.

Why do guides say anything from 6 GB to 30 GB?

They count different things. The text encoder is bigger than the image model, so whether it stays on the GPU decides most of the answer: with the encoder on the CPU, a Q4_K_M DiT ran 1024×1024 images in 6.1 GB on an RTX 3090; with everything in BF16 on the GPU, the weights alone are 30.8 GB.

Can Qwen-Image-2.1 run on a 12 GB or 16 GB GPU?

Yes, with quantized files and the text encoder unloaded after the prompt or kept on the CPU. An RTX 4070 12GB edited 2048×2048 images at 11.3 GB peak with the INT8 DiT and FP8 encoder in ComfyUI; an RTX 4060 Ti 16GB peaked near 15 GB with both in INT8.

Which text encoder does Qwen-Image-2.1 use?

Qwen3-VL-8B (Qwen3VLForConditionalGeneration), not Qwen2.5-VL. The 9B Qwen3.5 "prompt enhancer" files in the Comfy-Org repository are an optional extra model, not the text encoder, and are not counted here.

Is GB here GB or GiB?

GiB, the unit nvidia-smi and GPU memory sizes use, as everywhere on this site. The diffusers measurements were reported in GB by their authors and are shown as given.

More calculators

Updated