Qwen-Image-2.1 VRAM Calculator
Qwen-Image-2.1 is three models: a 7B diffusion transformer, the Qwen3-VL-8B text encoder and a VAE. Pick the file you run for each, where the text encoder lives, and the image size, and see the peak VRAM range and which GPUs fit.
– estimated peak VRAM
Weights on the GPU at once: – (exact file sizes). Working memory on top is the range measured on real GPUs below.
| GPU memory | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB |
|---|---|---|---|---|---|
| Result | – | – | – | – | – |
“Tight”: fits only at the low end of the range, as in ComfyUI, which can also move part of the model to system RAM (slower). Leave about 0.5–1 GB for the desktop if the GPU drives your screen.
Common setups at 1024×1024
| Setup | Weights on GPU | Peak range | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB |
|---|---|---|---|---|---|---|---|
| Full precision (diffusers default) | 30.8 GB | 32.9 GB–35.5 GB | No | No | No | No | No |
| ComfyUI INT8 ConvRot set | 16.1 GB | 18.2 GB–20.8 GB | No | No | No | Fits | Fits |
| ComfyUI INT8 set, encoder unloaded | 9.3 GB | 11.4 GB–14.0 GB | No | Tight | Fits | Fits | Fits |
| FP8 DiT + FP8 encoder (unsloth) | 16.0 GB | 18.1 GB–20.7 GB | No | No | No | Fits | Fits |
| GGUF Q8_0 + encoder Q8_0 | 15.9 GB | 18.0 GB–20.6 GB | No | No | No | Fits | Fits |
| GGUF Q4_K_M + encoder INT8 | 13.2 GB | 15.3 GB–17.9 GB | No | No | Tight | Fits | Fits |
| GGUF Q4_K_M + encoder Q4_K_M | 9.2 GB | 11.3 GB–13.9 GB | No | Tight | Fits | Fits | Fits |
| GGUF Q4_K_M, encoder on CPU | 4.5 GB | 6.1 GB–6.9 GB | Fits | Fits | Fits | Fits | Fits |
Low-VRAM setups for 8, 12 and 16 GB
Each row is a mix of the files above at 1024×1024. The three phase columns are the weights on the GPU while the prompt is encoded, while the image is denoised and while the VAE decodes it: exact file sizes. The peak adds the working memory measured in real runs, 1.6–2.4 GB with the text encoder on the CPU.
| Card | DiT + text encoder | Text encoder | Prompt | Denoise | VAE decode | Peak | Result | Try it |
|---|---|---|---|---|---|---|---|---|
| 8 GB | GGUF Q4_K_M + GGUF Q4_K_M + tiled VAE decode | On the CPU | 0.6 GB | 4.5 GB | 4.5 GB | 6.1 GB–6.9 GB | Fits | Try it |
| 8 GB | GGUF Q5_K_M + GGUF Q4_K_M + tiled VAE decode | On the CPU | 0.6 GB | 5.6 GB | 5.6 GB | 7.2 GB–8.0 GB | Tight | Try it |
| 8 GB | GGUF Q3_K_M + GGUF Q3_K_M An RTX 3070 8GB ran out of memory in VAE decode with this pair (both on the GPU). | Everything on the GPU | 7.4 GB | 7.4 GB | 7.4 GB | 9.5 GB–12.1 GB | No | Try it |
| 12 GB | GGUF Q8_0 + GGUF Q4_K_M + tiled VAE decode | On the CPU | 0.6 GB | 7.7 GB | 7.7 GB | 9.3 GB–10.1 GB | Fits | Try it |
| 12 GB | GGUF Q4_K_M + GGUF Q4_K_M | Unloaded after the prompt | 5.3 GB | 4.5 GB | 4.5 GB | 7.4 GB–10.0 GB | Fits | Try it |
| 12 GB | INT8 ConvRot (ComfyUI) + FP8 | Unloaded after the prompt | 9.4 GB | 7.4 GB | 7.4 GB | 11.5 GB–14.1 GB | Tight | Try it |
| 16 GB | INT8 ConvRot (ComfyUI) + INT8 ConvRot (ComfyUI) | Unloaded after the prompt | 9.3 GB | 7.4 GB | 7.4 GB | 11.4 GB–14.0 GB | Fits | Try it |
| 16 GB | GGUF Q6_K + GGUF Q4_K_M | Everything on the GPU | 11.2 GB | 11.2 GB | 11.2 GB | 13.3 GB–15.9 GB | Fits | Try it |
| 16 GB | GGUF Q4_K_M + INT8 ConvRot (ComfyUI) | Everything on the GPU | 13.2 GB | 13.2 GB | 13.2 GB | 15.3 GB–17.9 GB | Tight | Try it |
On Windows, “runs on 8 GB” can mean 18 GB
NVIDIA’s Windows driver moves what does not fit into shared system memory instead of failing (System Memory Fallback, since driver 536.40), so a job finishes but runs from system RAM. One RTX 3070 8GB owner saw the INT8 set use 7.3 GB of dedicated VRAM plus 11.2 GB of shared memory. On the same card, the Q4_K_M DiT on the GPU with the Q4_K_M encoder on the CPU peaked at 6.9 GB, and Q3_K_M files for both, together 6.8 GB, ran out of memory while decoding. If Task Manager shows “Shared GPU memory” climbing during a run, it does not fit; setting “CUDA - Sysmem Fallback Policy” to “Prefer No Sysmem Fallback” in the NVIDIA Control Panel makes that an out-of-memory error instead.
On a Mac (unified memory)
A public report from an M5 Max 48 GB shows Unsloth Desktop on macOS using about 34 GB after loading and about 42 GB while generating with the Q4_K_M DiT (image size not stated), far more than the CUDA figures on this page. The load figure is close to the files with the DiT held at BF16 (32.4 GB with a BF16 encoder and VAE, against 22.4 GB with the Q4_K_M DiT kept quantized, in decimal GB), so that path likely dequantizes it; this has not been confirmed. The ranges here come from CUDA runs (ComfyUI and diffusers) and may underestimate Mac paths such as MPS or MLX. macOS also lets the GPU use only part of the RAM by default, logged at about two thirds to three quarters, so leave headroom on a 48 GB Mac; in the same thread, the user’s 8-bit MLX build in MLX Serve peaked at about 20 GB. stable-diffusion.cpp runs GGUF files on Metal, but we have no Mac measurement of it.
Settings that do this
- ComfyUI:
--disable-dynamic-vram --lowvramruns the text encoders on the CPU (with dynamic VRAM on,--lowvramdoes nothing); the “VAE Decode (Tiled)” node decodes in tiles;--reserve-vram 1keeps 1 GB free for the desktop;--novramis the last resort when--lowvramis not enough. - stable-diffusion.cpp:
--backend te=cpuruns the text encoder on the CPU (it replaces the deprecated--clip-on-cpu);--vae-tilingdecodes in tiles;--diffusion-fauses flash attention in the DiT;--offload-to-cpukeeps the weights in system RAM and copies them to the GPU when needed;--max-vram cuda0=6caps what it plans to use.
Flags checked on 2026-09-29 in ComfyUI’s cli_args.py and nodes.py, and stable-diffusion.cpp’s performance, backend and Qwen Image 2.1 docs.
Peak VRAM people measured
| GPU | App | Files | Image | Peak | Source |
|---|---|---|---|---|---|
| RTX 3090 24GB | ComfyUI | DiT GGUF Q4_K_M + encoder INT8 ConvRot + VAE BF16 | 1024×1024 | 15.39 GiB | github.com 2026-09-21 |
| RTX 3090 24GB | ComfyUI | same | 2048×1152 | 16.83 GiB | github.com 2026-09-21 |
| RTX 3090 24GB | ComfyUI | same, image editing | 1024×1024 | 17.15 GiB | github.com 2026-09-21 |
| RTX 3090 24GB | ComfyUI --lowvram | same, encoder on CPU, tiled VAE | 1024×1024 | 6.14 GiB | github.com 2026-09-21 |
| RTX 4060 Ti 16GB | ComfyUI | DiT INT8 ConvRot + encoder INT8 + VAE BF16 | not stated | ~15 GB | huggingface.co 2026-09-22 |
| RTX 4070 12GB | ComfyUI | DiT INT8 ConvRot + encoder FP8 + VAE BF16, editing | 2048×2048 | 11.28 GiB | note.com 2026-09-24 |
| RTX 5090 32GB | diffusers | DiT + encoder FP8 (torchao), all resident | 1024×1024 | 21.3 GB | huggingface.co 2026-09-22 |
| RTX 5090 32GB | diffusers | same | 2048×2048 | 22.6 GB | huggingface.co 2026-09-22 |
| RTX 5090 32GB | diffusers | DiT + encoder NF4 (bitsandbytes) | 1024×1024 | 15.2 GB | huggingface.co 2026-09-22 |
| RTX 3070 8GB | not stated | DiT GGUF Q4_K_M on the GPU, encoder Q4_K_M on the CPU | 1024×1024, 40 steps | 6.9 GB | x.com 2026-09-26 |
| RTX 3070 8GB | not stated | INT8 set, spilling into shared system memory | not stated | 7.3 GB dedicated + 11.2 GB shared | x.com 2026-09-26 |
| RTX 3070 8GB | not stated | DiT GGUF Q3_K_M + encoder Q3_K_M, both on the GPU | not stated | out of memory in VAE decode | x.com 2026-09-26 |
| Apple M5 Max 48GB | Unsloth (macOS) | DiT GGUF Q4_K_M; encoder and VAE not stated; the DiT appears to be held at BF16 | not stated | ~34 GB loaded, ~42 GB generating | github.com 2026-09-29 |
File sizes read from Hugging Face on : Comfy-Org/
How to use
- Pick the diffusion model (DiT) file: BF16, FP8, INT8 or a GGUF quant.
- Pick the text encoder file and where it runs: on the GPU the whole time, unloaded after the prompt (ComfyUI does this when memory is short), or on the CPU.
- Choose the image size, or editing with reference images.
- Read the peak range and the GPU table. "Tight" means it fits only at the low end of the measured range.
Frequently asked questions
How much VRAM does Qwen-Image-2.1 need?
The full-precision files take 30.8 GB (BF16 DiT 13.3 GB, BF16 Qwen3-VL-8B encoder 16.3 GB, FP32 VAE 1.3 GB). ComfyUI's INT8 ConvRot set takes 16.1 GB, and a GGUF Q4_K_M DiT with the INT8 encoder and BF16 VAE 13.2 GB. Generating adds roughly 2–6 GB, depending on the image size and the app.
Why do guides say anything from 6 GB to 30 GB?
They count different things. The text encoder is bigger than the image model, so whether it stays on the GPU decides most of the answer: with the encoder on the CPU, a Q4_K_M DiT ran 1024×1024 images in 6.1 GB on an RTX 3090; with everything in BF16 on the GPU, the weights alone are 30.8 GB.
Can Qwen-Image-2.1 run on a 12 GB or 16 GB GPU?
Yes, with quantized files and the text encoder unloaded after the prompt or kept on the CPU. An RTX 4070 12GB edited 2048×2048 images at 11.3 GB peak with the INT8 DiT and FP8 encoder in ComfyUI; an RTX 4060 Ti 16GB peaked near 15 GB with both in INT8.
Which text encoder does Qwen-Image-2.1 use?
Qwen3-VL-8B (Qwen3VL
Is GB here GB or GiB?
GiB, the unit nvidia-smi and GPU memory sizes use, as everywhere on this site. The diffusers measurements were reported in GB by their authors and are shown as given.
More calculators
- Fine-Tuning VRAM Calculator GPU memory for full fine-tuning, LoRA and QLoRA in Transformers or Unsloth, checked against published runs. Open →
- vLLM KV Cache & Concurrency Calculator The KV cache pool vLLM allocates, in tokens, and how many requests fit at once, with the vllm serve command. Open →
- What Can My PC Run? Detects your GPU in the browser and lists the local LLMs it runs, with the best quantization and speed. Open →
- MiniMax H3 VRAM Calculator
(ComfyUI) VRAM and system RAM for MiniMax H3 video in ComfyUI: pruned, INT8, NVFP4 and GGUF files on 8–96 GB GPUs. Open → - MoE Offload Calculator
(--n-cpu-moe) The smallest llama.cpp --n-cpu-moe that fits your GPU, from real GGUF tensor sizes. Open → - Qwen3.8 27B GGUF Quants: Bonsai 2 vs GSQ-RCO vs UD Ternary Bonsai 2, GSQ-RCO and Unsloth UD files of Qwen3.8 27B: VRAM, quality, speed and the engine each needs. Open →
Updated