MiniMax H3 VRAM Calculator (ComfyUI)
MiniMax H3 is four models in a row: a 33B diffusion transformer, the Qwen3-VL-32B text encoder, a video VAE and an audio VAE. Pick your files, where the text encoder runs and the video size to see the peak VRAM, which GPUs hold it, and how much system RAM the part that does not fit needs.
On the GPU for the prompt, then swapped out (ComfyUI default)
864×480 (0.4 MP, template preview)
5 s (124 frames)
– estimated peak VRAM with every needed file on the GPU
No published run is this large, so there is no working-memory range. The file sizes below are still exact; try a shorter length or a smaller size.
Weights on the GPU at once: – of – in all files (exact sizes, audio VAE included).
| GPU memory | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB | 96 GB |
|---|---|---|---|---|---|---|---|
| Result | – | – | – | – | – | – | – |
| System RAM | – | – | – | – | – | – | – |
System RAM: the weights not on the GPU plus about 6 GB for ComfyUI. To keep a copy of every file in RAM, which ComfyUI does by default for fast reruns, plan for –. A 3090 with 32 GB of RAM was OOM-killed until --disable-pinned-memory.
“Fits”: every needed file and the top of the working-memory range fit. “Tight”: only the bottom does. “Streams”: the working memory fits but not every weight, so ComfyUI's dynamic VRAM keeps the rest in system RAM and moves it every step: it works, slower (a 4090 moved 9–10 GB per step in discussion #59). “Risky”: even the working memory may not fit. Leave about 0.5–1 GB for the desktop if the GPU drives your screen.
GB = GiB (1024³ bytes); measurements keep the units their source used.
Common setups at 864×480, 5 seconds
| Setup | Weights on GPU | All files | Peak range | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB | 96 GB |
|---|---|---|---|---|---|---|---|---|---|---|
| Full precision, everything on the GPU | 115.1 GB | 115.1 GB | 116.4 GB–121.8 GB | Streams | Streams | Streams | Streams | Streams | Streams | Streams |
| Full precision, encoder swapped out | 67.1 GB | 115.1 GB | 68.4 GB–73.8 GB | Streams | Streams | Streams | Streams | Streams | Streams | Fits |
| Full INT8 ConvRot set, everything on the GPU | 62.4 GB | 62.4 GB | 63.7 GB–69.1 GB | Streams | Streams | Streams | Streams | Streams | Streams | Fits |
| Pruned BF16 + INT8 encoder | 42.9 GB | 68.2 GB | 44.2 GB–49.6 GB | Streams | Streams | Streams | Streams | Streams | Tight | Fits |
| Pruned INT8 + NVFP4 encoder (Comfy-Org smallest set) | 24.9 GB | 39.6 GB | 26.2 GB–31.6 GB | Streams | Streams | Streams | Streams | Fits | Fits | Fits |
| Pruned FP8 + NVFP4 encoder | 24.9 GB | 39.5 GB | 26.2 GB–31.6 GB | Streams | Streams | Streams | Streams | Fits | Fits | Fits |
| GGUF Q8_0 + NVFP4 encoder + INT8 VAE | 23.3 GB | 37.9 GB | 24.6 GB–30.0 GB | Streams | Streams | Streams | Streams | Fits | Fits | Fits |
| W4A8 + NVFP4 encoder + INT8 VAE | 17.8 GB | 29.5 GB | 19.1 GB–24.5 GB | Streams | Streams | Streams | Tight | Fits | Fits | Fits |
| GGUF Q4_K_M + GGUF Q4_K_M encoder + INT8 VAE | 16.8 GB | 27.5 GB | 18.1 GB–23.5 GB | Streams | Streams | Streams | Fits | Fits | Fits | Fits |
| GGUF Q4_K_M, encoder on the CPU | 14.0 GB | 27.5 GB | 15.3 GB–20.7 GB | Streams | Streams | Tight | Fits | Fits | Fits | Fits |
| GGUF Q2_K, encoder on the CPU | 9.4 GB | 23.0 GB | 10.7 GB–16.1 GB | Streams | Tight | Tight | Fits | Fits | Fits | Fits |
What people measured
Units as each author gave them. Most runs streamed part of the weights, so the VRAM column is what the card held, not what the job would need with every file on it.
| GPU | App | Files | Video | VRAM | System RAM used | Time | Source |
|---|---|---|---|---|---|---|---|
| RTX PRO 6000 96GB RAM: not stated | ComfyUI 0.35.0 | Full INT8 ConvRot DiT + INT8 encoder + both VAEs; DiT and encoder on the card together | 864×480 · 5 s · 20 steps | 63.7 GB peak | – | 47 s | huggingface.co 2026-09-13 |
| RTX 5090 32GB RAM: 94 GB | ComfyUI 0.35.0 | same; 1–3 GB of the DiT crossed PCIe every step | 864×480 · 5 s · 20 steps | 31.8 GB | – | 69 s | huggingface.co 2026-09-13 |
| RTX 4090 24GB RAM: 86–108 GB | ComfyUI 0.35.0 | same; 9–10 GB crossed PCIe every step | 864×480 · 5 s · 20 steps | – | 68 GB peak | 92–93 s | huggingface.co 2026-09-13 |
| RTX 3090 24GB RAM: 32 GB | ComfyUI 0.30.1, --disable-pinned-memory | Pruned INT8 DiT + NVFP4 encoder + FP16 VAE | 832×480 · 124 frames · 20 steps | 23,716 MiB peak | – | 4 min 26 s | github.com 2026-08-04 |
| RTX 3090 24GB RAM: 32 GB | ComfyUI 0.30.1, --disable-pinned-memory | same | 832×480 · 362 frames · 20 steps | 18,884 MiB peak | 7.5 GB (29.9 GB and OOM-killed without the flag) | 23 min 17 s | github.com 2026-08-04 |
| RTX 5070 Ti 16GB RAM: 125 GB | ComfyUI 0.30.1 | Pruned INT8 DiT + NVFP4 encoder + both VAEs (42.5 GB) | 1344×768 · 5 s | 14,437 MiB peak | 45.4 GiB RSS; 12.6 GiB with --fast-disk | – | huggingface.co 2026-08-06 |
| RTX 5070 Ti 16GB RAM: 125 GB | ComfyUI 0.30.1 | same | 640×480 · 30 s | 14,197 MiB peak | same | – | huggingface.co 2026-08-06 |
| RTX 3090 24GB RAM: not stated | ComfyUI 0.34.0 | Pruned INT8 DiT + LightX2V 4-step LoRA | 1280×704 · 192 frames · 4 steps | 23,703 MiB at step 1 (stopped) | – | – | github.com 2026-09-09 |
| RTX 3090 24GB RAM: not stated | ComfyUI 0.32 | same LoRA family | 1344×768 · 192 frames | ~19,575 MiB peak | – | – | github.com 2026-09-09 |
| RTX 5070 Ti 16GB RAM: not stated | ComfyUI 0.30.2 | Full INT8 ConvRot DiT + INT8 encoder | 1280×736 · 362 frames · 20 steps | 15.1 GB average | – | 26 min 20 s (~79 s/step) | github.com 2026-08-16 |
| RTX 5070 Ti 16GB RAM: not stated | ComfyUI 0.33.1 | same | 1280×736 · 362 frames · 20 steps | 15,773 MB average | – | ~2 h estimated (341+ s/step), cancelled | github.com 2026-08-16 |
| RTX 5070 Ti 16GB RAM: 16 GB | ComfyUI 0.34.2 / 0.35.0 | Community hybrid INT8 turbo DiT (10Eros Max H3) + NVFP4 encoder + INT8 VAE | 0.6 MP · 6 s · 6 steps | ~74% (0.34.2); ~90% on 0.35.0 with --vram-headroom 1, which crashed the PC without it | – | 129–154 s | github.com 2026-09-09 |
| RTX 4060 Laptop 8GB RAM: 16 GB | ComfyUI 0.34.0 + PR #16148 | Pruned W4A8 DiT + NVFP4 encoder | not stated (38,968 tokens) · 20 steps | whole card | – | 20 min 51 s (72 s/step) | github.com 2026-09-06 |
| RTX 3060 12GB RAM: not stated | ComfyUI | 8-step Turbo LoRA, INT8 video VAE | 864×480 · 5 s · 8 steps | – | – | 4.5 min | huggingface.co 2026-08-06 |
| RTX 5090 32GB RAM: not stated | ComfyUI 0.33.0 nightly | Pruned INT8 DiT + NVFP4 encoder + 4-step Turbo LoRA | 1344×768 · 56 frames · 4 steps | – | – | 16.2 s | huggingface.co 2026-08-16 |
| Ryzen AI Max (Radeon 8060S), 48 GB as VRAM RAM: 15 GB | ComfyUI | Pruned INT8 DiT + NVFP4 encoder, both loaded completely | 480p · 5 s · 20 steps | all weights resident | – | 29 min 30 s (88.5 s/step) | huggingface.co 2026-08-05 |
| DGX Spark GB10, 128 GB unified RAM: shared | ComfyUI 0.30.1 | Pruned INT8 DiT + NVFP4 encoder + FP16 VAE | 864×480 · 39 / 56 / 124 frames · 20 steps | – | ~39 GB container memory at 124 frames | 88 / 122 / 326 s | github.com 2026-09-26 |
| H200 141GB RAM: not stated | ComfyUI | Pruned INT8 DiT, loaded completely | 1920×1088 · 10 s · 20 steps | – | – | 46 min 45 s (140 s/step) | huggingface.co 2026-08-06 |
File sizes read from the Hugging Face API on : Comfy-Org/MiniMax-H3, Abiray/
How to use
- Pick the diffusion model: full or pruned, BF16, INT8, FP8, W4A8, NVFP4 or a GGUF quant.
- Pick the text encoder, the video VAE and, if you use one, the Fun ControlNet patch.
- Choose where the text encoder runs. ComfyUI's default (0.37 and later, with dynamic VRAM) puts it on the GPU for the prompt and then swaps it out.
- Choose the resolution and length. Read the peak range, the GPU table and the RAM each card needs. "Streams" means ComfyUI keeps part of the weights in RAM, which works but is slower.
Frequently asked questions
How much VRAM does MiniMax H3 need?
The full-precision files take 115.1 GB (DiT 61.7 GB, Qwen3-VL-32B encoder 48.0 GB, VAEs 5.4 GB). Comfy-Org's smallest set (pruned INT8 DiT, NVFP4 encoder, FP16 VAE) is 39.6 GB, and ComfyUI never needs it all at once: the larger of DiT and encoder plus the VAEs is 24.9 GB, and a 864×480, 5-second clip adds 1.3–6.7 GB. Smaller GPUs still run it; ComfyUI streams the rest of the weights from system RAM.
How much system RAM do I need?
The weights that are not on the GPU plus about 6 GB. With the smallest set on a 24 GB card that is about 28 GB for a 5-second 864×480 clip. ComfyUI keeps a copy of every file by default, which takes about 46 GB: a 5070 Ti used 45.4 GiB, or 12.6 GiB with --fast-disk. A 3090 with 32 GB of RAM was killed at 29.9 GB until --disable-pinned-memory, which cut it to 7.5 GB.
Should I use the pruned or the full MiniMax H3 model?
For generating video, the pruned one: pruned INT8 is 19.5 GB instead of 31.7 GB, pruned BF16 37.5 GB instead of 61.7 GB. The model card says about 13B of the DiT's 33B parameters are AdaLN modulation branches whose outputs can be precomputed, and Comfy-Org replaced them with a lookup table "with no loss in output quality", per the ComfyUI blog. The full files are for fine-tuning.
Why did a ComfyUI update make H3 slower or crash?
Several versions changed how H3 is loaded. Issue #15665: from 0.32.0, a 1280×736, 362-frame job went from about 79 s to over 341 s per step on a 16 GB card. Issue #16150: 0.35.0 filled VRAM and crashed some PCs; --vram-headroom 1 and --disable-comfy-compiler helped. If a new version is slower on your card, run the same workflow once with --disable-dynamic-vram and compare the seconds per step.
What changed for the text encoder in ComfyUI 0.37?
PR #16374, "Always put text encoder on GPU when dynamic vram on", was merged on 2026-09-17 and shipped in v0.37.0. With dynamic VRAM on (the default), the Qwen3-VL-32B encoder now always runs on the GPU for the prompt and is then swapped out for the DiT, and --lowvram does nothing. To run the encoder on the CPU instead, start ComfyUI with --disable-dynamic-vram --lowvram. To see which is faster on your card, run the same workflow with and without --disable-dynamic-vram and compare the time per step and per prompt.
Is GB here GB or GiB?
GiB, the unit nvidia-smi and GPU memory sizes use, as everywhere on this site. The measurements are shown in the units their authors used.
More calculators
- Fine-Tuning VRAM Calculator GPU memory for full fine-tuning, LoRA and QLoRA in Transformers or Unsloth, checked against published runs. Open →
- vLLM KV Cache & Concurrency Calculator The KV cache pool vLLM allocates, in tokens, and how many requests fit at once, with the vllm serve command. Open →
- What Can My PC Run? Detects your GPU in the browser and lists the local LLMs it runs, with the best quantization and speed. Open →
- MoE Offload Calculator
(--n-cpu-moe) The smallest llama.cpp --n-cpu-moe that fits your GPU, from real GGUF tensor sizes. Open → - Qwen3.8 27B GGUF Quants: Bonsai 2 vs GSQ-RCO vs UD Ternary Bonsai 2, GSQ-RCO and Unsloth UD files of Qwen3.8 27B: VRAM, quality, speed and the engine each needs. Open →
- Qwen-Image-2.1 VRAM Calculator Peak VRAM for every DiT, text encoder and VAE combination, with real measurements. Open →
Updated