MiniMax H3 VRAM Calculator (ComfyUI)

MiniMax H3 is four models in a row: a 33B diffusion transformer, the Qwen3-VL-32B text encoder, a video VAE and an audio VAE. Pick your files, where the text encoder runs and the video size to see the peak VRAM, which GPUs hold it, and how much system RAM the part that does not fit needs.

On the GPU for the prompt, then swapped out (ComfyUI default)

864×480 (0.4 MP, template preview)

5 s (124 frames)

– estimated peak VRAM with every needed file on the GPU

Weights on the GPU at once: – of – in all files (exact sizes, audio VAE included).

GPU memory8 GB12 GB16 GB24 GB32 GB48 GB96 GB
Result–––––––
System RAM–––––––

System RAM: the weights not on the GPU plus about 6 GB for ComfyUI. To keep a copy of every file in RAM, which ComfyUI does by default for fast reruns, plan for –. A 3090 with 32 GB of RAM was OOM-killed until --disable-pinned-memory.

“Fits”: every needed file and the top of the working-memory range fit. “Tight”: only the bottom does. “Streams”: the working memory fits but not every weight, so ComfyUI's dynamic VRAM keeps the rest in system RAM and moves it every step: it works, slower (a 4090 moved 9–10 GB per step in discussion #59). “Risky”: even the working memory may not fit. Leave about 0.5–1 GB for the desktop if the GPU drives your screen.

GB = GiB (1024³ bytes); measurements keep the units their source used.

Common setups at 864×480, 5 seconds

SetupWeights on GPUAll filesPeak range8 GB12 GB16 GB24 GB32 GB48 GB96 GB
Full precision, everything on the GPU 115.1 GB 115.1 GB 116.4 GB–121.8 GB StreamsStreamsStreamsStreamsStreamsStreamsStreams
Full precision, encoder swapped out 67.1 GB 115.1 GB 68.4 GB–73.8 GB StreamsStreamsStreamsStreamsStreamsStreamsFits
Full INT8 ConvRot set, everything on the GPU 62.4 GB 62.4 GB 63.7 GB–69.1 GB StreamsStreamsStreamsStreamsStreamsStreamsFits
Pruned BF16 + INT8 encoder 42.9 GB 68.2 GB 44.2 GB–49.6 GB StreamsStreamsStreamsStreamsStreamsTightFits
Pruned INT8 + NVFP4 encoder (Comfy-Org smallest set) 24.9 GB 39.6 GB 26.2 GB–31.6 GB StreamsStreamsStreamsStreamsFitsFitsFits
Pruned FP8 + NVFP4 encoder 24.9 GB 39.5 GB 26.2 GB–31.6 GB StreamsStreamsStreamsStreamsFitsFitsFits
GGUF Q8_0 + NVFP4 encoder + INT8 VAE 23.3 GB 37.9 GB 24.6 GB–30.0 GB StreamsStreamsStreamsStreamsFitsFitsFits
W4A8 + NVFP4 encoder + INT8 VAE 17.8 GB 29.5 GB 19.1 GB–24.5 GB StreamsStreamsStreamsTightFitsFitsFits
GGUF Q4_K_M + GGUF Q4_K_M encoder + INT8 VAE 16.8 GB 27.5 GB 18.1 GB–23.5 GB StreamsStreamsStreamsFitsFitsFitsFits
GGUF Q4_K_M, encoder on the CPU 14.0 GB 27.5 GB 15.3 GB–20.7 GB StreamsStreamsTightFitsFitsFitsFits
GGUF Q2_K, encoder on the CPU 9.4 GB 23.0 GB 10.7 GB–16.1 GB StreamsTightTightFitsFitsFitsFits

What people measured

Units as each author gave them. Most runs streamed part of the weights, so the VRAM column is what the card held, not what the job would need with every file on it.

GPUAppFilesVideoVRAMSystem RAM usedTimeSource
RTX PRO 6000 96GB
RAM: not stated
ComfyUI 0.35.0 Full INT8 ConvRot DiT + INT8 encoder + both VAEs; DiT and encoder on the card together 864×480 · 5 s · 20 steps 63.7 GB peak – 47 s huggingface.co 2026-09-13
RTX 5090 32GB
RAM: 94 GB
ComfyUI 0.35.0 same; 1–3 GB of the DiT crossed PCIe every step 864×480 · 5 s · 20 steps 31.8 GB – 69 s huggingface.co 2026-09-13
RTX 4090 24GB
RAM: 86–108 GB
ComfyUI 0.35.0 same; 9–10 GB crossed PCIe every step 864×480 · 5 s · 20 steps – 68 GB peak 92–93 s huggingface.co 2026-09-13
RTX 3090 24GB
RAM: 32 GB
ComfyUI 0.30.1, --disable-pinned-memory Pruned INT8 DiT + NVFP4 encoder + FP16 VAE 832×480 · 124 frames · 20 steps 23,716 MiB peak – 4 min 26 s github.com 2026-08-04
RTX 3090 24GB
RAM: 32 GB
ComfyUI 0.30.1, --disable-pinned-memory same 832×480 · 362 frames · 20 steps 18,884 MiB peak 7.5 GB (29.9 GB and OOM-killed without the flag) 23 min 17 s github.com 2026-08-04
RTX 5070 Ti 16GB
RAM: 125 GB
ComfyUI 0.30.1 Pruned INT8 DiT + NVFP4 encoder + both VAEs (42.5 GB) 1344×768 · 5 s 14,437 MiB peak 45.4 GiB RSS; 12.6 GiB with --fast-disk – huggingface.co 2026-08-06
RTX 5070 Ti 16GB
RAM: 125 GB
ComfyUI 0.30.1 same 640×480 · 30 s 14,197 MiB peak same – huggingface.co 2026-08-06
RTX 3090 24GB
RAM: not stated
ComfyUI 0.34.0 Pruned INT8 DiT + LightX2V 4-step LoRA 1280×704 · 192 frames · 4 steps 23,703 MiB at step 1 (stopped) – – github.com 2026-09-09
RTX 3090 24GB
RAM: not stated
ComfyUI 0.32 same LoRA family 1344×768 · 192 frames ~19,575 MiB peak – – github.com 2026-09-09
RTX 5070 Ti 16GB
RAM: not stated
ComfyUI 0.30.2 Full INT8 ConvRot DiT + INT8 encoder 1280×736 · 362 frames · 20 steps 15.1 GB average – 26 min 20 s (~79 s/step) github.com 2026-08-16
RTX 5070 Ti 16GB
RAM: not stated
ComfyUI 0.33.1 same 1280×736 · 362 frames · 20 steps 15,773 MB average – ~2 h estimated (341+ s/step), cancelled github.com 2026-08-16
RTX 5070 Ti 16GB
RAM: 16 GB
ComfyUI 0.34.2 / 0.35.0 Community hybrid INT8 turbo DiT (10Eros Max H3) + NVFP4 encoder + INT8 VAE 0.6 MP · 6 s · 6 steps ~74% (0.34.2); ~90% on 0.35.0 with --vram-headroom 1, which crashed the PC without it – 129–154 s github.com 2026-09-09
RTX 4060 Laptop 8GB
RAM: 16 GB
ComfyUI 0.34.0 + PR #16148 Pruned W4A8 DiT + NVFP4 encoder not stated (38,968 tokens) · 20 steps whole card – 20 min 51 s (72 s/step) github.com 2026-09-06
RTX 3060 12GB
RAM: not stated
ComfyUI 8-step Turbo LoRA, INT8 video VAE 864×480 · 5 s · 8 steps – – 4.5 min huggingface.co 2026-08-06
RTX 5090 32GB
RAM: not stated
ComfyUI 0.33.0 nightly Pruned INT8 DiT + NVFP4 encoder + 4-step Turbo LoRA 1344×768 · 56 frames · 4 steps – – 16.2 s huggingface.co 2026-08-16
Ryzen AI Max (Radeon 8060S), 48 GB as VRAM
RAM: 15 GB
ComfyUI Pruned INT8 DiT + NVFP4 encoder, both loaded completely 480p · 5 s · 20 steps all weights resident – 29 min 30 s (88.5 s/step) huggingface.co 2026-08-05
DGX Spark GB10, 128 GB unified
RAM: shared
ComfyUI 0.30.1 Pruned INT8 DiT + NVFP4 encoder + FP16 VAE 864×480 · 39 / 56 / 124 frames · 20 steps – ~39 GB container memory at 124 frames 88 / 122 / 326 s github.com 2026-09-26
H200 141GB
RAM: not stated
ComfyUI Pruned INT8 DiT, loaded completely 1920×1088 · 10 s · 20 steps – – 46 min 45 s (140 s/step) huggingface.co 2026-08-06

File sizes read from the Hugging Face API on : Comfy-Org/MiniMax-H3, Abiray/MiniMax-H3-Pruned-GGUF, koongrizzly/MiniMax_H3_int4_W4A8_ConvRot_Pruned, Abiray/Minimax-H3-nvfp4-INT4-INT8-Convrot, unsloth/MiniMax-H3-GGUF, Abiray/MiniMax-H3-GGUF. For image models, see the Qwen-Image 2.1 VRAM calculator; for language models, the LLM VRAM calculator.

How to use

  1. Pick the diffusion model: full or pruned, BF16, INT8, FP8, W4A8, NVFP4 or a GGUF quant.
  2. Pick the text encoder, the video VAE and, if you use one, the Fun ControlNet patch.
  3. Choose where the text encoder runs. ComfyUI's default (0.37 and later, with dynamic VRAM) puts it on the GPU for the prompt and then swaps it out.
  4. Choose the resolution and length. Read the peak range, the GPU table and the RAM each card needs. "Streams" means ComfyUI keeps part of the weights in RAM, which works but is slower.

Frequently asked questions

How much VRAM does MiniMax H3 need?

The full-precision files take 115.1 GB (DiT 61.7 GB, Qwen3-VL-32B encoder 48.0 GB, VAEs 5.4 GB). Comfy-Org's smallest set (pruned INT8 DiT, NVFP4 encoder, FP16 VAE) is 39.6 GB, and ComfyUI never needs it all at once: the larger of DiT and encoder plus the VAEs is 24.9 GB, and a 864×480, 5-second clip adds 1.3–6.7 GB. Smaller GPUs still run it; ComfyUI streams the rest of the weights from system RAM.

How much system RAM do I need?

The weights that are not on the GPU plus about 6 GB. With the smallest set on a 24 GB card that is about 28 GB for a 5-second 864×480 clip. ComfyUI keeps a copy of every file by default, which takes about 46 GB: a 5070 Ti used 45.4 GiB, or 12.6 GiB with --fast-disk. A 3090 with 32 GB of RAM was killed at 29.9 GB until --disable-pinned-memory, which cut it to 7.5 GB.

Should I use the pruned or the full MiniMax H3 model?

For generating video, the pruned one: pruned INT8 is 19.5 GB instead of 31.7 GB, pruned BF16 37.5 GB instead of 61.7 GB. The model card says about 13B of the DiT's 33B parameters are AdaLN modulation branches whose outputs can be precomputed, and Comfy-Org replaced them with a lookup table "with no loss in output quality", per the ComfyUI blog. The full files are for fine-tuning.

Why did a ComfyUI update make H3 slower or crash?

Several versions changed how H3 is loaded. Issue #15665: from 0.32.0, a 1280×736, 362-frame job went from about 79 s to over 341 s per step on a 16 GB card. Issue #16150: 0.35.0 filled VRAM and crashed some PCs; --vram-headroom 1 and --disable-comfy-compiler helped. If a new version is slower on your card, run the same workflow once with --disable-dynamic-vram and compare the seconds per step.

What changed for the text encoder in ComfyUI 0.37?

PR #16374, "Always put text encoder on GPU when dynamic vram on", was merged on 2026-09-17 and shipped in v0.37.0. With dynamic VRAM on (the default), the Qwen3-VL-32B encoder now always runs on the GPU for the prompt and is then swapped out for the DiT, and --lowvram does nothing. To run the encoder on the CPU instead, start ComfyUI with --disable-dynamic-vram --lowvram. To see which is faster on your card, run the same workflow with and without --disable-dynamic-vram and compare the time per step and per prompt.

Is GB here GB or GiB?

GiB, the unit nvidia-smi and GPU memory sizes use, as everywhere on this site. The measurements are shown in the units their authors used.

More calculators

Updated