Datasheet for local AI49 models98 GPUsData read 2026-09-25
No adsNo tracking
Check my GPU →

MiniMax H3 Pruned VRAM requirements

An official, smaller cut of MiniMax H3 published by Comfy-Org, made for cards that cannot hold the full 33B model.

Released 2026-07Licence: MiniMax H3 Community License (custom; commercial-use restrictions, see license)Steps: 20–40Text encoder: Qwen3-VL 32B (27.1 GB as INT8)
TypeVideo
Parameterspruned 33B
Fits entirely from15 GB
8-bit or better from27 GB

01The files, and how much VRAM each needs

FileSizeNeededMin. VRAMQualitySource
16-bit40.2 GB46.0 GB47 GBthe original weightsComfy-Org/MiniMax-H3 →
Q8_021.6 GB27.4 GB28 GBpractically identical to the originalAbiray/MiniMax-H3-Pruned-GGUF →
FP821.0 GB26.8 GB27 GBpractically identical to the originalComfy-Org/MiniMax-H3 →
INT821.0 GB26.8 GB27 GBpractically identical to the originalComfy-Org/MiniMax-H3 →
Q6_K16.7 GB22.5 GB23 GBvery close to the originalAbiray/MiniMax-H3-Pruned-GGUF →
Q5_K_M14.1 GB19.9 GB20 GBclose; small differences in fine detailAbiray/MiniMax-H3-Pruned-GGUF →
Q4_K_M11.6 GB17.4 GB18 GBgood; some loss in fine detail and textAbiray/MiniMax-H3-Pruned-GGUF →
Q3_K_M8.9 GB14.7 GB15 GBnoticeable loss of detailAbiray/MiniMax-H3-Pruned-GGUF →

“Needed” = file + 5 GB working memory + 0.8 GB system reserve.

03Best GPU for MiniMax H3 Pruned

The cheapest cards (by launch price) that run it well, and every card sorted by memory: best GPU for MiniMax H3 Pruned → Planning bigger images or longer clips? Open the calculator →

04By graphics card

GPUVRAMVerdictBest fileNeeded
Desktop graphics cards
RTX 2060 6 GB6 GBOffload onlyQ4_K_M17.4 GB
RTX 3050 6 GB6 GBOffload onlyQ4_K_M17.4 GB
RTX 2070 Super 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 2080 Super 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 3050 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 3060 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 3060 Ti 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 3070 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 3070 Ti 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 4060 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 4060 Ti 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 5050 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 5060 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 5060 Ti 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RX 7600 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RX 9050 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RX 9060 XT 8 GB8 GBOffload onlyQ4_K_M17.4 GB
Arc B570 10 GB10 GBOffload onlyQ4_K_M17.4 GB
RTX 3080 10 GB10 GBOffload onlyQ4_K_M17.4 GB
RTX 2080 Ti 11 GB11 GBOffload onlyQ4_K_M17.4 GB
Arc B580 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RTX 2060 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RTX 3060 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RTX 3080 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RTX 3080 Ti 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RTX 4070 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RTX 4070 Super 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RTX 4070 Ti 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RTX 5070 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RX 7700 XT 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RX 9070 GRE 12 GB12 GBOffload onlyQ4_K_M17.4 GB
Arc A770 16 GB16 GBTightQ3_K_M14.7 GB
RTX 4060 Ti 16 GB16 GBTightQ3_K_M14.7 GB
RTX 4070 Ti Super 16 GB16 GBTightQ3_K_M14.7 GB
RTX 4080 16 GB16 GBTightQ3_K_M14.7 GB
RTX 4080 Super 16 GB16 GBTightQ3_K_M14.7 GB
RTX 5060 Ti 16 GB16 GBTightQ3_K_M14.7 GB
RTX 5070 Ti 16 GB16 GBTightQ3_K_M14.7 GB
RTX 5080 16 GB16 GBTightQ3_K_M14.7 GB
RX 7600 XT 16 GB16 GBTightQ3_K_M14.7 GB
RX 7800 XT 16 GB16 GBTightQ3_K_M14.7 GB
RX 7900 GRE 16 GB16 GBTightQ3_K_M14.7 GB
RX 9060 XT 16 GB16 GBTightQ3_K_M14.7 GB
RX 9070 16 GB16 GBTightQ3_K_M14.7 GB
RX 9070 XT 16 GB16 GBTightQ3_K_M14.7 GB
RX 7900 XT 20 GB20 GBRunsQ5_K_M19.9 GB
Arc Pro B60 24 GB24 GBRunsQ6_K22.5 GB
RTX 3090 24 GB24 GBRunsQ6_K22.5 GB
RTX 3090 Ti 24 GB24 GBRunsQ6_K22.5 GB
RTX 4090 24 GB24 GBRunsQ6_K22.5 GB
RX 7900 XTX 24 GB24 GBRunsQ6_K22.5 GB
Arc Pro B70 32 GB32 GBRuns wellQ8_027.4 GB
RTX 5090 32 GB32 GBRuns wellFP826.8 GB
Laptop GPUs
RTX 3050 Laptop 4 GB4 GBNot practicalQ3_K_M14.7 GB
RTX 3050 Ti Laptop 4 GB4 GBNot practicalQ3_K_M14.7 GB
RTX 2060 Laptop 6 GB6 GBOffload onlyQ4_K_M17.4 GB
RTX 3050 Laptop 6 GB6 GBOffload onlyQ4_K_M17.4 GB
RTX 3060 Laptop 6 GB6 GBOffload onlyQ4_K_M17.4 GB
RTX 4050 Laptop 6 GB6 GBOffload onlyQ4_K_M17.4 GB
RTX 2070 Laptop 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 2070 Super Laptop 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 2080 Laptop 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 2080 Super Laptop 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 3070 Laptop 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 3070 Ti Laptop 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 3080 Laptop 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 4060 Laptop 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 4070 Laptop 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 5050 Laptop 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 5060 Laptop 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 5070 Laptop 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RX 7600M 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RX 7600M XT 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RX 7600S 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RX 7700S 8 GB8 GBOffload onlyQ4_K_M17.4 GB
RTX 4080 Laptop 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RTX 5070 Laptop 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RTX 5070 Ti Laptop 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RX 7800M 12 GB12 GBOffload onlyQ4_K_M17.4 GB
RTX 3080 Laptop 16 GB16 GBTightQ3_K_M14.7 GB
RTX 3080 Ti Laptop 16 GB16 GBTightQ3_K_M14.7 GB
RTX 4090 Laptop 16 GB16 GBTightQ3_K_M14.7 GB
RTX 5080 Laptop 16 GB16 GBTightQ3_K_M14.7 GB
RX 7900M 16 GB16 GBTightQ3_K_M14.7 GB
RTX 5090 Laptop 24 GB24 GBRunsQ6_K22.5 GB
Unified memory
Radeon 8060S (Strix Halo) 96 GB96 GBRuns well16-bit46.0 GB
Apple Silicon Macs (by memory)
Mac 16 GB12.7 GBOffload onlyQ4_K_M17.4 GB
Mac 18 GB14.4 GBOffload onlyQ3_K_M14.7 GB
Mac 24 GB19.6 GBRunsQ4_K_M17.4 GB
Mac 32 GB26.8 GBRunsQ6_K22.5 GB
Mac 36 GB30.2 GBRuns wellQ8_027.4 GB
Mac 48 GB40.2 GBRuns wellQ8_027.4 GB
Mac 64 GB55.7 GBRuns well16-bit46.0 GB
Mac 96 GB85 GBRuns well16-bit46.0 GB
Mac 128 GB115.4 GBRuns well16-bit46.0 GB
Mac 192 GB175.4 GBRuns well16-bit46.0 GB
Mac 256 GB236.9 GBRuns well16-bit46.0 GB
Mac 512 GB498.1 GBRuns well16-bit46.0 GB

On RTX 40/50 GPUs the FP8 file is preferred over Q8_0 when both fit (hardware FP8). All verdicts are calculated; see how the numbers work.

05Text encoder, VAE and other files

FileFolderSizeWhen
Qwen3-VL 32B NVFP4 (AWQ)
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
models/text_encoders15.7 GBtext encoder · smaller, recommendedDownload →
Qwen3-VL 32B INT8 (convrot)
qwen3vl_32b_minimax_h3_int8_convrot.safetensors
models/text_encoders27.1 GBtext encoder · alternativeDownload →
Qwen3-VL 32B BF16
qwen3vl_32b_minimax_h3_bf16.safetensors
models/text_encoders51.5 GBtext encoder · alternativeDownload →
MiniMax H3 video VAE FP16
minimax_h3_video_vae_fp16.safetensors
models/vae5.2 GBrequiredDownload →
MiniMax H3 audio VAE FP32
minimax_h3_audio_vae_fp32.safetensors
models/vae0.6 GBrequiredDownload →

The files the official ComfyUI workflows load next to the model. Sizes read from Hugging Face (2026-09-25). Every GPU page for this model lists the exact set to download for that card, with the total.

Qwen3-VL 32B: 51.5 GB as 16-bit, 27.1 GB as INT8, 15.7 GB as NVFP4. ComfyUI encodes the prompt first and can push the encoder out of VRAM before sampling, so it does not have to fit together with the model. On 16 GB the encoder itself is too big for VRAM, so let it run from system RAM: slower prompt encoding, same images.

06Where the files go in ComfyUI

FileFolderLoader node
Diffusion model (.safetensors: 16-bit, FP8, INT8)ComfyUI/models/diffusion_modelsLoad Diffusion Model
GGUF file (.gguf)ComfyUI/models/unetUnet Loader (GGUF) — from the ComfyUI-GGUF node pack
Text encoderComfyUI/models/text_encodersLoad CLIP / DualCLIPLoader (or the GGUF versions)
VAEComfyUI/models/vaeLoad VAE

Standard ComfyUI folders. After copying files, press R in ComfyUI (or restart it) to refresh the lists. Some uploads need their uploader's own loader node — see the notes above.

07AMD, Intel and NVIDIA: which file types are fast

File typeRTX 50RTX 40RTX 30 / 20RX 9000RX 7000/6000 · Strix HaloIntel Arc
16-bitRunsRunsRunsRunsRunsRuns
FP8Native FP8Native FP8No FP8 speed-upNative FP8No FP8 speed-upNo FP8 speed-up
INT8Native INT8Native INT8Native INT8Native INT8Native INT8No INT8 speed-up
GGUFRuns (GGUF node)Runs (GGUF node)Runs (GGUF node)Runs (GGUF node)Runs (GGUF node)Runs (GGUF node)

Every file type loads on every listed GPU, so the memory verdicts apply to all of them. What differs is speed: FP8 maths needs RTX 40/50 or RX 9000 (with ROCm 6.4+ and PyTorch 2.7+); ComfyUI's INT8 maths runs on NVIDIA and AMD, not on Intel; GGUF is unpacked on the fly on any GPU, which costs some speed. NVFP4 files are fast only on RTX 50. AMD runs ComfyUI on Windows through ROCm, Intel through PyTorch XPU; some custom nodes are NVIDIA-only. Source: ComfyUI model_management.py. AMD and Intel guide →

08Measured and reported results

LabelGPUSetupResultPeak VRAMDateSource
reportedRTX 4060 8 GBFL2VA pruned 20B INT8 ConvRot + Q2_K GGUF text encoder · 832x480 · 20 steps
“Denoising time: About 5.5 minutes”
Desktop RTX 4060 8GB, 96 GB RAM, WSL; 5.5 min denoise only; "Total first-run time: Just under 8 minutes"; a later 15-second run took about 15 minutes
330 s / clip (107 frames)—2026-08mountainmeadowsystems.com →
reportedRTX 4070 Laptop 8 GBpruned int8 convrot (FL2VA/Ref2VA) · 832x480
“On a laptop with an RTX 4070 (8GB VRAM) and 32GB of RAM, we generated one 832×480, 15-second clip in 45 minutes.”
15-second clip, 24 fps, with audio; ComfyUI I2V/R2V workflows (which one not stated); 45 min = 2700 s; steps not stated
2700 s / clip—2026-08-04metallab.ai →
reportedRTX 4090 24 GBminimax_h3_fl2va_pruned_int8_convrot + qwen3vl_32b nvfp4_awq · 832x480 · 8 steps
“5 s clip (832×480, 8 steps): ~7 min”
~7 min approx = 420 s for 5 s clip; VRAM ~6.6 GB during sampling, ~22.5 GB spike at model load; 15 s clip ~25-30 min; no date shown
420 s / clip22.5 GB—github.com →
reportedRTX 5090 32 GBminimax_h3_fl2va_pruned_nvfp4 · 864x480 · 10 steps
“175 s for a 864×480 ten-second clip, 26,914 MiB peak VRAM.”
ComfyUI 0.30.1, 500 W power cap; 26,914 MiB = 26.28 GiB; INT8-ConvRot = 185 s / 28,581 MiB
175 s / clip (243 frames)26.28 GB2026-08-04ai-muninn.com →

Reported results are other people's numbers, copied as published, with a link. Settings, drivers and ComfyUI versions differ, so compare them with care. Send yours.

09Training a LoRA for MiniMax H3 Pruned

No trainer documentation with a VRAM figure for this model was found yet. What the trainers say for other models →

10Every file tracked for MiniMax H3 Pruned

FileTypeSizeRepo
MiniMax-H3-FL2VA-Pruned-Q8_0.gguf (fl2va-pruned)Q8_021.58 GBAbiray/MiniMax-H3-Pruned-GGUF →
MiniMax-H3-Ref2VA-Pruned-Q8_0.gguf (ref2va-pruned)Q8_021.58 GBAbiray/MiniMax-H3-Pruned-GGUF →
MiniMax-H3-FL2VA-Pruned-Q6_K.gguf (fl2va-pruned)Q6_K16.73 GBAbiray/MiniMax-H3-Pruned-GGUF →
MiniMax-H3-Ref2VA-Pruned-Q6_K.gguf (ref2va-pruned)Q6_K16.73 GBAbiray/MiniMax-H3-Pruned-GGUF →
MiniMax-H3-FL2VA-Pruned-Q5_K_M.gguf (fl2va-pruned)Q5_K_M14.07 GBAbiray/MiniMax-H3-Pruned-GGUF →
MiniMax-H3-FL2VA-Pruned-Q5_K_S.gguf (fl2va-pruned)Q5_K_S14.07 GBAbiray/MiniMax-H3-Pruned-GGUF →
MiniMax-H3-Ref2VA-Pruned-Q5_K_M.gguf (ref2va-pruned)Q5_K_M14.07 GBAbiray/MiniMax-H3-Pruned-GGUF →
MiniMax-H3-Ref2VA-Pruned-Q5_K_S.gguf (ref2va-pruned)Q5_K_S14.07 GBAbiray/MiniMax-H3-Pruned-GGUF →
MiniMax-H3-FL2VA-Pruned-Q4_K_M.gguf (fl2va-pruned)Q4_K_M11.56 GBAbiray/MiniMax-H3-Pruned-GGUF →
MiniMax-H3-FL2VA-Pruned-Q4_K_S.gguf (fl2va-pruned)Q4_K_S11.56 GBAbiray/MiniMax-H3-Pruned-GGUF →
MiniMax-H3-Ref2VA-Pruned-Q4_K_M.gguf (ref2va-pruned)Q4_K_M11.56 GBAbiray/MiniMax-H3-Pruned-GGUF →
MiniMax-H3-Ref2VA-Pruned-Q4_K_S.gguf (ref2va-pruned)Q4_K_S11.56 GBAbiray/MiniMax-H3-Pruned-GGUF →
MiniMax-H3-FL2VA-Pruned-Q3_K_M.gguf (fl2va-pruned)Q3_K_M8.90 GBAbiray/MiniMax-H3-Pruned-GGUF →
MiniMax-H3-Ref2VA-Pruned-Q3_K_M.gguf (ref2va-pruned)Q3_K_M8.90 GBAbiray/MiniMax-H3-Pruned-GGUF →
minimax_h3_fl2va_pruned_bf16.safetensors (fl2va-pruned)BF1640.23 GBComfy-Org/MiniMax-H3 →
minimax_h3_ref2va_pruned_bf16.safetensors (ref2va-pruned)BF1640.23 GBComfy-Org/MiniMax-H3 →
minimax_h3_fl2va_pruned_int8_convrot.safetensors (fl2va-pruned)INT820.97 GBComfy-Org/MiniMax-H3 →
minimax_h3_ref2va_pruned_int8_convrot.safetensors (ref2va-pruned)INT820.97 GBComfy-Org/MiniMax-H3 →
minimax_h3_fl2va_pruned_fp8_scaled.safetensors (fl2va-pruned)FP820.96 GBComfy-Org/MiniMax-H3 →
minimax_h3_ref2va_pruned_fp8_scaled.safetensors (ref2va-pruned)FP820.96 GBComfy-Org/MiniMax-H3 →

33B (HF API 33.1B) omni transformer generating video WITH native stereo audio, 4-15 s, 768p native (2K via upscale pass). Two denoisers: fl2va (text/first/last-frame-to-video+audio) and ref2va (up to 9 images/3 videos/3 audio refs); load one. 'Pruned' variants are smaller official Comfy-Org cuts (40.2 GB bf16 vs 66.3 GB). VAEs: minimax_h3_video_vae_fp16 5207808496 + minimax_h3_audio_vae_fp32 605254808. Native ComfyUI support (day-0, Aug 2026 blog); Comfy blog says smallest variants total 42.5 GB (from 123.6 GB). unsloth/MiniMax-H3-GGUF (~1.18M downloads, pruned only, quants named Q3_K/Q4_K/UD-*) is documented for stable-diffusion.cpp, not ComfyUI - not listed.