Datasheet for local AI49 models98 GPUsData read 2026-09-25
No adsNo tracking
Check my GPU →

Qwen-Image VRAM requirements

Alibaba's 20B text-to-image model, known for rendering long text and posters well. Big: the 8-bit file alone is over 20 GB.

Released 2025-08Licence: Apache-2.0Steps: 20–50Text encoder: Qwen2.5-VL 7B (9.4 GB as FP8)
TypeImage
Parameters20B
Fits entirely from10 GB
8-bit or better from24 GB

01The files, and how much VRAM each needs

FileSizeNeededMin. VRAMQualitySource
16-bit40.9 GB43.7 GB44 GBthe original weightsComfy-Org/Qwen-Image_ComfyUI →
Q8_021.8 GB24.6 GB25 GBpractically identical to the originalcity96/Qwen-Image-gguf →
FP820.4 GB23.2 GB24 GBpractically identical to the originalComfy-Org/Qwen-Image_ComfyUI →
Q6_K16.8 GB19.6 GB20 GBvery close to the originalcity96/Qwen-Image-gguf →
Q5_K_M14.9 GB17.7 GB18 GBclose; small differences in fine detailcity96/Qwen-Image-gguf →
Q4_K_M13.1 GB15.9 GB16 GBgood; some loss in fine detail and textcity96/Qwen-Image-gguf →
Q3_K_M9.7 GB12.5 GB13 GBnoticeable loss of detailcity96/Qwen-Image-gguf →
Q2_K7.1 GB9.9 GB10 GBheavy loss; a last resortcity96/Qwen-Image-gguf →

“Needed” = file + 2 GB working memory + 0.8 GB system reserve.

03Best GPU for Qwen-Image

The cheapest cards (by launch price) that run it well, and every card sorted by memory: best GPU for Qwen-Image → Planning bigger images or longer clips? Open the calculator →

04By graphics card

GPUVRAMVerdictBest fileNeeded
Desktop graphics cards
RTX 2060 6 GB6 GBNot practicalQ2_K9.9 GB
RTX 3050 6 GB6 GBNot practicalQ2_K9.9 GB
RTX 2070 Super 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 2080 Super 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 3050 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 3060 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 3060 Ti 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 3070 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 3070 Ti 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 4060 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 4060 Ti 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 5050 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 5060 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 5060 Ti 8 GB8 GBOffload onlyQ2_K9.9 GB
RX 7600 8 GB8 GBOffload onlyQ2_K9.9 GB
RX 9050 8 GB8 GBOffload onlyQ2_K9.9 GB
RX 9060 XT 8 GB8 GBOffload onlyQ2_K9.9 GB
Arc B570 10 GB10 GBTightQ2_K9.9 GB
RTX 3080 10 GB10 GBTightQ2_K9.9 GB
RTX 2080 Ti 11 GB11 GBTightQ2_K9.9 GB
Arc B580 12 GB12 GBTightQ2_K9.9 GB
RTX 2060 12 GB12 GBTightQ2_K9.9 GB
RTX 3060 12 GB12 GBTightQ2_K9.9 GB
RTX 3080 12 GB12 GBTightQ2_K9.9 GB
RTX 3080 Ti 12 GB12 GBTightQ2_K9.9 GB
RTX 4070 12 GB12 GBTightQ2_K9.9 GB
RTX 4070 Super 12 GB12 GBTightQ2_K9.9 GB
RTX 4070 Ti 12 GB12 GBTightQ2_K9.9 GB
RTX 5070 12 GB12 GBTightQ2_K9.9 GB
RX 7700 XT 12 GB12 GBTightQ2_K9.9 GB
RX 9070 GRE 12 GB12 GBTightQ2_K9.9 GB
Arc A770 16 GB16 GBRunsQ4_K_M15.9 GB
RTX 4060 Ti 16 GB16 GBRunsQ4_K_M15.9 GB
RTX 4070 Ti Super 16 GB16 GBRunsQ4_K_M15.9 GB
RTX 4080 16 GB16 GBRunsQ4_K_M15.9 GB
RTX 4080 Super 16 GB16 GBRunsQ4_K_M15.9 GB
RTX 5060 Ti 16 GB16 GBRunsQ4_K_M15.9 GB
RTX 5070 Ti 16 GB16 GBRunsQ4_K_M15.9 GB
RTX 5080 16 GB16 GBRunsQ4_K_M15.9 GB
RX 7600 XT 16 GB16 GBRunsQ4_K_M15.9 GB
RX 7800 XT 16 GB16 GBRunsQ4_K_M15.9 GB
RX 7900 GRE 16 GB16 GBRunsQ4_K_M15.9 GB
RX 9060 XT 16 GB16 GBRunsQ4_K_M15.9 GB
RX 9070 16 GB16 GBRunsQ4_K_M15.9 GB
RX 9070 XT 16 GB16 GBRunsQ4_K_M15.9 GB
RX 7900 XT 20 GB20 GBRunsQ6_K19.6 GB
Arc Pro B60 24 GB24 GBRuns wellFP823.2 GB
RTX 3090 24 GB24 GBRuns wellFP823.2 GB
RTX 3090 Ti 24 GB24 GBRuns wellFP823.2 GB
RTX 4090 24 GB24 GBRuns wellFP823.2 GB
RX 7900 XTX 24 GB24 GBRuns wellFP823.2 GB
Arc Pro B70 32 GB32 GBRuns wellQ8_024.6 GB
RTX 5090 32 GB32 GBRuns wellFP823.2 GB
Laptop GPUs
RTX 3050 Laptop 4 GB4 GBNot practicalQ2_K9.9 GB
RTX 3050 Ti Laptop 4 GB4 GBNot practicalQ2_K9.9 GB
RTX 2060 Laptop 6 GB6 GBNot practicalQ2_K9.9 GB
RTX 3050 Laptop 6 GB6 GBNot practicalQ2_K9.9 GB
RTX 3060 Laptop 6 GB6 GBNot practicalQ2_K9.9 GB
RTX 4050 Laptop 6 GB6 GBNot practicalQ2_K9.9 GB
RTX 2070 Laptop 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 2070 Super Laptop 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 2080 Laptop 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 2080 Super Laptop 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 3070 Laptop 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 3070 Ti Laptop 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 3080 Laptop 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 4060 Laptop 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 4070 Laptop 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 5050 Laptop 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 5060 Laptop 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 5070 Laptop 8 GB8 GBOffload onlyQ2_K9.9 GB
RX 7600M 8 GB8 GBOffload onlyQ2_K9.9 GB
RX 7600M XT 8 GB8 GBOffload onlyQ2_K9.9 GB
RX 7600S 8 GB8 GBOffload onlyQ2_K9.9 GB
RX 7700S 8 GB8 GBOffload onlyQ2_K9.9 GB
RTX 4080 Laptop 12 GB12 GBTightQ2_K9.9 GB
RTX 5070 Laptop 12 GB12 GBTightQ2_K9.9 GB
RTX 5070 Ti Laptop 12 GB12 GBTightQ2_K9.9 GB
RX 7800M 12 GB12 GBTightQ2_K9.9 GB
RTX 3080 Laptop 16 GB16 GBRunsQ4_K_M15.9 GB
RTX 3080 Ti Laptop 16 GB16 GBRunsQ4_K_M15.9 GB
RTX 4090 Laptop 16 GB16 GBRunsQ4_K_M15.9 GB
RTX 5080 Laptop 16 GB16 GBRunsQ4_K_M15.9 GB
RX 7900M 16 GB16 GBRunsQ4_K_M15.9 GB
RTX 5090 Laptop 24 GB24 GBRuns wellFP823.2 GB
Unified memory
Radeon 8060S (Strix Halo) 96 GB96 GBRuns well16-bit43.7 GB
Apple Silicon Macs (by memory)
Mac 16 GB12.7 GBTightQ3_K_M12.5 GB
Mac 18 GB14.4 GBTightQ3_K_M12.5 GB
Mac 24 GB19.6 GBRunsQ5_K_M17.7 GB
Mac 32 GB26.8 GBRuns wellQ8_024.6 GB
Mac 36 GB30.2 GBRuns wellQ8_024.6 GB
Mac 48 GB40.2 GBRuns wellQ8_024.6 GB
Mac 64 GB55.7 GBRuns well16-bit43.7 GB
Mac 96 GB85 GBRuns well16-bit43.7 GB
Mac 128 GB115.4 GBRuns well16-bit43.7 GB
Mac 192 GB175.4 GBRuns well16-bit43.7 GB
Mac 256 GB236.9 GBRuns well16-bit43.7 GB
Mac 512 GB498.1 GBRuns well16-bit43.7 GB

On RTX 40/50 GPUs the FP8 file is preferred over Q8_0 when both fit (hardware FP8). All verdicts are calculated; see how the numbers work.

05Text encoder, VAE and other files

FileFolderSizeWhen
Qwen2.5-VL 7B FP8
qwen_2.5_vl_7b_fp8_scaled.safetensors
models/text_encoders9.4 GBtext encoder · smaller, recommendedDownload →
Qwen2.5-VL 7B BF16
qwen_2.5_vl_7b.safetensors
models/text_encoders16.6 GBtext encoder · alternativeDownload →
Qwen-Image VAE
qwen_image_vae.safetensors
models/vae0.3 GBrequiredDownload →

The files the official ComfyUI workflows load next to the model. Sizes read from Hugging Face (2026-09-25). Every GPU page for this model lists the exact set to download for that card, with the total.

Qwen2.5-VL 7B: 16.6 GB as 16-bit, 9.4 GB as FP8. ComfyUI encodes the prompt first and can push the encoder out of VRAM before sampling, so it does not have to fit together with the model. On 16 GB the FP8 encoder fits on its own, so prompt encoding stays fast.

06Where the files go in ComfyUI

FileFolderLoader node
Diffusion model (.safetensors: 16-bit, FP8, INT8)ComfyUI/models/diffusion_modelsLoad Diffusion Model
GGUF file (.gguf)ComfyUI/models/unetUnet Loader (GGUF) — from the ComfyUI-GGUF node pack
Text encoderComfyUI/models/text_encodersLoad CLIP / DualCLIPLoader (or the GGUF versions)
VAEComfyUI/models/vaeLoad VAE

Standard ComfyUI folders. After copying files, press R in ComfyUI (or restart it) to refresh the lists. Some uploads need their uploader's own loader node — see the notes above.

07AMD, Intel and NVIDIA: which file types are fast

File typeRTX 50RTX 40RTX 30 / 20RX 9000RX 7000/6000 · Strix HaloIntel Arc
16-bitRunsRunsRunsRunsRunsRuns
FP8Native FP8Native FP8No FP8 speed-upNative FP8No FP8 speed-upNo FP8 speed-up
GGUFRuns (GGUF node)Runs (GGUF node)Runs (GGUF node)Runs (GGUF node)Runs (GGUF node)Runs (GGUF node)

Every file type loads on every listed GPU, so the memory verdicts apply to all of them. What differs is speed: FP8 maths needs RTX 40/50 or RX 9000 (with ROCm 6.4+ and PyTorch 2.7+); ComfyUI's INT8 maths runs on NVIDIA and AMD, not on Intel; GGUF is unpacked on the fly on any GPU, which costs some speed. NVFP4 files are fast only on RTX 50. AMD runs ComfyUI on Windows through ROCm, Intel through PyTorch XPU; some custom nodes are NVIDIA-only. Source: ComfyUI model_management.py. AMD and Intel guide →

08Measured and reported results

LabelGPUSetupResultPeak VRAMDateSource
reportedRadeon 8060S (Strix Halo) 96 GBQwen-Image-2512 BF16 + 4-step Lightning LoRA · 1328x1328 · 4 steps
“"workflow": "Qwen-Image-2512-BF16-4-Step-LoRA.json", ... "duration_seconds": 75.37661480903625”
kyuz0 Strix Halo ComfyUI toolbox benchmark (Ryzen AI Max, ROCm); cold run incl. model load, flags --disable-mmap --gpu-only --disable-smart-memory --cache-none; resolution from ben
75.38 s / image—2026-02-13raw.githubusercontent.com →
reportedRTX 3060 12 GB20 steps
“Qwen-ImageがRTX 3060(12GB)で動くと聞いて、早速ComfyUI版をお試し。確かに問題なく動いて、20stepでちょうど5分。”
'ちょうど5分' = exactly 5 minutes (converted to 300 s); ComfyUI version, file/resolution not stated; quote taken from search index (x.com not fetchable); date from tweet ID
300 s / image—2025-08-05x.com →

Reported results are other people's numbers, copied as published, with a link. Settings, drivers and ComfyUI versions differ, so compare them with care. Send yours.

09Training a LoRA for Qwen-Image

TrainerVRAMTypeSettings and quoteSource
musubi-tuner12 GBstated minimum1024x1024, batch 1, bf16 mixed precision, gradient checkpointing, xformers, --fp8_base --fp8_scaled + --blocks_to_swap 45; 64GB RAM recommended. Table: none 42GB, fp8 30GB, +swap16 24GB
“+ --blocks_to_swap 45|12GB”
github.com →
OneTrainer16 GBexample runOfficial preset: 512px, batch 2, transformer fp8, TE fp8, layer offload fraction 0.5 (24GB preset uses 0.1)
“#qwen LoRA 16GB.json”
github.com →
ai-toolkit24 GBexample runOfficial example config: 3-bit (uint3) base with accuracy recovery adapter, fp8 TE, cached text embeddings, low_vram, rank 16, batch 1, gradient checkpointing
“# 3bit is required for 24GB”
github.com →
diffusion-pipe24 GBstated minimumExample config: fp8 transformer, blocks_to_swap 8, rank 32, activation checkpointing, 640px dataset suggested, expandable_segments
“You will need block swapping. See the [example 24GB VRAM config]”
github.com →
SimpleTuner24 GBstated minimumint2-quanto or nf4-bnb base, batch 1, gradient checkpointing, LoRA rank 1-8, 512-768px start; 40GB+ strongly recommended
“A 24GB GPU is the absolute minimum, and even then you'll need extensive quantization and careful configuration.”
github.com →

reported Figures as stated by each trainer's own documentation or official example configs, read 2026-09-25. “Stated minimum” = the docs call it a minimum; “example run” = a config or measured run at that size. They differ a lot because of settings: an 8-bit or 4-bit base model, block swapping and lower resolution all cut memory. All models →

10Every file tracked for Qwen-Image

FileTypeSizeRepo
qwen_image_bf16.safetensors (original)BF1640.86 GBComfy-Org/Qwen-Image_ComfyUI →
qwen_image_2512_bf16.safetensors (2512)BF1640.86 GBComfy-Org/Qwen-Image_ComfyUI →
qwen_image_fp8_hq.safetensors (original)FP822.74 GBComfy-Org/Qwen-Image_ComfyUI →
qwen_image_fp8mixed.safetensors (original)FP820.53 GBComfy-Org/Qwen-Image_ComfyUI →
qwen_image_2512_fp8_e4m3fn.safetensors (2512)FP820.43 GBComfy-Org/Qwen-Image_ComfyUI →
qwen_image_fp8_e4m3fn.safetensors (original)FP820.43 GBComfy-Org/Qwen-Image_ComfyUI →
qwen_image_nvfp4.safetensors (original)NVFP419.77 GBComfy-Org/Qwen-Image_ComfyUI →
qwen-image-BF16.ggufBF1640.87 GBcity96/Qwen-Image-gguf →
qwen-image-Q8_0.ggufQ8_021.76 GBcity96/Qwen-Image-gguf →
qwen-image-Q6_K.ggufQ6_K16.82 GBcity96/Qwen-Image-gguf →
qwen-image-Q5_1.ggufQ5_115.39 GBcity96/Qwen-Image-gguf →
qwen-image-Q5_K_M.ggufQ5_K_M14.93 GBcity96/Qwen-Image-gguf →
qwen-image-Q5_0.ggufQ5_014.40 GBcity96/Qwen-Image-gguf →
qwen-image-Q5_K_S.ggufQ5_K_S14.12 GBcity96/Qwen-Image-gguf →
qwen-image-Q4_K_M.ggufQ4_K_M13.07 GBcity96/Qwen-Image-gguf →
qwen-image-Q4_1.ggufQ4_112.84 GBcity96/Qwen-Image-gguf →
qwen-image-Q4_K_S.ggufQ4_K_S12.14 GBcity96/Qwen-Image-gguf →
qwen-image-Q4_0.ggufQ4_011.85 GBcity96/Qwen-Image-gguf →
qwen-image-Q3_K_M.ggufQ3_K_M9.68 GBcity96/Qwen-Image-gguf →
qwen-image-Q3_K_S.ggufQ3_K_S8.95 GBcity96/Qwen-Image-gguf →
qwen-image-Q2_K.ggufQ2_K7.06 GBcity96/Qwen-Image-gguf →

MMDiT 20B. Text encoder Qwen2.5-VL-7B (see qwen2.5-vl-7b entry). Qwen-Image-2512 is an updated checkpoint of the same architecture (same sizes). nvfp4 needs RTX 50-series for speedup.