Datasheet for local AI49 models98 GPUsData read 2026-09-25
No adsNo tracking
Check my GPU →

RTX 4090 24 GB for local AI

Which image and video models run on the RTX 4090 24 GB, which file to download for each, and how much VRAM they need.

NVIDIA24 GB GDDR6X · 384-bit · 1008 GB/s · Ada LovelaceData 2026-09-25
VRAM24 GB
MemoryGDDR6X
Bus384-bit
Bandwidth1008 GB/s
ArchitectureAda Lovelace
Launched2022-10
Launch price$1,599
Runs well43 of 49

Specs: www.nvidia.com · launch: en.wikipedia.org

01What runs on it

ModelVerdictBest fileSizeNeeded
image models
FLUX.1 [dev]Runs wellFP811.9 GB14.2 GB
FLUX.1 [schnell]Runs wellFP811.9 GB14.2 GB
FLUX.1 Kontext [dev]Runs wellFP811.9 GB14.5 GB
FLUX.1 Krea [dev]Runs wellFP811.9 GB14.2 GB
FLUX.1 Fill [dev]Runs wellQ8_012.7 GB15.3 GB
FLUX.2 [dev]RunsQ4_K_M20.1 GB23.4 GB
FLUX.2 [klein] 9BRuns well16-bit18.2 GB20.5 GB
FLUX.2 [klein] 4BRuns well16-bit7.8 GB9.6 GB
Krea 2 (Turbo)Runs wellFP813.1 GB15.7 GB
Qwen-ImageRuns wellFP820.4 GB23.2 GB
Qwen-Image-Edit (2511)Runs wellFP820.5 GB23.5 GB
Qwen-Image 2.1Runs well16-bit14.2 GB16.5 GB
Z-Image TurboRuns well16-bit12.3 GB14.3 GB
Z-Image (base)Runs well16-bit12.3 GB14.3 GB
Ideogram 4Runs wellFP89.3 GB20.9 GB
Boogu-Image (Turbo)Runs well16-bit20.6 GB23.2 GB
ERNIE-Image (Turbo)Runs well16-bit16.1 GB18.4 GB
HiDream-O1-ImageRuns well16-bit16.4 GB19.2 GB
Mage-Flow (Microsoft)Runs well16-bit8.2 GB10.0 GB
Lumina Image 2.0Runs well16-bit5.2 GB7.0 GB
HiDream-I1 (Full)Runs wellFP817.1 GB19.9 GB
HiDream-I1 (Dev)Runs wellFP817.1 GB19.9 GB
Stable Diffusion 3.5 LargeRuns well16-bit16.5 GB18.8 GB
Stable Diffusion 3.5 MediumRuns well16-bit5.1 GB6.9 GB
Chroma1-HDRuns well16-bit17.8 GB20.1 GB
SDXL 1.0Runs well16-bit6.9 GB7.1 GB
Illustrious XL / Pony (SDXL anime)Runs well16-bit6.9 GB7.1 GB
Stable Diffusion 1.5Runs well16-bit2.1 GB3.3 GB
HunyuanImage 2.1Runs wellFP817.4 GB20.7 GB
video models
Wan 2.1 T2V 14BRuns wellFP814.3 GB18.6 GB
Wan 2.1 T2V 1.3BRuns well16-bit2.8 GB5.6 GB
Wan 2.1 I2V 14B 480PRuns wellFP816.4 GB20.7 GB
Wan 2.1 I2V 14B 720PRuns wellFP816.4 GB23.2 GB
Wan 2.1 VACE 14BRuns wellQ8_018.7 GB23.5 GB
Wan 2.2 T2V A14BRuns wellFP814.3 GB18.6 GB
Wan 2.2 I2V A14BRuns wellFP814.3 GB18.6 GB
Wan 2.2 TI2V 5BRuns well16-bit10.0 GB13.8 GB
Wan 2.2 Animate 14BRuns wellFP817.3 GB22.6 GB
Wan Animate 2 (14B)Runs wellQ8_018.1 GB23.4 GB
Wan 2.2 S2V 14BRuns wellFP816.4 GB21.2 GB
SCAIL-2 (character animation)Runs wellFP817.7 GB23.0 GB
HunyuanVideo (13B, original)Runs wellFP813.2 GB17.5 GB
LTX-Video 13B (0.9.8)Runs wellFP815.7 GB20.0 GB
HunyuanVideo 1.5Runs well16-bit16.7 GB21.0 GB
LTX-2 (19B)RunsQ6_K16.0 GB20.8 GB
LTX-2.3 (22B)RunsQ6_K17.8 GB22.6 GB
LTX-2.5 (22B)RunsQ6_K18.7 GB23.5 GB
MiniMax H3 (33B)TightQ3_K_M15.6 GB21.4 GB
MiniMax H3 PrunedRunsQ6_K16.7 GB22.5 GB

Calculated from real file sizes plus working memory. How the numbers work.

02Good to know

The RTX 4090 24 GB is an Ada Lovelace GPU with hardware FP8, so ComfyUI can compute Comfy-Org's FP8 files natively: small and fast. (Plain FP8 files use FP8 maths with the --fast fp8_matrix_mult option.)

03Measured and reported results

LabelModelSetupResultPeak VRAMDateSource
reportedFLUX.1 devfp8 (ComfyUI template)
“Prompt executed in 11.28 seconds”
ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread
11.28 s / image—2025-12-28github.com →
reportedFLUX.1 devQ8_0 GGUF · 1024x1024 · 20 steps
“15 seconds at the fastest to 17 seconds at the slowest on my RTX 4090 with Euler 20 Steps for 1024x1024 images”
stated range 15-17 s; city96 ComfyUI-GGUF Q8
15 s / image—2024-08-25github.com →
reportedFLUX.1 devfp8 (--fast) · 1024x1024 · 20 steps
“Prompt executed in 10.01 seconds”
ComfyUI 'RTX 4090 benchmarks - FLUX model' thread; OP settings 1024x1024, 20 steps; Aug 2024 ComfyUI/PyTorch 2.5 dev; FP8 with --fast, GPU at 2.52 GHz/875mV undervolt (9.07 s at 2.
10.01 s / image—2024-08-26github.com →
reportedHunyuanVideo 1.5720p model · 848x480
“The 5 seconds video took 297s to generate so barely longer than on your end”
Replicated the 5090 poster's ComfyUI workflow (720p model at 848x480, 5 s @24fps); card undervolted (~5% slower per poster)
297 s / clip—2025-12-05huggingface.co →
reportedIllustrious / PonyIllustrious-XL-v2.0 · 1024x1024
“NVIDIA RTX 4090 | 7.00it/s | 0.14s/it | CUDA 12.9 | ComfyUI (Unknown) | Windows 11 24H2”
Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section
0.14 s/it—2025-11-29huggingface.co →
reportedMiniMax H3 Prunedminimax_h3_fl2va_pruned_int8_convrot + qwen3vl_32b nvfp4_awq · 832x480 · 8 steps
“5 s clip (832×480, 8 steps): ~7 min”
~7 min approx = 420 s for 5 s clip; VRAM ~6.6 GB during sampling, ~22.5 GB spike at model load; 15 s clip ~25-30 min; no date shown
420 s / clip22.5 GB—github.com →
reportedWan 2.1 I2V 480P30 steps
“100%|███| 30/30 [06:13<00:00, 12.46s/it]”
Kijai WanVideoWrapper wanvideo_480p_I2V_example_01.json; 32GB system RAM, process later 'Killed' (RAM); resolution/frames not stated
12.46 s/it—2025-02-26github.com →
reportedWan 2.2 I2VWan2.2-I2V-A14B High/Low Q6_K GGUF · 800x448 · 8 steps
“RTX 4090なら1分半ほど、RTX 5080はほぼ2分で480p解像度を5秒出力できます。”
Chimolog GPU review; ComfyUI 0.3.5x, Kijai-based workflow, 2+2+4 steps with Lightx2v/Lightning LoRAs; '1分半ほど' = about 1.5 min (converted to 90 s, approximate); exact values only in
90 s / clip (81 frames)—2025-08-28chimolog.co →

Reported results are other people's numbers, copied as published, with a link. Settings, drivers and ComfyUI versions differ, so compare them with care. Send yours.