RTX 4090 24 GB for local AI
Which image and video models run on the RTX 4090 24 GB, which file to download for each, and how much VRAM they need.
VRAM24 GB
MemoryGDDR6X
Bus384-bit
Bandwidth1008 GB/s
ArchitectureAda Lovelace
Launched2022-10
Launch price$1,599
Runs well43 of 49
Specs: www.nvidia.com · launch: en.wikipedia.org
01What runs on it
| Model | Verdict | Best file | Size | Needed |
|---|---|---|---|---|
| image models | ||||
| FLUX.1 [dev] | Runs well | FP8 | 11.9 GB | 14.2 GB |
| FLUX.1 [schnell] | Runs well | FP8 | 11.9 GB | 14.2 GB |
| FLUX.1 Kontext [dev] | Runs well | FP8 | 11.9 GB | 14.5 GB |
| FLUX.1 Krea [dev] | Runs well | FP8 | 11.9 GB | 14.2 GB |
| FLUX.1 Fill [dev] | Runs well | Q8_0 | 12.7 GB | 15.3 GB |
| FLUX.2 [dev] | Runs | Q4_K_M | 20.1 GB | 23.4 GB |
| FLUX.2 [klein] 9B | Runs well | 16-bit | 18.2 GB | 20.5 GB |
| FLUX.2 [klein] 4B | Runs well | 16-bit | 7.8 GB | 9.6 GB |
| Krea 2 (Turbo) | Runs well | FP8 | 13.1 GB | 15.7 GB |
| Qwen-Image | Runs well | FP8 | 20.4 GB | 23.2 GB |
| Qwen-Image-Edit (2511) | Runs well | FP8 | 20.5 GB | 23.5 GB |
| Qwen-Image 2.1 | Runs well | 16-bit | 14.2 GB | 16.5 GB |
| Z-Image Turbo | Runs well | 16-bit | 12.3 GB | 14.3 GB |
| Z-Image (base) | Runs well | 16-bit | 12.3 GB | 14.3 GB |
| Ideogram 4 | Runs well | FP8 | 9.3 GB | 20.9 GB |
| Boogu-Image (Turbo) | Runs well | 16-bit | 20.6 GB | 23.2 GB |
| ERNIE-Image (Turbo) | Runs well | 16-bit | 16.1 GB | 18.4 GB |
| HiDream-O1-Image | Runs well | 16-bit | 16.4 GB | 19.2 GB |
| Mage-Flow (Microsoft) | Runs well | 16-bit | 8.2 GB | 10.0 GB |
| Lumina Image 2.0 | Runs well | 16-bit | 5.2 GB | 7.0 GB |
| HiDream-I1 (Full) | Runs well | FP8 | 17.1 GB | 19.9 GB |
| HiDream-I1 (Dev) | Runs well | FP8 | 17.1 GB | 19.9 GB |
| Stable Diffusion 3.5 Large | Runs well | 16-bit | 16.5 GB | 18.8 GB |
| Stable Diffusion 3.5 Medium | Runs well | 16-bit | 5.1 GB | 6.9 GB |
| Chroma1-HD | Runs well | 16-bit | 17.8 GB | 20.1 GB |
| SDXL 1.0 | Runs well | 16-bit | 6.9 GB | 7.1 GB |
| Illustrious XL / Pony (SDXL anime) | Runs well | 16-bit | 6.9 GB | 7.1 GB |
| Stable Diffusion 1.5 | Runs well | 16-bit | 2.1 GB | 3.3 GB |
| HunyuanImage 2.1 | Runs well | FP8 | 17.4 GB | 20.7 GB |
| video models | ||||
| Wan 2.1 T2V 14B | Runs well | FP8 | 14.3 GB | 18.6 GB |
| Wan 2.1 T2V 1.3B | Runs well | 16-bit | 2.8 GB | 5.6 GB |
| Wan 2.1 I2V 14B 480P | Runs well | FP8 | 16.4 GB | 20.7 GB |
| Wan 2.1 I2V 14B 720P | Runs well | FP8 | 16.4 GB | 23.2 GB |
| Wan 2.1 VACE 14B | Runs well | Q8_0 | 18.7 GB | 23.5 GB |
| Wan 2.2 T2V A14B | Runs well | FP8 | 14.3 GB | 18.6 GB |
| Wan 2.2 I2V A14B | Runs well | FP8 | 14.3 GB | 18.6 GB |
| Wan 2.2 TI2V 5B | Runs well | 16-bit | 10.0 GB | 13.8 GB |
| Wan 2.2 Animate 14B | Runs well | FP8 | 17.3 GB | 22.6 GB |
| Wan Animate 2 (14B) | Runs well | Q8_0 | 18.1 GB | 23.4 GB |
| Wan 2.2 S2V 14B | Runs well | FP8 | 16.4 GB | 21.2 GB |
| SCAIL-2 (character animation) | Runs well | FP8 | 17.7 GB | 23.0 GB |
| HunyuanVideo (13B, original) | Runs well | FP8 | 13.2 GB | 17.5 GB |
| LTX-Video 13B (0.9.8) | Runs well | FP8 | 15.7 GB | 20.0 GB |
| HunyuanVideo 1.5 | Runs well | 16-bit | 16.7 GB | 21.0 GB |
| LTX-2 (19B) | Runs | Q6_K | 16.0 GB | 20.8 GB |
| LTX-2.3 (22B) | Runs | Q6_K | 17.8 GB | 22.6 GB |
| LTX-2.5 (22B) | Runs | Q6_K | 18.7 GB | 23.5 GB |
| MiniMax H3 (33B) | Tight | Q3_K_M | 15.6 GB | 21.4 GB |
| MiniMax H3 Pruned | Runs | Q6_K | 16.7 GB | 22.5 GB |
Calculated from real file sizes plus working memory. How the numbers work.
02Good to know
The RTX 4090 24 GB is an Ada Lovelace GPU with hardware FP8, so ComfyUI can compute Comfy-Org's FP8 files natively: small and fast. (Plain FP8 files use FP8 maths with the --fast fp8_matrix_mult option.)
03Measured and reported results
| Label | Model | Setup | Result | Peak VRAM | Date | Source |
|---|---|---|---|---|---|---|
| reported | FLUX.1 dev | fp8 (ComfyUI template) “Prompt executed in 11.28 seconds” ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread | 11.28 s / image | — | 2025-12-28 | github.com → |
| reported | FLUX.1 dev | Q8_0 GGUF · 1024x1024 · 20 steps “15 seconds at the fastest to 17 seconds at the slowest on my RTX 4090 with Euler 20 Steps for 1024x1024 images” stated range 15-17 s; city96 ComfyUI-GGUF Q8 | 15 s / image | — | 2024-08-25 | github.com → |
| reported | FLUX.1 dev | fp8 (--fast) · 1024x1024 · 20 steps “Prompt executed in 10.01 seconds” ComfyUI 'RTX 4090 benchmarks - FLUX model' thread; OP settings 1024x1024, 20 steps; Aug 2024 ComfyUI/PyTorch 2.5 dev; FP8 with --fast, GPU at 2.52 GHz/875mV undervolt (9.07 s at 2. | 10.01 s / image | — | 2024-08-26 | github.com → |
| reported | HunyuanVideo 1.5 | 720p model · 848x480 “The 5 seconds video took 297s to generate so barely longer than on your end” Replicated the 5090 poster's ComfyUI workflow (720p model at 848x480, 5 s @24fps); card undervolted (~5% slower per poster) | 297 s / clip | — | 2025-12-05 | huggingface.co → |
| reported | Illustrious / Pony | Illustrious-XL-v2.0 · 1024x1024 “NVIDIA RTX 4090 | 7.00it/s | 0.14s/it | CUDA 12.9 | ComfyUI (Unknown) | Windows 11 24H2” Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section | 0.14 s/it | — | 2025-11-29 | huggingface.co → |
| reported | MiniMax H3 Pruned | minimax_h3_fl2va_pruned_int8_convrot + qwen3vl_32b nvfp4_awq · 832x480 · 8 steps “5 s clip (832×480, 8 steps): ~7 min” ~7 min approx = 420 s for 5 s clip; VRAM ~6.6 GB during sampling, ~22.5 GB spike at model load; 15 s clip ~25-30 min; no date shown | 420 s / clip | 22.5 GB | — | github.com → |
| reported | Wan 2.1 I2V 480P | 30 steps “100%|███| 30/30 [06:13<00:00, 12.46s/it]” Kijai WanVideoWrapper wanvideo_480p_I2V_example_01.json; 32GB system RAM, process later 'Killed' (RAM); resolution/frames not stated | 12.46 s/it | — | 2025-02-26 | github.com → |
| reported | Wan 2.2 I2V | Wan2.2-I2V-A14B High/Low Q6_K GGUF · 800x448 · 8 steps “RTX 4090なら1分半ほど、RTX 5080はほぼ2分で480p解像度を5秒出力できます。” Chimolog GPU review; ComfyUI 0.3.5x, Kijai-based workflow, 2+2+4 steps with Lightx2v/Lightning LoRAs; '1分半ほど' = about 1.5 min (converted to 90 s, approximate); exact values only in | 90 s / clip (81 frames) | — | 2025-08-28 | chimolog.co → |
Reported results are other people's numbers, copied as published, with a link. Settings, drivers and ComfyUI versions differ, so compare them with care. Send yours.