From the RTX 4080 Super 16 GB to the RTX 5090 32 GB: what changes for local AI
+16 GB of VRAM (16 → 32 GB). Of 49 models, 22 newly run well, 16 get a better file or stop offloading, and 11 stay the same.
22newly run well
16better file
11no change
+16GB more VRAM
01Newly runs well
| Model | On the RTX 4080 Super 16 GB | On the RTX 5090 32 GB | Needed |
|---|---|---|---|
| Qwen-Image | Runs Q4_K_M | Runs well FP8 | 23.2 GB |
| Qwen-Image-Edit (2511) | Tight Q3_K_M | Runs well FP8 | 23.5 GB |
| Ideogram 4 | Runs Q4_1 | Runs well FP8 | 20.9 GB |
| HiDream-I1 (Full) | Runs Q5_K_M | Runs well FP8 | 19.9 GB |
| HiDream-I1 (Dev) | Runs Q5_K_M | Runs well FP8 | 19.9 GB |
| HunyuanImage 2.1 | Runs Q4_K_M | Runs well FP8 | 20.7 GB |
| Wan 2.1 T2V 14B | Runs Q5_K_M | Runs well FP8 | 18.6 GB |
| Wan 2.1 I2V 14B 480P | Runs Q4_K_M | Runs well FP8 | 20.7 GB |
| Wan 2.1 I2V 14B 720P | Tight Q3_K_M | Runs well FP8 | 23.2 GB |
| Wan 2.1 VACE 14B | Tight Q3_K_S | Runs well Q8_0 | 23.5 GB |
| Wan 2.2 T2V A14B | Runs Q5_K_M | Runs well FP8 | 18.6 GB |
| Wan 2.2 I2V A14B | Runs Q5_K_M | Runs well FP8 | 18.6 GB |
| Wan 2.2 Animate 14B | Tight Q3_K_M | Runs well FP8 | 22.6 GB |
| Wan Animate 2 (14B) | Tight Q3_K_M | Runs well Q8_0 | 23.4 GB |
| Wan 2.2 S2V 14B | Tight Q2_K | Runs well FP8 | 21.2 GB |
| SCAIL-2 (character animation) | Tight Q3_K_M | Runs well FP8 | 23.0 GB |
| HunyuanVideo (13B, original) | Runs Q6_K | Runs well 16-bit | 29.9 GB |
| LTX-Video 13B (0.9.8) | Runs Q6_K | Runs well FP8 | 20.0 GB |
| LTX-2 (19B) | Tight Q3_K_M | Runs well Q8_0 | 25.2 GB |
| LTX-2.3 (22B) | Tight Q3_K_M | Runs well Q8_0 | 27.6 GB |
| LTX-2.5 (22B) | Tight Q2_K | Runs well Q8_0 | 28.4 GB |
| MiniMax H3 Pruned | Tight Q3_K_M | Runs well FP8 | 26.8 GB |
“Runs well” means a 16-bit or 8-bit file fits entirely in VRAM.
02A better file, or less offloading
| Model | On the RTX 4080 Super 16 GB | On the RTX 5090 32 GB | Needed |
|---|---|---|---|
| FLUX.2 [dev] | Offload only Q2_K | Runs Q6_K | 30.7 GB |
| MiniMax H3 (33B) | Offload only Q4_K_M | Runs Q5_K_M | 29.7 GB |
| FLUX.1 [dev] | Runs well FP8 | Runs well 16-bit | 26.1 GB |
| FLUX.1 [schnell] | Runs well FP8 | Runs well 16-bit | 26.1 GB |
| FLUX.1 Kontext [dev] | Runs well FP8 | Runs well 16-bit | 26.4 GB |
| FLUX.1 Krea [dev] | Runs well FP8 | Runs well 16-bit | 26.1 GB |
| FLUX.1 Fill [dev] | Runs well Q8_0 | Runs well 16-bit | 26.4 GB |
| FLUX.2 [klein] 9B | Runs well FP8 | Runs well 16-bit | 20.5 GB |
| Krea 2 (Turbo) | Runs well FP8 | Runs well 16-bit | 28.9 GB |
| Qwen-Image 2.1 | Runs well Q8_0 | Runs well 16-bit | 16.5 GB |
| Boogu-Image (Turbo) | Runs well FP8 | Runs well 16-bit | 23.2 GB |
| ERNIE-Image (Turbo) | Runs well Q8_0 | Runs well 16-bit | 18.4 GB |
| HiDream-O1-Image | Runs well FP8 | Runs well 16-bit | 19.2 GB |
| Stable Diffusion 3.5 Large | Runs well Q8_0 | Runs well 16-bit | 18.8 GB |
| Chroma1-HD | Runs well FP8 | Runs well 16-bit | 20.1 GB |
| HunyuanVideo 1.5 | Runs well FP8 | Runs well 16-bit | 21.0 GB |
03No change in what fits
FLUX.2 klein 4B Runs wellZ-Image Turbo Runs wellZ-Image Runs wellMage-Flow Runs wellLumina 2.0 Runs wellSD 3.5 Medium Runs wellSDXL Runs wellIllustrious / Pony Runs wellSD 1.5 Runs wellWan 2.1 1.3B Runs wellWan 2.2 5B Runs well
Same verdict and same file on both cards. Speed can still differ.
04Speed and features
Memory bandwidth goes from 736 to 1792 GB/s (×2.43). For models that fit on both cards, bandwidth and compute decide the speed; this ratio is a rough first guide, not a benchmark. It also adds FP4 (NVFP4) hardware, for models that ship NVFP4 files.