Datasheet for local AI49 models98 GPUsData read 2026-09-25
No adsNo tracking
Check my GPU →

From the RTX 4090 24 GB to the RTX 5090 32 GB: what changes for local AI

+8 GB of VRAM (24 → 32 GB). Of 49 models, 4 newly run well, 9 get a better file or stop offloading, and 36 stay the same.

24 GB GDDR6X · 384-bit · 1008 GB/s · Ada Lovelace32 GB GDDR7 · 512-bit · 1792 GB/s · BlackwellData 2026-09-25
4newly run well
9better file
36no change
+8GB more VRAM

01Newly runs well

ModelOn the RTX 4090 24 GBOn the RTX 5090 32 GBNeeded
LTX-2 (19B)Runs Q6_KRuns well Q8_025.2 GB
LTX-2.3 (22B)Runs Q6_KRuns well Q8_027.6 GB
LTX-2.5 (22B)Runs Q6_KRuns well Q8_028.4 GB
MiniMax H3 PrunedRuns Q6_KRuns well FP826.8 GB

“Runs well” means a 16-bit or 8-bit file fits entirely in VRAM.

02A better file, or less offloading

ModelOn the RTX 4090 24 GBOn the RTX 5090 32 GBNeeded
FLUX.1 [dev]Runs well FP8Runs well 16-bit26.1 GB
FLUX.1 [schnell]Runs well FP8Runs well 16-bit26.1 GB
FLUX.1 Kontext [dev]Runs well FP8Runs well 16-bit26.4 GB
FLUX.1 Krea [dev]Runs well FP8Runs well 16-bit26.1 GB
FLUX.1 Fill [dev]Runs well Q8_0Runs well 16-bit26.4 GB
FLUX.2 [dev]Runs Q4_K_MRuns Q6_K30.7 GB
Krea 2 (Turbo)Runs well FP8Runs well 16-bit28.9 GB
HunyuanVideo (13B, original)Runs well FP8Runs well 16-bit29.9 GB
MiniMax H3 (33B)Tight Q3_K_MRuns Q5_K_M29.7 GB

04Speed and features

Memory bandwidth goes from 1008 to 1792 GB/s (×1.78). For models that fit on both cards, bandwidth and compute decide the speed; this ratio is a rough first guide, not a benchmark. It also adds FP4 (NVFP4) hardware, for models that ship NVFP4 files.