From the RX 7900 XTX 24 GB to the RTX 5090 32 GB: what changes for local AI
+8 GB of VRAM (24 → 32 GB). Of 49 models, 4 newly run well, 9 get a better file or stop offloading, and 36 stay the same.
01Newly runs well
| Model | On the RX 7900 XTX 24 GB | On the RTX 5090 32 GB | Needed |
|---|---|---|---|
| LTX-2 (19B) | Runs Q6_K | Runs well Q8_0 | 25.2 GB |
| LTX-2.3 (22B) | Runs Q6_K | Runs well Q8_0 | 27.6 GB |
| LTX-2.5 (22B) | Runs Q6_K | Runs well Q8_0 | 28.4 GB |
| MiniMax H3 Pruned | Runs Q6_K | Runs well FP8 | 26.8 GB |
“Runs well” means a 16-bit or 8-bit file fits entirely in VRAM.
02A better file, or less offloading
| Model | On the RX 7900 XTX 24 GB | On the RTX 5090 32 GB | Needed |
|---|---|---|---|
| FLUX.1 [dev] | Runs well Q8_0 | Runs well 16-bit | 26.1 GB |
| FLUX.1 [schnell] | Runs well Q8_0 | Runs well 16-bit | 26.1 GB |
| FLUX.1 Kontext [dev] | Runs well Q8_0 | Runs well 16-bit | 26.4 GB |
| FLUX.1 Krea [dev] | Runs well Q8_0 | Runs well 16-bit | 26.1 GB |
| FLUX.1 Fill [dev] | Runs well Q8_0 | Runs well 16-bit | 26.4 GB |
| FLUX.2 [dev] | Runs Q4_K_M | Runs Q6_K | 30.7 GB |
| Krea 2 (Turbo) | Runs well Q8_0 | Runs well 16-bit | 28.9 GB |
| HunyuanVideo (13B, original) | Runs well Q8_0 | Runs well 16-bit | 29.9 GB |
| MiniMax H3 (33B) | Tight Q3_K_M | Runs Q5_K_M | 29.7 GB |
03No change in what fits
Same verdict and same file on both cards. Speed can still differ.
04Speed and features
Memory bandwidth goes from 960 to 1792 GB/s (×1.87). For models that fit on both cards, bandwidth and compute decide the speed; this ratio is a rough first guide, not a benchmark. The RTX 5090 32 GB computes FP8 natively, so FP8 files run faster there than on the RX 7900 XTX 24 GB, which only uses them to save memory. It also adds FP4 (NVFP4) hardware, for models that ship NVFP4 files. Switching from AMD to NVIDIA also changes the software route (CUDA, ROCm or XPU) and which custom nodes work; see the AMD and Intel guide.