Datasheet for local AI49 models98 GPUsData read 2026-09-25
No adsNo tracking
Check my GPU →

From the RTX 4070 Super 12 GB to the RTX 3090 24 GB: what changes for local AI

+12 GB of VRAM (12 → 24 GB). Of 49 models, 26 newly run well, 15 get a better file or stop offloading, and 8 stay the same.

12 GB GDDR6X · 192-bit · 504 GB/s · Ada Lovelace24 GB GDDR6X · 384-bit · 936 GB/s · AmpereData 2026-09-25
26newly run well
15better file
8no change
+12GB more VRAM

01Newly runs well

ModelOn the RTX 4070 Super 12 GBOn the RTX 3090 24 GBNeeded
FLUX.1 [dev]Runs Q5_K_SRuns well Q8_015.0 GB
FLUX.1 [schnell]Runs Q5_K_SRuns well Q8_015.0 GB
FLUX.1 Kontext [dev]Runs Q5_K_MRuns well Q8_015.3 GB
FLUX.1 Krea [dev]Runs Q5_K_MRuns well Q8_015.0 GB
FLUX.1 Fill [dev]Runs Q5_K_SRuns well Q8_015.3 GB
Krea 2 (Turbo)Runs Q5_K_MRuns well Q8_016.3 GB
Qwen-ImageTight Q2_KRuns well FP823.2 GB
Qwen-Image-Edit (2511)Tight Q2_KRuns well FP823.5 GB
Ideogram 4Offload only Q4_1Runs well Q8_022.6 GB
Boogu-Image (Turbo)Runs Q5_1Runs well 16-bit23.2 GB
HiDream-I1 (Full)Tight Q3_K_MRuns well Q8_021.5 GB
HiDream-I1 (Dev)Tight Q3_K_MRuns well Q8_021.5 GB
HunyuanImage 2.1Tight Q2_KRuns well Q8_023.1 GB
Wan 2.1 T2V 14BTight Q3_K_MRuns well Q8_020.2 GB
Wan 2.1 I2V 14B 480POffload only Q3_K_MRuns well Q8_022.4 GB
Wan 2.1 I2V 14B 720POffload only Q4_K_MRuns well FP823.2 GB
Wan 2.1 VACE 14BOffload only Q3_K_SRuns well Q8_023.5 GB
Wan 2.2 T2V A14BTight Q3_K_MRuns well Q8_019.7 GB
Wan 2.2 I2V A14BTight Q3_K_MRuns well Q8_019.7 GB
Wan 2.2 Animate 14BTight Q2_KRuns well FP822.6 GB
Wan Animate 2 (14B)Tight Q2_KRuns well Q8_023.4 GB
Wan 2.2 S2V 14BOffload only Q4_K_MRuns well FP821.2 GB
SCAIL-2 (character animation)Offload only Q2_KRuns well Q8_023.4 GB
HunyuanVideo (13B, original)Tight Q3_K_MRuns well Q8_018.3 GB
LTX-Video 13B (0.9.8)Tight Q3_K_MRuns well Q8_018.3 GB
HunyuanVideo 1.5Runs Q6_KRuns well 16-bit21.0 GB

“Runs well” means a 16-bit or 8-bit file fits entirely in VRAM.

02A better file, or less offloading

ModelOn the RTX 4070 Super 12 GBOn the RTX 3090 24 GBNeeded
FLUX.2 [dev]Offload only Q4_K_MRuns Q4_K_M23.4 GB
LTX-2 (19B)Offload only Q2_KRuns Q6_K20.8 GB
LTX-2.3 (22B)Offload only Q2_KRuns Q6_K22.6 GB
LTX-2.5 (22B)Offload only Q2_KRuns Q6_K23.5 GB
MiniMax H3 (33B)Offload only Q4_K_MTight Q3_K_M21.4 GB
MiniMax H3 PrunedOffload only Q4_K_MRuns Q6_K22.5 GB
FLUX.2 [klein] 9BRuns well FP8Runs well 16-bit20.5 GB
Qwen-Image 2.1Runs well Q8_0Runs well 16-bit16.5 GB
Z-Image TurboRuns well Q8_0Runs well 16-bit14.3 GB
Z-Image (base)Runs well Q8_0Runs well 16-bit14.3 GB
ERNIE-Image (Turbo)Runs well Q8_0Runs well 16-bit18.4 GB
HiDream-O1-ImageRuns well FP8Runs well 16-bit19.2 GB
Stable Diffusion 3.5 LargeRuns well Q8_0Runs well 16-bit18.8 GB
Chroma1-HDRuns well FP8Runs well 16-bit20.1 GB
Wan 2.2 TI2V 5BRuns well Q8_0Runs well 16-bit13.8 GB

04Speed and features

Memory bandwidth goes from 504 to 936 GB/s (×1.86). For models that fit on both cards, bandwidth and compute decide the speed; this ratio is a rough first guide, not a benchmark.