From the RTX 3060 12 GB to the RTX 4070 Ti Super 16 GB: what changes for local AI
+4 GB of VRAM (12 → 16 GB). Of 49 models, 8 newly run well, 26 get a better file or stop offloading, and 15 stay the same.
8newly run well
26better file
15no change
+4GB more VRAM
01Newly runs well
| Model | On the RTX 3060 12 GB | On the RTX 4070 Ti Super 16 GB | Needed |
|---|---|---|---|
| FLUX.1 [dev] | Runs Q5_K_S | Runs well FP8 | 14.2 GB |
| FLUX.1 [schnell] | Runs Q5_K_S | Runs well FP8 | 14.2 GB |
| FLUX.1 Kontext [dev] | Runs Q5_K_M | Runs well FP8 | 14.5 GB |
| FLUX.1 Krea [dev] | Runs Q5_K_M | Runs well FP8 | 14.2 GB |
| FLUX.1 Fill [dev] | Runs Q5_K_S | Runs well Q8_0 | 15.3 GB |
| Krea 2 (Turbo) | Runs Q5_K_M | Runs well FP8 | 15.7 GB |
| Boogu-Image (Turbo) | Runs Q5_1 | Runs well FP8 | 12.9 GB |
| HunyuanVideo 1.5 | Runs Q6_K | Runs well FP8 | 12.6 GB |
“Runs well” means a 16-bit or 8-bit file fits entirely in VRAM.
02A better file, or less offloading
| Model | On the RTX 3060 12 GB | On the RTX 4070 Ti Super 16 GB | Needed |
|---|---|---|---|
| Ideogram 4 | Offload only Q4_1 | Runs Q4_1 | 14.7 GB |
| Wan 2.1 I2V 14B 480P | Offload only Q3_K_M | Runs Q4_K_M | 15.6 GB |
| Wan 2.1 I2V 14B 720P | Offload only Q4_K_M | Tight Q3_K_M | 15.4 GB |
| Wan 2.1 VACE 14B | Offload only Q3_K_S | Tight Q3_K_S | 12.6 GB |
| Wan 2.2 S2V 14B | Offload only Q4_K_M | Tight Q2_K | 14.3 GB |
| SCAIL-2 (character animation) | Offload only Q2_K | Tight Q3_K_M | 14.4 GB |
| LTX-2 (19B) | Offload only Q2_K | Tight Q3_K_M | 14.9 GB |
| LTX-2.3 (22B) | Offload only Q2_K | Tight Q3_K_M | 15.6 GB |
| LTX-2.5 (22B) | Offload only Q2_K | Tight Q2_K | 13.6 GB |
| MiniMax H3 Pruned | Offload only Q4_K_M | Tight Q3_K_M | 14.7 GB |
| FLUX.2 [dev] | Offload only Q4_K_M | Offload only Q2_K | 16.2 GB |
| Qwen-Image | Tight Q2_K | Runs Q4_K_M | 15.9 GB |
| Qwen-Image-Edit (2511) | Tight Q2_K | Tight Q3_K_M | 12.9 GB |
| Z-Image Turbo | Runs well Q8_0 | Runs well 16-bit | 14.3 GB |
| Z-Image (base) | Runs well Q8_0 | Runs well 16-bit | 14.3 GB |
| HiDream-I1 (Full) | Tight Q3_K_M | Runs Q5_K_M | 15.8 GB |
| HiDream-I1 (Dev) | Tight Q3_K_M | Runs Q5_K_M | 15.8 GB |
| HunyuanImage 2.1 | Tight Q2_K | Runs Q4_K_M | 14.6 GB |
| Wan 2.1 T2V 14B | Tight Q3_K_M | Runs Q5_K_M | 15.6 GB |
| Wan 2.2 T2V A14B | Tight Q3_K_M | Runs Q5_K_M | 15.1 GB |
| Wan 2.2 I2V A14B | Tight Q3_K_M | Runs Q5_K_M | 15.1 GB |
| Wan 2.2 TI2V 5B | Runs well Q8_0 | Runs well 16-bit | 13.8 GB |
| Wan 2.2 Animate 14B | Tight Q2_K | Tight Q3_K_M | 13.9 GB |
| Wan Animate 2 (14B) | Tight Q2_K | Tight Q3_K_M | 13.9 GB |
| HunyuanVideo (13B, original) | Tight Q3_K_M | Runs Q6_K | 15.3 GB |
| LTX-Video 13B (0.9.8) | Tight Q3_K_M | Runs Q6_K | 15.2 GB |
03No change in what fits
FLUX.2 klein 9B Runs wellFLUX.2 klein 4B Runs wellQwen-Image 2.1 Runs wellERNIE-Image Runs wellHiDream-O1 Runs wellMage-Flow Runs wellLumina 2.0 Runs wellSD 3.5 Large Runs wellSD 3.5 Medium Runs wellChroma1-HD Runs wellSDXL Runs wellIllustrious / Pony Runs wellSD 1.5 Runs wellWan 2.1 1.3B Runs wellMiniMax H3 Offload only
Same verdict and same file on both cards. Speed can still differ.
04Speed and features
Memory bandwidth goes from 360 to 672 GB/s (×1.87). For models that fit on both cards, bandwidth and compute decide the speed; this ratio is a rough first guide, not a benchmark. The RTX 4070 Ti Super 16 GB computes FP8 natively, so FP8 files run faster there than on the RTX 3060 12 GB, which only uses them to save memory.