Datasheet for local AI49 models98 GPUsData read 2026-09-25
No adsNo tracking
Check my GPU →

From the RTX 3060 12 GB to the RTX 5060 Ti 16 GB: what changes for local AI

+4 GB of VRAM (12 → 16 GB). Of 49 models, 8 newly run well, 26 get a better file or stop offloading, and 15 stay the same.

12 GB GDDR6 · 192-bit · 360 GB/s · Ampere16 GB GDDR7 · 128-bit · 448 GB/s · BlackwellData 2026-09-25
8newly run well
26better file
15no change
+4GB more VRAM

01Newly runs well

ModelOn the RTX 3060 12 GBOn the RTX 5060 Ti 16 GBNeeded
FLUX.1 [dev]Runs Q5_K_SRuns well FP814.2 GB
FLUX.1 [schnell]Runs Q5_K_SRuns well FP814.2 GB
FLUX.1 Kontext [dev]Runs Q5_K_MRuns well FP814.5 GB
FLUX.1 Krea [dev]Runs Q5_K_MRuns well FP814.2 GB
FLUX.1 Fill [dev]Runs Q5_K_SRuns well Q8_015.3 GB
Krea 2 (Turbo)Runs Q5_K_MRuns well FP815.7 GB
Boogu-Image (Turbo)Runs Q5_1Runs well FP812.9 GB
HunyuanVideo 1.5Runs Q6_KRuns well FP812.6 GB

“Runs well” means a 16-bit or 8-bit file fits entirely in VRAM.

02A better file, or less offloading

ModelOn the RTX 3060 12 GBOn the RTX 5060 Ti 16 GBNeeded
Ideogram 4Offload only Q4_1Runs Q4_114.7 GB
Wan 2.1 I2V 14B 480POffload only Q3_K_MRuns Q4_K_M15.6 GB
Wan 2.1 I2V 14B 720POffload only Q4_K_MTight Q3_K_M15.4 GB
Wan 2.1 VACE 14BOffload only Q3_K_STight Q3_K_S12.6 GB
Wan 2.2 S2V 14BOffload only Q4_K_MTight Q2_K14.3 GB
SCAIL-2 (character animation)Offload only Q2_KTight Q3_K_M14.4 GB
LTX-2 (19B)Offload only Q2_KTight Q3_K_M14.9 GB
LTX-2.3 (22B)Offload only Q2_KTight Q3_K_M15.6 GB
LTX-2.5 (22B)Offload only Q2_KTight Q2_K13.6 GB
MiniMax H3 PrunedOffload only Q4_K_MTight Q3_K_M14.7 GB
FLUX.2 [dev]Offload only Q4_K_MOffload only Q2_K16.2 GB
Qwen-ImageTight Q2_KRuns Q4_K_M15.9 GB
Qwen-Image-Edit (2511)Tight Q2_KTight Q3_K_M12.9 GB
Z-Image TurboRuns well Q8_0Runs well 16-bit14.3 GB
Z-Image (base)Runs well Q8_0Runs well 16-bit14.3 GB
HiDream-I1 (Full)Tight Q3_K_MRuns Q5_K_M15.8 GB
HiDream-I1 (Dev)Tight Q3_K_MRuns Q5_K_M15.8 GB
HunyuanImage 2.1Tight Q2_KRuns Q4_K_M14.6 GB
Wan 2.1 T2V 14BTight Q3_K_MRuns Q5_K_M15.6 GB
Wan 2.2 T2V A14BTight Q3_K_MRuns Q5_K_M15.1 GB
Wan 2.2 I2V A14BTight Q3_K_MRuns Q5_K_M15.1 GB
Wan 2.2 TI2V 5BRuns well Q8_0Runs well 16-bit13.8 GB
Wan 2.2 Animate 14BTight Q2_KTight Q3_K_M13.9 GB
Wan Animate 2 (14B)Tight Q2_KTight Q3_K_M13.9 GB
HunyuanVideo (13B, original)Tight Q3_K_MRuns Q6_K15.3 GB
LTX-Video 13B (0.9.8)Tight Q3_K_MRuns Q6_K15.2 GB

04Speed and features

Memory bandwidth goes from 360 to 448 GB/s (×1.24). For models that fit on both cards, bandwidth and compute decide the speed; this ratio is a rough first guide, not a benchmark. The RTX 5060 Ti 16 GB computes FP8 natively, so FP8 files run faster there than on the RTX 3060 12 GB, which only uses them to save memory. It also adds FP4 (NVFP4) hardware, for models that ship NVFP4 files.