Datasheet for local AI49 models98 GPUsData read 2026-09-25
No adsNo tracking
Check my GPU →

From the RTX 3080 10 GB to the RTX 5070 Ti 16 GB: what changes for local AI

+6 GB of VRAM (10 → 16 GB). Of 49 models, 13 newly run well, 27 get a better file or stop offloading, and 9 stay the same.

10 GB GDDR6X · 320-bit · 760 GB/s · Ampere16 GB GDDR7 · 256-bit · 896 GB/s · BlackwellData 2026-09-25
13newly run well
27better file
9no change
+6GB more VRAM

01Newly runs well

ModelOn the RTX 3080 10 GBOn the RTX 5070 Ti 16 GBNeeded
FLUX.1 [dev]Runs Q4_K_SRuns well FP814.2 GB
FLUX.1 [schnell]Runs Q4_K_SRuns well FP814.2 GB
FLUX.1 Kontext [dev]Runs Q4_K_MRuns well FP814.5 GB
FLUX.1 Krea [dev]Runs Q4_K_MRuns well FP814.2 GB
FLUX.1 Fill [dev]Runs Q4_K_SRuns well Q8_015.3 GB
FLUX.2 [klein] 9BRuns Q5_K_MRuns well FP811.7 GB
Krea 2 (Turbo)Tight Q3_K_MRuns well FP815.7 GB
Boogu-Image (Turbo)Offload only Q5_1Runs well FP812.9 GB
ERNIE-Image (Turbo)Runs Q6_KRuns well Q8_011.0 GB
HiDream-O1-ImageOffload only FP8Runs well FP810.9 GB
Stable Diffusion 3.5 LargeRuns Q5_1Runs well Q8_011.1 GB
Chroma1-HDRuns Q6_KRuns well FP811.5 GB
HunyuanVideo 1.5Runs Q4_K_MRuns well FP812.6 GB

“Runs well” means a 16-bit or 8-bit file fits entirely in VRAM.

02A better file, or less offloading

ModelOn the RTX 3080 10 GBOn the RTX 5070 Ti 16 GBNeeded
Qwen-Image-Edit (2511)Offload only Q2_KTight Q3_K_M12.9 GB
Ideogram 4Offload only Q4_1Runs Q4_114.7 GB
HunyuanImage 2.1Offload only Q2_KRuns Q4_K_M14.6 GB
Wan 2.1 T2V 14BOffload only Q3_K_MRuns Q5_K_M15.6 GB
Wan 2.1 I2V 14B 480POffload only Q4_K_MRuns Q4_K_M15.6 GB
Wan 2.1 I2V 14B 720POffload only Q4_K_MTight Q3_K_M15.4 GB
Wan 2.1 VACE 14BOffload only Q4_K_MTight Q3_K_S12.6 GB
Wan 2.2 Animate 14BOffload only Q2_KTight Q3_K_M13.9 GB
Wan Animate 2 (14B)Offload only Q2_KTight Q3_K_M13.9 GB
Wan 2.2 S2V 14BOffload only Q4_K_MTight Q2_K14.3 GB
SCAIL-2 (character animation)Offload only Q4_K_MTight Q3_K_M14.4 GB
HunyuanVideo (13B, original)Offload only Q3_K_MRuns Q6_K15.3 GB
LTX-2 (19B)Offload only Q4_K_MTight Q3_K_M14.9 GB
LTX-2.3 (22B)Offload only Q4_K_MTight Q3_K_M15.6 GB
LTX-2.5 (22B)Offload only Q4_K_MTight Q2_K13.6 GB
MiniMax H3 PrunedOffload only Q4_K_MTight Q3_K_M14.7 GB
FLUX.2 [dev]Not practical Q2_KOffload only Q2_K16.2 GB
Qwen-ImageTight Q2_KRuns Q4_K_M15.9 GB
Z-Image TurboRuns well Q8_0Runs well 16-bit14.3 GB
Z-Image (base)Runs well Q8_0Runs well 16-bit14.3 GB
Mage-Flow (Microsoft)Runs well INT8Runs well 16-bit10.0 GB
HiDream-I1 (Full)Tight Q2_KRuns Q5_K_M15.8 GB
HiDream-I1 (Dev)Tight Q2_KRuns Q5_K_M15.8 GB
Wan 2.2 T2V A14BTight Q2_KRuns Q5_K_M15.1 GB
Wan 2.2 I2V A14BTight Q2_KRuns Q5_K_M15.1 GB
Wan 2.2 TI2V 5BRuns well Q8_0Runs well 16-bit13.8 GB
LTX-Video 13B (0.9.8)Tight Q2_KRuns Q6_K15.2 GB

04Speed and features

Memory bandwidth goes from 760 to 896 GB/s (×1.18). For models that fit on both cards, bandwidth and compute decide the speed; this ratio is a rough first guide, not a benchmark. The RTX 5070 Ti 16 GB computes FP8 natively, so FP8 files run faster there than on the RTX 3080 10 GB, which only uses them to save memory. It also adds FP4 (NVFP4) hardware, for models that ship NVFP4 files.