Datasheet for local AI49 models98 GPUsData read 2026-09-25
No adsNo tracking
Check my GPU →

What AI models run on 16 GB of VRAM?

25 of 49 image and video models run at 8-bit quality or better on 16 GB. Here is every one, with the file to download.

01Every model on 16 GB

ModelVerdictBest fileSizeNeeded
image models
FLUX.1 [dev]Runs wellQ8_012.7 GB15.0 GB
FLUX.1 [schnell]Runs wellQ8_012.7 GB15.0 GB
FLUX.1 Kontext [dev]Runs wellQ8_012.7 GB15.3 GB
FLUX.1 Krea [dev]Runs wellQ8_012.7 GB15.0 GB
FLUX.1 Fill [dev]Runs wellQ8_012.7 GB15.3 GB
FLUX.2 [dev]Offload onlyQ2_K12.9 GB16.2 GB
FLUX.2 [klein] 9BRuns wellQ8_010.0 GB12.3 GB
FLUX.2 [klein] 4BRuns well16-bit7.8 GB9.6 GB
Krea 2 (Turbo)Runs wellFP813.1 GB15.7 GB
Qwen-ImageRunsQ4_K_M13.1 GB15.9 GB
Qwen-Image-Edit (2511)TightQ3_K_M9.9 GB12.9 GB
Qwen-Image 2.1Runs wellQ8_07.6 GB9.9 GB
Z-Image TurboRuns well16-bit12.3 GB14.3 GB
Z-Image (base)Runs well16-bit12.3 GB14.3 GB
Ideogram 4RunsQ4_16.2 GB14.7 GB
Boogu-Image (Turbo)Runs wellQ8_011.6 GB14.2 GB
ERNIE-Image (Turbo)Runs wellQ8_08.7 GB11.0 GB
HiDream-O1-ImageRuns wellFP88.1 GB10.9 GB
Mage-Flow (Microsoft)Runs well16-bit8.2 GB10.0 GB
Lumina Image 2.0Runs well16-bit5.2 GB7.0 GB
HiDream-I1 (Full)RunsQ5_K_M13.0 GB15.8 GB
HiDream-I1 (Dev)RunsQ5_K_M13.0 GB15.8 GB
Stable Diffusion 3.5 LargeRuns wellQ8_08.8 GB11.1 GB
Stable Diffusion 3.5 MediumRuns well16-bit5.1 GB6.9 GB
Chroma1-HDRuns wellQ8_09.7 GB12.0 GB
SDXL 1.0Runs well16-bit6.9 GB7.1 GB
Illustrious XL / Pony (SDXL anime)Runs well16-bit6.9 GB7.1 GB
Stable Diffusion 1.5Runs well16-bit2.1 GB3.3 GB
HunyuanImage 2.1RunsQ4_K_M11.3 GB14.6 GB
video models
Wan 2.1 T2V 14BRunsQ5_K_M11.3 GB15.6 GB
Wan 2.1 T2V 1.3BRuns well16-bit2.8 GB5.6 GB
Wan 2.1 I2V 14B 480PRunsQ4_K_M11.3 GB15.6 GB
Wan 2.1 I2V 14B 720PTightQ3_K_M8.6 GB15.4 GB
Wan 2.1 VACE 14BTightQ3_K_S7.8 GB12.6 GB
Wan 2.2 T2V A14BRunsQ5_K_M10.8 GB15.1 GB
Wan 2.2 I2V A14BRunsQ5_K_M10.8 GB15.1 GB
Wan 2.2 TI2V 5BRuns well16-bit10.0 GB13.8 GB
Wan 2.2 Animate 14BTightQ3_K_M8.6 GB13.9 GB
Wan Animate 2 (14B)TightQ3_K_M8.6 GB13.9 GB
Wan 2.2 S2V 14BTightQ2_K9.5 GB14.3 GB
SCAIL-2 (character animation)TightQ3_K_M9.1 GB14.4 GB
HunyuanVideo (13B, original)RunsQ6_K11.0 GB15.3 GB
LTX-Video 13B (0.9.8)RunsQ6_K10.9 GB15.2 GB
HunyuanVideo 1.5Runs wellQ8_09.0 GB13.3 GB
LTX-2 (19B)TightQ3_K_M10.1 GB14.9 GB
LTX-2.3 (22B)TightQ3_K_M10.8 GB15.6 GB
LTX-2.5 (22B)TightQ2_K8.8 GB13.6 GB
MiniMax H3 (33B)Offload onlyQ4_K_M19.9 GB25.7 GB
MiniMax H3 PrunedTightQ3_K_M8.9 GB14.7 GB

Assumes a GPU without FP8 compute (Q8_0 preferred for 8-bit). On RTX 40/50, FP8 files are used where available. How the numbers work.