RTX 5060 Ti 16 GB for local AI
Which image and video models run on the RTX 5060 Ti 16 GB, which file to download for each, and how much VRAM they need.
VRAM16 GB
MemoryGDDR7
Bus128-bit
Bandwidth448 GB/s
ArchitectureBlackwell
Launched2025-04
Launch price$429
Runs well25 of 49
Specs: www.nvidia.com · launch: www.nvidia.com
01What runs on it
| Model | Verdict | Best file | Size | Needed |
|---|---|---|---|---|
| image models | ||||
| FLUX.1 [dev] | Runs well | FP8 | 11.9 GB | 14.2 GB |
| FLUX.1 [schnell] | Runs well | FP8 | 11.9 GB | 14.2 GB |
| FLUX.1 Kontext [dev] | Runs well | FP8 | 11.9 GB | 14.5 GB |
| FLUX.1 Krea [dev] | Runs well | FP8 | 11.9 GB | 14.2 GB |
| FLUX.1 Fill [dev] | Runs well | Q8_0 | 12.7 GB | 15.3 GB |
| FLUX.2 [dev] | Offload only | Q2_K | 12.9 GB | 16.2 GB |
| FLUX.2 [klein] 9B | Runs well | FP8 | 9.4 GB | 11.7 GB |
| FLUX.2 [klein] 4B | Runs well | 16-bit | 7.8 GB | 9.6 GB |
| Krea 2 (Turbo) | Runs well | FP8 | 13.1 GB | 15.7 GB |
| Qwen-Image | Runs | Q4_K_M | 13.1 GB | 15.9 GB |
| Qwen-Image-Edit (2511) | Tight | Q3_K_M | 9.9 GB | 12.9 GB |
| Qwen-Image 2.1 | Runs well | Q8_0 | 7.6 GB | 9.9 GB |
| Z-Image Turbo | Runs well | 16-bit | 12.3 GB | 14.3 GB |
| Z-Image (base) | Runs well | 16-bit | 12.3 GB | 14.3 GB |
| Ideogram 4 | Runs | Q4_1 | 6.2 GB | 14.7 GB |
| Boogu-Image (Turbo) | Runs well | FP8 | 10.3 GB | 12.9 GB |
| ERNIE-Image (Turbo) | Runs well | Q8_0 | 8.7 GB | 11.0 GB |
| HiDream-O1-Image | Runs well | FP8 | 8.1 GB | 10.9 GB |
| Mage-Flow (Microsoft) | Runs well | 16-bit | 8.2 GB | 10.0 GB |
| Lumina Image 2.0 | Runs well | 16-bit | 5.2 GB | 7.0 GB |
| HiDream-I1 (Full) | Runs | Q5_K_M | 13.0 GB | 15.8 GB |
| HiDream-I1 (Dev) | Runs | Q5_K_M | 13.0 GB | 15.8 GB |
| Stable Diffusion 3.5 Large | Runs well | Q8_0 | 8.8 GB | 11.1 GB |
| Stable Diffusion 3.5 Medium | Runs well | 16-bit | 5.1 GB | 6.9 GB |
| Chroma1-HD | Runs well | FP8 | 9.2 GB | 11.5 GB |
| SDXL 1.0 | Runs well | 16-bit | 6.9 GB | 7.1 GB |
| Illustrious XL / Pony (SDXL anime) | Runs well | 16-bit | 6.9 GB | 7.1 GB |
| Stable Diffusion 1.5 | Runs well | 16-bit | 2.1 GB | 3.3 GB |
| HunyuanImage 2.1 | Runs | Q4_K_M | 11.3 GB | 14.6 GB |
| video models | ||||
| Wan 2.1 T2V 14B | Runs | Q5_K_M | 11.3 GB | 15.6 GB |
| Wan 2.1 T2V 1.3B | Runs well | 16-bit | 2.8 GB | 5.6 GB |
| Wan 2.1 I2V 14B 480P | Runs | Q4_K_M | 11.3 GB | 15.6 GB |
| Wan 2.1 I2V 14B 720P | Tight | Q3_K_M | 8.6 GB | 15.4 GB |
| Wan 2.1 VACE 14B | Tight | Q3_K_S | 7.8 GB | 12.6 GB |
| Wan 2.2 T2V A14B | Runs | Q5_K_M | 10.8 GB | 15.1 GB |
| Wan 2.2 I2V A14B | Runs | Q5_K_M | 10.8 GB | 15.1 GB |
| Wan 2.2 TI2V 5B | Runs well | 16-bit | 10.0 GB | 13.8 GB |
| Wan 2.2 Animate 14B | Tight | Q3_K_M | 8.6 GB | 13.9 GB |
| Wan Animate 2 (14B) | Tight | Q3_K_M | 8.6 GB | 13.9 GB |
| Wan 2.2 S2V 14B | Tight | Q2_K | 9.5 GB | 14.3 GB |
| SCAIL-2 (character animation) | Tight | Q3_K_M | 9.1 GB | 14.4 GB |
| HunyuanVideo (13B, original) | Runs | Q6_K | 11.0 GB | 15.3 GB |
| LTX-Video 13B (0.9.8) | Runs | Q6_K | 10.9 GB | 15.2 GB |
| HunyuanVideo 1.5 | Runs well | FP8 | 8.3 GB | 12.6 GB |
| LTX-2 (19B) | Tight | Q3_K_M | 10.1 GB | 14.9 GB |
| LTX-2.3 (22B) | Tight | Q3_K_M | 10.8 GB | 15.6 GB |
| LTX-2.5 (22B) | Tight | Q2_K | 8.8 GB | 13.6 GB |
| MiniMax H3 (33B) | Offload only | Q4_K_M | 19.9 GB | 25.7 GB |
| MiniMax H3 Pruned | Tight | Q3_K_M | 8.9 GB | 14.7 GB |
Calculated from real file sizes plus working memory. How the numbers work.
02Good to know
The RTX 5060 Ti 16 GB is a Blackwell GPU with FP8 and FP4 hardware: ComfyUI computes Comfy-Org's FP8 files natively here, and NVFP4 files (where a model offers them) are faster still.
03Measured and reported results
| Label | Model | Setup | Result | Peak VRAM | Date | Source |
|---|---|---|---|---|---|---|
| reported | FLUX.1 dev | fp8 (ComfyUI template) “Prompt executed in 25.71 seconds” ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread; poster's log shows 16311 | 25.71 s / image | — | 2025-08-04 | github.com → |
| reported | FLUX.2 dev | flux2_dev_fp8mixed + mistral_3_small_flux2_fp8 · 1024x1024 · 20 steps “FP8 t2i 142秒程度(20Steps)” ComfyUI t2i at 1MP; approximate (程度) | 142 s / image | — | 2025-11-28 | note.com → |
| reported | FLUX.2 dev | GGUF Q4_0 (flux2-dev-q4_0.gguf) · 1024x1024 · 20 steps “GGUF Q4_0 t2i 180秒程度(20Steps)” ComfyUI t2i at 1MP; approximate (程度); author notes GGUF slower than fp8 when spilling VRAM | 180 s / image | — | 2025-11-28 | note.com → |
| reported | Illustrious / Pony | Illustrious-XL-v2.0 · 1024x1024 “NVIDIA RTX 5060 Ti | 2.60it/s | 0.39s/it | CUDA 12.9 | ComfyUI (Unknown) | Arch Linux” Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section; row says 'RTX 5060 Ti', co | 0.39 s/it | — | 2025-11-29 | huggingface.co → |
| reported | Wan 2.2 I2V | video_wan2_2_14B_i2v template (quant not stated) · 720p “RTX5060ti 16GBだと165.9秒なので、2〜3倍時間が必要ですが、落ちずに720p動画が生成できる事はたいしたものです。” Comparison figure given in the RTX 3060 post, presumably same 720p/53-frame test (not explicitly restated); steps not stated | 165.9 s / clip (53 frames) | — | 2025-11-26 | note.com → |
| reported | Z-Image Turbo | z_image_turbo_bf16.safetensors · 1328x1328 · 8 steps “約 35秒かかりました。一度モデルを VRAM にロードした後、プロンプトを変えての再実行だと約 22秒で生成できます。” ComfyUI; ~35 s first run incl. load, ~22 s warm; approximate ('約') | 22 s / image | — | 2025-11-30 | iwannacreateapps.com → |
Reported results are other people's numbers, copied as published, with a link. Settings, drivers and ComfyUI versions differ, so compare them with care. Send yours.