The RTX 5060 Ti 16 GB is one of the most popular cards for local AI right now: 16 GB of VRAM, FP8 and FP4 hardware, and a lower launch price than the RTX 4060 Ti 16 GB it replaced. It is also the card I use every day. Here is every model on this site, checked against it, with real timings where I have measured them.
The short answer
- 26 models run a top-quality file (16-bit or 8-bit) entirely in VRAM.
- 11 more run entirely in VRAM with a compressed GGUF file.
- 13 only fit heavily compressed or with part of the model in system RAM.
For most people the bigger limit is system RAM, not VRAM: with 32 GB of RAM, big video models like Wan 2.2 14B fill it completely. 64 GB makes life easier.
Every model
| Model | Verdict | Best file | VRAM needed | Measured |
|---|---|---|---|---|
| Boogu-Image image | Runs well | FP8 | 12.9 GB | — |
| Chroma1-HD image | Runs well | FP8 | 11.5 GB | — |
| ERNIE-Image image | Runs well | Q8_0 | 11.0 GB | — |
| FLUX.1 Fill image | Runs well | Q8_0 | 15.3 GB | — |
| FLUX.1 Kontext image | Runs well | FP8 | 14.5 GB | — |
| FLUX.1 Krea image | Runs well | FP8 | 14.2 GB | — |
| FLUX.1 dev image | Runs well | FP8 | 14.2 GB | 36.4 s / image |
| FLUX.1 schnell image | Runs well | FP8 | 14.2 GB | 9.4 s / image |
| FLUX.2 klein 4B image | Runs well | 16-bit | 9.6 GB | 2.9 s / image |
| FLUX.2 klein 9B image | Runs well | FP8 | 11.7 GB | 8.5 s / image |
| HiDream-O1 image | Runs well | FP8 | 10.9 GB | — |
| Illustrious / Pony image | Runs well | 16-bit | 7.1 GB | — |
| Krea 2 image | Runs well | FP8 | 15.7 GB | 10.9 s / image |
| Lumina 2.0 image | Runs well | 16-bit | 7.0 GB | — |
| Mage-Flow image | Runs well | 16-bit | 10.0 GB | — |
| Ming-Image image | Runs well | 16-bit | 14.6 GB | 7.4 s / image |
| Qwen-Image 2.1 image | Runs well | INT8 | 9.6 GB | 17.2 s / image |
| SD 1.5 image | Runs well | 16-bit | 3.3 GB | — |
| SD 3.5 Large image | Runs well | Q8_0 | 11.1 GB | — |
| SD 3.5 Medium image | Runs well | 16-bit | 6.9 GB | — |
| SDXL image | Runs well | 16-bit | 7.1 GB | — |
| Z-Image image | Runs well | 16-bit | 14.3 GB | — |
| Z-Image Turbo image | Runs well | 16-bit | 14.3 GB | 11.0 s / image |
| HunyuanVideo 1.5 video | Runs well | FP8 | 12.6 GB | — |
| Wan 2.1 1.3B video | Runs well | 16-bit | 5.6 GB | — |
| Wan 2.2 5B video | Runs well | 16-bit | 13.8 GB | 8.6 min / clip |
| HiDream-I1 image | Runs | Q5_K_M | 15.8 GB | — |
| HiDream-I1 Full image | Runs | Q5_K_M | 15.8 GB | — |
| HunyuanImage 2.1 image | Runs | Q4_K_M | 14.6 GB | — |
| Ideogram 4 image | Runs | Q4_1 | 14.7 GB | — |
| Qwen-Image image | Runs | Q4_K_M | 15.9 GB | — |
| HunyuanVideo 13B video | Runs | Q6_K | 15.3 GB | — |
| LTX-Video 13B video | Runs | Q6_K | 15.2 GB | — |
| Wan 2.1 14B video | Runs | Q5_K_M | 15.6 GB | — |
| Wan 2.1 I2V 480P video | Runs | Q4_K_M | 15.6 GB | — |
| Wan 2.2 I2V video | Runs | Q5_K_M | 15.1 GB | — |
| Wan 2.2 T2V video | Runs | Q5_K_M | 15.1 GB | 19.1 min / clip |
| Qwen-Image-Edit image | Tight | Q3_K_M | 12.9 GB | — |
| LTX-2 video | Tight | Q3_K_M | 14.9 GB | — |
| LTX-2.3 video | Tight | Q3_K_M | 15.6 GB | — |
| LTX-2.5 video | Tight | Q2_K | 13.6 GB | — |
| MiniMax H3 Pruned video | Tight | Q3_K_M | 14.7 GB | — |
| SCAIL-2 video | Tight | Q3_K_M | 14.4 GB | — |
| Wan 2.1 I2V 720P video | Tight | Q3_K_M | 15.4 GB | — |
| Wan 2.2 Animate video | Tight | Q3_K_M | 13.9 GB | — |
| Wan 2.2 S2V video | Tight | Q2_K | 14.3 GB | — |
| Wan Animate 2 video | Tight | Q3_K_M | 13.9 GB | — |
| Wan VACE 14B video | Tight | Q3_K_S | 12.6 GB | — |
| FLUX.2 dev image | Offload only | Q2_K | 16.2 GB | — |
| MiniMax H3 video | Offload only | Q4_K_M | 25.7 GB | — |
“VRAM needed” is calculated from the real file size, the model's working memory at about one megapixel (or a short clip for video) and a small reserve for Windows. “Measured” is the fastest file I timed on my own RTX 5060 Ti 16 GB with 32 GB of RAM. Click a model for every file, the download list and the RAM you need.
What I learned from timing it
- Use FP8 where it exists. The 50 series computes FP8 in hardware. FLUX.1 dev FP8 was faster than GGUF Q8_0, and Wan 2.2 14B FP8 was faster than a GGUF that fits. FLUX numbers · Wan numbers.
- Small, distilled models fly. Z-Image Turbo and FLUX.1 schnell take around 10 seconds per 1024×1024 image.
- Video is slow but works. A 5-second Wan 2.2 clip takes 8 to 20+ minutes depending on the model.
Want to run them yourself? Download the workflows I timed — free, with every file linked.
Full card details: RTX 5060 Ti 16 GB. Thinking about a different card? See what an upgrade would change.