RTX 4060 Ti 16 GB for local AI
Which image and video models run on the RTX 4060 Ti 16 GB, which file to download for each, and how much VRAM they need.
VRAM16 GB
MemoryGDDR6
Bus128-bit
Bandwidth288 GB/s
ArchitectureAda Lovelace
Launched2023-07
Launch price$499
Runs well25 of 49
Specs: www.nvidia.com · launch: en.wikipedia.org
01What runs on it
| Model | Verdict | Best file | Size | Needed |
|---|---|---|---|---|
| image models | ||||
| FLUX.1 [dev] | Runs well | FP8 | 11.9 GB | 14.2 GB |
| FLUX.1 [schnell] | Runs well | FP8 | 11.9 GB | 14.2 GB |
| FLUX.1 Kontext [dev] | Runs well | FP8 | 11.9 GB | 14.5 GB |
| FLUX.1 Krea [dev] | Runs well | FP8 | 11.9 GB | 14.2 GB |
| FLUX.1 Fill [dev] | Runs well | Q8_0 | 12.7 GB | 15.3 GB |
| FLUX.2 [dev] | Offload only | Q2_K | 12.9 GB | 16.2 GB |
| FLUX.2 [klein] 9B | Runs well | FP8 | 9.4 GB | 11.7 GB |
| FLUX.2 [klein] 4B | Runs well | 16-bit | 7.8 GB | 9.6 GB |
| Krea 2 (Turbo) | Runs well | FP8 | 13.1 GB | 15.7 GB |
| Qwen-Image | Runs | Q4_K_M | 13.1 GB | 15.9 GB |
| Qwen-Image-Edit (2511) | Tight | Q3_K_M | 9.9 GB | 12.9 GB |
| Qwen-Image 2.1 | Runs well | Q8_0 | 7.6 GB | 9.9 GB |
| Z-Image Turbo | Runs well | 16-bit | 12.3 GB | 14.3 GB |
| Z-Image (base) | Runs well | 16-bit | 12.3 GB | 14.3 GB |
| Ideogram 4 | Runs | Q4_1 | 6.2 GB | 14.7 GB |
| Boogu-Image (Turbo) | Runs well | FP8 | 10.3 GB | 12.9 GB |
| ERNIE-Image (Turbo) | Runs well | Q8_0 | 8.7 GB | 11.0 GB |
| HiDream-O1-Image | Runs well | FP8 | 8.1 GB | 10.9 GB |
| Mage-Flow (Microsoft) | Runs well | 16-bit | 8.2 GB | 10.0 GB |
| Lumina Image 2.0 | Runs well | 16-bit | 5.2 GB | 7.0 GB |
| HiDream-I1 (Full) | Runs | Q5_K_M | 13.0 GB | 15.8 GB |
| HiDream-I1 (Dev) | Runs | Q5_K_M | 13.0 GB | 15.8 GB |
| Stable Diffusion 3.5 Large | Runs well | Q8_0 | 8.8 GB | 11.1 GB |
| Stable Diffusion 3.5 Medium | Runs well | 16-bit | 5.1 GB | 6.9 GB |
| Chroma1-HD | Runs well | FP8 | 9.2 GB | 11.5 GB |
| SDXL 1.0 | Runs well | 16-bit | 6.9 GB | 7.1 GB |
| Illustrious XL / Pony (SDXL anime) | Runs well | 16-bit | 6.9 GB | 7.1 GB |
| Stable Diffusion 1.5 | Runs well | 16-bit | 2.1 GB | 3.3 GB |
| HunyuanImage 2.1 | Runs | Q4_K_M | 11.3 GB | 14.6 GB |
| video models | ||||
| Wan 2.1 T2V 14B | Runs | Q5_K_M | 11.3 GB | 15.6 GB |
| Wan 2.1 T2V 1.3B | Runs well | 16-bit | 2.8 GB | 5.6 GB |
| Wan 2.1 I2V 14B 480P | Runs | Q4_K_M | 11.3 GB | 15.6 GB |
| Wan 2.1 I2V 14B 720P | Tight | Q3_K_M | 8.6 GB | 15.4 GB |
| Wan 2.1 VACE 14B | Tight | Q3_K_S | 7.8 GB | 12.6 GB |
| Wan 2.2 T2V A14B | Runs | Q5_K_M | 10.8 GB | 15.1 GB |
| Wan 2.2 I2V A14B | Runs | Q5_K_M | 10.8 GB | 15.1 GB |
| Wan 2.2 TI2V 5B | Runs well | 16-bit | 10.0 GB | 13.8 GB |
| Wan 2.2 Animate 14B | Tight | Q3_K_M | 8.6 GB | 13.9 GB |
| Wan Animate 2 (14B) | Tight | Q3_K_M | 8.6 GB | 13.9 GB |
| Wan 2.2 S2V 14B | Tight | Q2_K | 9.5 GB | 14.3 GB |
| SCAIL-2 (character animation) | Tight | Q3_K_M | 9.1 GB | 14.4 GB |
| HunyuanVideo (13B, original) | Runs | Q6_K | 11.0 GB | 15.3 GB |
| LTX-Video 13B (0.9.8) | Runs | Q6_K | 10.9 GB | 15.2 GB |
| HunyuanVideo 1.5 | Runs well | FP8 | 8.3 GB | 12.6 GB |
| LTX-2 (19B) | Tight | Q3_K_M | 10.1 GB | 14.9 GB |
| LTX-2.3 (22B) | Tight | Q3_K_M | 10.8 GB | 15.6 GB |
| LTX-2.5 (22B) | Tight | Q2_K | 8.8 GB | 13.6 GB |
| MiniMax H3 (33B) | Offload only | Q4_K_M | 19.9 GB | 25.7 GB |
| MiniMax H3 Pruned | Tight | Q3_K_M | 8.9 GB | 14.7 GB |
Calculated from real file sizes plus working memory. How the numbers work.
02Good to know
The RTX 4060 Ti 16 GB is an Ada Lovelace GPU with hardware FP8, so ComfyUI can compute Comfy-Org's FP8 files natively: small and fast. (Plain FP8 files use FP8 maths with the --fast fp8_matrix_mult option.)
03Measured and reported results
| Label | Model | Setup | Result | Peak VRAM | Date | Source |
|---|---|---|---|---|---|---|
| reported | FLUX.1 dev | FP8 “Flux-Dev(FP8) … 2回目は72秒から51秒とあまり効果が感じられませんでした。” ComfyUI; same PC before/after swapping RTX 4060 Ti 8GB -> 16GB; 2nd-run times; resolution/steps not stated; 51 s = 16GB card | 51 s / image | — | 2025-01-03 | note.com → |
| reported | FLUX.1 schnell | FP8 “Flux-Schnell(FP8) … 2回目は24秒から11秒です。” ComfyUI; same PC before/after swapping RTX 4060 Ti 8GB -> 16GB; 2nd-run times; resolution/steps not stated; 11 s = 16GB card | 11 s / image | — | 2025-01-03 | note.com → |
| reported | Qwen-Image 2.1 | Qwen Image 2.1 INT8 Convrot (+Qwen3 VL 8B INT8 TE) “Image generation: ~20 seconds ... Peak VRAM: ~15 GB” ComfyUI, no CPU/disk offload; approximate values ('~'); editing ~60 s; resolution/steps not stated; date derived from '1 day ago' on 2026-09-25 | 20 s / image | 15 GB | 2026-09-24 | huggingface.co → |
Reported results are other people's numbers, copied as published, with a link. Settings, drivers and ComfyUI versions differ, so compare them with care. Send yours.