RTX 4060 Laptop 8 GB for local AI
Which image and video models run on the RTX 4060 Laptop 8 GB, which file to download for each, and how much VRAM they need.
Specs: www.nvidia.com · launch: www.nvidia.com
01What runs on it
| Model | Verdict | Best file | Size | Needed |
|---|---|---|---|---|
| image models | ||||
| FLUX.1 [dev] | Tight | Q3_K_S | 5.2 GB | 7.5 GB |
| FLUX.1 [schnell] | Tight | Q3_K_S | 5.2 GB | 7.5 GB |
| FLUX.1 Kontext [dev] | Tight | Q3_K_M | 5.4 GB | 8.0 GB |
| FLUX.1 Krea [dev] | Tight | Q3_K_M | 5.4 GB | 7.7 GB |
| FLUX.1 Fill [dev] | Tight | Q3_K_S | 5.2 GB | 7.8 GB |
| FLUX.2 [dev] | Not practical | Q2_K | 12.9 GB | 16.2 GB |
| FLUX.2 [klein] 9B | Tight | Q3_K_M | 4.8 GB | 7.1 GB |
| FLUX.2 [klein] 4B | Runs well | FP8 | 4.1 GB | 5.9 GB |
| Krea 2 (Turbo) | Tight | Q2_K | 4.9 GB | 7.5 GB |
| Qwen-Image | Offload only | Q2_K | 7.1 GB | 9.9 GB |
| Qwen-Image-Edit (2511) | Offload only | Q4_K_M | 13.2 GB | 16.2 GB |
| Qwen-Image 2.1 | Runs | Q5_K_M | 5.0 GB | 7.3 GB |
| Z-Image Turbo | Runs | Q6_K | 5.9 GB | 7.9 GB |
| Z-Image (base) | Runs | Q5_K_M | 5.6 GB | 7.6 GB |
| Ideogram 4 | Offload only | Q4_1 | 6.2 GB | 14.7 GB |
| Boogu-Image (Turbo) | Offload only | Q4_1 | 7.4 GB | 10.0 GB |
| ERNIE-Image (Turbo) | Runs | Q4_K_M | 5.0 GB | 7.3 GB |
| HiDream-O1-Image | Offload only | FP8 | 8.1 GB | 10.9 GB |
| Mage-Flow (Microsoft) | Runs well | INT8 | 4.2 GB | 6.0 GB |
| Lumina Image 2.0 | Runs well | 16-bit | 5.2 GB | 7.0 GB |
| HiDream-I1 (Full) | Offload only | Q2_K | 6.6 GB | 9.4 GB |
| HiDream-I1 (Dev) | Offload only | Q2_K | 6.6 GB | 9.4 GB |
| Stable Diffusion 3.5 Large | Runs | Q4_1 | 5.3 GB | 7.6 GB |
| Stable Diffusion 3.5 Medium | Runs well | 16-bit | 5.1 GB | 6.9 GB |
| Chroma1-HD | Runs | Q4_K_M | 5.6 GB | 7.9 GB |
| SDXL 1.0 | Runs well | 16-bit | 6.9 GB | 7.1 GB |
| Illustrious XL / Pony (SDXL anime) | Runs well | 16-bit | 6.9 GB | 7.1 GB |
| Stable Diffusion 1.5 | Runs well | 16-bit | 2.1 GB | 3.3 GB |
| HunyuanImage 2.1 | Offload only | Q4_K_M | 11.3 GB | 14.6 GB |
| video models | ||||
| Wan 2.1 T2V 14B | Offload only | Q4_K_M | 10.1 GB | 14.4 GB |
| Wan 2.1 T2V 1.3B | Runs well | 16-bit | 2.8 GB | 5.6 GB |
| Wan 2.1 I2V 14B 480P | Offload only | Q4_K_M | 11.3 GB | 15.6 GB |
| Wan 2.1 I2V 14B 720P | Offload only | Q4_K_M | 11.3 GB | 18.1 GB |
| Wan 2.1 VACE 14B | Offload only | Q4_K_M | 11.6 GB | 16.4 GB |
| Wan 2.2 T2V A14B | Offload only | Q2_K | 5.3 GB | 9.6 GB |
| Wan 2.2 I2V A14B | Offload only | Q2_K | 5.3 GB | 9.6 GB |
| Wan 2.2 TI2V 5B | Runs | Q5_K_M | 3.8 GB | 7.6 GB |
| Wan 2.2 Animate 14B | Offload only | Q4_K_M | 11.5 GB | 16.8 GB |
| Wan Animate 2 (14B) | Offload only | Q4_K_M | 11.3 GB | 16.6 GB |
| Wan 2.2 S2V 14B | Offload only | Q4_K_M | 13.9 GB | 18.7 GB |
| SCAIL-2 (character animation) | Offload only | Q4_K_M | 11.5 GB | 16.8 GB |
| HunyuanVideo (13B, original) | Offload only | Q4_K_M | 7.9 GB | 12.2 GB |
| LTX-Video 13B (0.9.8) | Offload only | Q2_K | 4.7 GB | 9.0 GB |
| HunyuanVideo 1.5 | Offload only | Q4_K_M | 5.1 GB | 9.4 GB |
| LTX-2 (19B) | Offload only | Q4_K_M | 12.8 GB | 17.6 GB |
| LTX-2.3 (22B) | Offload only | Q4_K_M | 14.3 GB | 19.1 GB |
| LTX-2.5 (22B) | Offload only | Q4_K_M | 15.1 GB | 19.9 GB |
| MiniMax H3 (33B) | Not practical | Q3_K_M | 15.6 GB | 21.4 GB |
| MiniMax H3 Pruned | Offload only | Q4_K_M | 11.6 GB | 17.4 GB |
Calculated from real file sizes plus working memory. How the numbers work.
02Good to know
The RTX 4060 Laptop 8 GB is an Ada Lovelace GPU with hardware FP8, so ComfyUI can compute Comfy-Org's FP8 files natively: small and fast. (Plain FP8 files use FP8 maths with the --fast fp8_matrix_mult option.) As a laptop GPU it runs at a lower power limit than desktop cards (35–115 W depending on the laptop). The memory verdicts are the same; speed depends heavily on how much power the laptop maker allows.
03Measured and reported results
| Label | Model | Setup | Result | Peak VRAM | Date | Source |
|---|---|---|---|---|---|---|
| reported | Illustrious / Pony | WAI-Illustrious SDXL v16.0 · 1024x1024 · 20 steps “1024x1024 | 1.47 | 13s | 15.81s | ~5.6GB” ComfyUI, euler_ancestral/Karras CFG 5; columns it/s | KSampler | total | VRAM; no --lowvram; VRAM approximate | 15.81 s / image · 1.47 it/s | 5.6 GB | 2026-02-26 | lilting.ch → |
| reported | Wan 2.2 5B | Wan2_2-TI2V-5B_fp8_e4m3fn_scaled_KJ · 480x480 · 30 steps “30/30 [00:57<00:00, 1.93s/it] ... Prompt executed in 94.93 seconds” Article says "RTX 4060 (8GB VRAM)", 32 GB RAM, Win 11; same author (lilting) documents this machine as an RTX 4060 Laptop in other posts; I2V; frames not stated; 50 steps = 113.93 | 94.93 s / clip · 1.93 s/it | — | 2026-03-06 | lilting.ch → |
| reported | Wan 2.2 I2V | WAN 2.2 14B Rapid distilled (all-in-one) · 480x480 · 4 steps “4/4 [00:45<00:00, 11.46s/it] ... Prompt executed in 111.41 seconds” Same machine note as above (8 GB, likely Laptop); 4,569 MB loaded on GPU with 11,067 MB offloaded (peak_vram = loaded weights, not measured peak) | 111.41 s / clip (33 frames) · 11.46 s/it | 4.46 GB | 2026-03-06 | lilting.ch → |
Reported results are other people's numbers, copied as published, with a link. Settings, drivers and ComfyUI versions differ, so compare them with care. Send yours.