RTX 4070 Super 12 GB for local AI
Which image and video models run on the RTX 4070 Super 12 GB, which file to download for each, and how much VRAM they need.
VRAM12 GB
MemoryGDDR6X
Bus192-bit
Bandwidth504 GB/s
ArchitectureAda Lovelace
Launched2024-01
Launch price$599
Runs well17 of 49
Specs: www.nvidia.com · launch: www.nvidia.com
01What runs on it
| Model | Verdict | Best file | Size | Needed |
|---|---|---|---|---|
| image models | ||||
| FLUX.1 [dev] | Runs | Q5_K_S | 8.3 GB | 10.6 GB |
| FLUX.1 [schnell] | Runs | Q5_K_S | 8.3 GB | 10.6 GB |
| FLUX.1 Kontext [dev] | Runs | Q5_K_M | 8.4 GB | 11.0 GB |
| FLUX.1 Krea [dev] | Runs | Q5_K_M | 8.4 GB | 10.7 GB |
| FLUX.1 Fill [dev] | Runs | Q5_K_S | 8.3 GB | 10.9 GB |
| FLUX.2 [dev] | Offload only | Q4_K_M | 20.1 GB | 23.4 GB |
| FLUX.2 [klein] 9B | Runs well | FP8 | 9.4 GB | 11.7 GB |
| FLUX.2 [klein] 4B | Runs well | 16-bit | 7.8 GB | 9.6 GB |
| Krea 2 (Turbo) | Runs | Q5_K_M | 8.9 GB | 11.5 GB |
| Qwen-Image | Tight | Q2_K | 7.1 GB | 9.9 GB |
| Qwen-Image-Edit (2511) | Tight | Q2_K | 7.5 GB | 10.5 GB |
| Qwen-Image 2.1 | Runs well | Q8_0 | 7.6 GB | 9.9 GB |
| Z-Image Turbo | Runs well | Q8_0 | 7.2 GB | 9.2 GB |
| Z-Image (base) | Runs well | Q8_0 | 7.2 GB | 9.2 GB |
| Ideogram 4 | Offload only | Q4_1 | 6.2 GB | 14.7 GB |
| Boogu-Image (Turbo) | Runs | Q5_1 | 8.6 GB | 11.2 GB |
| ERNIE-Image (Turbo) | Runs well | Q8_0 | 8.7 GB | 11.0 GB |
| HiDream-O1-Image | Runs well | FP8 | 8.1 GB | 10.9 GB |
| Mage-Flow (Microsoft) | Runs well | 16-bit | 8.2 GB | 10.0 GB |
| Lumina Image 2.0 | Runs well | 16-bit | 5.2 GB | 7.0 GB |
| HiDream-I1 (Full) | Tight | Q3_K_M | 8.8 GB | 11.6 GB |
| HiDream-I1 (Dev) | Tight | Q3_K_M | 8.8 GB | 11.6 GB |
| Stable Diffusion 3.5 Large | Runs well | Q8_0 | 8.8 GB | 11.1 GB |
| Stable Diffusion 3.5 Medium | Runs well | 16-bit | 5.1 GB | 6.9 GB |
| Chroma1-HD | Runs well | FP8 | 9.2 GB | 11.5 GB |
| SDXL 1.0 | Runs well | 16-bit | 6.9 GB | 7.1 GB |
| Illustrious XL / Pony (SDXL anime) | Runs well | 16-bit | 6.9 GB | 7.1 GB |
| Stable Diffusion 1.5 | Runs well | 16-bit | 2.1 GB | 3.3 GB |
| HunyuanImage 2.1 | Tight | Q2_K | 7.3 GB | 10.6 GB |
| video models | ||||
| Wan 2.1 T2V 14B | Tight | Q3_K_M | 7.6 GB | 11.9 GB |
| Wan 2.1 T2V 1.3B | Runs well | 16-bit | 2.8 GB | 5.6 GB |
| Wan 2.1 I2V 14B 480P | Offload only | Q3_K_M | 8.6 GB | 12.9 GB |
| Wan 2.1 I2V 14B 720P | Offload only | Q4_K_M | 11.3 GB | 18.1 GB |
| Wan 2.1 VACE 14B | Offload only | Q3_K_S | 7.8 GB | 12.6 GB |
| Wan 2.2 T2V A14B | Tight | Q3_K_M | 7.2 GB | 11.5 GB |
| Wan 2.2 I2V A14B | Tight | Q3_K_M | 7.2 GB | 11.5 GB |
| Wan 2.2 TI2V 5B | Runs well | Q8_0 | 5.4 GB | 9.2 GB |
| Wan 2.2 Animate 14B | Tight | Q2_K | 6.5 GB | 11.8 GB |
| Wan Animate 2 (14B) | Tight | Q2_K | 6.5 GB | 11.8 GB |
| Wan 2.2 S2V 14B | Offload only | Q4_K_M | 13.9 GB | 18.7 GB |
| SCAIL-2 (character animation) | Offload only | Q2_K | 7.3 GB | 12.6 GB |
| HunyuanVideo (13B, original) | Tight | Q3_K_M | 6.2 GB | 10.5 GB |
| LTX-Video 13B (0.9.8) | Tight | Q3_K_M | 6.5 GB | 10.8 GB |
| HunyuanVideo 1.5 | Runs | Q6_K | 7.0 GB | 11.3 GB |
| LTX-2 (19B) | Offload only | Q2_K | 8.1 GB | 12.9 GB |
| LTX-2.3 (22B) | Offload only | Q2_K | 8.3 GB | 13.1 GB |
| LTX-2.5 (22B) | Offload only | Q2_K | 8.8 GB | 13.6 GB |
| MiniMax H3 (33B) | Offload only | Q4_K_M | 19.9 GB | 25.7 GB |
| MiniMax H3 Pruned | Offload only | Q4_K_M | 11.6 GB | 17.4 GB |
Calculated from real file sizes plus working memory. How the numbers work.
02Good to know
The RTX 4070 Super 12 GB is an Ada Lovelace GPU with hardware FP8, so ComfyUI can compute Comfy-Org's FP8 files natively: small and fast. (Plain FP8 files use FP8 maths with the --fast fp8_matrix_mult option.)
03Measured and reported results
| Label | Model | Setup | Result | Peak VRAM | Date | Source |
|---|---|---|---|---|---|---|
| reported | FLUX.1 dev | Q4_0 GGUF “1.9s/it with Q4_0” RTX 4070 Super 12GB; same post: 2.6s/it with Q5_1, 1.3s/it with NF4; resolution not stated; early (Aug 2024) ComfyUI-GGUF | 1.9 s/it | — | 2024-08-17 | huggingface.co → |
| reported | FLUX.2 klein 4B | flux-2-klein-4b (distilled) + qwen_3_4b “5 秒でした。” ComfyUI 4B T2I template (distilled, 4 steps default); resolution not stated; base 4B T2I 37 s incl. model load; distilled edit 14-17 s | 5 s / image | — | 2026-01-20 | note.com → |
Reported results are other people's numbers, copied as published, with a link. Settings, drivers and ComfyUI versions differ, so compare them with care. Send yours.