Measured and reported results
69 real-world results for models on specific GPUs: seconds per image or clip, iterations per second, peak VRAM. Each one is copied as published, with a link to the source.
| Label | Model / GPU | Setup | Result | Peak VRAM | Date | Source |
|---|---|---|---|---|---|---|
| reported | FLUX.1 dev on Arc A770 16 GB | fp8 (ComfyUI template) · 20 steps “20/20 [00:46<00:00, 2.33s/it] Prompt executed in 47.13 seconds” 'GPU Benchmark Flux DEV fp8' thread; Intel A770 on Fedora Linux, PyTorch 2.3.110+xpu (53.26 s with PyTorch nightly); resolution not stated | 47.13 s / image · 2.33 s/it | — | 2025-07-24 | github.com → |
| reported | SDXL on Arc A770 16 GB | SDXL 1.0 base · 1024x1024 · 20 steps “20/20 [00:14<00:00, 1.36it/s] ... Prompt executed in 23.44 seconds” ACER A770 16GB; thread benchmark = SDXL 1024x1024 20 steps seed 1; later runs 1.53 it/s / 14.28 s | 23.44 s / image · 1.36 it/s | — | 2025-07-10 | github.com → |
| reported | FLUX.1 dev on Arc B580 12 GB | fp8 · 1024x1024 · 20 steps “Flux1 dev fp8 (20step): GOOD (35s)” Intel Arc B580 12GB, PyTorch 2.8 XPU in Docker (YanWenKun); single 1024x1024 image, pre-warmed inference only (model load excluded) | 35 s / image | — | 2025-08-03 | github.com → |
| reported | FLUX.1 Krea on Arc B580 12 GB | Flux1 Krea dev (precision not stated) · 1024x1024 · 20 steps “Flux1 Krea dev (20step): OK (46s)” Intel Arc B580 12GB, PyTorch 2.8 XPU in Docker (YanWenKun); single 1024x1024 image, pre-warmed inference only (model load excluded) | 46 s / image | — | 2025-08-03 | github.com → |
| reported | FLUX.1 schnell on Arc B580 12 GB | Flux1 schnell (template) · 1024x1024 · 4 steps “Flux1 schnell (4step): GOOD (8s)” Intel Arc B580 12GB, PyTorch 2.8 XPU in Docker (YanWenKun); single 1024x1024 image, pre-warmed inference only (model load excluded) | 8 s / image | — | 2025-08-03 | github.com → |
| reported | SD 3.5 Large on Arc B580 12 GB | fp8 · 1024x1024 · 20 steps “SD 3.5 large fp8 (20step): GOOD (26s)” Intel Arc B580 12GB, PyTorch 2.8 XPU in Docker (YanWenKun); single 1024x1024 image, pre-warmed inference only (model load excluded) | 26 s / image | — | 2025-08-03 | github.com → |
| reported | SDXL on Arc B580 12 GB | sd_xl_base_1.0 · 1024x1024 · 20 steps “100%|██| 20/20 [00:05<00:00, 3.96it/s]” ComfyUI 'GPU Benchmark' thread: default workflow, SDXL 1.0 base, 1024x1024, seed 1, second run; Intel B580 Steel Legend OC 12 GB | 3.96 it/s | — | 2025-05-12 | github.com → |
| reported | Wan 2.2 5B on Arc B580 12 GB | Wan 2.2 5B T2V (template) · 640x352 “Wan 2.2 5B Text to Video (640x352, 121 frames): FAST (60s, very poor quality)” Arc B580 12GB, PyTorch 2.8 XPU; same post: 960x544/121f = 181s, 1280x704 OOM; steps not stated | 60 s / clip (121 frames) | — | 2025-08-03 | github.com → |
| reported | FLUX.1 dev on Radeon 8060S (Strix Halo) 96 GB | flux1-dev full precision (23.8 GB) + t5xxl_fp16 · 1024x1024 · 20 steps “Steady state | 3.64 s/it | 77.56 s” AMD ROCm blog, Ryzen AI Max+ 395 / Radeon 8060S, 128 GB unified, Windows ComfyUI; first run 103.89 s; vendor-published measurement | 77.56 s / image · 3.64 s/it | — | 2026-07-14 | rocm.blogs.amd.com → |
| reported | LTX-2 on Radeon 8060S (Strix Halo) 96 GB | LTX-2 BF16 (T2V) · 1280x720 “"workflow": "LTX2-T2V-BF16.json", ... "duration_seconds": 615.0017409324646” kyuz0 Strix Halo toolbox benchmark; cold run incl. model load; resolution/frames from benchmark page; I2V = 616.16 s; steps not stated | 615 s / clip (121 frames) | — | 2026-02-13 | raw.githubusercontent.com → |
| reported | Qwen-Image on Radeon 8060S (Strix Halo) 96 GB | Qwen-Image-2512 BF16 + 4-step Lightning LoRA · 1328x1328 · 4 steps “"workflow": "Qwen-Image-2512-BF16-4-Step-LoRA.json", ... "duration_seconds": 75.37661480903625” kyuz0 Strix Halo ComfyUI toolbox benchmark (Ryzen AI Max, ROCm); cold run incl. model load, flags --disable-mmap --gpu-only --disable-smart-memory --cache-none; resolution from ben | 75.38 s / image | — | 2026-02-13 | raw.githubusercontent.com → |
| reported | Qwen-Image-Edit on Radeon 8060S (Strix Halo) 96 GB | Qwen-Image-Edit-2511 BF16 + 4-step LoRA · ~1.6MP (dynamic) · 4 steps “"workflow": "Qwen-Image-Edit-2511-BF16-4-Step-LoRA.json", ... "duration_seconds": 112.70737218856812” kyuz0 Strix Halo toolbox benchmark; cold run incl. model load; 20-step run = 667.24 s; date from run timestamp | 112.71 s / image | — | 2026-02-13 | raw.githubusercontent.com → |
| reported | Wan 2.2 T2V on Radeon 8060S (Strix Halo) 96 GB | Wan 2.2 T2V A14B (high+low noise experts) · 640x640 · 4 steps “Steady state | 197 s/it | 186 s/it | 26 min 51 s” AMD ROCm blog, Ryzen AI Max+ 395; s/it for high-noise / low-noise expert; 26 min 51 s = 1611 s; first run 36 min 14 s; vendor-published | 1611 s / clip (81 frames) | — | 2026-07-14 | rocm.blogs.amd.com → |
| reported | SDXL on RTX 2070 Laptop 8 GB | SDXL 1.0 base · 1024x1024 · 20 steps “2070 rtx mobile 20/20 [00:19<00:00, 1.05it/s] ... Prompt executed in 24.22 seconds” Thread benchmark SDXL 1024x1024 20 steps | 24.22 s / image · 1.05 it/s | — | 2024-03-04 | github.com → |
| reported | Qwen-Image on RTX 3060 12 GB | 20 steps “Qwen-ImageがRTX 3060(12GB)で動くと聞いて、早速ComfyUI版をお試し。確かに問題なく動いて、20stepでちょうど5分。” 'ちょうど5分' = exactly 5 minutes (converted to 300 s); ComfyUI version, file/resolution not stated; quote taken from search index (x.com not fetchable); date from tweet ID | 300 s / image | — | 2025-08-05 | x.com → |
| reported | Wan 2.2 I2V on RTX 3060 12 GB | video_wan2_2_14B_i2v template (quant not stated) · 720p “53frame/16fpsで513.5秒” ComfyUI on native Linux, PyTorch 2.8; 10GB offloaded via MultiGPU node; steps/LoRA not stated | 513.5 s / clip (53 frames) | — | 2025-11-26 | note.com → |
| reported | Z-Image Turbo on RTX 3060 12 GB | 1024x1024 “私の環境(グラボ:RTX 3060 12GB)では解像度1024*1024pxの画像を生成するのに大体45秒程度かかりました” ComfyUI; approximate ('大体…程度'); file/steps not stated (article notes ~12GB model) | 45 s / image | — | 2025-12-05 | kurokumasoft.com → |
| reported | Illustrious / Pony on RTX 3060 Ti 8 GB | Illustrious-XL-v2.0 · 1024x1024 “NVIDIA RTX 3060 Ti | 1.96it/s | 0.51s/it | CUDA 12.9 | ComfyUI (Unknown) | Ubuntu Server 24.04.2 LTS” Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section | 0.51 s/it | — | 2025-11-29 | huggingface.co → |
| reported | Krea 2 on RTX 3060 Ti 8 GB | fp8 “3060 ti here, 8GB VRAM, fp8 quant gets about 3.9 seconds per iteration” Posted on Krea-2-Turbo repo; ComfyUI ('run ... in Comfy now'); resolution/steps not stated; year inferred from Krea 2 release (Jun 2026) | 3.9 s/it | — | 2026-07-01 | huggingface.co → |
| reported | SDXL on RTX 3070 Laptop 8 GB | SDXL 1.0 base · 1024x1024 · 20 steps “20/20 [00:11<00:00, 1.73it/s] ... Prompt executed in 15.44 seconds” RTX 3070 Laptop GPU, Asus ZenBook Duo; thread benchmark SDXL 1024x1024 20 steps | 15.44 s / image · 1.73 it/s | — | 2024-04-22 | github.com → |
| reported | SDXL on RTX 3080 Ti 12 GB | sd_xl_base_1.0 · 1024x1024 · 20 steps “20/20 [00:04<00:00, 4.01it/s] Prompt executed in 5.60 second” ComfyUI 'GPU Benchmark' thread: default workflow, SDXL 1.0 base, 1024x1024, seed 1, second run | 5.6 s / image · 4.01 it/s | — | 2025-04-30 | github.com → |
| reported | FLUX.1 dev on RTX 3090 24 GB | fp8 (ComfyUI template) “Nvidia 3090: 26s” ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread | 26 s / image | — | 2025-07-22 | github.com → |
| reported | HiDream-I1 on RTX 3090 24 GB | HiDream Full fp8 (T5 fp8, Llama3.1 fp8 scaled) “ComfyUI Full example from their page at https://comfyanonymous.github.io/ComfyUI_examples/hidream/ runs at 115.09s on my RTX 3090, from cold start.” HiDream-I1 Full via ComfyUI example workflow, Windows portable; cold start incl. load; resolution/steps not restated | 115.09 s / image | — | 2025-04-18 | huggingface.co → |
| reported | Illustrious / Pony on RTX 3090 24 GB | Illustrious-XL-v2.0 · 1024x1024 “NVIDIA RTX 3090 | 4.00it/s | 0.25s/it | CUDA 12.9 | ComfyUI (Unknown) | Arch Linux” Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section | 0.25 s/it | — | 2025-11-29 | huggingface.co → |
| reported | MiniMax H3 on RTX 3090 24 GB | transformer variant not stated; NVFP4 text encoder · 832x480 “render 23 min 17 s ... peak VRAM ~19.8 GB of 24 GB” 15.08 s clip at 24 fps with audio; needs --disable-pinned-memory on 32 GB RAM; 23 min 17 s = 1397 s; no date shown | 1397 s / clip (362 frames) | 19.8 GB | — | github.com → |
| reported | SDXL on RTX 3090 24 GB | sd_xl_base_1.0 · 1024x1024 “Prompt executed in 6.16 seconds” ComfyUI 'GPU Benchmark' thread: default workflow, SDXL 1.0 base, 1024x1024, seed 1, second run | 6.16 s / image | — | 2024-03-06 | github.com → |
| reported | Z-Image Turbo on RTX 3090 24 GB | Tongyi-MAI/Z-Image-Turbo · 1024x1024 “NVIDIA RTX 3090 | 0.50it/s | 2.01s/it | CUDA 12.6 | ComfyUI (5151cff) | Arch Linux” Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; Z-Image 1024px section; CFG 8 (above Turbo's us | 2.01 s/it | — | 2025-11-29 | huggingface.co → |
| reported | MiniMax H3 Pruned on RTX 4060 8 GB | FL2VA pruned 20B INT8 ConvRot + Q2_K GGUF text encoder · 832x480 · 20 steps “Denoising time: About 5.5 minutes” Desktop RTX 4060 8GB, 96 GB RAM, WSL; 5.5 min denoise only; "Total first-run time: Just under 8 minutes"; a later 15-second run took about 15 minutes | 330 s / clip (107 frames) | — | 2026-08 | mountainmeadowsystems.com → |
| reported | Illustrious / Pony on RTX 4060 Laptop 8 GB | WAI-Illustrious SDXL v16.0 · 1024x1024 · 20 steps “1024x1024 | 1.47 | 13s | 15.81s | ~5.6GB” ComfyUI, euler_ancestral/Karras CFG 5; columns it/s | KSampler | total | VRAM; no --lowvram; VRAM approximate | 15.81 s / image · 1.47 it/s | 5.6 GB | 2026-02-26 | lilting.ch → |
| reported | Wan 2.2 5B on RTX 4060 Laptop 8 GB | Wan2_2-TI2V-5B_fp8_e4m3fn_scaled_KJ · 480x480 · 30 steps “30/30 [00:57<00:00, 1.93s/it] ... Prompt executed in 94.93 seconds” Article says "RTX 4060 (8GB VRAM)", 32 GB RAM, Win 11; same author (lilting) documents this machine as an RTX 4060 Laptop in other posts; I2V; frames not stated; 50 steps = 113.93 | 94.93 s / clip · 1.93 s/it | — | 2026-03-06 | lilting.ch → |
| reported | Wan 2.2 I2V on RTX 4060 Laptop 8 GB | WAN 2.2 14B Rapid distilled (all-in-one) · 480x480 · 4 steps “4/4 [00:45<00:00, 11.46s/it] ... Prompt executed in 111.41 seconds” Same machine note as above (8 GB, likely Laptop); 4,569 MB loaded on GPU with 11,067 MB offloaded (peak_vram = loaded weights, not measured peak) | 111.41 s / clip (33 frames) · 11.46 s/it | 4.46 GB | 2026-03-06 | lilting.ch → |
| reported | FLUX.1 dev on RTX 4060 Ti 16 GB | FP8 “Flux-Dev(FP8) … 2回目は72秒から51秒とあまり効果が感じられませんでした。” ComfyUI; same PC before/after swapping RTX 4060 Ti 8GB -> 16GB; 2nd-run times; resolution/steps not stated; 51 s = 16GB card | 51 s / image | — | 2025-01-03 | note.com → |
| reported | FLUX.1 schnell on RTX 4060 Ti 16 GB | FP8 “Flux-Schnell(FP8) … 2回目は24秒から11秒です。” ComfyUI; same PC before/after swapping RTX 4060 Ti 8GB -> 16GB; 2nd-run times; resolution/steps not stated; 11 s = 16GB card | 11 s / image | — | 2025-01-03 | note.com → |
| reported | Qwen-Image 2.1 on RTX 4060 Ti 16 GB | Qwen Image 2.1 INT8 Convrot (+Qwen3 VL 8B INT8 TE) “Image generation: ~20 seconds ... Peak VRAM: ~15 GB” ComfyUI, no CPU/disk offload; approximate values ('~'); editing ~60 s; resolution/steps not stated; date derived from '1 day ago' on 2026-09-25 | 20 s / image | 15 GB | 2026-09-24 | huggingface.co → |
| reported | FLUX.1 dev on RTX 4060 Ti 8 GB | FP8 “Flux-Dev(FP8) … 2回目は72秒から51秒とあまり効果が感じられませんでした。” ComfyUI; same PC before/after swapping RTX 4060 Ti 8GB -> 16GB; 2nd-run times; resolution/steps not stated; 72 s = 8GB card | 72 s / image | — | 2025-01-03 | note.com → |
| reported | FLUX.1 Krea on RTX 4060 Ti 8 GB | Flux.1 Krea Dev CLIP+VAE FP8 (~20GB checkpoint) · 768x1024 · 20 steps “すべて768x1024または1024x768、20ステップ、seed固定。… 合計(2回目以降) | 約40秒” ComfyUI with async weight offloading to system RAM; sampling ~35 s; first run ~130 s; approximate ('約'); 768x1024 or 1024x768 | 40 s / image | — | 2026-05-06 | hide10.com → |
| reported | FLUX.1 schnell on RTX 4060 Ti 8 GB | FP8 “Flux-Schnell(FP8) … 2回目は24秒から11秒です。” ComfyUI; same PC before/after swapping RTX 4060 Ti 8GB -> 16GB; 2nd-run times; resolution/steps not stated; 24 s = 8GB card | 24 s / image | — | 2025-01-03 | note.com → |
| reported | Qwen-Image 2.1 on RTX 4070 12 GB | Qwen-Image-2.1 int8 convrot + qwen3vl_8b_int8 convrot TE + VAE bf16 · 1024x1024 · 25 steps “1024x1024・25ステップ(公式ワークフロー相当) | 14.0 秒” ComfyUI, RTX 4070 12GB with partial offload (17GB weights); same table: 12 steps 9.6 s, 2048x2048 20 steps 90.5 s; repo date not shown (Qwen-Image 2.1 era, 2026) | 14 s / image | — | — | github.com → |
| reported | Qwen-Image 2.1 on RTX 4070 12 GB | Qwen Image 2.1 INT8 ConvRot · 832x1248 · 25 steps “「全体の所要時間」で見ると、平均で32.63秒から19.83秒へ(約39.2%短縮)… GPU全体のVRAMピーク:11.19 GiB → 11.28 GiB” ComfyUI, RTX 4070 12GB, Euler/simple CFG 1; 32.63 s = standard total time (19.83 s with Spectrum speedup node); VRAM in GiB | 32.63 s / image | 11.19 GB | 2026-09-24 | note.com → |
| reported | SDXL on RTX 4070 12 GB | sd_xl_base_1.0 · 1024x1024 · 20 steps “20/20 [00:06<00:00, 3.21it/s] Prompt executed in 7.13 seconds” ComfyUI 'GPU Benchmark' thread: default workflow, SDXL 1.0 base, 1024x1024, seed 1, second run; GPU stated as 'RTX 4070 12Gb' | 7.13 s / image · 3.21 it/s | — | 2024-03-21 | github.com → |
| reported | MiniMax H3 Pruned on RTX 4070 Laptop 8 GB | pruned int8 convrot (FL2VA/Ref2VA) · 832x480 “On a laptop with an RTX 4070 (8GB VRAM) and 32GB of RAM, we generated one 832×480, 15-second clip in 45 minutes.” 15-second clip, 24 fps, with audio; ComfyUI I2V/R2V workflows (which one not stated); 45 min = 2700 s; steps not stated | 2700 s / clip | — | 2026-08-04 | metallab.ai → |
| reported | FLUX.1 dev on RTX 4070 Super 12 GB | Q4_0 GGUF “1.9s/it with Q4_0” RTX 4070 Super 12GB; same post: 2.6s/it with Q5_1, 1.3s/it with NF4; resolution not stated; early (Aug 2024) ComfyUI-GGUF | 1.9 s/it | — | 2024-08-17 | huggingface.co → |
| reported | FLUX.2 klein 4B on RTX 4070 Super 12 GB | flux-2-klein-4b (distilled) + qwen_3_4b “5 秒でした。” ComfyUI 4B T2I template (distilled, 4 steps default); resolution not stated; base 4B T2I 37 s incl. model load; distilled edit 14-17 s | 5 s / image | — | 2026-01-20 | note.com → |
| reported | FLUX.1 dev on RTX 4080 16 GB | fp8_e4m3fn weight_dtype · 1024x1024 · 20 steps “weight_dtype (fp8_e4m3fn) with --fast (13sec)” ComfyUI 'RTX 4090 benchmarks - FLUX model' thread; OP settings 1024x1024, 20 steps; Aug 2024 ComfyUI/PyTorch 2.5 dev; with --fast flag (same poster: 19sec without --fast, 28sec def | 13 s / image | — | 2024-08-23 | github.com → |
| reported | FLUX.1 dev on RTX 4090 24 GB | fp8 (ComfyUI template) “Prompt executed in 11.28 seconds” ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread | 11.28 s / image | — | 2025-12-28 | github.com → |
| reported | FLUX.1 dev on RTX 4090 24 GB | Q8_0 GGUF · 1024x1024 · 20 steps “15 seconds at the fastest to 17 seconds at the slowest on my RTX 4090 with Euler 20 Steps for 1024x1024 images” stated range 15-17 s; city96 ComfyUI-GGUF Q8 | 15 s / image | — | 2024-08-25 | github.com → |
| reported | FLUX.1 dev on RTX 4090 24 GB | fp8 (--fast) · 1024x1024 · 20 steps “Prompt executed in 10.01 seconds” ComfyUI 'RTX 4090 benchmarks - FLUX model' thread; OP settings 1024x1024, 20 steps; Aug 2024 ComfyUI/PyTorch 2.5 dev; FP8 with --fast, GPU at 2.52 GHz/875mV undervolt (9.07 s at 2. | 10.01 s / image | — | 2024-08-26 | github.com → |
| reported | HunyuanVideo 1.5 on RTX 4090 24 GB | 720p model · 848x480 “The 5 seconds video took 297s to generate so barely longer than on your end” Replicated the 5090 poster's ComfyUI workflow (720p model at 848x480, 5 s @24fps); card undervolted (~5% slower per poster) | 297 s / clip | — | 2025-12-05 | huggingface.co → |
| reported | Illustrious / Pony on RTX 4090 24 GB | Illustrious-XL-v2.0 · 1024x1024 “NVIDIA RTX 4090 | 7.00it/s | 0.14s/it | CUDA 12.9 | ComfyUI (Unknown) | Windows 11 24H2” Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section | 0.14 s/it | — | 2025-11-29 | huggingface.co → |
| reported | MiniMax H3 Pruned on RTX 4090 24 GB | minimax_h3_fl2va_pruned_int8_convrot + qwen3vl_32b nvfp4_awq · 832x480 · 8 steps “5 s clip (832×480, 8 steps): ~7 min” ~7 min approx = 420 s for 5 s clip; VRAM ~6.6 GB during sampling, ~22.5 GB spike at model load; 15 s clip ~25-30 min; no date shown | 420 s / clip | 22.5 GB | — | github.com → |
| reported | Wan 2.1 I2V 480P on RTX 4090 24 GB | 30 steps “100%|███| 30/30 [06:13<00:00, 12.46s/it]” Kijai WanVideoWrapper wanvideo_480p_I2V_example_01.json; 32GB system RAM, process later 'Killed' (RAM); resolution/frames not stated | 12.46 s/it | — | 2025-02-26 | github.com → |
| reported | Wan 2.2 I2V on RTX 4090 24 GB | Wan2.2-I2V-A14B High/Low Q6_K GGUF · 800x448 · 8 steps “RTX 4090なら1分半ほど、RTX 5080はほぼ2分で480p解像度を5秒出力できます。” Chimolog GPU review; ComfyUI 0.3.5x, Kijai-based workflow, 2+2+4 steps with Lightx2v/Lightning LoRAs; '1分半ほど' = about 1.5 min (converted to 90 s, approximate); exact values only in | 90 s / clip (81 frames) | — | 2025-08-28 | chimolog.co → |
| reported | FLUX.1 dev on RTX 5060 Ti 16 GB | fp8 (ComfyUI template) “Prompt executed in 25.71 seconds” ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread; poster's log shows 16311 | 25.71 s / image | — | 2025-08-04 | github.com → |
| reported | FLUX.2 dev on RTX 5060 Ti 16 GB | flux2_dev_fp8mixed + mistral_3_small_flux2_fp8 · 1024x1024 · 20 steps “FP8 t2i 142秒程度(20Steps)” ComfyUI t2i at 1MP; approximate (程度) | 142 s / image | — | 2025-11-28 | note.com → |
| reported | FLUX.2 dev on RTX 5060 Ti 16 GB | GGUF Q4_0 (flux2-dev-q4_0.gguf) · 1024x1024 · 20 steps “GGUF Q4_0 t2i 180秒程度(20Steps)” ComfyUI t2i at 1MP; approximate (程度); author notes GGUF slower than fp8 when spilling VRAM | 180 s / image | — | 2025-11-28 | note.com → |
| reported | Illustrious / Pony on RTX 5060 Ti 16 GB | Illustrious-XL-v2.0 · 1024x1024 “NVIDIA RTX 5060 Ti | 2.60it/s | 0.39s/it | CUDA 12.9 | ComfyUI (Unknown) | Arch Linux” Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section; row says 'RTX 5060 Ti', co | 0.39 s/it | — | 2025-11-29 | huggingface.co → |
| reported | Wan 2.2 I2V on RTX 5060 Ti 16 GB | video_wan2_2_14B_i2v template (quant not stated) · 720p “RTX5060ti 16GBだと165.9秒なので、2〜3倍時間が必要ですが、落ちずに720p動画が生成できる事はたいしたものです。” Comparison figure given in the RTX 3060 post, presumably same 720p/53-frame test (not explicitly restated); steps not stated | 165.9 s / clip (53 frames) | — | 2025-11-26 | note.com → |
| reported | Z-Image Turbo on RTX 5060 Ti 16 GB | z_image_turbo_bf16.safetensors · 1328x1328 · 8 steps “約 35秒かかりました。一度モデルを VRAM にロードした後、プロンプトを変えての再実行だと約 22秒で生成できます。” ComfyUI; ~35 s first run incl. load, ~22 s warm; approximate ('約') | 22 s / image | — | 2025-11-30 | iwannacreateapps.com → |
| reported | Krea 2 on RTX 5070 Ti 16 GB | Krea 2 Turbo (ComfyUI release, 17.8GB total download) · 1024x1024 · 8 steps “generating a 1024x1024 pixel image took 16-25 seconds when the prompts were rewritten, and approximately 9-15 seconds when the same prompts were reused” GIGAZINE review, ComfyUI; range only (9-15 s same prompt, 16-25 s new prompt incl. text encoding) | 9–15 s / image | — | 2026-06-24 | gigazine.net → |
| reported | FLUX.1 dev on RTX 5090 32 GB | fp8 (ComfyUI template) “Getting 8.78s at 2.38it/s for 3 runs.” ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread; Inno3D RTX 5090 X3 OC | 8.78 s / image · 2.38 it/s | — | 2025-08-05 | github.com → |
| reported | HunyuanVideo 1.5 on RTX 5090 32 GB | 720p model · 848x480 “I was able to generate 5 seconds video in 284s at 24fps using the 720p model at 848*480” ComfyUI; 5 s video at 24 fps; steps not stated | 284 s / clip | — | 2025-11-22 | huggingface.co → |
| reported | Illustrious / Pony on RTX 5090 32 GB | Illustrious-XL-v2.0 · 1024x1024 “NVIDIA RTX 5090 | 8.95it/s | 0.11s/it | CUDA 12.8 | ComfyUI (Unknown) | Windows 11 24H2” Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section | 0.11 s/it | — | 2025-11-29 | huggingface.co → |
| reported | LTX-2 on RTX 5090 32 GB | LTX-2 19B NVFP4 · 720p “generating 143 frames at 720p takes about 66 seconds end‑to‑end in ComfyUI” Issue says this is ~30-40% slower than NVIDIA's expected 40-45 s; steps not stated | 66 s / clip (143 frames) | — | 2026-01-08 | github.com → |
| reported | MiniMax H3 Pruned on RTX 5090 32 GB | minimax_h3_fl2va_pruned_nvfp4 · 864x480 · 10 steps “175 s for a 864×480 ten-second clip, 26,914 MiB peak VRAM.” ComfyUI 0.30.1, 500 W power cap; 26,914 MiB = 26.28 GiB; INT8-ConvRot = 185 s / 28,581 MiB | 175 s / clip (243 frames) | 26.28 GB | 2026-08-04 | ai-muninn.com → |
| reported | Z-Image Turbo on RTX 5090 32 GB | fp8_e4m3fn (per log) “without SA: 0.95 seg” 'seg' = segundos (seconds) per image; SageAttention variants gave 0.87-0.91 s; resolution/steps not stated | 0.95 s / image | — | 2025-11-27 | github.com → |
| reported | SDXL on RX 7600 XT 16 GB | SDXL 1.0 base · 1024x1024 · 20 steps “20/20 [00:17<00:00, 1.13it/s] Prompt executed in 20.52 seconds” ComfyUI 0.3.29 Zluda, Windows 11, driver 25.3.1; thread benchmark SDXL 1024x1024 20 steps | 20.52 s / image · 1.13 it/s | — | 2025-04-18 | github.com → |
| reported | LTX-2.3 on RX 7900 XTX 24 GB | ltx-2.3-22b-dev-fp8 + gemma_3_12B_it_fp4_mixed (I2V template) · 768x1280 “[INFO] Prompt executed in 197.38 seconds” User QualiaSG, 64 GB RAM, ROCm 7.13, with --disable-dynamic-vram --disable-smart-memory --disable-pinned-memory (stalls without them); I2V template (distilled), 24 fps; comment dat | 197.38 s / clip (121 frames) | — | — | github.com → |
| reported | Wan 2.1 I2V 480P on RX 7900 XTX 24 GB | wan2.1_i2v_480p_14B_fp8_scaled · 480x704 · 25 steps “Wan2.1 i2v | 480×704 | 81 | 25 | ~40 min” Windows 11, native ROCm 7.1, PYTORCH_NO_HIP_MEMORY_CACHING=1; approximate (~40 min = 2400 s); author notes slow vs CUDA | 2400 s / clip (81 frames) | — | 2026-04-05 | joshwaamein.github.io → |
| reported | SDXL on RX 9070 XT 16 GB | SDXL · 1024x1024 · 20 steps “| SDXL | 20 | 1024x1024 | 1.5it/s | 27.76s | Manual tiled VAE decoder |” Ubuntu 24.04, ROCm 6.4.1, PyTorch nightly, TUNABLEOP + AOTriton experimental, --use-pytorch-cross-attention; without tiled VAE 1.49 it/s / 34.56 s; default VAE decode could OOM | 27.76 s / image · 1.5 it/s | — | 2025-05-30 | github.com → |
Reported results are other people's numbers, copied as published, with a link. Settings, drivers and ComfyUI versions differ, so compare them with care. Send yours.