Datasheet for local AI49 models98 GPUsData read 2026-09-25
No adsNo tracking
Check my GPU →

Measured and reported results

69 real-world results for models on specific GPUs: seconds per image or clip, iterations per second, peak VRAM. Each one is copied as published, with a link to the source.

LabelModel / GPUSetupResultPeak VRAMDateSource
reportedFLUX.1 dev on Arc A770 16 GBfp8 (ComfyUI template) · 20 steps
“20/20 [00:46<00:00, 2.33s/it] Prompt executed in 47.13 seconds”
'GPU Benchmark Flux DEV fp8' thread; Intel A770 on Fedora Linux, PyTorch 2.3.110+xpu (53.26 s with PyTorch nightly); resolution not stated
47.13 s / image · 2.33 s/it—2025-07-24github.com →
reportedSDXL on Arc A770 16 GBSDXL 1.0 base · 1024x1024 · 20 steps
“20/20 [00:14<00:00, 1.36it/s] ... Prompt executed in 23.44 seconds”
ACER A770 16GB; thread benchmark = SDXL 1024x1024 20 steps seed 1; later runs 1.53 it/s / 14.28 s
23.44 s / image · 1.36 it/s—2025-07-10github.com →
reportedFLUX.1 dev on Arc B580 12 GBfp8 · 1024x1024 · 20 steps
“Flux1 dev fp8 (20step): GOOD (35s)”
Intel Arc B580 12GB, PyTorch 2.8 XPU in Docker (YanWenKun); single 1024x1024 image, pre-warmed inference only (model load excluded)
35 s / image—2025-08-03github.com →
reportedFLUX.1 Krea on Arc B580 12 GBFlux1 Krea dev (precision not stated) · 1024x1024 · 20 steps
“Flux1 Krea dev (20step): OK (46s)”
Intel Arc B580 12GB, PyTorch 2.8 XPU in Docker (YanWenKun); single 1024x1024 image, pre-warmed inference only (model load excluded)
46 s / image—2025-08-03github.com →
reportedFLUX.1 schnell on Arc B580 12 GBFlux1 schnell (template) · 1024x1024 · 4 steps
“Flux1 schnell (4step): GOOD (8s)”
Intel Arc B580 12GB, PyTorch 2.8 XPU in Docker (YanWenKun); single 1024x1024 image, pre-warmed inference only (model load excluded)
8 s / image—2025-08-03github.com →
reportedSD 3.5 Large on Arc B580 12 GBfp8 · 1024x1024 · 20 steps
“SD 3.5 large fp8 (20step): GOOD (26s)”
Intel Arc B580 12GB, PyTorch 2.8 XPU in Docker (YanWenKun); single 1024x1024 image, pre-warmed inference only (model load excluded)
26 s / image—2025-08-03github.com →
reportedSDXL on Arc B580 12 GBsd_xl_base_1.0 · 1024x1024 · 20 steps
“100%|██| 20/20 [00:05<00:00, 3.96it/s]”
ComfyUI 'GPU Benchmark' thread: default workflow, SDXL 1.0 base, 1024x1024, seed 1, second run; Intel B580 Steel Legend OC 12 GB
3.96 it/s—2025-05-12github.com →
reportedWan 2.2 5B on Arc B580 12 GBWan 2.2 5B T2V (template) · 640x352
“Wan 2.2 5B Text to Video (640x352, 121 frames): FAST (60s, very poor quality)”
Arc B580 12GB, PyTorch 2.8 XPU; same post: 960x544/121f = 181s, 1280x704 OOM; steps not stated
60 s / clip (121 frames)—2025-08-03github.com →
reportedFLUX.1 dev on Radeon 8060S (Strix Halo) 96 GBflux1-dev full precision (23.8 GB) + t5xxl_fp16 · 1024x1024 · 20 steps
“Steady state | 3.64 s/it | 77.56 s”
AMD ROCm blog, Ryzen AI Max+ 395 / Radeon 8060S, 128 GB unified, Windows ComfyUI; first run 103.89 s; vendor-published measurement
77.56 s / image · 3.64 s/it—2026-07-14rocm.blogs.amd.com →
reportedLTX-2 on Radeon 8060S (Strix Halo) 96 GBLTX-2 BF16 (T2V) · 1280x720
“"workflow": "LTX2-T2V-BF16.json", ... "duration_seconds": 615.0017409324646”
kyuz0 Strix Halo toolbox benchmark; cold run incl. model load; resolution/frames from benchmark page; I2V = 616.16 s; steps not stated
615 s / clip (121 frames)—2026-02-13raw.githubusercontent.com →
reportedQwen-Image on Radeon 8060S (Strix Halo) 96 GBQwen-Image-2512 BF16 + 4-step Lightning LoRA · 1328x1328 · 4 steps
“"workflow": "Qwen-Image-2512-BF16-4-Step-LoRA.json", ... "duration_seconds": 75.37661480903625”
kyuz0 Strix Halo ComfyUI toolbox benchmark (Ryzen AI Max, ROCm); cold run incl. model load, flags --disable-mmap --gpu-only --disable-smart-memory --cache-none; resolution from ben
75.38 s / image—2026-02-13raw.githubusercontent.com →
reportedQwen-Image-Edit on Radeon 8060S (Strix Halo) 96 GBQwen-Image-Edit-2511 BF16 + 4-step LoRA · ~1.6MP (dynamic) · 4 steps
“"workflow": "Qwen-Image-Edit-2511-BF16-4-Step-LoRA.json", ... "duration_seconds": 112.70737218856812”
kyuz0 Strix Halo toolbox benchmark; cold run incl. model load; 20-step run = 667.24 s; date from run timestamp
112.71 s / image—2026-02-13raw.githubusercontent.com →
reportedWan 2.2 T2V on Radeon 8060S (Strix Halo) 96 GBWan 2.2 T2V A14B (high+low noise experts) · 640x640 · 4 steps
“Steady state | 197 s/it | 186 s/it | 26 min 51 s”
AMD ROCm blog, Ryzen AI Max+ 395; s/it for high-noise / low-noise expert; 26 min 51 s = 1611 s; first run 36 min 14 s; vendor-published
1611 s / clip (81 frames)—2026-07-14rocm.blogs.amd.com →
reportedSDXL on RTX 2070 Laptop 8 GBSDXL 1.0 base · 1024x1024 · 20 steps
“2070 rtx mobile 20/20 [00:19<00:00, 1.05it/s] ... Prompt executed in 24.22 seconds”
Thread benchmark SDXL 1024x1024 20 steps
24.22 s / image · 1.05 it/s—2024-03-04github.com →
reportedQwen-Image on RTX 3060 12 GB20 steps
“Qwen-ImageがRTX 3060(12GB)で動くと聞いて、早速ComfyUI版をお試し。確かに問題なく動いて、20stepでちょうど5分。”
'ちょうど5分' = exactly 5 minutes (converted to 300 s); ComfyUI version, file/resolution not stated; quote taken from search index (x.com not fetchable); date from tweet ID
300 s / image—2025-08-05x.com →
reportedWan 2.2 I2V on RTX 3060 12 GBvideo_wan2_2_14B_i2v template (quant not stated) · 720p
“53frame/16fpsで513.5秒”
ComfyUI on native Linux, PyTorch 2.8; 10GB offloaded via MultiGPU node; steps/LoRA not stated
513.5 s / clip (53 frames)—2025-11-26note.com →
reportedZ-Image Turbo on RTX 3060 12 GB1024x1024
“私の環境(グラボ:RTX 3060 12GB)では解像度1024*1024pxの画像を生成するのに大体45秒程度かかりました”
ComfyUI; approximate ('大体…程度'); file/steps not stated (article notes ~12GB model)
45 s / image—2025-12-05kurokumasoft.com →
reportedIllustrious / Pony on RTX 3060 Ti 8 GBIllustrious-XL-v2.0 · 1024x1024
“NVIDIA RTX 3060 Ti | 1.96it/s | 0.51s/it | CUDA 12.9 | ComfyUI (Unknown) | Ubuntu Server 24.04.2 LTS”
Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section
0.51 s/it—2025-11-29huggingface.co →
reportedKrea 2 on RTX 3060 Ti 8 GBfp8
“3060 ti here, 8GB VRAM, fp8 quant gets about 3.9 seconds per iteration”
Posted on Krea-2-Turbo repo; ComfyUI ('run ... in Comfy now'); resolution/steps not stated; year inferred from Krea 2 release (Jun 2026)
3.9 s/it—2026-07-01huggingface.co →
reportedSDXL on RTX 3070 Laptop 8 GBSDXL 1.0 base · 1024x1024 · 20 steps
“20/20 [00:11<00:00, 1.73it/s] ... Prompt executed in 15.44 seconds”
RTX 3070 Laptop GPU, Asus ZenBook Duo; thread benchmark SDXL 1024x1024 20 steps
15.44 s / image · 1.73 it/s—2024-04-22github.com →
reportedSDXL on RTX 3080 Ti 12 GBsd_xl_base_1.0 · 1024x1024 · 20 steps
“20/20 [00:04<00:00, 4.01it/s] Prompt executed in 5.60 second”
ComfyUI 'GPU Benchmark' thread: default workflow, SDXL 1.0 base, 1024x1024, seed 1, second run
5.6 s / image · 4.01 it/s—2025-04-30github.com →
reportedFLUX.1 dev on RTX 3090 24 GBfp8 (ComfyUI template)
“Nvidia 3090: 26s”
ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread
26 s / image—2025-07-22github.com →
reportedHiDream-I1 on RTX 3090 24 GBHiDream Full fp8 (T5 fp8, Llama3.1 fp8 scaled)
“ComfyUI Full example from their page at https://comfyanonymous.github.io/ComfyUI_examples/hidream/ runs at 115.09s on my RTX 3090, from cold start.”
HiDream-I1 Full via ComfyUI example workflow, Windows portable; cold start incl. load; resolution/steps not restated
115.09 s / image—2025-04-18huggingface.co →
reportedIllustrious / Pony on RTX 3090 24 GBIllustrious-XL-v2.0 · 1024x1024
“NVIDIA RTX 3090 | 4.00it/s | 0.25s/it | CUDA 12.9 | ComfyUI (Unknown) | Arch Linux”
Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section
0.25 s/it—2025-11-29huggingface.co →
reportedMiniMax H3 on RTX 3090 24 GBtransformer variant not stated; NVFP4 text encoder · 832x480
“render 23 min 17 s ... peak VRAM ~19.8 GB of 24 GB”
15.08 s clip at 24 fps with audio; needs --disable-pinned-memory on 32 GB RAM; 23 min 17 s = 1397 s; no date shown
1397 s / clip (362 frames)19.8 GB—github.com →
reportedSDXL on RTX 3090 24 GBsd_xl_base_1.0 · 1024x1024
“Prompt executed in 6.16 seconds”
ComfyUI 'GPU Benchmark' thread: default workflow, SDXL 1.0 base, 1024x1024, seed 1, second run
6.16 s / image—2024-03-06github.com →
reportedZ-Image Turbo on RTX 3090 24 GBTongyi-MAI/Z-Image-Turbo · 1024x1024
“NVIDIA RTX 3090 | 0.50it/s | 2.01s/it | CUDA 12.6 | ComfyUI (5151cff) | Arch Linux”
Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; Z-Image 1024px section; CFG 8 (above Turbo's us
2.01 s/it—2025-11-29huggingface.co →
reportedMiniMax H3 Pruned on RTX 4060 8 GBFL2VA pruned 20B INT8 ConvRot + Q2_K GGUF text encoder · 832x480 · 20 steps
“Denoising time: About 5.5 minutes”
Desktop RTX 4060 8GB, 96 GB RAM, WSL; 5.5 min denoise only; "Total first-run time: Just under 8 minutes"; a later 15-second run took about 15 minutes
330 s / clip (107 frames)—2026-08mountainmeadowsystems.com →
reportedIllustrious / Pony on RTX 4060 Laptop 8 GBWAI-Illustrious SDXL v16.0 · 1024x1024 · 20 steps
“1024x1024 | 1.47 | 13s | 15.81s | ~5.6GB”
ComfyUI, euler_ancestral/Karras CFG 5; columns it/s | KSampler | total | VRAM; no --lowvram; VRAM approximate
15.81 s / image · 1.47 it/s5.6 GB2026-02-26lilting.ch →
reportedWan 2.2 5B on RTX 4060 Laptop 8 GBWan2_2-TI2V-5B_fp8_e4m3fn_scaled_KJ · 480x480 · 30 steps
“30/30 [00:57<00:00, 1.93s/it] ... Prompt executed in 94.93 seconds”
Article says "RTX 4060 (8GB VRAM)", 32 GB RAM, Win 11; same author (lilting) documents this machine as an RTX 4060 Laptop in other posts; I2V; frames not stated; 50 steps = 113.93
94.93 s / clip · 1.93 s/it—2026-03-06lilting.ch →
reportedWan 2.2 I2V on RTX 4060 Laptop 8 GBWAN 2.2 14B Rapid distilled (all-in-one) · 480x480 · 4 steps
“4/4 [00:45<00:00, 11.46s/it] ... Prompt executed in 111.41 seconds”
Same machine note as above (8 GB, likely Laptop); 4,569 MB loaded on GPU with 11,067 MB offloaded (peak_vram = loaded weights, not measured peak)
111.41 s / clip (33 frames) · 11.46 s/it4.46 GB2026-03-06lilting.ch →
reportedFLUX.1 dev on RTX 4060 Ti 16 GBFP8
“Flux-Dev(FP8) … 2回目は72秒から51秒とあまり効果が感じられませんでした。”
ComfyUI; same PC before/after swapping RTX 4060 Ti 8GB -> 16GB; 2nd-run times; resolution/steps not stated; 51 s = 16GB card
51 s / image—2025-01-03note.com →
reportedFLUX.1 schnell on RTX 4060 Ti 16 GBFP8
“Flux-Schnell(FP8) … 2回目は24秒から11秒です。”
ComfyUI; same PC before/after swapping RTX 4060 Ti 8GB -> 16GB; 2nd-run times; resolution/steps not stated; 11 s = 16GB card
11 s / image—2025-01-03note.com →
reportedQwen-Image 2.1 on RTX 4060 Ti 16 GBQwen Image 2.1 INT8 Convrot (+Qwen3 VL 8B INT8 TE)
“Image generation: ~20 seconds ... Peak VRAM: ~15 GB”
ComfyUI, no CPU/disk offload; approximate values ('~'); editing ~60 s; resolution/steps not stated; date derived from '1 day ago' on 2026-09-25
20 s / image15 GB2026-09-24huggingface.co →
reportedFLUX.1 dev on RTX 4060 Ti 8 GBFP8
“Flux-Dev(FP8) … 2回目は72秒から51秒とあまり効果が感じられませんでした。”
ComfyUI; same PC before/after swapping RTX 4060 Ti 8GB -> 16GB; 2nd-run times; resolution/steps not stated; 72 s = 8GB card
72 s / image—2025-01-03note.com →
reportedFLUX.1 Krea on RTX 4060 Ti 8 GBFlux.1 Krea Dev CLIP+VAE FP8 (~20GB checkpoint) · 768x1024 · 20 steps
“すべて768x1024または1024x768、20ステップ、seed固定。… 合計(2回目以降) | 約40秒”
ComfyUI with async weight offloading to system RAM; sampling ~35 s; first run ~130 s; approximate ('約'); 768x1024 or 1024x768
40 s / image—2026-05-06hide10.com →
reportedFLUX.1 schnell on RTX 4060 Ti 8 GBFP8
“Flux-Schnell(FP8) … 2回目は24秒から11秒です。”
ComfyUI; same PC before/after swapping RTX 4060 Ti 8GB -> 16GB; 2nd-run times; resolution/steps not stated; 24 s = 8GB card
24 s / image—2025-01-03note.com →
reportedQwen-Image 2.1 on RTX 4070 12 GBQwen-Image-2.1 int8 convrot + qwen3vl_8b_int8 convrot TE + VAE bf16 · 1024x1024 · 25 steps
“1024x1024・25ステップ(公式ワークフロー相当) | 14.0 秒”
ComfyUI, RTX 4070 12GB with partial offload (17GB weights); same table: 12 steps 9.6 s, 2048x2048 20 steps 90.5 s; repo date not shown (Qwen-Image 2.1 era, 2026)
14 s / image——github.com →
reportedQwen-Image 2.1 on RTX 4070 12 GBQwen Image 2.1 INT8 ConvRot · 832x1248 · 25 steps
“「全体の所要時間」で見ると、平均で32.63秒から19.83秒へ(約39.2%短縮)… GPU全体のVRAMピーク:11.19 GiB → 11.28 GiB”
ComfyUI, RTX 4070 12GB, Euler/simple CFG 1; 32.63 s = standard total time (19.83 s with Spectrum speedup node); VRAM in GiB
32.63 s / image11.19 GB2026-09-24note.com →
reportedSDXL on RTX 4070 12 GBsd_xl_base_1.0 · 1024x1024 · 20 steps
“20/20 [00:06<00:00, 3.21it/s] Prompt executed in 7.13 seconds”
ComfyUI 'GPU Benchmark' thread: default workflow, SDXL 1.0 base, 1024x1024, seed 1, second run; GPU stated as 'RTX 4070 12Gb'
7.13 s / image · 3.21 it/s—2024-03-21github.com →
reportedMiniMax H3 Pruned on RTX 4070 Laptop 8 GBpruned int8 convrot (FL2VA/Ref2VA) · 832x480
“On a laptop with an RTX 4070 (8GB VRAM) and 32GB of RAM, we generated one 832×480, 15-second clip in 45 minutes.”
15-second clip, 24 fps, with audio; ComfyUI I2V/R2V workflows (which one not stated); 45 min = 2700 s; steps not stated
2700 s / clip—2026-08-04metallab.ai →
reportedFLUX.1 dev on RTX 4070 Super 12 GBQ4_0 GGUF
“1.9s/it with Q4_0”
RTX 4070 Super 12GB; same post: 2.6s/it with Q5_1, 1.3s/it with NF4; resolution not stated; early (Aug 2024) ComfyUI-GGUF
1.9 s/it—2024-08-17huggingface.co →
reportedFLUX.2 klein 4B on RTX 4070 Super 12 GBflux-2-klein-4b (distilled) + qwen_3_4b
“5 秒でした。”
ComfyUI 4B T2I template (distilled, 4 steps default); resolution not stated; base 4B T2I 37 s incl. model load; distilled edit 14-17 s
5 s / image—2026-01-20note.com →
reportedFLUX.1 dev on RTX 4080 16 GBfp8_e4m3fn weight_dtype · 1024x1024 · 20 steps
“weight_dtype (fp8_e4m3fn) with --fast (13sec)”
ComfyUI 'RTX 4090 benchmarks - FLUX model' thread; OP settings 1024x1024, 20 steps; Aug 2024 ComfyUI/PyTorch 2.5 dev; with --fast flag (same poster: 19sec without --fast, 28sec def
13 s / image—2024-08-23github.com →
reportedFLUX.1 dev on RTX 4090 24 GBfp8 (ComfyUI template)
“Prompt executed in 11.28 seconds”
ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread
11.28 s / image—2025-12-28github.com →
reportedFLUX.1 dev on RTX 4090 24 GBQ8_0 GGUF · 1024x1024 · 20 steps
“15 seconds at the fastest to 17 seconds at the slowest on my RTX 4090 with Euler 20 Steps for 1024x1024 images”
stated range 15-17 s; city96 ComfyUI-GGUF Q8
15 s / image—2024-08-25github.com →
reportedFLUX.1 dev on RTX 4090 24 GBfp8 (--fast) · 1024x1024 · 20 steps
“Prompt executed in 10.01 seconds”
ComfyUI 'RTX 4090 benchmarks - FLUX model' thread; OP settings 1024x1024, 20 steps; Aug 2024 ComfyUI/PyTorch 2.5 dev; FP8 with --fast, GPU at 2.52 GHz/875mV undervolt (9.07 s at 2.
10.01 s / image—2024-08-26github.com →
reportedHunyuanVideo 1.5 on RTX 4090 24 GB720p model · 848x480
“The 5 seconds video took 297s to generate so barely longer than on your end”
Replicated the 5090 poster's ComfyUI workflow (720p model at 848x480, 5 s @24fps); card undervolted (~5% slower per poster)
297 s / clip—2025-12-05huggingface.co →
reportedIllustrious / Pony on RTX 4090 24 GBIllustrious-XL-v2.0 · 1024x1024
“NVIDIA RTX 4090 | 7.00it/s | 0.14s/it | CUDA 12.9 | ComfyUI (Unknown) | Windows 11 24H2”
Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section
0.14 s/it—2025-11-29huggingface.co →
reportedMiniMax H3 Pruned on RTX 4090 24 GBminimax_h3_fl2va_pruned_int8_convrot + qwen3vl_32b nvfp4_awq · 832x480 · 8 steps
“5 s clip (832×480, 8 steps): ~7 min”
~7 min approx = 420 s for 5 s clip; VRAM ~6.6 GB during sampling, ~22.5 GB spike at model load; 15 s clip ~25-30 min; no date shown
420 s / clip22.5 GB—github.com →
reportedWan 2.1 I2V 480P on RTX 4090 24 GB30 steps
“100%|███| 30/30 [06:13<00:00, 12.46s/it]”
Kijai WanVideoWrapper wanvideo_480p_I2V_example_01.json; 32GB system RAM, process later 'Killed' (RAM); resolution/frames not stated
12.46 s/it—2025-02-26github.com →
reportedWan 2.2 I2V on RTX 4090 24 GBWan2.2-I2V-A14B High/Low Q6_K GGUF · 800x448 · 8 steps
“RTX 4090なら1分半ほど、RTX 5080はほぼ2分で480p解像度を5秒出力できます。”
Chimolog GPU review; ComfyUI 0.3.5x, Kijai-based workflow, 2+2+4 steps with Lightx2v/Lightning LoRAs; '1分半ほど' = about 1.5 min (converted to 90 s, approximate); exact values only in
90 s / clip (81 frames)—2025-08-28chimolog.co →
reportedFLUX.1 dev on RTX 5060 Ti 16 GBfp8 (ComfyUI template)
“Prompt executed in 25.71 seconds”
ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread; poster's log shows 16311
25.71 s / image—2025-08-04github.com →
reportedFLUX.2 dev on RTX 5060 Ti 16 GBflux2_dev_fp8mixed + mistral_3_small_flux2_fp8 · 1024x1024 · 20 steps
“FP8 t2i 142秒程度(20Steps)”
ComfyUI t2i at 1MP; approximate (程度)
142 s / image—2025-11-28note.com →
reportedFLUX.2 dev on RTX 5060 Ti 16 GBGGUF Q4_0 (flux2-dev-q4_0.gguf) · 1024x1024 · 20 steps
“GGUF Q4_0 t2i 180秒程度(20Steps)”
ComfyUI t2i at 1MP; approximate (程度); author notes GGUF slower than fp8 when spilling VRAM
180 s / image—2025-11-28note.com →
reportedIllustrious / Pony on RTX 5060 Ti 16 GBIllustrious-XL-v2.0 · 1024x1024
“NVIDIA RTX 5060 Ti | 2.60it/s | 0.39s/it | CUDA 12.9 | ComfyUI (Unknown) | Arch Linux”
Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section; row says 'RTX 5060 Ti', co
0.39 s/it—2025-11-29huggingface.co →
reportedWan 2.2 I2V on RTX 5060 Ti 16 GBvideo_wan2_2_14B_i2v template (quant not stated) · 720p
“RTX5060ti 16GBだと165.9秒なので、2〜3倍時間が必要ですが、落ちずに720p動画が生成できる事はたいしたものです。”
Comparison figure given in the RTX 3060 post, presumably same 720p/53-frame test (not explicitly restated); steps not stated
165.9 s / clip (53 frames)—2025-11-26note.com →
reportedZ-Image Turbo on RTX 5060 Ti 16 GBz_image_turbo_bf16.safetensors · 1328x1328 · 8 steps
“約 35秒かかりました。一度モデルを VRAM にロードした後、プロンプトを変えての再実行だと約 22秒で生成できます。”
ComfyUI; ~35 s first run incl. load, ~22 s warm; approximate ('約')
22 s / image—2025-11-30iwannacreateapps.com →
reportedKrea 2 on RTX 5070 Ti 16 GBKrea 2 Turbo (ComfyUI release, 17.8GB total download) · 1024x1024 · 8 steps
“generating a 1024x1024 pixel image took 16-25 seconds when the prompts were rewritten, and approximately 9-15 seconds when the same prompts were reused”
GIGAZINE review, ComfyUI; range only (9-15 s same prompt, 16-25 s new prompt incl. text encoding)
9–15 s / image—2026-06-24gigazine.net →
reportedFLUX.1 dev on RTX 5090 32 GBfp8 (ComfyUI template)
“Getting 8.78s at 2.38it/s for 3 runs.”
ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread; Inno3D RTX 5090 X3 OC
8.78 s / image · 2.38 it/s—2025-08-05github.com →
reportedHunyuanVideo 1.5 on RTX 5090 32 GB720p model · 848x480
“I was able to generate 5 seconds video in 284s at 24fps using the 720p model at 848*480”
ComfyUI; 5 s video at 24 fps; steps not stated
284 s / clip—2025-11-22huggingface.co →
reportedIllustrious / Pony on RTX 5090 32 GBIllustrious-XL-v2.0 · 1024x1024
“NVIDIA RTX 5090 | 8.95it/s | 0.11s/it | CUDA 12.8 | ComfyUI (Unknown) | Windows 11 24H2”
Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section
0.11 s/it—2025-11-29huggingface.co →
reportedLTX-2 on RTX 5090 32 GBLTX-2 19B NVFP4 · 720p
“generating 143 frames at 720p takes about 66 seconds end‑to‑end in ComfyUI”
Issue says this is ~30-40% slower than NVIDIA's expected 40-45 s; steps not stated
66 s / clip (143 frames)—2026-01-08github.com →
reportedMiniMax H3 Pruned on RTX 5090 32 GBminimax_h3_fl2va_pruned_nvfp4 · 864x480 · 10 steps
“175 s for a 864×480 ten-second clip, 26,914 MiB peak VRAM.”
ComfyUI 0.30.1, 500 W power cap; 26,914 MiB = 26.28 GiB; INT8-ConvRot = 185 s / 28,581 MiB
175 s / clip (243 frames)26.28 GB2026-08-04ai-muninn.com →
reportedZ-Image Turbo on RTX 5090 32 GBfp8_e4m3fn (per log)
“without SA: 0.95 seg”
'seg' = segundos (seconds) per image; SageAttention variants gave 0.87-0.91 s; resolution/steps not stated
0.95 s / image—2025-11-27github.com →
reportedSDXL on RX 7600 XT 16 GBSDXL 1.0 base · 1024x1024 · 20 steps
“20/20 [00:17<00:00, 1.13it/s] Prompt executed in 20.52 seconds”
ComfyUI 0.3.29 Zluda, Windows 11, driver 25.3.1; thread benchmark SDXL 1024x1024 20 steps
20.52 s / image · 1.13 it/s—2025-04-18github.com →
reportedLTX-2.3 on RX 7900 XTX 24 GBltx-2.3-22b-dev-fp8 + gemma_3_12B_it_fp4_mixed (I2V template) · 768x1280
“[INFO] Prompt executed in 197.38 seconds”
User QualiaSG, 64 GB RAM, ROCm 7.13, with --disable-dynamic-vram --disable-smart-memory --disable-pinned-memory (stalls without them); I2V template (distilled), 24 fps; comment dat
197.38 s / clip (121 frames)——github.com →
reportedWan 2.1 I2V 480P on RX 7900 XTX 24 GBwan2.1_i2v_480p_14B_fp8_scaled · 480x704 · 25 steps
“Wan2.1 i2v | 480×704 | 81 | 25 | ~40 min”
Windows 11, native ROCm 7.1, PYTORCH_NO_HIP_MEMORY_CACHING=1; approximate (~40 min = 2400 s); author notes slow vs CUDA
2400 s / clip (81 frames)—2026-04-05joshwaamein.github.io →
reportedSDXL on RX 9070 XT 16 GBSDXL · 1024x1024 · 20 steps
“| SDXL | 20 | 1024x1024 | 1.5it/s | 27.76s | Manual tiled VAE decoder |”
Ubuntu 24.04, ROCm 6.4.1, PyTorch nightly, TUNABLEOP + AOTriton experimental, --use-pytorch-cross-attention; without tiled VAE 1.49 it/s / 34.56 s; default VAE decode could OOM
27.76 s / image · 1.5 it/s—2025-05-30github.com →

Reported results are other people's numbers, copied as published, with a link. Settings, drivers and ComfyUI versions differ, so compare them with care. Send yours.