Qwen-Image 2.1, Krea 2 and Ming-Image all came out in 2026, and Comfy-Org publishes each of them in a new 8-bit format, int8_convrot. I downloaded the main files and timed them on my RTX 5060 Ti 16 GB with ComfyUI 0.38.1.
The result
Seconds per 1024×1024 image, at each model's usual step count (Qwen-Image 2.1: 25, Krea 2 Turbo: 8, Ming-Image: 12). Shorter is better. INT8 files are highlighted.
| Model | File | Size | Resolution, steps | Per image | Peak VRAM | Peak RAM |
|---|---|---|---|---|---|---|
| Qwen-Image 2.1 | qwen_image_2.1_int8_convrot.safetensors | 7.3 GB | 1024×1024, 25 | 17.2 s | 14.3 GB | 22.2 GB |
| Qwen-Image 2.1 | qwen_image_2.1_bf16.safetensors | 14.2 GB | 1024×1024, 25 | 39.7 s | 15.5 GB | 24.5 GB |
| Qwen-Image 2.1 | qwen_image_2.1_Q8_0.gguf | 7.6 GB | 1024×1024, 25 | 43.6 s | 14.8 GB | 21.9 GB |
| Krea 2 Turbo | krea2_turbo_int8_convrot.safetensors | 13.5 GB | 1024×1024, 8 | 10.9 s | 15.3 GB | 22.8 GB |
| Krea 2 Turbo | krea2_turbo_fp8_scaled.safetensors | 13.1 GB | 1024×1024, 8 | 17.7 s | 14.9 GB | 23.0 GB |
| Ming-Image | ming_image_0.1_design_int8_convrot.safetensors | 6.2 GB | 1024×1024, 12 | 7.4 s | 13.6 GB | 28.4 GB |
| Ming-Image | ming_image_0.1_design_bf16.safetensors | 12.3 GB | 1024×1024, 12 | 17.4 s | 15.3 GB | 31.3 GB |
| Ming-Image | ming_image_0.1_design_int8_convrot.safetensors | 6.2 GB | 2048×2048, 12 | 44.9 s | 14.4 GB | 26.9 GB |
Average of 2–3 timed runs after 1 warm-up run; the runs were within 0.05 s of each other. Peak VRAM is the memory in use on the whole card, including the text encoder and Windows. Peak RAM is system RAM in use on the whole PC (32 GB).
What it means
- Take the INT8 file. On every model it was the fastest. Qwen-Image 2.1: INT8 17.2 s, 16-bit 39.7 s, GGUF Q8_0 43.6 s. Krea 2: INT8 10.9 s, FP8 17.7 s. Ming-Image: INT8 7.4 s, 16-bit 17.4 s.
- GGUF Q8_0 was the slowest of all. The Qwen-Image 2.1 Q8_0 GGUF is about the same size as the INT8 file but took 43.6 s instead of 17.2 s: GGUF weights are unpacked at every step. On an NVIDIA card with current ComfyUI, use GGUF only when no native 8-bit file fits. Because of this result, the calculator on this site now suggests INT8 before Q8_0 on RTX 40 and RTX 50 cards.
- Qwen-Image 2.1 in 16-bit still works on 16 GB. The 14.2 GB file does not fit next to the text encoder, so ComfyUI streams part of it from system RAM — and it was still faster than the GGUF.
- Ming-Image needs RAM more than VRAM. Its text encoder is a whole language model: 19.5 GB even in INT8. With the 16-bit image model, system RAM peaked at 31.3 GB of my 32 GB; with INT8 at 28.4 GB. With 32 GB of RAM, close other programs; with 16 GB of RAM, expect heavy swapping.
- Ming-Image at 2048×2048 runs on 16 GB. The INT8 file made a 2048×2048 poster in 44.9 s — about 6 times the 1024 time for 4 times the pixels.
My setup
- GPU: NVIDIA GeForce RTX 5060 Ti 16 GB. System RAM: 32 GB. ComfyUI 0.38.1 portable, PyTorch 2.12 with CUDA 13.0, default options.
- Qwen-Image 2.1: text encoder Qwen3-VL 8B INT8, Qwen-Image 2.1 VAE, euler / simple, CFG 1.
- Krea 2 Turbo: text encoder Qwen3-VL 4B FP8, Qwen-Image VAE, euler / simple, CFG 1.
- Ming-Image 0.1 Design: text encoder Ling-Mini-2.0 INT8, Ming-Image VAE, euler / simple.
- Files: Comfy-Org (16-bit, FP8, INT8), Abiray (Qwen-Image 2.1 GGUF). Time is sampling plus VAE decode, measured through the ComfyUI API. Date: 2026-10-04.
The exact workflows I timed, with file links: free ComfyUI workflows.
Every model on this card: what runs on an RTX 5060 Ti 16 GB. More: Qwen-Image 2.1 · Krea 2 · Ming-Image · INT8, NVFP4, MXFP8 explained.