FLUX.2 klein is the small, fast branch of FLUX.2: a 4B model under the Apache 2.0 licence and a 9B model under a non-commercial licence, both distilled to about 4 steps. I timed both on my RTX 5060 Ti 16 GB.
The result
Seconds per 1024×1024 image, 4 steps. Shorter is better.
klein 4B FP82.9 s
klein 4B 16-bit3.9 s
klein 4B Q8_04.3 s
klein 9B Q8_08.5 s
klein 9B Q4_K_M9.8 s
| Model and file | File | Size | Per image | Peak VRAM |
|---|---|---|---|---|
| klein 4B FP8 | flux-2-klein-4b-fp8.safetensors | 4.1 GB | 2.9 s | 12.2 GB |
| klein 4B 16-bit | flux-2-klein-4b.safetensors | 7.8 GB | 3.9 s | 14.1 GB |
| klein 4B Q8_0 | flux-2-klein-4b-Q8_0.gguf | 4.3 GB | 4.3 s | 12.4 GB |
| klein 9B Q8_0 | flux-2-klein-9b-Q8_0.gguf | 10.0 GB | 8.5 s | 13.5 GB |
| klein 9B Q4_K_M | flux-2-klein-9b-Q4_K_M.gguf | 5.9 GB | 9.8 s | 13.5 GB |
Average of 3 timed runs after 1 warm-up run. Peak VRAM is the memory in use on the whole card during the run, including the text encoder and Windows.
What it means
- klein 4B is extremely fast. 2.9 seconds per image in FP8. That is fast enough to try prompts almost live.
- FP8 beats 16-bit and GGUF on RTX 40/50. FP8 was the fastest of the three 4B files; Q8_0 GGUF was the slowest, because GGUF weights are unpacked at every step.
- For the 9B, Q4 was slower than Q8. Q4_K_M took 9.8 s against 8.5 s for Q8_0. A smaller GGUF only helps when the bigger one does not fit — on 16 GB, Q8_0 fits, so take it.
- Licence matters for selling images. klein 4B is Apache 2.0. klein 9B uses the FLUX non-commercial licence; check it before using results commercially.
My setup
- GPU: NVIDIA GeForce RTX 5060 Ti 16 GB. System RAM: 32 GB. ComfyUI 0.25.0 portable, PyTorch 2.12 with CUDA 13.0, default options.
- Text encoders: Qwen3 4B FP8 (for 4B) and Qwen3 8B FP8 (for 9B). VAE: flux2-vae. Sampler euler with the Flux2Scheduler, CFG 1.
- Files: Comfy-Org (16-bit 4B), black-forest-labs (FP8 4B), unsloth (GGUF).
- Date: 2026-10-01.
The exact workflows I timed, with file links: free ComfyUI workflows.
Every model on this card: what runs on an RTX 5060 Ti 16 GB. More: FLUX.2 klein 4B · FLUX.2 klein 9B.