Datasheet for local AI50 models99 GPUsData read 2026-10-01
No adsNo tracking
Check my GPU →

FLUX.2 klein 4B and 9B on a 16 GB RTX 5060 Ti: timed

FLUX.2 klein 4B in FP8 made a 1024×1024 image in 2.9 s on my RTX 5060 Ti. The 9B model took 8.5 s as GGUF Q8_0 — and Q4 was slower, not faster.

G.04MeasuredUpdated 2026-10-01

FLUX.2 klein is the small, fast branch of FLUX.2: a 4B model under the Apache 2.0 licence and a 9B model under a non-commercial licence, both distilled to about 4 steps. I timed both on my RTX 5060 Ti 16 GB.

The result

Seconds per 1024×1024 image, 4 steps. Shorter is better.

klein 4B FP82.9 s
klein 4B 16-bit3.9 s
klein 4B Q8_04.3 s
klein 9B Q8_08.5 s
klein 9B Q4_K_M9.8 s
Model and fileFileSizePer imagePeak VRAM
klein 4B FP8flux-2-klein-4b-fp8.safetensors4.1 GB2.9 s12.2 GB
klein 4B 16-bitflux-2-klein-4b.safetensors7.8 GB3.9 s14.1 GB
klein 4B Q8_0flux-2-klein-4b-Q8_0.gguf4.3 GB4.3 s12.4 GB
klein 9B Q8_0flux-2-klein-9b-Q8_0.gguf10.0 GB8.5 s13.5 GB
klein 9B Q4_K_Mflux-2-klein-9b-Q4_K_M.gguf5.9 GB9.8 s13.5 GB

Average of 3 timed runs after 1 warm-up run. Peak VRAM is the memory in use on the whole card during the run, including the text encoder and Windows.

What it means

  • klein 4B is extremely fast. 2.9 seconds per image in FP8. That is fast enough to try prompts almost live.
  • FP8 beats 16-bit and GGUF on RTX 40/50. FP8 was the fastest of the three 4B files; Q8_0 GGUF was the slowest, because GGUF weights are unpacked at every step.
  • For the 9B, Q4 was slower than Q8. Q4_K_M took 9.8 s against 8.5 s for Q8_0. A smaller GGUF only helps when the bigger one does not fit — on 16 GB, Q8_0 fits, so take it.
  • Licence matters for selling images. klein 4B is Apache 2.0. klein 9B uses the FLUX non-commercial licence; check it before using results commercially.

My setup

  • GPU: NVIDIA GeForce RTX 5060 Ti 16 GB. System RAM: 32 GB. ComfyUI 0.25.0 portable, PyTorch 2.12 with CUDA 13.0, default options.
  • Text encoders: Qwen3 4B FP8 (for 4B) and Qwen3 8B FP8 (for 9B). VAE: flux2-vae. Sampler euler with the Flux2Scheduler, CFG 1.
  • Files: Comfy-Org (16-bit 4B), black-forest-labs (FP8 4B), unsloth (GGUF).
  • Date: 2026-10-01.

The exact workflows I timed, with file links: free ComfyUI workflows.

Every model on this card: what runs on an RTX 5060 Ti 16 GB. More: FLUX.2 klein 4B · FLUX.2 klein 9B.