Datasheet for local AI49 models98 GPUsData read 2026-09-25
No adsNo tracking
Check my GPU →

FLUX.1 [dev] VRAM requirements

The 12B text-to-image model from Black Forest Labs that most local workflows are built around. Strong prompt following and readable text.

Released 2024-08Licence: FLUX.1 [dev] Non-Commercial LicenseSteps: 20–30Text encoder: T5-XXL (4.9 GB as FP8)
TypeImage
Parameters12B
Fits entirely from7 GB
8-bit or better from15 GB

01The files, and how much VRAM each needs

FileSizeNeededMin. VRAMQualitySource
16-bit23.8 GB26.1 GB27 GBthe original weightsblack-forest-labs/FLUX.1-dev →
Q8_012.7 GB15.0 GB16 GBpractically identical to the originalcity96/FLUX.1-dev-gguf →
FP811.9 GB14.2 GB15 GBpractically identical to the originalKijai/flux-fp8 →
Q6_K9.9 GB12.2 GB13 GBvery close to the originalcity96/FLUX.1-dev-gguf →
Q5_K_S8.3 GB10.6 GB11 GBclose; small differences in fine detailcity96/FLUX.1-dev-gguf →
Q4_K_S6.8 GB9.1 GB10 GBgood; some loss in fine detail and textcity96/FLUX.1-dev-gguf →
Q3_K_S5.2 GB7.5 GB8 GBnoticeable loss of detailcity96/FLUX.1-dev-gguf →
Q2_K4.0 GB6.3 GB7 GBheavy loss; a last resortcity96/FLUX.1-dev-gguf →

“Needed” = file + 1.5 GB working memory + 0.8 GB system reserve.

03Best GPU for FLUX.1 dev

The cheapest cards (by launch price) that run it well, and every card sorted by memory: best GPU for FLUX.1 dev → Planning bigger images or longer clips? Open the calculator →

04By graphics card

GPUVRAMVerdictBest fileNeeded
Desktop graphics cards
RTX 2060 6 GB6 GBOffload onlyQ3_K_S7.5 GB
RTX 3050 6 GB6 GBOffload onlyQ3_K_S7.5 GB
RTX 2070 Super 8 GB8 GBTightQ3_K_S7.5 GB
RTX 2080 Super 8 GB8 GBTightQ3_K_S7.5 GB
RTX 3050 8 GB8 GBTightQ3_K_S7.5 GB
RTX 3060 8 GB8 GBTightQ3_K_S7.5 GB
RTX 3060 Ti 8 GB8 GBTightQ3_K_S7.5 GB
RTX 3070 8 GB8 GBTightQ3_K_S7.5 GB
RTX 3070 Ti 8 GB8 GBTightQ3_K_S7.5 GB
RTX 4060 8 GB8 GBTightQ3_K_S7.5 GB
RTX 4060 Ti 8 GB8 GBTightQ3_K_S7.5 GB
RTX 5050 8 GB8 GBTightQ3_K_S7.5 GB
RTX 5060 8 GB8 GBTightQ3_K_S7.5 GB
RTX 5060 Ti 8 GB8 GBTightQ3_K_S7.5 GB
RX 7600 8 GB8 GBTightQ3_K_S7.5 GB
RX 9050 8 GB8 GBTightQ3_K_S7.5 GB
RX 9060 XT 8 GB8 GBTightQ3_K_S7.5 GB
Arc B570 10 GB10 GBRunsQ4_K_S9.1 GB
RTX 3080 10 GB10 GBRunsQ4_K_S9.1 GB
RTX 2080 Ti 11 GB11 GBRunsQ5_K_S10.6 GB
Arc B580 12 GB12 GBRunsQ5_K_S10.6 GB
RTX 2060 12 GB12 GBRunsQ5_K_S10.6 GB
RTX 3060 12 GB12 GBRunsQ5_K_S10.6 GB
RTX 3080 12 GB12 GBRunsQ5_K_S10.6 GB
RTX 3080 Ti 12 GB12 GBRunsQ5_K_S10.6 GB
RTX 4070 12 GB12 GBRunsQ5_K_S10.6 GB
RTX 4070 Super 12 GB12 GBRunsQ5_K_S10.6 GB
RTX 4070 Ti 12 GB12 GBRunsQ5_K_S10.6 GB
RTX 5070 12 GB12 GBRunsQ5_K_S10.6 GB
RX 7700 XT 12 GB12 GBRunsQ5_K_S10.6 GB
RX 9070 GRE 12 GB12 GBRunsQ5_K_S10.6 GB
Arc A770 16 GB16 GBRuns wellQ8_015.0 GB
RTX 4060 Ti 16 GB16 GBRuns wellFP814.2 GB
RTX 4070 Ti Super 16 GB16 GBRuns wellFP814.2 GB
RTX 4080 16 GB16 GBRuns wellFP814.2 GB
RTX 4080 Super 16 GB16 GBRuns wellFP814.2 GB
RTX 5060 Ti 16 GB16 GBRuns wellFP814.2 GB
RTX 5070 Ti 16 GB16 GBRuns wellFP814.2 GB
RTX 5080 16 GB16 GBRuns wellFP814.2 GB
RX 7600 XT 16 GB16 GBRuns wellQ8_015.0 GB
RX 7800 XT 16 GB16 GBRuns wellQ8_015.0 GB
RX 7900 GRE 16 GB16 GBRuns wellQ8_015.0 GB
RX 9060 XT 16 GB16 GBRuns wellFP814.2 GB
RX 9070 16 GB16 GBRuns wellFP814.2 GB
RX 9070 XT 16 GB16 GBRuns wellFP814.2 GB
RX 7900 XT 20 GB20 GBRuns wellQ8_015.0 GB
Arc Pro B60 24 GB24 GBRuns wellQ8_015.0 GB
RTX 3090 24 GB24 GBRuns wellQ8_015.0 GB
RTX 3090 Ti 24 GB24 GBRuns wellQ8_015.0 GB
RTX 4090 24 GB24 GBRuns wellFP814.2 GB
RX 7900 XTX 24 GB24 GBRuns wellQ8_015.0 GB
Arc Pro B70 32 GB32 GBRuns well16-bit26.1 GB
RTX 5090 32 GB32 GBRuns well16-bit26.1 GB
Laptop GPUs
RTX 3050 Laptop 4 GB4 GBOffload onlyQ4_K_S9.1 GB
RTX 3050 Ti Laptop 4 GB4 GBOffload onlyQ4_K_S9.1 GB
RTX 2060 Laptop 6 GB6 GBOffload onlyQ3_K_S7.5 GB
RTX 3050 Laptop 6 GB6 GBOffload onlyQ3_K_S7.5 GB
RTX 3060 Laptop 6 GB6 GBOffload onlyQ3_K_S7.5 GB
RTX 4050 Laptop 6 GB6 GBOffload onlyQ3_K_S7.5 GB
RTX 2070 Laptop 8 GB8 GBTightQ3_K_S7.5 GB
RTX 2070 Super Laptop 8 GB8 GBTightQ3_K_S7.5 GB
RTX 2080 Laptop 8 GB8 GBTightQ3_K_S7.5 GB
RTX 2080 Super Laptop 8 GB8 GBTightQ3_K_S7.5 GB
RTX 3070 Laptop 8 GB8 GBTightQ3_K_S7.5 GB
RTX 3070 Ti Laptop 8 GB8 GBTightQ3_K_S7.5 GB
RTX 3080 Laptop 8 GB8 GBTightQ3_K_S7.5 GB
RTX 4060 Laptop 8 GB8 GBTightQ3_K_S7.5 GB
RTX 4070 Laptop 8 GB8 GBTightQ3_K_S7.5 GB
RTX 5050 Laptop 8 GB8 GBTightQ3_K_S7.5 GB
RTX 5060 Laptop 8 GB8 GBTightQ3_K_S7.5 GB
RTX 5070 Laptop 8 GB8 GBTightQ3_K_S7.5 GB
RX 7600M 8 GB8 GBTightQ3_K_S7.5 GB
RX 7600M XT 8 GB8 GBTightQ3_K_S7.5 GB
RX 7600S 8 GB8 GBTightQ3_K_S7.5 GB
RX 7700S 8 GB8 GBTightQ3_K_S7.5 GB
RTX 4080 Laptop 12 GB12 GBRunsQ5_K_S10.6 GB
RTX 5070 Laptop 12 GB12 GBRunsQ5_K_S10.6 GB
RTX 5070 Ti Laptop 12 GB12 GBRunsQ5_K_S10.6 GB
RX 7800M 12 GB12 GBRunsQ5_K_S10.6 GB
RTX 3080 Laptop 16 GB16 GBRuns wellQ8_015.0 GB
RTX 3080 Ti Laptop 16 GB16 GBRuns wellQ8_015.0 GB
RTX 4090 Laptop 16 GB16 GBRuns wellFP814.2 GB
RTX 5080 Laptop 16 GB16 GBRuns wellFP814.2 GB
RX 7900M 16 GB16 GBRuns wellQ8_015.0 GB
RTX 5090 Laptop 24 GB24 GBRuns wellFP814.2 GB
Unified memory
Radeon 8060S (Strix Halo) 96 GB96 GBRuns well16-bit26.1 GB
Apple Silicon Macs (by memory)
Mac 16 GB12.7 GBRunsQ6_K12.2 GB
Mac 18 GB14.4 GBRunsQ6_K12.2 GB
Mac 24 GB19.6 GBRuns wellQ8_015.0 GB
Mac 32 GB26.8 GBRuns well16-bit26.1 GB
Mac 36 GB30.2 GBRuns well16-bit26.1 GB
Mac 48 GB40.2 GBRuns well16-bit26.1 GB
Mac 64 GB55.7 GBRuns well16-bit26.1 GB
Mac 96 GB85 GBRuns well16-bit26.1 GB
Mac 128 GB115.4 GBRuns well16-bit26.1 GB
Mac 192 GB175.4 GBRuns well16-bit26.1 GB
Mac 256 GB236.9 GBRuns well16-bit26.1 GB
Mac 512 GB498.1 GBRuns well16-bit26.1 GB

On RTX 40/50 GPUs the FP8 file is preferred over Q8_0 when both fit (hardware FP8). All verdicts are calculated; see how the numbers work.

05Text encoder, VAE and other files

FileFolderSizeWhen
CLIP-L
clip_l.safetensors
models/text_encoders0.2 GBrequiredDownload →
T5-XXL FP8
t5xxl_fp8_e4m3fn.safetensors
models/text_encoders4.9 GBtext encoder · smaller, recommendedDownload →
T5-XXL FP16
t5xxl_fp16.safetensors
models/text_encoders9.8 GBtext encoder · alternativeDownload →
FLUX.1 VAE (ae)
ae.safetensors
models/vae0.3 GBrequiredDownload →

The files the official ComfyUI workflows load next to the model. Sizes read from Hugging Face (2026-09-25). Every GPU page for this model lists the exact set to download for that card, with the total.

T5-XXL: 9.8 GB as 16-bit, 4.9 GB as FP8, 2.9 GB as GGUF Q4_K_M (plus CLIP-L (0.25 GB)). ComfyUI encodes the prompt first and can push the encoder out of VRAM before sampling, so it does not have to fit together with the model. On 16 GB the FP8 encoder fits on its own, so prompt encoding stays fast.

06Where the files go in ComfyUI

FileFolderLoader node
Diffusion model (.safetensors: 16-bit, FP8, INT8)ComfyUI/models/diffusion_modelsLoad Diffusion Model
GGUF file (.gguf)ComfyUI/models/unetUnet Loader (GGUF) — from the ComfyUI-GGUF node pack
Text encoderComfyUI/models/text_encodersLoad CLIP / DualCLIPLoader (or the GGUF versions)
VAEComfyUI/models/vaeLoad VAE

Standard ComfyUI folders. After copying files, press R in ComfyUI (or restart it) to refresh the lists. Some uploads need their uploader's own loader node — see the notes above.

07AMD, Intel and NVIDIA: which file types are fast

File typeRTX 50RTX 40RTX 30 / 20RX 9000RX 7000/6000 · Strix HaloIntel Arc
16-bitRunsRunsRunsRunsRunsRuns
FP8Native FP8Native FP8No FP8 speed-upNative FP8No FP8 speed-upNo FP8 speed-up
GGUFRuns (GGUF node)Runs (GGUF node)Runs (GGUF node)Runs (GGUF node)Runs (GGUF node)Runs (GGUF node)

Every file type loads on every listed GPU, so the memory verdicts apply to all of them. What differs is speed: FP8 maths needs RTX 40/50 or RX 9000 (with ROCm 6.4+ and PyTorch 2.7+); ComfyUI's INT8 maths runs on NVIDIA and AMD, not on Intel; GGUF is unpacked on the fly on any GPU, which costs some speed. NVFP4 files are fast only on RTX 50. AMD runs ComfyUI on Windows through ROCm, Intel through PyTorch XPU; some custom nodes are NVIDIA-only. Source: ComfyUI model_management.py. AMD and Intel guide →

08Measured and reported results

LabelGPUSetupResultPeak VRAMDateSource
reportedArc A770 16 GBfp8 (ComfyUI template) · 20 steps
“20/20 [00:46<00:00, 2.33s/it] Prompt executed in 47.13 seconds”
'GPU Benchmark Flux DEV fp8' thread; Intel A770 on Fedora Linux, PyTorch 2.3.110+xpu (53.26 s with PyTorch nightly); resolution not stated
47.13 s / image · 2.33 s/it—2025-07-24github.com →
reportedArc B580 12 GBfp8 · 1024x1024 · 20 steps
“Flux1 dev fp8 (20step): GOOD (35s)”
Intel Arc B580 12GB, PyTorch 2.8 XPU in Docker (YanWenKun); single 1024x1024 image, pre-warmed inference only (model load excluded)
35 s / image—2025-08-03github.com →
reportedRadeon 8060S (Strix Halo) 96 GBflux1-dev full precision (23.8 GB) + t5xxl_fp16 · 1024x1024 · 20 steps
“Steady state | 3.64 s/it | 77.56 s”
AMD ROCm blog, Ryzen AI Max+ 395 / Radeon 8060S, 128 GB unified, Windows ComfyUI; first run 103.89 s; vendor-published measurement
77.56 s / image · 3.64 s/it—2026-07-14rocm.blogs.amd.com →
reportedRTX 3090 24 GBfp8 (ComfyUI template)
“Nvidia 3090: 26s”
ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread
26 s / image—2025-07-22github.com →
reportedRTX 4060 Ti 16 GBFP8
“Flux-Dev(FP8) … 2回目は72秒から51秒とあまり効果が感じられませんでした。”
ComfyUI; same PC before/after swapping RTX 4060 Ti 8GB -> 16GB; 2nd-run times; resolution/steps not stated; 51 s = 16GB card
51 s / image—2025-01-03note.com →
reportedRTX 4060 Ti 8 GBFP8
“Flux-Dev(FP8) … 2回目は72秒から51秒とあまり効果が感じられませんでした。”
ComfyUI; same PC before/after swapping RTX 4060 Ti 8GB -> 16GB; 2nd-run times; resolution/steps not stated; 72 s = 8GB card
72 s / image—2025-01-03note.com →
reportedRTX 4070 Super 12 GBQ4_0 GGUF
“1.9s/it with Q4_0”
RTX 4070 Super 12GB; same post: 2.6s/it with Q5_1, 1.3s/it with NF4; resolution not stated; early (Aug 2024) ComfyUI-GGUF
1.9 s/it—2024-08-17huggingface.co →
reportedRTX 4080 16 GBfp8_e4m3fn weight_dtype · 1024x1024 · 20 steps
“weight_dtype (fp8_e4m3fn) with --fast (13sec)”
ComfyUI 'RTX 4090 benchmarks - FLUX model' thread; OP settings 1024x1024, 20 steps; Aug 2024 ComfyUI/PyTorch 2.5 dev; with --fast flag (same poster: 19sec without --fast, 28sec def
13 s / image—2024-08-23github.com →
reportedRTX 4090 24 GBfp8 (ComfyUI template)
“Prompt executed in 11.28 seconds”
ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread
11.28 s / image—2025-12-28github.com →
reportedRTX 4090 24 GBQ8_0 GGUF · 1024x1024 · 20 steps
“15 seconds at the fastest to 17 seconds at the slowest on my RTX 4090 with Euler 20 Steps for 1024x1024 images”
stated range 15-17 s; city96 ComfyUI-GGUF Q8
15 s / image—2024-08-25github.com →
reportedRTX 4090 24 GBfp8 (--fast) · 1024x1024 · 20 steps
“Prompt executed in 10.01 seconds”
ComfyUI 'RTX 4090 benchmarks - FLUX model' thread; OP settings 1024x1024, 20 steps; Aug 2024 ComfyUI/PyTorch 2.5 dev; FP8 with --fast, GPU at 2.52 GHz/875mV undervolt (9.07 s at 2.
10.01 s / image—2024-08-26github.com →
reportedRTX 5060 Ti 16 GBfp8 (ComfyUI template)
“Prompt executed in 25.71 seconds”
ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread; poster's log shows 16311
25.71 s / image—2025-08-04github.com →
reportedRTX 5090 32 GBfp8 (ComfyUI template)
“Getting 8.78s at 2.38it/s for 3 runs.”
ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread; Inno3D RTX 5090 X3 OC
8.78 s / image · 2.38 it/s—2025-08-05github.com →

Reported results are other people's numbers, copied as published, with a link. Settings, drivers and ComfyUI versions differ, so compare them with care. Send yours.

09Training a LoRA for FLUX.1 dev

TrainerVRAMTypeSettings and quoteSource
sd-scripts8 GBstated minimum--fp8_base, --blocks_to_swap 28, fp8 T5XXL recommended; table also: 24GB batch 2, 16GB batch 1 + swap, 12GB swap 16 + AdamW8bit, 10GB swap 22
“8GB VRAM: Use --blocks_to_swap 28, recommend fp8 format for T5XXL”
github.com →
SimpleTuner10 GBstated minimumRank-16 LoRA; NF4 base ~9 GB (int8 ~18 GB, int4 ~13 GB, unquantised ~30 GB); lowest config 512px, batch 1, Lion8bit paged; 1024px needs >=12 GB. Guide example uses FLUX.1 Krea
“the absolute minimum is a single 3080 10G”
github.com →
fluxgym12 GBstated minimum12G preset: kohya sd-scripts backend, --fp8_base, Adafactor, --split_mode, train_blocks=single, gradient checkpointing, cached TE outputs; default 512px, rank 4
“Dead simple web UI for training FLUX LoRA with LOW VRAM (12GB/16GB/20GB) support.”
github.com →
OneTrainer12 GBstated minimumWiki Flux page; no specific settings given (NF4/fp8 mentioned nearby); 8GB possible with GGUF
“It is possible to train a Flux Lora on a GPU with 12GB of VRAM.”
github.com →
ai-toolkit24 GBstated minimumHistorical README requirement (removed from README on 2026-03-31 in commit ad474e3); 8-bit quantized base, low_vram flag if GPU drives monitors
“You currently need a GPU with at least 24GB of VRAM to train FLUX.1.”
github.com →
ai-toolkit24 GBexample runCurrent official example config: rank 16, batch 1, 512/768/1024 buckets, gradient checkpointing, quantize: true (8-bit); README still points to this file
“Copy the example config file located at config/examples/train_lora_flux_24gb.yaml”
github.com →

reported Figures as stated by each trainer's own documentation or official example configs, read 2026-09-25. “Stated minimum” = the docs call it a minimum; “example run” = a config or measured run at that size. They differ a lot because of settings: an 8-bit or 4-bit base model, block swapping and lower resolution all cut memory. All models →

10Every file tracked for FLUX.1 dev

FileTypeSizeRepo
flux1-dev-fp8-e4m3fn.safetensorsFP811.90 GBKijai/flux-fp8 →
flux1-dev-fp8-e5m2.safetensorsFP811.90 GBKijai/flux-fp8 →
flux1-dev.safetensorsBF1623.80 GBblack-forest-labs/FLUX.1-dev →
flux1-dev-F16.ggufF1623.80 GBcity96/FLUX.1-dev-gguf →
flux1-dev-Q8_0.ggufQ8_012.71 GBcity96/FLUX.1-dev-gguf →
flux1-dev-Q6_K.ggufQ6_K9.86 GBcity96/FLUX.1-dev-gguf →
flux1-dev-Q5_1.ggufQ5_19.01 GBcity96/FLUX.1-dev-gguf →
flux1-dev-Q5_K_S.ggufQ5_K_S8.29 GBcity96/FLUX.1-dev-gguf →
flux1-dev-Q5_0.ggufQ5_08.27 GBcity96/FLUX.1-dev-gguf →
flux1-dev-Q4_1.ggufQ4_17.53 GBcity96/FLUX.1-dev-gguf →
flux1-dev-Q4_K_S.ggufQ4_K_S6.81 GBcity96/FLUX.1-dev-gguf →
flux1-dev-Q4_0.ggufQ4_06.79 GBcity96/FLUX.1-dev-gguf →
flux1-dev-Q3_K_S.ggufQ3_K_S5.23 GBcity96/FLUX.1-dev-gguf →
flux1-dev-Q2_K.ggufQ2_K4.03 GBcity96/FLUX.1-dev-gguf →

Text encoders: clip_l + T5-XXL. VAE ae.safetensors 335304388 bytes (BFL repo).