SDXL 1.0 VRAM requirements
The 2023 classic. One 6.9 GB checkpoint with the text encoders and VAE inside, and the biggest library of LoRAs and fine-tunes.
01The files, and how much VRAM each needs
| File | Size | Needed | Min. VRAM | Quality | Source |
|---|---|---|---|---|---|
| 16-bit | 6.9 GB | 7.1 GB | 8 GB | the original weights | stabilityai/stable-diffusion-xl-base-1.0 → |
“Needed” = file + 1.2 GB working memory + 0.8 GB system reserve. The 6.94 GB checkpoint also holds the two text encoders and the VAE. During sampling only the UNet (about 5.1 GB at 16-bit, 2.6B parameters) has to sit in VRAM, so that is the figure used for the 16-bit file.
02By amount of VRAM
03Best GPU for SDXL
The cheapest cards (by launch price) that run it well, and every card sorted by memory: best GPU for SDXL → Planning bigger images or longer clips? Open the calculator →
04By graphics card
| GPU | VRAM | Verdict | Best file | Needed |
|---|---|---|---|---|
| Desktop graphics cards | ||||
| RTX 2060 6 GB | 6 GB | Offload only | 16-bit | 7.1 GB |
| RTX 3050 6 GB | 6 GB | Offload only | 16-bit | 7.1 GB |
| RTX 2070 Super 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 2080 Super 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3050 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3060 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3060 Ti 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3070 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3070 Ti 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4060 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4060 Ti 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5050 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5060 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5060 Ti 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RX 7600 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RX 9050 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RX 9060 XT 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| Arc B570 10 GB | 10 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3080 10 GB | 10 GB | Runs well | 16-bit | 7.1 GB |
| RTX 2080 Ti 11 GB | 11 GB | Runs well | 16-bit | 7.1 GB |
| Arc B580 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RTX 2060 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3060 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3080 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3080 Ti 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4070 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4070 Super 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4070 Ti 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5070 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RX 7700 XT 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RX 9070 GRE 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| Arc A770 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4060 Ti 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4070 Ti Super 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4080 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4080 Super 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5060 Ti 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5070 Ti 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5080 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RX 7600 XT 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RX 7800 XT 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RX 7900 GRE 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RX 9060 XT 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RX 9070 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RX 9070 XT 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RX 7900 XT 20 GB | 20 GB | Runs well | 16-bit | 7.1 GB |
| Arc Pro B60 24 GB | 24 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3090 24 GB | 24 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3090 Ti 24 GB | 24 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4090 24 GB | 24 GB | Runs well | 16-bit | 7.1 GB |
| RX 7900 XTX 24 GB | 24 GB | Runs well | 16-bit | 7.1 GB |
| Arc Pro B70 32 GB | 32 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5090 32 GB | 32 GB | Runs well | 16-bit | 7.1 GB |
| Laptop GPUs | ||||
| RTX 3050 Laptop 4 GB | 4 GB | Offload only | 16-bit | 7.1 GB |
| RTX 3050 Ti Laptop 4 GB | 4 GB | Offload only | 16-bit | 7.1 GB |
| RTX 2060 Laptop 6 GB | 6 GB | Offload only | 16-bit | 7.1 GB |
| RTX 3050 Laptop 6 GB | 6 GB | Offload only | 16-bit | 7.1 GB |
| RTX 3060 Laptop 6 GB | 6 GB | Offload only | 16-bit | 7.1 GB |
| RTX 4050 Laptop 6 GB | 6 GB | Offload only | 16-bit | 7.1 GB |
| RTX 2070 Laptop 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 2070 Super Laptop 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 2080 Laptop 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 2080 Super Laptop 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3070 Laptop 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3070 Ti Laptop 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3080 Laptop 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4060 Laptop 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4070 Laptop 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5050 Laptop 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5060 Laptop 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5070 Laptop 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RX 7600M 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RX 7600M XT 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RX 7600S 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RX 7700S 8 GB | 8 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4080 Laptop 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5070 Laptop 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5070 Ti Laptop 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RX 7800M 12 GB | 12 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3080 Laptop 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RTX 3080 Ti Laptop 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RTX 4090 Laptop 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5080 Laptop 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RX 7900M 16 GB | 16 GB | Runs well | 16-bit | 7.1 GB |
| RTX 5090 Laptop 24 GB | 24 GB | Runs well | 16-bit | 7.1 GB |
| Unified memory | ||||
| Radeon 8060S (Strix Halo) 96 GB | 96 GB | Runs well | 16-bit | 7.1 GB |
| Apple Silicon Macs (by memory) | ||||
| Mac 16 GB | 12.7 GB | Runs well | 16-bit | 7.1 GB |
| Mac 18 GB | 14.4 GB | Runs well | 16-bit | 7.1 GB |
| Mac 24 GB | 19.6 GB | Runs well | 16-bit | 7.1 GB |
| Mac 32 GB | 26.8 GB | Runs well | 16-bit | 7.1 GB |
| Mac 36 GB | 30.2 GB | Runs well | 16-bit | 7.1 GB |
| Mac 48 GB | 40.2 GB | Runs well | 16-bit | 7.1 GB |
| Mac 64 GB | 55.7 GB | Runs well | 16-bit | 7.1 GB |
| Mac 96 GB | 85 GB | Runs well | 16-bit | 7.1 GB |
| Mac 128 GB | 115.4 GB | Runs well | 16-bit | 7.1 GB |
| Mac 192 GB | 175.4 GB | Runs well | 16-bit | 7.1 GB |
| Mac 256 GB | 236.9 GB | Runs well | 16-bit | 7.1 GB |
| Mac 512 GB | 498.1 GB | Runs well | 16-bit | 7.1 GB |
On RTX 40/50 GPUs the FP8 file is preferred over Q8_0 when both fit (hardware FP8). All verdicts are calculated; see how the numbers work.
05Text encoder, VAE and other files
The text encoder is inside the checkpoint, so there is nothing extra to download.
06Where the files go in ComfyUI
| File | Folder | Loader node |
|---|---|---|
| The checkpoint (.safetensors) | ComfyUI/models/checkpoints | Load Checkpoint |
Standard ComfyUI folders. After copying files, press R in ComfyUI (or restart it) to refresh the lists. Some uploads need their uploader's own loader node — see the notes above.
07AMD, Intel and NVIDIA: which file types are fast
| File type | RTX 50 | RTX 40 | RTX 30 / 20 | RX 9000 | RX 7000/6000 · Strix Halo | Intel Arc |
|---|---|---|---|---|---|---|
| 16-bit | Runs | Runs | Runs | Runs | Runs | Runs |
Every file type loads on every listed GPU, so the memory verdicts apply to all of them. What differs is speed: FP8 maths needs RTX 40/50 or RX 9000 (with ROCm 6.4+ and PyTorch 2.7+); ComfyUI's INT8 maths runs on NVIDIA and AMD, not on Intel; GGUF is unpacked on the fly on any GPU, which costs some speed. NVFP4 files are fast only on RTX 50. AMD runs ComfyUI on Windows through ROCm, Intel through PyTorch XPU; some custom nodes are NVIDIA-only. Source: ComfyUI model_management.py. AMD and Intel guide →
08Measured and reported results
| Label | GPU | Setup | Result | Peak VRAM | Date | Source |
|---|---|---|---|---|---|---|
| reported | Arc A770 16 GB | SDXL 1.0 base · 1024x1024 · 20 steps “20/20 [00:14<00:00, 1.36it/s] ... Prompt executed in 23.44 seconds” ACER A770 16GB; thread benchmark = SDXL 1024x1024 20 steps seed 1; later runs 1.53 it/s / 14.28 s | 23.44 s / image · 1.36 it/s | — | 2025-07-10 | github.com → |
| reported | Arc B580 12 GB | sd_xl_base_1.0 · 1024x1024 · 20 steps “100%|██| 20/20 [00:05<00:00, 3.96it/s]” ComfyUI 'GPU Benchmark' thread: default workflow, SDXL 1.0 base, 1024x1024, seed 1, second run; Intel B580 Steel Legend OC 12 GB | 3.96 it/s | — | 2025-05-12 | github.com → |
| reported | RTX 2070 Laptop 8 GB | SDXL 1.0 base · 1024x1024 · 20 steps “2070 rtx mobile 20/20 [00:19<00:00, 1.05it/s] ... Prompt executed in 24.22 seconds” Thread benchmark SDXL 1024x1024 20 steps | 24.22 s / image · 1.05 it/s | — | 2024-03-04 | github.com → |
| reported | RTX 3070 Laptop 8 GB | SDXL 1.0 base · 1024x1024 · 20 steps “20/20 [00:11<00:00, 1.73it/s] ... Prompt executed in 15.44 seconds” RTX 3070 Laptop GPU, Asus ZenBook Duo; thread benchmark SDXL 1024x1024 20 steps | 15.44 s / image · 1.73 it/s | — | 2024-04-22 | github.com → |
| reported | RTX 3080 Ti 12 GB | sd_xl_base_1.0 · 1024x1024 · 20 steps “20/20 [00:04<00:00, 4.01it/s] Prompt executed in 5.60 second” ComfyUI 'GPU Benchmark' thread: default workflow, SDXL 1.0 base, 1024x1024, seed 1, second run | 5.6 s / image · 4.01 it/s | — | 2025-04-30 | github.com → |
| reported | RTX 3090 24 GB | sd_xl_base_1.0 · 1024x1024 “Prompt executed in 6.16 seconds” ComfyUI 'GPU Benchmark' thread: default workflow, SDXL 1.0 base, 1024x1024, seed 1, second run | 6.16 s / image | — | 2024-03-06 | github.com → |
| reported | RTX 4070 12 GB | sd_xl_base_1.0 · 1024x1024 · 20 steps “20/20 [00:06<00:00, 3.21it/s] Prompt executed in 7.13 seconds” ComfyUI 'GPU Benchmark' thread: default workflow, SDXL 1.0 base, 1024x1024, seed 1, second run; GPU stated as 'RTX 4070 12Gb' | 7.13 s / image · 3.21 it/s | — | 2024-03-21 | github.com → |
| reported | RX 7600 XT 16 GB | SDXL 1.0 base · 1024x1024 · 20 steps “20/20 [00:17<00:00, 1.13it/s] Prompt executed in 20.52 seconds” ComfyUI 0.3.29 Zluda, Windows 11, driver 25.3.1; thread benchmark SDXL 1024x1024 20 steps | 20.52 s / image · 1.13 it/s | — | 2025-04-18 | github.com → |
| reported | RX 9070 XT 16 GB | SDXL · 1024x1024 · 20 steps “| SDXL | 20 | 1024x1024 | 1.5it/s | 27.76s | Manual tiled VAE decoder |” Ubuntu 24.04, ROCm 6.4.1, PyTorch nightly, TUNABLEOP + AOTriton experimental, --use-pytorch-cross-attention; without tiled VAE 1.49 it/s / 34.56 s; default VAE decode could OOM | 27.76 s / image · 1.5 it/s | — | 2025-05-30 | github.com → |
Reported results are other people's numbers, copied as published, with a link. Settings, drivers and ComfyUI versions differ, so compare them with care. Send yours.
09Training a LoRA for SDXL
| Trainer | VRAM | Type | Settings and quote | Source |
|---|---|---|---|---|
| OneTrainer | 5.8 GB | example run | Community test (efhosci): fp8 weights, 1024px, batch 1, latent caching, SDP, gradient checkpointing; ~8.6 GB with fp16 weights “with resolution set to 1024 the VRAM usage peaked around 5.8 GB” | github.com → |
| sd-scripts | 8 GB | stated minimum | 1024px default; U-Net only, gradient checkpointing, cached TE outputs + latents, 8-bit optimizer or Adafactor, dim 4-8 for 8GB; 10GB recommended “The LoRA training can be done with 8GB GPU memory (10GB recommended).” | github.com → |
reported Figures as stated by each trainer's own documentation or official example configs, read 2026-09-25. “Stated minimum” = the docs call it a minimum; “example run” = a config or measured run at that size. They differ a lot because of settings: an 8-bit or 4-bit base model, block swapping and lower resolution all cut memory. All models →
10Every file tracked for SDXL
| File | Type | Size | Repo |
|---|---|---|---|
| sd_xl_base_1.0.safetensors | 16-bit | 6.94 GB | stabilityai/stable-diffusion-xl-base-1.0 → all-in-one checkpoint: includes text encoder(s) and/or VAE, not diffusion-model-only (UNet 2.6B + CLIP-L + OpenCLIP-G + VAE) |
Single checkpoint file loaded with Load Checkpoint. params_b is total incl. text encoders (UNet alone ~2.6B).