| File | umt5_xxl_fp8_e4m3fn_scaled.safetensors |
| What it is | UMT5-XXL FP8 (scaled) (text encoder) |
| Size | 6.7 GB |
| Folder | ComfyUI/models/text_encoders/ |
| Loader node | a text-encoder loader: Load CLIP (CLIPLoader) for one encoder, DualCLIPLoader when the model takes two (FLUX.1: CLIP-L + T5-XXL) |
| Source | Comfy-Org/Wan_2.1_ComfyUI_repackaged on Hugging Face |
Download umt5_xxl_fp8_e4m3fn_scaled.safetensors (6.7 GB)
Where to put it
Put the file in ComfyUI/models/text_encoders/ (in the portable version: ComfyUI_windows_portable/ComfyUI/models/text_encoders/). Older guides say models/clip: that folder still works. Then press R in ComfyUI, or restart it, so the file shows up in a text-encoder loader: Load CLIP (CLIPLoader) for one encoder, DualCLIPLoader when the model takes two (FLUX.1: CLIP-L + T5-XXL).
If a workflow says “Value not in list” for this node, the file name in the workflow is different from yours, or the file is in another folder. Click the node and pick the file from the list.
What it does
A text encoder turns your prompt into numbers the model understands. It runs once per prompt, then ComfyUI can move it out of VRAM to make room for the image model — so it costs load time and system RAM more than speed.
This is a smaller FP8 version. It loads faster and needs about half the RAM of the 16-bit file; the difference in the pictures is usually hard to see. It is the version this site lists by default.
FP8 files load on every NVIDIA card in ComfyUI. On a Mac (Apple Silicon) FP8 does not load — use the 16-bit file there.
Models that use umt5_xxl_fp8_e4m3fn_scaled.safetensors
| Model | Type | Fits from | Runs well from |
|---|---|---|---|
| SCAIL-2 (character animation) | video | 13 GB | 22 GB |
| Wan 2.1 I2V 14B 480P | video | 13 GB | 21 GB |
| Wan 2.1 I2V 14B 720P | video | 16 GB | 24 GB |
| Wan 2.1 T2V 1.3B | video | 4 GB | 5 GB |
| Wan 2.1 T2V 14B | video | 12 GB | 19 GB |
| Wan 2.1 VACE 14B | video | 13 GB | 24 GB |
| Wan 2.2 Animate 14B | video | 12 GB | 23 GB |
| Wan 2.2 I2V A14B | video | 10 GB | 19 GB |
| Wan 2.2 S2V 14B | video | 15 GB | 22 GB |
| Wan 2.2 T2V A14B | video | 10 GB | 19 GB |
| Wan 2.2 TI2V 5B | video | 6 GB | 10 GB |
| Wan Animate 2 (14B) | video | 12 GB | 22 GB |
VRAM of the graphics card, for the model with the file this site picks for that card. Open a model to see every file and every GPU.
Other versions of this text encoder
| File | Precision | Size | Used by |
|---|---|---|---|
| umt5_xxl_fp8_e4m3fn_scaled.safetensors | FP8 | 6.7 GB | 12 models |
| umt5_xxl_fp16.safetensors | FP16 | 11.4 GB | 12 models |
Same encoder, different precision. Any of them works in the same loader node; pick one and select it in the node.
Tested with this file
I use this exact file in 3 of my free workflows, measured on an RTX 5060 Ti 16 GB:
- Wan 2.2 TI2V 5B — text to video — 9 min for a 121-frame 1280×704 clip
- Wan 2.2 T2V 14B — FP8 — 19 min for an 81-frame 832×480 clip
- Wan 2.2 T2V 14B — GGUF Q5_K_M — 26 min for an 81-frame 832×480 clip
Not sure your card can run these models? Check your GPU. All shared files: ComfyUI model files.
Size and source read from Hugging Face (2026-10-01).