| File | clip_vision_h.safetensors |
| What it is | CLIP Vision H (other) |
| Size | 1.3 GB |
| Folder | ComfyUI/models/clip_vision/ |
| Loader node | Load CLIP Vision (CLIPVisionLoader) |
| Source | Comfy-Org/Wan_2.1_ComfyUI_repackaged on Hugging Face |
Download clip_vision_h.safetensors (1.3 GB)
Where to put it
Put the file in ComfyUI/models/clip_vision/ (in the portable version: ComfyUI_windows_portable/ComfyUI/models/clip_vision/). Then press R in ComfyUI, or restart it, so the file shows up in Load CLIP Vision (CLIPVisionLoader).
If a workflow says “Value not in list” for this node, the file name in the workflow is different from yours, or the file is in another folder. Click the node and pick the file from the list.
What it does
A vision encoder reads an input image, for example the start frame in image-to-video. It is only needed in workflows that take an image.
Models that use clip_vision_h.safetensors
| Model | Type | Fits from | Runs well from |
|---|---|---|---|
| SCAIL-2 (character animation) | video | 13 GB | 22 GB |
| Wan 2.1 I2V 14B 480P | video | 13 GB | 21 GB |
| Wan 2.1 I2V 14B 720P | video | 16 GB | 24 GB |
| Wan 2.2 Animate 14B | video | 12 GB | 23 GB |
| Wan Animate 2 (14B) | video | 12 GB | 22 GB |
VRAM of the graphics card, for the model with the file this site picks for that card. Open a model to see every file and every GPU.
Not sure your card can run these models? Check your GPU. All shared files: ComfyUI model files.
Size and source read from Hugging Face (2026-10-01).