Wan2.1 VACE 14B: All-in-One Video Creation and Editing
Wan2.1 VACE 14B is an open video creation and editing model from Alibaba's Wan team that unifies multiple video tasks in a single model. Rather than shipping separate checkpoints for generation, editing, and completion, VACE (Video All-in-one Creation and Editing) accepts a text prompt together with optional video, mask, and reference-image inputs and handles the full range of tasks from that one interface. The 14B model supports both 480P and 720P output, while the smaller 1.3B variant of the family targets 480P.
One Model, Many Tasks
VACE covers the major video creation and editing workflows through a unified conditioning interface:
- Text-to-Video: Generate video directly from a text prompt
- Reference-to-Video: Generate video guided by one or more reference images, preserving subject identity without any preprocessing step
- Video-to-Video Editing: Re-render an existing video under new conditions such as depth or pose control extracted from the source
- Masked Video Editing: Inpaint or outpaint regions of a video using mask inputs, enabling object replacement, removal, and canvas extension
- First-Last Frame Control: Generate the motion between provided start and end frames
- Task Composition: Combine inputs, for example a reference image plus a control video, in a single generation pass
Built on the Wan2.1 Foundation
VACE inherits the Wan2.1 video foundation stack:
- Wan-VAE: A causal 3D VAE that encodes and decodes high-resolution video of arbitrary length while preserving temporal information
- Diffusion Transformer: The Wan2.1 DiT backbone, extended with VACE context blocks that inject video, mask, and reference conditioning
- UMT5 Text Encoder: Multilingual prompt understanding covering both English and Chinese, including Wan2.1's distinctive ability to render legible text in generated video
Inputs of any resolution are accepted, with best results when the source material falls near the model's native 480P and 720P generation sizes.
ComfyUI Integration
VACE is supported natively in ComfyUI through the WanVaceToVideo core node. This template provisions the official ComfyUI workflow templates for the model covering text-to-video, reference-to-video, video-to-video with a control video, first-last-frame generation, inpainting, and outpainting, along with the model weights repackaged by Comfy-Org at Comfy-Org/Wan_2.1_ComfyUI_repackaged. The bundled workflows include distillation LoRAs that cut the number of denoising steps for faster iteration, and the inpainting workflow includes segmentation-assisted mask generation.
Use Cases
- Subject-consistent video generation from character or product reference images
- Object replacement and removal in existing footage
- Extending video canvases beyond their original framing
- Repose and restyle of existing clips via pose and depth control
- Animating between keyframes with first-last frame control
- Standard text-to-video generation
Deploy Wan2.1 VACE 14B on Vast.ai with ComfyUI to run the full creation and editing suite from a browser-based node interface.