Model Library/Wan2.2 S2V 14B (FP8)

Alibaba logoWan2.2 S2V 14B (FP8)

Video
ComfyUI

Wan2.2 S2V generates audio-driven cinematic video from a reference image, an audio clip, and a text prompt

On-Demand Dedicated 1xRTX 5090

Details

Modalities

video

Version

2.2

Recommended Hardware

1xRTX 5090

Estimated Price

Loading...

Provider

Alibaba

Family

Wan

Parameters

14B

License

Apache 2.0

Wan2.2 S2V 14B: Audio-Driven Cinematic Video Generation

Wan2.2 S2V 14B is an open-source speech-to-video generation model from the Wan2.2 family that produces audio-synchronized, film-quality video from a reference image, an audio clip, and an optional text prompt. Rather than limiting itself to talking-head animation, S2V targets full cinematic performance: nuanced character interactions, realistic body movement, and dynamic camera work driven by the rhythm and content of the input audio.

How It Works

S2V conditions video generation on an audio signal. A Wav2Vec2-based audio encoder extracts features from the input speech or singing, and the diffusion model generates frames whose lip motion, facial expression, and body language follow the audio. The reference image establishes character identity and scene composition, while the text prompt steers content, style, and camera behavior. Video length adapts automatically to the duration of the input audio.

Key Capabilities

  • Speech- and singing-driven animation: Characters speak or sing in sync with the provided audio track, with expression and gesture matched to delivery
  • Reference-image identity: A single input image fixes the character, wardrobe, and setting for the generated clip
  • Text-guided direction: Optional prompts control scene content, mood, and camera movement on top of the audio conditioning
  • Pose-driven performance: The model additionally supports pose-sequence guidance for choreographed motion synchronized to audio
  • Long-form generation: The underlying approach extends to longer clips and precise lip-sync editing
  • Resolution support: Generates at 480P and 720P, with aspect ratio following the reference image

Wan2.2 Foundation

S2V builds on the Wan2.2 generation of video diffusion models, trained on a substantially expanded dataset compared to Wan2.1, with 65.6% more images and 83.2% more videos, plus curated aesthetic data labeled for lighting, composition, contrast, and color tone. This training recipe underpins the family's cinematic output quality and improved handling of complex motion. In benchmark evaluations reported by the Wan team, the S2V approach outperforms comparable audio-driven human animation systems, and the broader Wan2.2 family achieves superior performance against leading closed-source commercial models on Wan-Bench 2.0.

ComfyUI Workflow

This template deploys the model through ComfyUI using the official Wan2.2 S2V workflow. The workflow loads the FP8-quantized S2V diffusion model together with the UMT5-XXL text encoder, the Wan2.1 VAE, and the Wav2Vec2 audio encoder, and includes a distillation LoRA for fast four-step sampling. Upload an audio file and a reference image, adjust the prompt, and queue the graph to generate synchronized video. All required weights are downloaded automatically at first boot.

Use Cases

  • Virtual presenters and spokespeople animated from a single photo and a voice track
  • Music videos with characters singing in sync to a recorded performance
  • Dialogue scenes and previsualization for film and animation pipelines
  • Localized marketing content where one visual identity speaks many languages
  • Educational and explainer videos narrated by a consistent on-screen character
  • Social media content that turns portraits and voice notes into performed clips

Deploy Wan2.2 S2V 14B on Vast.ai to generate audio-driven video with ComfyUI.

Quick Start Guide

Choose a model and click 'Deploy' above to find available GPUs recommended for this model.

Rent your dedicated instance preconfigured with the model you've selected.

Start sending requests to your model instance and getting responses right now.