Alibaba Cloud's Wan-Video (万相) is not a single model — it is a complete visual generation ecosystem. Backed by Alibaba's Tongyi Lab, the Wan family covers both image generation and video generation, built to deliver end-to-end creation from text, to pictures, to video.
What makes Wan-Video stand out is its twin-track strategy of comprehensive openness and commercial production: it publishes open-source weights and source code on GitHub for researchers and developers to download, deploy, and customize freely, while simultaneously offering the Wan 3.0-Video cloud API through Alibaba Cloud's Bailian (百炼) platform for enterprises that need 30-second long-form video and multimodal reference workflows.
Whether you are a developer running models on a consumer GPU or a creative team building ad campaigns through an API, this guide takes you from the open-source fundamentals to the commercial frontier of the Wan family.
Generate AI video with the Wan series on FuseAITools: Wan Video Hub.
I. What Is Wan-Video?
Wan-Video is the advanced multimodal visual generation model family from Alibaba's Tongyi Lab, covering two domains — image generation and video generation. Positioned as a comprehensive and open collection of video foundation models, its core mission is to push the boundaries of video generation and provide a full-chain generation capability that starts from text or images and ends in video.
Two features define Wan-Video more than any benchmark score:
- Scale flexibility: models at different sizes serve different efficiency-versus-quality needs, from a 1.3B parameter model for consumer GPUs to a 14B flagship and cloud-native production tiers.
- Openness: source code and all model weights are publicly available on GitHub, actively fueling research and the video generation community — an unusual move for a major cloud vendor.
The open-source community line (Wan 2.1, Wan 2.2) and the commercial cloud API line (Wan 3.0-Video) complement each other: community users deploy open models locally for research and customization, while enterprise users call the Bailian cloud API for commercial capabilities such as 30-second long video generation and multimodal reference generation.
II. Core Model Evolution: From Open-Source Foundation to All-in-One Production
The Wan family has iterated through several generations, each with a clear positioning — from the open-source foundation models to the all-around production model.
1. Open-Source Line: Wan 2.1 and Wan 2.2
Wan 2.1 established the technical roadmap of the series. Built on the diffusion-transformer (DiT) paradigm, it achieved a major breakthrough in video generation through a new VAE, a scalable pre-training strategy, and large-scale data curation. Wan 2.1 ships in two sizes:
- 14B (14 billion parameters): trained on a massive dataset of billions of images and videos; outperforms existing open models and several commercial solutions across internal and external benchmarks.
- 1.3B (1.3 billion parameters): built for resource efficiency — it runs in only 8.19 GB of VRAM and is compatible with a wide range of consumer GPUs.
Wan 2.2 extends Wan 2.1 with multi-GPU inference, FP8 quantization, and LoRA training support. Its headline addition is the Speech-to-Video (S2V) task, which generates video directly from speech and pairs with CosyVoice text-to-speech to enable one-step audio-video creation.
2. Commercialization Line: Wan 2.5 to Wan 3.0
Starting with Wan 2.5, the series kept its open-source ecosystem while moving toward productivity tools — Wan 2.5 introduced audio-supported short video generation and 480P test capabilities. Wan 2.6 and Wan 2.7 pushed multi-shot narrative and native audio-video sync: the former uses more explicit shot control, and the latter supports multi-shot transitions described in natural language.
Wan 3.0 entered public beta in August 2026 as an All-in-One reference-based video generation model built for production environments. It generates up to 30 seconds of 1080P video in a single pass and supports complex camera language such as continuous camera movement and one-take (一镜到底) shots.
3. Version Comparison at a Glance
| Version | Key Features | Best For |
|---|---|---|
| Wan 2.2 | Open-source base model; 1.3B/14B dual versions; runs in 8.19GB VRAM | Local deployment, academic research, dev testing |
| Wan 2.5 | Audio support; 480P short video | Quick visual drafts, fast testing |
| Wan 2.6 | Explicit multi-shot control; native audio | Structured storyboard narratives |
| Wan 2.7 | Natural-language multi-shot control; first/last-frame control | Scene-sequence descriptive creation |
| Wan 3.0 | 30s native duration; all-modality reference; document parsing | Commercial production, ad marketing, short dramas |
III. Core Capabilities Explained
1. Text-to-Video
Describe a scene in text and the model generates the corresponding video. Wan 3.0 outputs up to 30 seconds of 1080P HD video in one generation, with deeply optimized prompt adherence that faithfully reproduces complex motion trajectories and physical dynamics.
For prompts, structured formulas work best:
- Basic formula — Subject + Scene + Motion: ideal for first-time users. The more accurate and rich the description, the higher the output quality.
- Advanced formula — Subject description + Scene description + Motion description + Aesthetic control + Stylization: for experienced users, adding richer detail on top of the basics boosts texture quality and storytelling.
Try text-to-video on FuseAITools: Wan Text to Video.
2. Image-to-Video
Upload a static image as reference, and the model continues it into a motion video. Wan 3.0 delivers more stable, natural motion while keeping the subject's identity and visual style highly consistent with the source image. Image-to-video prompts should focus on describing the motion to be added, not re-describing what already exists in the image. The formula is Motion + Camera movement — control shots with instructions like "camera pushes in" or "camera moves left."
Animate a still image on FuseAITools: Wan Image to Video.
3. First-and-Last-Frame Control
Provide both a start frame and an end frame; the model generates the transition video between them. This is for scenes that need precisely controlled narrative start and end points and controllable transitions — for example, a product's form change or a character moving between scenes. Wan 2.5 and above support this feature steadily.
4. Reference-Based Video Generation
Wan 3.0 supports mixing multiple types of reference material in a single task. Per task you can feed up to:
- 10 images: lock character design, art style, prop style, and scene mood.
- 5 videos: provide reference for camera language and action rhythm.
- 5 audio clips: control voice style, dialogue tone, and background music mood.
Images, videos, text, and audio can be combined freely; the model absorbs the key information from all references to keep the entire video stylistically unified.
Generate from references on FuseAITools: Wan Reference-to-Video (R2V).
5. Document & Webpage Parsing
A signature capability that sets Wan 3.0 apart from most video models — it is the first to accept office documents and webpage links as input references:
- Supported formats: doc, xls, ppt, pdf, md.
- Limits: a single file or link up to 100 MB / 50 pages.
The model parses the document's hierarchy, layout, and data relationships, converting product specs, selling-point structures, and image-text layouts into video frames. Upload a product introduction PPT, for example, and Wan extracts the key information automatically to produce a promotional short film.
6. Native Audio-Visual Sync
Wan 3.0 supports native audio-visual sync generation: the model can automatically match background music and sound effects to the visuals, or accept external audio as a driver to align a character's lip movements and actions with the audio rhythm. Wan 2.5 and above all support audio-aware generation.
In the prompt, sound descriptions sit alongside visual descriptions. The formula is Subject + Scene + Motion + Sound description (voice / sound effects / background music), structured as:
- Voice = spoken content + emotion + tone + speaking speed + timbre + accent
- Sound effects = description of the concrete sound event
- Background music = musical style and mood
7. Multi-Shot Coherent Narrative
Wan 2.6 and Wan 2.7 introduced multi-shot coherent narrative generation. Users control shot structure, camera position, and timing through prompts while keeping key elements — subject, scene, atmosphere — consistent across shots.
The multi-shot prompt structure is Overall description + Shot number + Timestamp + Shot content:
- Overall description: briefly summarize the story theme, narrative style, and core emotion.
- Shot number / timestamp: define where each shot starts and ends.
- Shot content: describe that shot's visuals, action, and sound.
Wan 3.0 raises the multi-shot ceiling to 30 seconds, making a complete one-take or continuously-moving shot narrative feasible in a single generation.
Build first/last-frame stories on FuseAITools: Wan v2.7 Image to Video.
IV. Usage Tutorial
1. Preparation: Two Paths Into Wan-Video
Path A — Open-source local deployment (Wan 2.1 / Wan 2.2):
- Hardware: the 1.3B model needs 8.19GB VRAM and runs on consumer GPUs; the 14B model needs higher specs.
- Software: Python environment with PyTorch >= 2.4.0.
- Steps:
git clone https://github.com/Wan-Video/Wan2.2.git, thenpip install -r requirements.txt; installrequirements_s2v.txtfor speech synthesis; finally download model weights from Hugging Face or ModelScope.
Path B — Cloud API (Wan 3.0-Video):
- Obtain an API key from Alibaba Cloud's Bailian (百炼) platform.
- Currently in invitation-based beta; some users can apply for public beta access.
- No local GPU required — call through the API.
2. Step One: Choose a Generation Mode
| Mode | Input | Use Case |
|---|---|---|
| Text-to-Video | Text description | Ideation from scratch |
| Image-to-Video | Image + text | Animating static assets |
| First/Last-Frame Transition | Start image + end image + text | Precise control of start & end states |
| Reference Generation | Mixed image/video/audio + text | Multi-element consistent creation |
3. Step Two: Write Structured Prompts
Text-to-video (advanced formula) example:
- Subject description: a black-haired Miao ethnic-minority girl in traditional costume.
- Scene description: terraced rice fields wrapped in morning mist at dawn, layered mountains in the distance.
- Motion description: the girl slowly turns; her hair lifts in the breeze; the hem of her clothes sways gently.
- Aesthetic control: soft morning light from the side, medium shot, slow push-in.
- Stylization: photorealistic cinematic style, teal-and-warm color palette.
Image-to-video: after uploading the image, describe only the motion — "The person smiles and waves at the camera, slowly raising an arm and swaying it side to side; a breeze rustles the rice paddies in the background. Fixed camera."
Multimodal prompt with audio: "A man speaks in a dim recording studio. He says: 'Welcome to today's sharing,' with a calm tone, medium speed, deep voice. Minimal electronic ambient music plays quietly in the background."
4. Step Three: Configure Generation Parameters
- Duration & resolution: Wan 3.0 supports 480P, 720P, and 1080P, with up to 30 seconds per generation.
- Reference material: upload reference images, videos, or audio to keep characters or styles consistent.
- Extension: Wan 3.0 supports smart duration recommendations and video extension on top of an already-generated clip.
5. Step Four: Generate and Iterate
Submit the task and the model begins processing. Generation time depends on video length, resolution, and complexity. Preview the result when done; if it is not ideal, adjust the prompt and regenerate, or use editing features for targeted fixes.
V. Prompt Techniques and Examples
1. Five Core Techniques
Technique 1 — Use the structured formula. Organize prompts with the advanced formula "subject + scene + motion + aesthetic control + stylization" so the model understands every dimension precisely.
Technique 2 — For image-to-video, focus on motion. After uploading an input image, do not repeat what is already in it; only describe the motion and changes you want to add.
Technique 3 — Describe dialogue and audio separately. For prompts containing human voice, explicitly state the spoken content, emotion, tone, and speaking speed.
Technique 4 — Label multi-shot prompts with numbers and timestamps. Wan 2.6 and above support multi-shot control via structures like "Shot 1 (0-4s)" and "Shot 2 (5-8s)."
Technique 5 — Make good use of prompt expansion. Wan models support automatic prompt extension to raise generation quality, especially for shorter prompts.
2. Example Prompts
Nature scene: "Aerial view of Iceland's black sand beach; white waves crash against the black volcanic coastline while low clouds cast moving shadows. The camera pushes in slowly. Sound: continuous wave crashes and high-altitude wind."
Product showcase: "Subject: a deep-blue ceramic coffee cup with fine matte texture. Scene: pure white background, soft overhead lighting. Motion: the cup slowly rotates for a 360-degree showcase. Aesthetic: clean commercial lighting, product photography style."
Talking-head: "A woman in her early 30s sits facing the camera in a home studio and says: 'Today I'll show you three tips you can use right away,' with a relaxed, natural tone and medium speed. Low-volume Lo-Fi beats in the background. Warm key light, soft background blur."
Multi-shot narrative (Wan 2.6 / 2.7): "Overall: a girl searches the forest for her lost necklace. Shot 1 (0-4s): wide shot — the girl walks through the trees, looking around. Shot 2 (5-9s): medium close-up — the girl bends down, picks up the necklace from the ground, and lights up with joy. Shot 3 (10-14s): close-up — the pendant glints in the sunlight as the girl fastens the necklace."
VI. Use Cases and Limitations
1. Ideal Use Cases
Short-drama and short-video production. Wan 3.0's 30-second native long-form generation completes a full micro narrative in one pass, reducing visual discontinuities from multi-segment stitching, quickly producing script-matched draft footage, and cutting live-shoot costs.
Ad marketing and brand promotion. Feed product images as references to generate dynamic product demo videos for short-video distribution channels; document parsing converts a product PPT into video material with one click.
UI and product demos. Document parsing is especially friendly to UI demos and data-chart animation, preserving chart and interface structure without garbled, broken text.
Local development and academic research. The open-source line offers 1.3B and 14B models; the 1.3B runs in only 8.19GB of VRAM and fits consumer GPUs.
Character-consistency-critical projects. Wan 3.0's all-around reference mode accepts up to 10 images, 5 videos, and 5 audio clips as reference material, keeping character design, prop details, and spatial relationships consistent across long videos.
2. Limitations
Clear functional boundaries. Wan 3.0 does not support: Function Calling (tool invocation), web search, model fine-tuning (SFT), context caching, or batch asynchronous inference endpoints.
No video continuation. You cannot append new footage to the end of an existing video. Full films longer than 30 seconds must be generated in multiple passes and stitched on your side.
Long-video generation is still evolving. Wan 3.0 supports 30-second generation, but long-form output remains under continuous optimization — its benchmark long-narrative score is 64.3/100, meaning multi-scene coherent storytelling still has room to grow.
Invitation-stage restrictions. Wan 3.0-Video is still in invitation-based beta and is not yet open to all users for public or commercial use.
VII. FAQ
Q1: Is Wan-Video open source?
A: The Wan series follows a parallel open-source-plus-commercial model. The source code and model weights of Wan 2.1 and Wan 2.2 are fully open on GitHub for free download, deployment, and customization. Wan 3.0-Video, meanwhile, is offered commercially as an API service through Alibaba Cloud's Bailian platform.
Q2: What hardware does the Wan 2.1 1.3B model require?
A: The 1.3B version is extremely resource-efficient — it needs only 8.19GB of VRAM and is compatible with a wide range of consumer GPUs, making high-quality local video generation feasible on your own machine.
Q3: How long a video can Wan 3.0 generate?
A: Wan 3.0 generates up to 30 seconds of 1080P video in a single pass, supporting continuous camera movement, one-take shots, and multi-shot narratives. It does not support video continuation, so content beyond 30 seconds must be generated in segments and stitched together.
Q4: What input modalities does Wan 3.0 support?
A: Wan 3.0 accepts text, images, video, audio, documents (doc/xls/ppt/pdf/md), and webpage links. Each task accepts up to 10 images, 5 videos, and 5 audio clips, and the model automatically fuses multimodal reference information during generation.
Q5: Does Wan support native audio-visual sync?
A: Yes. Wan 2.5 and above support audio-aware video generation. Wan 3.0 can both auto-match background music and sound effects to the visuals, and accept external audio to drive lip and action alignment. A sound-description field in the prompt gives precise control over the audio.
Q6: Which versions support multi-shot control?
A: Wan 2.6 and Wan 2.7 support multi-shot coherent narrative generation. Wan 2.6 uses explicit shot control via a shot_type parameter; Wan 2.7 supports shot transitions described in natural language. Wan 3.0 supports one-take or multi-shot narratives within a single 30-second pass.
Q7: What is the pricing for the Wan 3.0 API?
A: Wan 3.0-Video is billed per second, with different unit prices across the three resolutions — 480P, 720P, and 1080P. For exact pricing, always refer to the latest announcement on Alibaba Cloud's Bailian platform.
VIII. Conclusion
Wan-Video represents a critical step in AI video generation moving from "toy" toward "production tool." From the open-source Wan 2.1 to the production-oriented Wan 3.0, the Wan series answers the question with disciplined version iteration.
Wan 3.0 is the latest result of that trajectory. It pushes generation length to 30 seconds, making one-take continuous narratives possible. It expands the input range beyond text, image, audio, and video to include office documents and webpage links — moving the starting point of video generation from "one sentence" to "an entire brief." And it uses an all-around reference mode to solve character consistency, the key pain point of commercial production.
Of course, Wan 3.0 is not the end point. It is still in invitation-stage beta, does not support video continuation or web search, and its long-narrative capability still has room to improve. But the path Wan-Video has charted — nurturing an open-source community while steadily evolving toward commercial production — offers a reference model worth watching for the whole AI video industry.
For developers, content creators, and enterprises alike, Wan's open-source ecosystem and cloud API are complementary: local deployment suits technical exploration and custom development; cloud calls suit large-scale commercial production. Understanding its capability boundaries and choosing the right access path releases far more value than simply "trying it out."
Generate with the Wan series on FuseAITools: Wan Text to Video, Wan Image to Video, Wan Video to Video, Wan v2.7 Image to Video, and Wan v2.7 Reference-to-Video — find the AI video tool that fits your creative workflow.
