Alibaba | Wan | 3.0 | Reference to Video
Alibaba Wan 3.0 Reference-to-Video creates AI video from prompts plus reference images, videos, audio, files, or links with guided duration controls.
- Runtime (p50)
- 1m
- Estimated price
- From $0.05
Overview
Alibaba | Wan | 3.0 | Reference to Video Overview
Alibaba | Wan | 3.0 | Reference to Video is a multimodal AI video generation model that turns prompts plus reference media into coherent clips up to 30 seconds long, with native audio and rich camera motion. Built by Alibaba Tongyi Lab as part of the Wan video family, it extends earlier Wan 2.x models with an omni-reference workflow: users can combine text, images, videos, audio and even documents or web pages as creative guidance in a single generation. The primary differentiator of Alibaba | Wan | 3.0 | Reference to Video is its ability to parse static, text-heavy inputs such as PDFs, slides, spreadsheets and URLs into “reality-grade” video narratives, making it uniquely suited for data-heavy, explainer and presentation-style content. On each::labs, this model focuses specifically on reference-to-video tasks, letting teams drive structure, style and motion from existing assets while still benefiting from powerful generative capabilities.
Capabilities
Capabilities
- Generates up to 30-second videos in a single pass with native synchronized audio, removing the need for separate audio workflows for many scenarios.
- Supports reference-to-video using mixed inputs: text prompts plus images, short video clips, audio snippets, office documents and public web pages in one request.
- Provides flexible aspect ratios including widescreen, vertical and square formats, enabling direct output for social media, presentations and product pages.
- Offers first-frame and first-and-last-frame control modes, letting users lock identity and composition while the model fills in motion between frames.
- Maintains strong texture and character detail with a focus on “reality-grade” rendering, especially for product shots and human subjects.
- Parses document and webpage references to build narrative structure around static charts, bullet lists and written explanations, ideal for turning data-heavy content into video.
- Handles up to ~10 image references and multiple short video or audio clips, so creators can combine style boards, motion references and sound cues in the same generation.
- Exposed via Alibaba | Wan | 3.0 | Reference to Video API endpoints on Alibaba Cloud Model Studio, making it accessible for programmatic workflows and integration into apps and pipelines via each::labs.
Use cases
Use Cases for Alibaba | Wan | 3.0 | Reference to Video
Creators and designers: Use Alibaba | Wan | 3.0 | Reference to Video to turn style boards and keyframes into polished motion for intros, music visuals or short cinematic scenes. For example: “Use these three concept art images as references for a 15-second fantasy fly-through, smooth drone-style camera, epic orchestral score, high-detail volumetric lighting.”
Marketers and product teams: Convert product shots and pitch decks into launch videos that highlight features and metrics. The document-to-video capability lets teams feed a PDF or slide deck and get a concise 20–30 second product story. Example: “Turn this product brochure PDF into a 25-second launch video, hero shots of the device, animated callouts of specs, upbeat electronic soundtrack.”
Developers: Integrate Alibaba | Wan | 3.0 | Reference to Video API into onboarding flows to auto-generate tutorials from help center articles or documentation. Example: “Use the linked API quick-start page as reference, generate a 30-second step-by-step onboarding video with UI-style motion graphics and calm narration-ready audio.”
Educators and analysts: Transform reports and dashboards into explainer clips, using charts and tables as visual anchors. Example: “Visualize the Q2 sales spreadsheet as a 20-second data story, bar charts animating in, key numbers highlighted, corporate presentation style, subtle background music.”
Tips & tricks
Tips and Tricks
Effective prompting for Alibaba | Wan | 3.0 | Reference to Video combines seven core elements: subject, action, environment, camera movement, lighting, atmosphere and visual style. Begin by anchoring the scene in the reference asset—describe how the subject should move relative to the provided image, clip or document content, then add one clear camera move and a concise lighting description. Keep prompts under roughly 50–80 words and avoid complex sequences of actions; focus instead on a single, continuous motion across the full duration. When using document or URL references, explicitly state what sections matter (“focus on the sales growth chart” or “visualize the hero slide”), which helps the model prioritize key content. For iteration, generate short 4–8 second drafts at 480p or 720p to validate motion and framing before committing to 1080p 30-second outputs.
Example prompts:
“Use the attached product image as the main subject. The camera slowly orbits around the bottle on a glossy black surface, droplets sliding down, cinematic studio lighting, subtle ambient music.”
“Turn the linked PDF pitch deck into a 20-second explainer video: visualize key market growth charts and product mockups, modern flat motion graphics, smooth camera pans, upbeat tech conference background audio.”
“Extend this five-second reference clip into a full 25-second sequence: keep the main character’s face and outfit consistent, continue the walk through the neon street, slow tracking shot, rain reflections, atmospheric cyberpunk film look.”
Technical spec
Technical Specifications
- Model family: Wan 3.0 (Tongyi Wanxiang video line), Alibaba Cloud / Tongyi Lab.
- Generation modes: Text-to-video, image-to-video (first frame / first & last frame), reference-to-video, video editing; this page focuses on reference-to-video.
- Max duration: Up to 30 seconds per single-pass generation; with video reference, input plus output must stay within about 30 seconds combined.
- Resolution: 480p, 720p and 1080p output, with 1080p commonly used as the default in API and playground examples.
- Aspect ratios: Adaptive plus standard ratios such as 16:9, 4:3, 1:1, 3:4 and 9:16 for vertical content.
- Inputs: Text prompts, up to ~10 images, up to ~5 short videos, up to ~5 audio clips, plus office documents (doc/xls/ppt/pdf/txt) and public web URLs as references.
- Outputs: MP4 video files with synchronized audio rendered in a single pass.
- Typical API pattern: Asynchronous REST endpoint in Alibaba Cloud Model Studio under IDs such as
wan3.0-videofor video synthesis. - Processing time: Varies by provider, but public beta examples report multi-second to tens-of-seconds latency for full 30-second clips, depending on resolution and load.
Things to be aware of
Things to Be Aware Of
While Alibaba | Wan | 3.0 | Reference to Video supports up to 30-second outputs, longer clips can exhibit pacing or continuity issues if prompts attempt complex multi-step actions instead of one continuous motion. Identity stability for human subjects can still drift across frames, especially when combining multiple reference images or when faces are small in frame; users should explicitly constrain identity and test shorter durations first. Text and UI elements rendered inside the video (labels, screen content, charts) may not be perfectly accurate and often require manual verification or overlay in post-production. Document and URL parsing depends on content clarity, so heavily formatted or very long files can lead to selective or incomplete visualization. Access to Alibaba | Wan | 3.0 | Reference to Video API may be gated in some regions or remain in preview, meaning quota and pricing terms can change during public beta.
Key considerations
Key Considerations
Alibaba | Wan | 3.0 | Reference to Video is optimized for short, high-impact clips rather than long-form footage, so shot planning should stay within the 2–30 second range. Users need at minimum a clear text prompt and at least one reference asset (image, short video, audio, document or URL) to fully leverage the reference-to-video mode. This model is best suited for explainer videos, product showcases, social clips and data-driven narratives where reference visuals or documents provide structure, while alternatives may be preferable for longer or purely cinematic films. Pricing on Alibaba | Wan | 3.0 | Reference to Video API is typically per-second and varies by resolution tier, so higher resolutions and full 30-second runs carry noticeably higher cost and should be reserved for final renders.
Limitations
Limitations
Alibaba | Wan | 3.0 | Reference to Video currently caps outputs at around 1080p and 30 seconds per pass, so it is not a solution for native 4K or multi-minute continuous scenes. The reference-to-video pipeline cannot always mix all reference types at once; some combinations of first-frame, last-frame and multiple video references are restricted per generation. Audio quality, while synchronized, may require editing or replacement for professional voiceovers or complex sound design, and the model does not provide fine-grained control over dialogue or lyrics. Open-source weights for Wan 3.0 have not been officially confirmed, so usage typically relies on cloud APIs rather than local deployment, which makes performance subject to provider infrastructure and regional availability.

