Alibaba | Wan | 3.0 | Prime | Reference to Video
Alibaba Wan 3.0 Prime Reference-to-Video creates AI video from prompts plus reference images, videos, audio, files, or web links for drafts.
- Runtime (p50)
- 1m
- Estimated price
- From $0.068
Overview
Alibaba | Wan | 3.0 | Prime | Reference to Video Overview
Alibaba | Wan | 3.0 | Prime | Reference to Video is a multimodal AI video generation model that turns prompts plus reference assets into polished short clips with native audio. It sits within Alibaba’s Wan 3.0 video family, which is designed for all‑in‑one text-to-video, image-to-video, reference-to-video, and video editing workflows through a single API endpoint. The primary differentiator of Alibaba | Wan | 3.0 | Prime | Reference to Video is its omni-reference capability: you can drive a single 30‑second shot from mixed inputs including images, video clips, audio files, documents, and even web pages, while keeping subjects, style, and scene layout consistent across the entire clip. On each::labs, this variant focuses specifically on reference-to-video, making it ideal for users who want to transform existing creative assets and data into high-quality video drafts.
Capabilities
Capabilities
- Generates up to 30‑second reference-to-video clips in a single pass, avoiding stitching and preserving camera movement and scene continuity.
- Accepts multimodal references—images, video clips, audio files, office documents, markdown, and live web pages—within one request to drive subject, style, narrative, and sound.
- Keeps character identity, product appearance, location, and overall visual style consistent across the entire shot when guided by coherent references.
- Produces native audio (speech, singing, ambience, and effects) together with visuals, enabling synchronized sound without a separate audio generation step.
- Supports multiple generation modes (text-to-video, image-to-video, reference-to-video, and video editing) through a unified API, automatically inferring the desired workflow from attached media.
- Allows document-to-video generation, turning static, text-heavy inputs such as PDFs, spreadsheets, slides, and web pages into reality-grade video content.
- Provides configurable resolution tiers (480p, 720p, 1080p) and common social/media aspect ratios including vertical, horizontal, and square formats.
- Enables fine-grained control over duration, aspect ratio, audio on/off, and sometimes “thinking” or high-quality modes through Alibaba | Wan | 3.0 | Prime | Reference to Video API parameters.
Use cases
Use Cases for Alibaba | Wan | 3.0 | Prime | Reference to Video
For creators and filmmakers, Alibaba | Wan | 3.0 | Prime | Reference to Video can turn concept art, animatics, and motion references into 10–30 second visual tests. By uploading character images and a short motion clip, then prompting “create a 12‑second tracking shot of the hero sprinting through a rainy alley, following the style of Image 1 and motion from Video A,” you get a coherent draft with native ambience and effects.
Marketers can feed product photos, brand guideline PDFs, and a landing page URL into the Alibaba | Wan | 3.0 | Prime | Reference to Video API to rapidly prototype social promos. A prompt like “use the attached brochure PDF and homepage to generate a 20‑second 9:16 ad for our fitness app, upbeat music, clean typography and UI screens animated in sequence” leverages document-to-video and omni-reference capabilities.
Designers and product teams can upload UI mockups and presentation decks to create motion design previews, while developers integrate Alibaba reference-to-video into automated workflows—such as turning weekly reports or dashboards into short video summaries—by passing structured documents and a template prompt through each::labs.
Tips & tricks
Tips and Tricks
To get the most out of Alibaba | Wan | 3.0 | Prime | Reference to Video, treat each request as a single cinematic shot rather than a loose montage. Reviews and prompt guides recommend writing the prompt like a shot description: specify the subject, action, setting, lighting, camera movement, and pacing in one connected paragraph, and call out how each reference asset should be used. Address references by order or name (“Image 1,” “Video A,” “Audio 3”) so the model can align subject appearance, motion style, and sound. Keep reference clips within the allowed total duration (typically up to 15 seconds of source video within a 30‑second output window) and avoid references with burnt‑in captions or watermarks, which can be reproduced in the generated video. When available, use “thinking” or high‑quality modes to improve motion continuity and reduce artifacts, especially at 1080p.
Example prompts:
- “Use Image 1 as the main character and Video A as motion style to create a 20‑second cinematic shot of a woman walking through a neon city at night, steady dolly camera, soft rim lighting, synchronized footsteps and ambient street noise.”
- “From the attached product photos and website URL, generate a 15‑second 9:16 promo video showing the gadget rotating on a clean studio table, macro camera moves, bright key light, and upbeat electronic music referencing Audio 1.”
- “Transform the uploaded presentation PDF into a 30‑second explainer video: animated charts from Slide 3 and 4, voiceover tone similar to Audio 2, minimalistic blue-and-white style, smooth camera pans over key bullet points.”
Technical spec
Technical Specifications
- Max duration: Generates videos up to 30 seconds in a single pass; typical duration range is 2–30 seconds per request.
- Resolution options: 480p, 720p, and 1080p output, with 1080p commonly documented as the default tier.
- Aspect ratios: Supports at least 16:9, 9:16, 1:1, 4:3, and 3:4, plus an “intelligent” or adaptive ratio mode in some API surfaces.
- Frame rate: Wan 3.0 reference docs indicate up to 30 fps for video generation.
- Input formats: Text prompts; image files; video clips (commonly MP4/MOV with H.264/H.265); audio files; office documents (PDF, PPT, DOC, XLS); markdown; and web page URLs, subject to size and page-count limits (for example, ~100 MB or ~50 pages in Alibaba documentation).
- Output formats: MP4 video with an embedded, natively generated audio track.
- Processing time: Live tests and reviews report turnarounds of tens of seconds for 5–15 second clips, and up to around a minute for full 30‑second generations, depending on resolution and queue load.
- Architecture: Closed, proprietary video diffusion/transformer architecture from Alibaba Tongyi Lab; weights are not publicly released.
Things to be aware of
Things to Be Aware Of
Alibaba | Wan | 3.0 | Prime | Reference to Video is a closed model, so you cannot fine-tune weights or run it locally; you interact only through cloud APIs. Official specs and third‑party guides consistently indicate a 1080p resolution ceiling, despite some early marketing content claiming native 4K, so treat 4K claims with caution unless the provider explicitly documents upscaling as a separate feature. Reference assets must respect size, duration, and count limits (commonly around tens of megabytes per file, up to ~20 references total), and the combined source video length cannot exceed the 30‑second output window. Users report that cluttered prompts, conflicting references, and assets with heavy overlays or captions can degrade subject consistency and may cause unwanted text to appear in the generated video.
Key considerations
Key Considerations
Alibaba | Wan | 3.0 | Prime | Reference to Video works best when you supply clear reference assets that express subject identity, style, motion, and sound, then bind them together with a focused prompt describing a single continuous shot. Because the model is closed and offered as a metered API, you access it via providers like Alibaba Cloud Model Studio or integrated platforms such as each::labs, and you should account for per‑second pricing that scales with resolution (commonly cited ranges of roughly $0.05–$0.20 per second across 480p–1080p). Use this model when you need 2–30 second, reference-driven clips with native audio from mixed inputs; for purely synthetic text-only video or still images, Wan 3.0’s other modes or different models may be more cost‑effective.
Limitations
Limitations
Alibaba | Wan | 3.0 | Prime | Reference to Video cannot currently generate beyond ~30 seconds per pass, and it is broadly documented as topping out at 1080p rather than true native 4K. The model depends heavily on the quality and coherence of references; noisy, mismatched, or low‑resolution inputs can lead to artifacts, identity drift, or unstable motion. It is not an open‑weights system, so you cannot train custom versions or inspect internals, and pricing tied to duration and resolution makes very high‑volume or experimentation-heavy workflows more expensive than lighter, text-only generation. Finally, document-to-video performance varies with layout complexity, so dense, unstructured reports may require prompt and asset curation to avoid chaotic visual results.

