Alibaba | Wan | 3.0 | Prime | Reference to Video

Video·wan-3.0·by Alibaba

Alibaba Wan 3.0 Prime Reference-to-Video creates AI video from prompts plus reference images, videos, audio, files, or web links for drafts.

Runtime (p50)
1m
Estimated price
From $0.068
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "alibaba-wan-3-0-prime-reference-to-video",
    "version": "0.0.1",
    "input": {
        "audio": true,
        "ratio": "1:1",
        "prompt": "Start with @Image1. The woman stands in front of the ornate gold mirror, phone raised in her right hand, her left hand tucked into her jeans pocket. She looks at the camera and gives a small smile. She shifts her weight slightly and her hair moves.\n\nShe pulls her left hand out of her pocket, raises it beside her face, and snaps her fingers once.\n\nOn the snap, her outfit changes instantly. She is now wearing a fitted black evening gown, floor-length and elegant, with silver stiletto heels and a small silver clutch bag held in her left hand. She straightens her posture, lifts her chin, and turns slightly to one side to show the silhouette of the dress.\n\nShe raises her hand and snaps her fingers a second time.\n\nThe outfit changes instantly again. She now wears a sharply tailored orange blazer suit, matching orange trousers, black stiletto heels, and a structured black handbag on her shoulder. She adjusts the lapel of the blazer with her free hand and gives a confident look at the camera.\n\nShe raises her hand and snaps her fingers a third time.\n\nThe outfit changes instantly one final time. She now wears an oversized cherry-red knit sweater, loose black trousers, and black sneakers. No bag. She relaxes her shoulders, shifts her weight onto one hip, laughs, and gives a small wave at the mirror.\n\nThroughout the entire video, the following must stay completely unchanged: her face, her hair and hairstyle, her skin tone, the phone in her right hand, the ornate gold mirror frame, the cream panelled walls, the white armchair, the marble side table with the vase, the patterned rug, the wooden floor, the daylight coming from the right, and the camera position. The camera does not move, zoom or pan at any point.\n\nOnly the clothing, shoes and bag change. Each change happens in a single instant on the finger snap, with no fade, no dissolve, no morph and no transition effect. The old outfit does not blend into the new one; it is replaced in one frame.\n\nPhotorealistic, natural window light, subtle film grain, unretouched. Sound of three crisp finger snaps and quiet room tone.",
        "duration": "15",
        "resolution": "1080P",
        "prompt_extend": true,
        "reference_images": [
            "https://cdn-us.eachlabs.ai/defaults/4aa8f8c3467642a89be28e5b49d136c2.png"
        ]
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    Alibaba | Wan | 3.0 | Prime | Reference to Video Overview

    Alibaba | Wan | 3.0 | Prime | Reference to Video is a multimodal AI video generation model that turns prompts plus reference assets into polished short clips with native audio. It sits within Alibaba’s Wan 3.0 video family, which is designed for all‑in‑one text-to-video, image-to-video, reference-to-video, and video editing workflows through a single API endpoint. The primary differentiator of Alibaba | Wan | 3.0 | Prime | Reference to Video is its omni-reference capability: you can drive a single 30‑second shot from mixed inputs including images, video clips, audio files, documents, and even web pages, while keeping subjects, style, and scene layout consistent across the entire clip. On each::labs, this variant focuses specifically on reference-to-video, making it ideal for users who want to transform existing creative assets and data into high-quality video drafts.

  • Capabilities

    Capabilities

    • Generates up to 30‑second reference-to-video clips in a single pass, avoiding stitching and preserving camera movement and scene continuity.
    • Accepts multimodal references—images, video clips, audio files, office documents, markdown, and live web pages—within one request to drive subject, style, narrative, and sound.
    • Keeps character identity, product appearance, location, and overall visual style consistent across the entire shot when guided by coherent references.
    • Produces native audio (speech, singing, ambience, and effects) together with visuals, enabling synchronized sound without a separate audio generation step.
    • Supports multiple generation modes (text-to-video, image-to-video, reference-to-video, and video editing) through a unified API, automatically inferring the desired workflow from attached media.
    • Allows document-to-video generation, turning static, text-heavy inputs such as PDFs, spreadsheets, slides, and web pages into reality-grade video content.
    • Provides configurable resolution tiers (480p, 720p, 1080p) and common social/media aspect ratios including vertical, horizontal, and square formats.
    • Enables fine-grained control over duration, aspect ratio, audio on/off, and sometimes “thinking” or high-quality modes through Alibaba | Wan | 3.0 | Prime | Reference to Video API parameters.
  • Use cases

    Use Cases for Alibaba | Wan | 3.0 | Prime | Reference to Video

    For creators and filmmakers, Alibaba | Wan | 3.0 | Prime | Reference to Video can turn concept art, animatics, and motion references into 10–30 second visual tests. By uploading character images and a short motion clip, then prompting “create a 12‑second tracking shot of the hero sprinting through a rainy alley, following the style of Image 1 and motion from Video A,” you get a coherent draft with native ambience and effects.

    Marketers can feed product photos, brand guideline PDFs, and a landing page URL into the Alibaba | Wan | 3.0 | Prime | Reference to Video API to rapidly prototype social promos. A prompt like “use the attached brochure PDF and homepage to generate a 20‑second 9:16 ad for our fitness app, upbeat music, clean typography and UI screens animated in sequence” leverages document-to-video and omni-reference capabilities.

    Designers and product teams can upload UI mockups and presentation decks to create motion design previews, while developers integrate Alibaba reference-to-video into automated workflows—such as turning weekly reports or dashboards into short video summaries—by passing structured documents and a template prompt through each::labs.

  • Tips & tricks

    Tips and Tricks

    To get the most out of Alibaba | Wan | 3.0 | Prime | Reference to Video, treat each request as a single cinematic shot rather than a loose montage. Reviews and prompt guides recommend writing the prompt like a shot description: specify the subject, action, setting, lighting, camera movement, and pacing in one connected paragraph, and call out how each reference asset should be used. Address references by order or name (“Image 1,” “Video A,” “Audio 3”) so the model can align subject appearance, motion style, and sound. Keep reference clips within the allowed total duration (typically up to 15 seconds of source video within a 30‑second output window) and avoid references with burnt‑in captions or watermarks, which can be reproduced in the generated video. When available, use “thinking” or high‑quality modes to improve motion continuity and reduce artifacts, especially at 1080p.

    Example prompts:

    • “Use Image 1 as the main character and Video A as motion style to create a 20‑second cinematic shot of a woman walking through a neon city at night, steady dolly camera, soft rim lighting, synchronized footsteps and ambient street noise.”
    • “From the attached product photos and website URL, generate a 15‑second 9:16 promo video showing the gadget rotating on a clean studio table, macro camera moves, bright key light, and upbeat electronic music referencing Audio 1.”
    • “Transform the uploaded presentation PDF into a 30‑second explainer video: animated charts from Slide 3 and 4, voiceover tone similar to Audio 2, minimalistic blue-and-white style, smooth camera pans over key bullet points.”
  • Technical spec

    Technical Specifications

    • Max duration: Generates videos up to 30 seconds in a single pass; typical duration range is 2–30 seconds per request.
    • Resolution options: 480p, 720p, and 1080p output, with 1080p commonly documented as the default tier.
    • Aspect ratios: Supports at least 16:9, 9:16, 1:1, 4:3, and 3:4, plus an “intelligent” or adaptive ratio mode in some API surfaces.
    • Frame rate: Wan 3.0 reference docs indicate up to 30 fps for video generation.
    • Input formats: Text prompts; image files; video clips (commonly MP4/MOV with H.264/H.265); audio files; office documents (PDF, PPT, DOC, XLS); markdown; and web page URLs, subject to size and page-count limits (for example, ~100 MB or ~50 pages in Alibaba documentation).
    • Output formats: MP4 video with an embedded, natively generated audio track.
    • Processing time: Live tests and reviews report turnarounds of tens of seconds for 5–15 second clips, and up to around a minute for full 30‑second generations, depending on resolution and queue load.
    • Architecture: Closed, proprietary video diffusion/transformer architecture from Alibaba Tongyi Lab; weights are not publicly released.
  • Things to be aware of

    Things to Be Aware Of

    Alibaba | Wan | 3.0 | Prime | Reference to Video is a closed model, so you cannot fine-tune weights or run it locally; you interact only through cloud APIs. Official specs and third‑party guides consistently indicate a 1080p resolution ceiling, despite some early marketing content claiming native 4K, so treat 4K claims with caution unless the provider explicitly documents upscaling as a separate feature. Reference assets must respect size, duration, and count limits (commonly around tens of megabytes per file, up to ~20 references total), and the combined source video length cannot exceed the 30‑second output window. Users report that cluttered prompts, conflicting references, and assets with heavy overlays or captions can degrade subject consistency and may cause unwanted text to appear in the generated video.

  • Key considerations

    Key Considerations

    Alibaba | Wan | 3.0 | Prime | Reference to Video works best when you supply clear reference assets that express subject identity, style, motion, and sound, then bind them together with a focused prompt describing a single continuous shot. Because the model is closed and offered as a metered API, you access it via providers like Alibaba Cloud Model Studio or integrated platforms such as each::labs, and you should account for per‑second pricing that scales with resolution (commonly cited ranges of roughly $0.05–$0.20 per second across 480p–1080p). Use this model when you need 2–30 second, reference-driven clips with native audio from mixed inputs; for purely synthetic text-only video or still images, Wan 3.0’s other modes or different models may be more cost‑effective.

  • Limitations

    Limitations

    Alibaba | Wan | 3.0 | Prime | Reference to Video cannot currently generate beyond ~30 seconds per pass, and it is broadly documented as topping out at 1080p rather than true native 4K. The model depends heavily on the quality and coherence of references; noisy, mismatched, or low‑resolution inputs can lead to artifacts, identity drift, or unstable motion. It is not an open‑weights system, so you cannot train custom versions or inspect internals, and pricing tied to duration and resolution makes very high‑volume or experimentation-heavy workflows more expensive than lighter, text-only generation. Finally, document-to-video performance varies with layout complexity, so dense, unstructured reports may require prompt and asset curation to avoid chaotic visual results.

Related models

4 models
* FAQ

About Alibaba | Wan | 3.0 | Prime | Reference to Video

01 / 03

What is Alibaba Wan 3.0 Prime Reference-to-Video?

Alibaba Wan 3.0 Prime Reference-to-Video is a reference-guided AI video model in the Wan 3.0 Prime line. It combines a written prompt with reference images, videos, audio, one supported file, or one public web link to guide the generated result.