Alibaba | Wan | 3.0 | Image to Video

Video·wan-3.0·by Alibaba

Alibaba Wan 3.0 Image-to-Video turns first-frame or first-and-last-frame images into AI video with motion, audio, duration, and resolution controls.

Runtime (p50)
1m
Estimated price
From $0.05
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "alibaba-wan-3-0-image-to-video",
    "version": "0.0.1",
    "input": {
        "audio": true,
        "ratio": "16:9",
        "prompt": "She carefully places the last piece of gold leaf onto the chocolate cake with the tweezers, then sets the tweezers down on the counter beside her. She straightens up, wipes her hands lightly, and looks directly into the camera with a warm smile.\n\nShe says: \"What do you think the secret to a beautiful cake is? Love, of course. Now let's deliver this beautiful chocolate cake to its owner.\"\n\nAs she finishes speaking she reaches down, slides both hands under the cake plate, and lifts it slightly, still smiling at the camera.\n\nHer face, her hair, the black patterned headband, the white chef's jacket, the mixing bowls, the bright kitchen window behind her and the whole room stay exactly the same throughout. Camera stays in place with a very slight natural drift, no cuts. Photorealistic, soft bright natural window light, subtle film grain. Warm intimate voice close to the microphone, quiet kitchen room tone, soft clink as the tweezers touch the counter.",
        "duration": "10",
        "resolution": "1080P",
        "first_frame": "https://cdn-us.eachlabs.ai/defaults/5b3f13fefafc42238966352a7f9cf013.png",
        "prompt_extend": true
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    Alibaba | Wan | 3.0 | Image to Video Overview

    Alibaba | Wan | 3.0 | Image to Video is an advanced image-to-video mode within Alibaba’s Wan 3.0 multimodal video generation family, designed to turn still images into continuous AI video with native audio in a single pass. It animates a first frame, or both a first and last frame, into directed clips of up to 30 seconds without stitching separate segments. The model’s primary differentiator is its combination of long single-pass duration, strong source fidelity for image-to-video, and synchronized audio, all controlled through prompts and reference inputs. Built by Alibaba’s Tongyi Lab and exposed via Wan 3.0 endpoints and partner platforms, it fits seamlessly into each::labs as a powerful option for creators who need cinematic motion, camera control, and sound from still imagery.

  • Capabilities

    Capabilities

    • Generates up to 30 seconds of continuous video from a single image, avoiding stitched multi-clip workflows.
    • Supports first-frame and first-and-last-frame image-to-video control, letting you define both the starting and ending visuals of a shot.
    • Produces synchronized native audio—voice, ambience, music, and effects—in the same pass as video, reducing post-production overhead.
    • Preserves subject identity, composition, and visual style from the source image, enabling high source fidelity for products and characters.
    • Provides flexible camera control via prompts, including push-ins, pullbacks, orbits, and follow shots.
    • Accepts multimodal references (images, video clips, audio, and documents such as PDFs or PPTs) to guide motion, pacing, and narrative in image-to-video workflows.
    • Offers multiple resolutions and aspect ratios for social, vertical, and cinematic formats, with up to 1080p confirmed output.
    • Integrates with Alibaba | Wan | 3.0 | Image to Video API endpoints, making it suitable for automation and embedding in each::labs pipelines.
  • Use cases

    Use Cases for Alibaba | Wan | 3.0 | Image to Video

    For creators and filmmakers, Alibaba | Wan | 3.0 | Image to Video can turn a single concept frame into a 20–30 second establishing shot with camera moves and atmosphere, using first-frame or first-and-last-frame control with native audio. Example prompt: “Animate this storyboard frame into a 25-second opening shot with a slow crane up, city ambience, and subtle musical swell.”

    Marketers and product teams can convert product images into promo clips that preserve packaging and branding while adding motion and sound. Example prompt: “Create a 10-second premium ad from this product photo. Keep logo and label exact, add rotating product motion, glossy reflections, and a clean sound bed.”

    Designers use image-to-video to explore motion studies from static compositions, leveraging camera control and source fidelity. Example prompt: “Turn this poster design into a 15-second motion study. Animate light streaks and parallax layers, keep typography unchanged, and add subtle synth ambience.”

    Developers integrate Alibaba | Wan | 3.0 | Image to Video API into pipelines to auto-generate explainer clips from UI mockups or document-based references. Example prompt: “From this dashboard screenshot, generate a 12-second tutorial clip with guided camera pans and voiceover explaining the main metrics.”

  • Tips & tricks

    Tips and Tricks

    To get the most from Alibaba | Wan | 3.0 | Image to Video, treat the source image as a locked first frame and describe motion, camera behavior, and atmosphere explicitly in your prompt. For character or product fidelity, supply focused reference images and keep prompts constrained to a few key actions rather than lengthy scripts. When using first-and-last-frame control, define where the shot should begin and end, then let Wan 3.0 fill the motion in between; this reduces unwanted cuts. Avoid overloading the model with dense crowds, complex hand interactions, or tiny UI text, which are common stress points. Start with 10–15 second clips at 720p or 1080p to balance speed and reliability, then extend to 30 seconds after you’re comfortable with its behavior.

    Example prompts:

    “Animate this product photo into a 12-second slow push-in video. Keep shape, label, and colors identical. Add soft studio lighting, subtle reflections, and calm background music. No on-screen text.”

    “Turn this character portrait into a 20-second cinematic shot. The camera circles slowly, hair and clothing move with a gentle breeze, and ambient city noise plays underneath.”

    “Use this first frame and last frame to create a 30-second continuous walking sequence, with smooth camera follow, consistent lighting, and natural footstep audio.”

  • Technical spec

    Technical Specifications

    • Generation length: Up to 30 seconds per clip in a single pass, with shorter durations available.
    • Resolutions: 480p, 720p, and 1080p output tiers; 1080p is the confirmed maximum in official documentation.
    • Aspect ratios: Common formats including 16:9, 9:16, 1:1, 3:4, and 4:3, with auto-selection based on prompt where supported.
    • Input types (image-to-video): Single starting image, optional last frame, plus reference images, videos, audio, and documents (doc, xls, ppt, pdf, txt, key, pages, md) depending on endpoint.
    • Output format: MP4 video with native audio track generated in the same pass.
    • Reference limits: Common configurations expose up to ~10 reference images and mixed media assets per generation, with a total duration cap of 30 seconds when video references are used.
    • Architecture: Closed, frontier-level multimodal video model; weights are not publicly released.
    • Processing time: Typically a single generation pass per clip; real-world reports note that longer, 30-second runs can be slower and may require retries.
  • Things to be aware of

    Things to Be Aware Of

    Reports on Wan 3.0 note that long 30-second runs can occasionally introduce uncommanded camera cuts or dissolves, so planning to trim or select the best segments is wise. Complex hand interactions, overlapping bodies, and dense crowds remain challenging and may show minor glitches. Small on-screen text, packaging copy, and UI elements can blur or drift over time, so critical typography should be composited later or kept large. Pricing for Alibaba | Wan | 3.0 | Image to Video API is typically per second and per resolution, which means 1080p 30-second clips can consume budget quickly. Because Wan 3.0 is closed and accessed via cloud endpoints, you should account for API quotas, rate limits, and potential latency when embedding it into each::labs workflows.

  • Key considerations

    Key Considerations

    Alibaba | Wan | 3.0 | Image to Video is best used when you need up to 30 seconds of continuous motion and synchronized audio from a still image, particularly for narrative or cinematic shots. Because Wan 3.0 is a closed model, you interact with it via Alibaba | Wan | 3.0 | Image to Video API endpoints or integrated platforms like each::labs rather than managing weights directly. Official specs cap output at 1080p, so ultra-high-resolution or true 4K workflows may require upscaling or alternative models. Longer clips increase the chance of minor visual or motion artifacts, so professional pipelines should plan for trimming, selective takes, and iterative prompting. For purely short loops or ultra-detailed typography, specialized image or video engines may be preferable.

  • Limitations

    Limitations

    Alibaba | Wan | 3.0 | Image to Video cannot currently output beyond 1080p, despite some early reports suggesting 4K; official specs and multiple technical deep dives confirm 480p, 720p, and 1080p as the real tiers. Long single-pass clips increase the probability of minor identity drift, motion artifacts, or unexpected cuts, especially in complex scenes. The model does not provide direct access to weights or on-prem deployment, limiting use to cloud APIs and integrated platforms. Fine-grained control over facial expressions, perfectly accurate small text, or fully custom avatars from scratch remains out of scope, and some cases may require additional specialist tools or post-processing.

Related models

4 models
* FAQ

About Alibaba | Wan | 3.0 | Image to Video

01 / 03

What is Alibaba Wan 3.0 Image-to-Video?

Alibaba Wan 3.0 Image-to-Video is a video generation model that turns a first-frame image into an AI video, with optional last-frame guidance for more controlled motion. It supports prompt guidance, duration control, optional audio, seed control, and 480P, 720P, or 1080P output.