Alibaba | Wan | 3.0 | Prime | Image to Video

Video·wan-3.0·by Alibaba

Alibaba Wan 3.0 Prime Image-to-Video turns source images into AI video with prompt guidance, audio, duration, and 480P-1080P output controls.

Runtime (p50)
1m
Estimated price
From $0.068
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "alibaba-wan-3-0-prime-image-to-video",
    "version": "0.0.1",
    "input": {
        "audio": true,
        "ratio": "16:9",
        "prompt": "The same woman in every scene, matching the reference image exactly: dark brown hair pulled back with loose strands falling around her face, freckles across her nose and cheeks, natural minimal makeup, small gold hoop earrings, a cream oversized blazer worn over a plain black t-shirt. The same modern office throughout: glass-walled meeting rooms, pale wood desks, a tall potted plant, polished concrete floors, bright daylight through large windows.\n\nScene 1 (0–7s). Selfie camera held at arm's length, her extended arm visible at the edge of the frame, slight wide-angle distortion and natural handheld movement. She stands in the office corridor beside the plant, looks into the lens and says brightly: \"Good morning! Come spend a day at the office with me.\" She smiles, turns and starts walking, the camera bobbing gently with her steps as the glass walls pass behind her.\n\nHARD CUT.\n\nScene 2 (7–17s). Fixed camera on her desk, facing her. She now wears clear-framed glasses. The footage is heavily sped up, timelapse style, roughly eight times normal speed. Her movements are fast and jittery: she types rapidly, scrolls, picks up a pen and writes, turns to speak to someone off-frame, drinks from a mug, pushes her glasses up, moves papers, leans back and stretches. Behind her the daylight shifts slowly from bright morning to warm afternoon while colleagues blur past.\n\nHARD CUT.\n\nScene 3 (17–25s). Normal speed. She sits sideways on her chair at the desk, glasses off, holding a cookie in one hand and the camera in the other. Late afternoon light comes in low and warm. She holds the cookie up near her face without eating it, looks into the lens and says: \"After hours of work, this tastes so good. A few more emails and then we're heading home.\" She smiles and laughs softly.\n\nHARD CUT.\n\nScene 4 (25–30s). Night. The office behind her is dark and empty, the monitors switched off, only a few dim ceiling lights on and city lights visible through the windows. She stands in the corridor holding the camera at arm's length, looks into the lens and says: \"Alright, let's go home.\" She smiles, and the video ends.\n\nPhotorealistic, handheld selfie camera with natural drift, subtle grain, unretouched vlog aesthetic. Natural daylight in the first three scenes, dark warm tones in the last. Ambient office sound, keyboard clicks, quiet room tone, and her voice close to the microphone.",
        "duration": "30",
        "resolution": "1080P",
        "first_frame": "https://cdn-us.eachlabs.ai/defaults/4944dc34a36648fdb16a9ab596a5214b.png",
        "prompt_extend": true
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    Alibaba | Wan | 3.0 | Prime | Image to Video Overview

    Alibaba | Wan | 3.0 | Prime | Image to Video turns a source image into a generated video with prompt-guided motion, composition control, and optional audio. It is part of Alibaba’s Wan 3.0 family and is positioned as a faster, high-quality image animation workflow for creators who need controlled movement from a still frame. Its main differentiator is the combination of first-frame image guidance, optional last-frame guidance, and flexible output control across 480p, 720p, and 1080p. The model is designed for short-form cinematic clips, product animations, social content, and prototype video generation inside an API workflow.

  • Capabilities

    Capabilities

    • Animates a still image into a generated video with prompt guidance.
    • Supports optional last-frame guidance for ending control.
    • Produces output at 480p, 720p, or 1080p.
    • Supports short clips from 2 seconds up to 30 seconds.
    • Can generate with optional audio enabled in the workflow.
    • Preserves subject identity and composition better when the source image is clean and well framed.
    • Fits API-driven pipelines for product demos, social content, and creative automation.
  • Use cases

    Use Cases for Alibaba | Wan | 3.0 | Prime | Image to Video

    Creators can turn a portrait or scene still into a polished motion clip by asking for subtle movement, camera drift, and atmospheric changes. Example: "Animate this portrait with gentle hair movement, soft blinking, and a slow cinematic push-in." This uses the model’s first-frame image anchoring.

    Marketers can create product teasers from a single hero image and keep the item centered while adding premium motion cues. Example: "Use this image as the first frame, add reflective light sweeps, a slow orbiting camera move, and a clean studio finish." This benefits from the 1080p output option.

    Designers can prototype concept shots before committing to full motion production. Example: "Transform this concept art into a 10-second scene with drifting clouds, subtle character movement, and a final frame matching the original composition." This makes use of the optional last-frame guidance.

    Developers can build automated image-to-video generation features inside a workflow. Example: "Generate a 720p clip from this first-frame image with mild environmental motion and no abrupt camera cuts." This aligns with the Alibaba | Wan | 3.0 | Prime | Image to Video API workflow.

  • Tips & tricks

    Tips and Tricks

    Write prompts around motion, not just style. Start with the subject, then describe the action, camera movement, environment, and ending state. Keep the image aligned with the desired aspect ratio so the model does not need to crop aggressively. Use the optional last-frame guidance only when the final pose or composition matters. For complex motion, enable deeper reasoning controls if available in your integration. Good prompts are specific and compact:

    "Animate this product photo with a slow camera push-in, subtle rotating reflections, and soft studio lighting, ending on a centered hero frame."

    "Turn this character portrait into a cinematic short clip with gentle wind movement, a slight head turn, and a steady left-to-right camera drift."

    "Use this image as the first frame, keep the subject identity stable, and add calm background motion with realistic lighting changes over 5 seconds."

  • Technical spec

    Technical Specifications

    • Model type: image-to-video generation.
    • Inputs: first-frame image, text prompt, optional last-frame image, optional audio, and optional thinking mode controls.
    • Output: generated video, with optional audio depending on settings.
    • Resolution support: 480p, 720p, and 1080p.
    • Duration: 2 to 30 seconds.
    • Aspect ratio: controlled by the source image and generation settings; keep the input image aligned with the intended frame.
    • Format: API-based JSON request flow with image attachment and video response.
    • Processing time: not officially standardized in the sources reviewed; async API workflows are commonly used for generation.
  • Things to be aware of

    Things to Be Aware Of

    Results depend heavily on the source image. Busy compositions, occluded subjects, and unclear framing can reduce motion quality or identity consistency. Prompts that request too many actions at once often produce weaker scene continuity. The model also works better when the intended aspect ratio already matches the input image, because mismatched framing can force unwanted crops. If you need custom audio or tightly synchronized audio behavior, confirm that the integration exposes those controls before production use.

  • Key considerations

    Key Considerations

    Alibaba | Wan | 3.0 | Prime | Image to Video works best when the source image has a clear subject, simple composition, and visible motion potential. It is a strong fit when you want controlled image animation rather than freeform text-to-video generation. The model is especially useful when you need longer clips than older image-to-video workflows and want to balance speed, quality, and cost by choosing 480p, 720p, or 1080p. For best results, use it when the opening frame matters most and when the desired motion can be described clearly in one prompt. The Alibaba | Wan | 3.0 | Prime | Image to Video API is most efficient for structured production pipelines rather than casual one-off experimentation.

  • Limitations

    Limitations

    Alibaba | Wan | 3.0 | Prime | Image to Video is still constrained by the quality and structure of the input image, so it cannot reliably invent complex new scenes from a weak source frame. It is not ideal for dense text rendering, crowded compositions, or highly intricate hand motion. Publicly reviewed sources also do not provide a single standardized processing-time benchmark, so latency can vary by resolution, duration, and API setup.

Related models

4 models
* FAQ

About Alibaba | Wan | 3.0 | Prime | Image to Video

01 / 03

What is Alibaba Wan 3.0 Prime Image-to-Video?

Alibaba Wan 3.0 Prime Image-to-Video is an image-to-video model in the Wan 3.0 Prime line. It turns a first-frame image, and optionally a last-frame image, into AI video with prompt guidance, duration control, optional audio, seed control, and selectable output resolution.