Alibaba | Wan | 3.0 | Prime | Text to Video

Video·wan-3.0·by Alibaba

Alibaba Wan 3.0 Prime Text-to-Video generates prompt-led AI video with aspect ratio, audio, seed, duration, and 480P-1080P output video controls.

Runtime (p50)
3m
Estimated price
From $0.068
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "alibaba-wan-3-0-prime-text-to-video",
    "version": "0.0.1",
    "input": {
        "audio": true,
        "ratio": "16:9",
        "prompt": "Scene 1 (0–8s). A woman lies asleep on her side in bed, face half buried in the pillow, soft morning light coming through the curtains. A large fat grey cat sits on the mattress beside her head. The cat lifts one front paw and taps her cheek. She doesn't move. The cat taps again, harder, and pats her nose. She frowns, groans, and slowly opens one eye, then the other. She pushes herself up onto her elbow, hair a mess, blinking, barely awake. The cat stares at her without blinking.\n\nHARD CUT.\n\nScene 2 (8–14s). Kitchen. Still in her pyjamas, hair messy, she crouches down and pours dry cat food from a small bag into a ceramic bowl on the floor. The kibble rattles as it fills the bowl. The fat grey cat pushes past her leg and starts eating immediately. She rubs her eyes with the back of her wrist and stands up.\n\nHARD CUT.\n\nScene 3 (14–22s). Living room. She picks up the remote from the coffee table and sits down on the sofa directly facing the television. The camera is positioned behind her, low over her shoulder, so that the television screen fills most of the frame and only the back of her head and shoulders are visible in the lower corner. She points the remote and presses the button.\n\nThe television turns on and the screen becomes the focus of the shot. A news broadcast is on air: a presenter sits at a desk in a studio, and across the lower third of the screen a bold red breaking news banner reads \"BREAKING NEWS — WAN 3.0 IS LIVE ON EACHLABS\" in clean white capital letters. The presenter looks into the camera and says clearly: \"Breaking news. Wan 3.0 is live on Eachlabs.\"\n\nShe sits up straight, suddenly fully awake. The fat grey cat jumps up onto the sofa beside her, also facing the screen.\n\nScene 4 (22–26s). The camera moves around to the front of the sofa, now facing the woman and the cat. She turns her head and looks down at the cat. The cat turns its head up and looks back at her. She breaks into a wide delighted smile and laughs, and the cat blinks slowly at her. She reaches out and scratches the top of the cat's head. They stay looking at each other as the shot holds.\n\nConsistency: the same woman throughout, the same fat grey cat, the same apartment, soft natural morning light in every scene. Photorealistic, handheld camera with slight natural drift, subtle film grain, unretouched. Sound of the cat's paw on skin, kibble hitting the bowl, the television clicking on, the presenter's voice, and quiet room tone at the end.",
        "duration": "26",
        "resolution": "1080P",
        "prompt_extend": true
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    Alibaba | Wan | 3.0 | Prime | Text to Video Overview

    Alibaba | Wan | 3.0 | Prime | Text to Video turns natural-language prompts into generated video clips for creators who need controlled, prompt-led motion without stitching together many short generations. It comes from Alibaba’s Wan family, and current documentation and API pages describe it as an accelerated text-to-video variant with flexible duration, aspect-ratio control, and optional audio generation. The clearest differentiator is its combination of longer single-pass output and richer prompt control, including a “thinking” mode for more deliberate interpretation of complex scenes. On each::labs, this model is useful when you want a production-oriented Alibaba text-to-video workflow that can produce cinematic clips from a single prompt while still exposing practical generation controls.

  • Capabilities

    Capabilities

    • Generates short-form video directly from a written prompt.
    • Supports 2–30 second clip lengths for flexible shot planning.
    • Offers selectable output quality including 480p, 720p, and 1080p in current API listings.
    • Provides aspect-ratio control for platform-specific framing.
    • Can generate optional audio along with the video in supported serving layers.
    • Includes deeper prompt interpretation controls through a “thinking” mode in current documentation.
    • Fits workflows that need cinematic composition, camera movement, and prompt-defined visual style.
    • Belongs to a broader Wan 3.0 family that is described as accepting multimodal references in the wider release coverage.
  • Use cases

    Use Cases for Alibaba | Wan | 3.0 | Prime | Text to Video

    Creators can use Alibaba | Wan | 3.0 | Prime | Text to Video to turn a scene idea into a draft clip for Reels, Shorts, or ads. A prompt like “A skateboarder jumps a curb at sunset, low-angle tracking shot, warm lens flare, energetic audio” leverages motion and camera control.

    Marketers can generate product teasers with a controlled studio look. A prompt like “A rotating perfume bottle on glossy black glass, soft rim light, slow push-in, premium mood” uses the model’s prompt-led framing and quality settings.

    Designers can prototype mood boards as moving scenes before final production. A prompt like “Minimal white room, floating architectural model, gentle camera drift, quiet ambient tone” works well when visual atmosphere matters more than dialogue.

    Developers building creative tools can expose the Alibaba | Wan | 3.0 | Prime | Text to Video API as a draft-generation step, letting users test low-cost versions first and then regenerate at higher quality once the prompt is refined.

  • Tips & tricks

    Tips and Tricks

    For Alibaba | Wan | 3.0 | Prime | Text to Video, prompt structure matters more than keyword stuffing. Use plain declarative prose, name the visible subject first, then describe the motion, camera treatment, lighting, and finally audio. Keep one primary action per shot, because overpacked prompts reduce consistency. If your scene includes dialogue, quote the line exactly and specify who is speaking; if you do not want captions burned into the frame, say so explicitly. If the output feels too loose, add a camera cue such as “slow push-in” or “locked-off shot.” Example prompts: “A product demo shot of a silver smartwatch on a black pedestal, slow orbiting camera, soft studio light, subtle ambient sound.” “A rainy neon street at night, one cyclist crosses frame, handheld tracking shot, reflective pavement, cinematic mood.” “A close-up of a chef plating pasta, gentle push-in, warm kitchen lighting, no on-screen text.”

  • Technical spec

    Technical Specifications

    • Model type: Text-to-video generation with prompt-led scene creation.
    • Duration: Flexible clip lengths from 2 to 30 seconds are listed in current API documentation.
    • Resolution: Output support includes 480p, 720p, and 1080p depending on the serving API.
    • Aspect ratios: Aspect-ratio selection is supported, though the exact preset set depends on the API layer.
    • Audio: Optional audio generation is supported in current Wan 3.0 Prime documentation.
    • Inputs: Text prompt is the core input; documented Wan 3.0 releases also describe multimodal inputs beyond text in the broader family.
    • Output: Video file output returned asynchronously through the API flow.
    • Processing: Generations are asynchronous; current API docs indicate you submit a request and poll for completion rather than holding a live connection.
  • Things to be aware of

    Things to Be Aware Of

    This model works best when each prompt describes one clear shot, not a full storyline with many events. If you overload it with too many subjects, actions, or style shifts, the output can drift. Audio is optional in the documented serving flow, so make sure your workflow matches the selected endpoint and settings. As with most AI video systems, hands, fast motion, crowded scenes, and dense visual detail can be harder to control than simple compositions. Because generations are asynchronous, you should plan for submit-and-poll workflow design rather than expecting instant frame delivery.

  • Key considerations

    Key Considerations

    Alibaba | Wan | 3.0 | Prime | Text to Video is best when you need a single prompt to produce a coherent short clip with cinematic framing and optional sound. It is a strong fit for concept previews, social video drafts, and prompt-driven shots that need camera direction and motion cues. For best results, you should write in clear prose and describe subject, action, setting, camera, and audio in that order. Cost varies by serving layer and resolution, so draft at lower settings first if you are testing ideas, then move to higher quality once the scene is locked. If you only need a finished marketing edit, this model is usually better as a raw generative shot source than as a complete end-to-end editor.

  • Limitations

    Limitations

    Alibaba | Wan | 3.0 | Prime | Text to Video is not a full video editor and does not replace timeline-based post-production. It is limited to short generated clips rather than long-form finished films. Current public coverage also notes that fine text rendering, dense scenes, and full reliability at the highest quality settings may still be imperfect. Availability, pricing, and exact parameter names can differ by hosting layer, so the Alibaba text-to-video experience on each::labs should be treated as API-defined rather than universally identical.

Related models

4 models
* FAQ

About Alibaba | Wan | 3.0 | Prime | Text to Video

01 / 03

What is Alibaba Wan 3.0 Prime Text-to-Video?

Alibaba Wan 3.0 Prime Text-to-Video is a prompt-based AI video model in the Wan 3.0 Prime line. It generates video from written descriptions and supports aspect ratio choices, duration control, optional audio, seed control, and 480P, 720P, or 1080P output.