Alibaba | Wan | 3.0 | Text to Video

Video·wan-3.0·by Alibaba

Alibaba Wan 3.0 Text-to-Video generates AI video from prompts with duration, aspect ratio, audio, seed, and 480P-1080P output controls for teams.

Runtime (p50)
3m
Estimated price
From $0.05
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "alibaba-wan-3-0-text-to-video",
    "version": "0.0.1",
    "input": {
        "audio": true,
        "ratio": "16:9",
        "prompt": "A young woman walks slowly along a rain-soaked Tokyo street at night, neon signs reflecting in the puddles around her. She holds a transparent umbrella, water running off its edges. She stops at a crossing, looks up at the glowing signs above her, and a faint smile appears. Then she turns to the camera and she says softly: \"This is my favorite time of the day.\" Behind her, the crossing light turns green and people begin to move past her in both directions. She steps forward into the crowd, and the camera holds as she disappears among the umbrellas.\n\nCinematic, shallow depth of field, cool blue and magenta neon palette, natural rain sound and distant city traffic. Steady handheld camera, slow forward drift.",
        "duration": "15",
        "resolution": "1080P",
        "prompt_extend": true
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    Alibaba | Wan | 3.0 | Text to Video Overview

    Alibaba | Wan | 3.0 | Text to Video is a multimodal AI video generation model from Alibaba’s Tongyi Lab that turns text prompts and reference media into up to 30‑second, sound‑enabled clips in a single pass. It is designed for teams that need fast, controllable short-form video without stitching multiple generations together, offering 480p, 720p, and 1080p output with synchronized audio. The primary differentiator of Alibaba | Wan | 3.0 | Text to Video is its unified pipeline: one endpoint accepts text, images, video, audio, and office documents and produces a coherent, narrated video clip. Through each::labs, teams can use this model for text-to-video campaigns, document-to-video explainers, or reference-driven edits while keeping control over duration, aspect ratio, seed, and resolution for repeatable, production-ready workflows.

  • Capabilities

    Capabilities

    • Generates up to 30-second videos from plain text prompts in a single pass, reducing the need to stitch shorter clips.
    • Supports 480p, 720p, and 1080p output, enabling quality scaling from rapid drafts to production-ready assets.
    • Provides intelligent duration and aspect ratio selection when those parameters are not specified, adapting video length and framing to the prompt.
    • Produces synchronized native audio with the video, including ambience and narrative voice, for complete short clips without manual sound design.
    • Maintains strong character and motion coherence across the full 30-second window, improving identity stability and physics compared with earlier Wan versions.
    • Underlying Wan 3.0 model accepts multimodal references such as images, videos, audio, and office documents (PPT, PDFs, spreadsheets, web pages) to drive generation, which can be exposed via compatible APIs.
    • Allows multiple reference assets in a single job (images, short clips, audio) with documented limits on counts and combined durations, supporting complex, guided edits.
    • Available via official Alibaba | Wan | 3.0 | Text to Video API endpoints with usage-based pricing and rate limits, making it suitable for integration into production pipelines on each::labs.
  • Use cases

    Use Cases for Alibaba | Wan | 3.0 | Text to Video

    For creators and filmmakers, Alibaba | Wan | 3.0 | Text to Video can deliver 30-second concept trailers or mood pieces from script-style prompts, taking advantage of its coherent motion and synchronized audio. A prompt might be: “30-second trailer of a sci-fi short, sweeping city vistas, close-ups on the main character, orchestral build.”

    For marketers, its controllable resolution and aspect ratio make it ideal for multi-platform campaigns; teams can generate 1080p landscape hero videos and vertical social shorts from similar prompts, such as: “20-second vertical launch teaser, product hero shots, bold typography, upbeat soundtrack.”

    For designers and product teams, the model can turn structured copy into interface walkthroughs or explainer clips, especially when paired with document-to-video pipelines; e.g.: “Explain a new dashboard layout in 25 seconds, clean UI mockups, calm narration, minimal transitions.”

    For developers, integration with the Alibaba | Wan | 3.0 | Text to Video API through each::labs enables automated generation of onboarding, release notes, or report summaries triggered from text or document inputs with consistent duration and format.

  • Tips & tricks

    Tips and Tricks

    Prompt engineering for Alibaba | Wan | 3.0 | Text to Video benefits from treating the prompt like a compact shot script: specify scene, characters, camera, lighting, and pacing in a few clear sentences. Use explicit duration and aspect ratio when you need tight control, but consider leaving them unset when you want the model’s “smart duration” and adaptive framing to choose a format from context. For consistent outputs, fix a seed value and reuse it across iterations to explore minor prompt tweaks while keeping composition stable. When targeting 1080p, favor clean compositions, moderate motion, and limited small text elements to reduce artifacts. Example prompts:

    “30-second 16:9 cinematic ad, a barista in a warm café, slow dolly-in, soft morning light, product close-up at the end, gentle piano soundtrack.”
    “20-second vertical explainer, minimalist icons and text, smooth transitions, narrator-style audio describing a fintech app onboarding flow.”
    “15-second action shot, futuristic city at night, single hero running across neon rooftops, dynamic tracking camera, energetic electronic music.”

  • Technical spec

    Technical Specifications

    • Provider / family: Alibaba Tongyi Lab, Wan 3.0 video generation family.
    • Generation type: Text-to-video (primary), with unified support for image and reference inputs in the underlying Wan 3.0 model.
    • Max duration: Up to 30 seconds per generation, with durations from ~2–30 seconds supported.
    • Frame rate: Up to 30 fps in the official Wan3.0 API preview.
    • Output resolutions: 480p, 720p, and 1080p, with 1080p typically the default tier.
    • Aspect ratios: Standard landscape and portrait ratios, with “intelligent” automatic aspect ratio selection when not set.
    • Input format (text-to-video): Natural language prompt, optional duration, aspect ratio, seed, and resolution parameters.
    • Output format: Encoded video file (commonly MP4) with integrated audio track.
    • Processing time: Typical runs complete in tens of seconds for 10–30s clips, depending on resolution and queue load.
  • Things to be aware of

    Things to Be Aware Of

    Despite its strengths, Alibaba | Wan | 3.0 | Text to Video can struggle with fine-grained in-frame text such as tiny labels, packaging, or subtitles; these elements may be inaccurate or soft. Long 30-second clips create more room for small visual drift, especially in complex multi-shot narratives, so users should review continuity and, when needed, regenerate shorter segments. Document or link inputs are subject to file and page limits (for example, around 100 MB and 50 pages per file in reported tests), which constrains very large slide decks or PDFs. As a preview or beta model, access can be gated, quotas can change, and independent benchmarks are still limited, so teams should validate quality against their own scenarios before large-scale deployment.

  • Key considerations

    Key Considerations

    Alibaba | Wan | 3.0 | Text to Video is optimized for short, high-fidelity clips up to 30 seconds, so it fits best for ads, social posts, trailers, and micro‑explainers rather than long-form content. The official APIs list it as a preview or public beta model, meaning behavior and pricing can evolve and production users should monitor version changes. Costs scale linearly with duration and resolution, with published per‑second rates at different tiers, so teams should balance 1080p quality against budget for large batches. It performs best when prompts and references are specific about characters, motion, and camera, using the model’s strengths in narrative coherence and synchronized audio, while each::labs can help coordinate seeds, aspect ratios, and retries across workflows.

  • Limitations

    Limitations

    Alibaba | Wan | 3.0 | Text to Video is focused on short clips up to 30 seconds and cannot produce long-form videos in a single generation, requiring external editing for minutes-long content. Maximum resolution is documented as 1080p rather than true 4K, despite some early rumors, so projects needing native 4K must upscale externally. Fine typography, dense UI, and crowded scenes can exhibit artifacts or misrendered details, making it less suitable for text-heavy instructional content where legibility is critical. API usage is subject to per-second pricing, rate limits, and concurrent job caps, which may restrict extremely high-volume real-time workloads without careful planning.

Related models

4 models
* FAQ

About Alibaba | Wan | 3.0 | Text to Video

01 / 03

What is Alibaba Wan 3.0 Text-to-Video?

Alibaba Wan 3.0 Text-to-Video is a prompt-based video generation model for creating short AI video clips from written descriptions. It supports Chinese and English prompts, aspect ratio choices, duration control, optional audio, seed control, and 480P, 720P, or 1080P output.