# Alibaba | Wan | 3.0 | Reference to Video Alibaba Wan 3.0 Reference-to-Video creates AI video from prompts plus reference images, videos, audio, files, or links with guided duration controls. ## API Information - **Model Slug:** alibaba-wan-3-0-reference-to-video - **Branded URL:** https://www.eachlabs.ai/alibaba/wan-3-0/alibaba-wan-3-0-reference-to-video - **Provider:** Alibaba - **Category:** Reference to Video - **Output Type:** video - **Status:** active - **Base Cost:** $0.05–$0.20 per second, depending on resolution - **Estimated Processing Time:** 60 seconds - **Interactive Demo:** https://www.eachlabs.ai/ai-models/alibaba-wan-3-0-reference-to-video ## Pricing - **Charge Type:** dynamic - **Estimate:** $0.05–$0.20 per second, depending on resolution - **Pricing Details:** 1080p: $0.2 - **Pricing Details:** 720p: $0.1 - **Pricing Details:** 480p: $0.05 - **Pricing Details:** default: $0.2 ## Input Schema | Parameter | Type | Required | Default | Constraints | Description | |-----------|------|----------|---------|-------------|-------------| | prompt | string | Yes | - | - | Text description of the desired video. Reference assets can be mentioned by order, such as Image1, Video1, and Audio1. | | reference_images | array | No | - | - | Reference image URLs. Maximum 10 images. Supported formats: JPEG, JPG, PNG without alpha, BMP, WEBP. Each side must be between 240 and 8000 pixels, aspect ratio must be between 1:8 and 8:1, and each file must be no larger than 20 MB. | | reference_videos | array | No | - | - | Reference video URLs. Maximum 5 videos. Supported formats: MP4 and MOV. Total reference video duration must not exceed 15 seconds. | | reference_audios | array | No | - | - | Reference audio URLs. Maximum 5 audio clips. Supported formats: WAV and MP3. Total reference audio duration must not exceed 15 seconds. | | file | string | No | - | - | Reference file URL. Only one file can be provided. File and link cannot be used together. | | link | string | No | - | - | Public web page URL used as reference. Only one link can be provided. Link and file cannot be used together. | | resolution | string | No | 1080P | 480P, 720P, 1080P | Output video resolution. | | ratio | string | No | adaptive | adaptive, 16:9, 9:16, 1:1, 4:3, 3:4 | Output video aspect ratio. Adaptive lets the model choose a suitable aspect ratio based on the reference content. | | duration | string | No | 5 | auto, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 | Generated video duration. Select Auto to let the model determine the duration automatically. | | audio | boolean | No | true | - | Whether to generate an audio track for the output video. | | prompt_extend | boolean | No | true | - | Enable automatic prompt rewriting. Enabled by default. Disabling it may reduce latency but can also reduce generation quality. | | seed | integer | No | - | 0–2147483647 | Random seed for reproducibility. The same seed may produce similar, but not identical, results. | ## Example Request ```bash curl -X POST https://api.eachlabs.ai/v1/prediction/ \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "alibaba-wan-3-0-reference-to-video", "input": { "prompt": "Start with @Image1 .The woman stands still against the blue backdrop, holding the small pink ice cream cup beside her, smiling at the camera. She shifts her weight slightly and her curls move.\n\nThe camera begins pushing slowly toward her. As it moves closer, she lifts the cup up toward the lens.\n\nAs the camera reaches her, fresh raspberries and chunks of dark chocolate burst outward through the frame in slow motion, scattering past the camera on both sides. Droplets of raspberry sauce fly through the air and catch the light. The berries tumble and rotate as they float around her.\n\nShe holds the cup steady, laughing, as the fruit and chocolate continue to drift and fall around her. The camera holds close on the cup with the raspberries still moving through the frame.\n\nCinematic commercial advertising style, powder-blue studio backdrop throughout, clean studio lighting, vibrant saturated colours, shallow depth of field. Slow motion on the burst. Sound of a soft whoosh and light impacts." } }' ``` ## Output Schema Response returned by `GET /v1/prediction/{id}` when the job completes: ```json { "status": "success", "predictionID": "string", "output": "string (URL of generated video)", "metrics": { "predict_time": "number (seconds)" } } ``` ## Polling ```bash curl https://api.eachlabs.ai/v1/prediction/{PREDICTION_ID} \ -H "Authorization: Bearer YOUR_API_KEY" ``` | Status | Meaning | |--------|---------| | `processing` | Still running — poll again | | `success` | Done — read `output` | | `error` | Failed — read `message` / `details` | ## Webhook (alternative to polling) Pass `"webhook_url": "https://your.host/path"` in the create request. Eachlabs POSTs this payload when the job ends: ```json { "exec_id": "prediction-uuid", "status": "succeeded", "output": "https://...", "error": "" } ``` `status` is `"succeeded"` or `"failed"`. `exec_id` equals the `predictionID` from create. Return 2xx within 30 seconds. ## Errors Error body: `{ "status": "error", "message": "...", "details": "..." }` | Code | Meaning | |------|---------| | `400` | Invalid input | | `401` | Missing / invalid `Authorization` bearer token | | `404` | Unknown model or prediction id | | `429` | Rate limit — 100 creates / min, 10 concurrent per key | | `5xx` | Retry with backoff | ## Overview **Alibaba | Wan | 3.0 | Reference to Video Overview** Alibaba | Wan | 3.0 | Reference to Video is a multimodal AI video generation model that turns prompts plus reference media into coherent clips up to 30 seconds long, with native audio and rich camera motion. Built by Alibaba Tongyi Lab as part of the Wan video family, it extends earlier Wan 2.x models with an **omni-reference** workflow: users can combine text, images, videos, audio and even documents or web pages as creative guidance in a single generation. The primary differentiator of Alibaba | Wan | 3.0 | Reference to Video is its ability to parse static, text-heavy inputs such as PDFs, slides, spreadsheets and URLs into “reality-grade” video narratives, making it uniquely suited for data-heavy, explainer and presentation-style content. On each::labs, this model focuses specifically on reference-to-video tasks, letting teams drive structure, style and motion from existing assets while still benefiting from powerful generative capabilities. ## Usage Notes - API Base URL: `https://api.eachlabs.ai/v1` - Authentication: send `Authorization: Bearer YOUR_API_KEY`. Generate a key from the Eachlabs dashboard at https://www.eachlabs.ai/dashboard/api-keys. - File-typed parameters (`*_url`, `image_url`, `video_url`, `audio_url`, etc.) accept publicly-reachable HTTPS URLs only. Upload your asset first (GCS / S3 / your CDN) and pass the resulting URL. Data-URIs and localhost URLs are rejected. - For structured parameters (arrays / objects) send real JSON values, not stringified payloads. - Monetary values are reported in USD; per-token / per-megapixel rates may be billed in micro-cents internally. - Prefer `webhook_url` over polling for long-running predictions — see the Webhook Callback section.