# Alibaba | Wan | 3.0 | Prime | Reference to Video Alibaba Wan 3.0 Prime Reference-to-Video creates AI video from prompts plus reference images, videos, audio, files, or web links for drafts. ## API Information - **Model Slug:** alibaba-wan-3-0-prime-reference-to-video - **Branded URL:** https://www.eachlabs.ai/alibaba/wan-3-0/alibaba-wan-3-0-prime-reference-to-video - **Provider:** Alibaba - **Category:** Reference to Video - **Output Type:** video - **Status:** active - **Base Cost:** $0.068–$0.28 per second, depending on resolution - **Estimated Processing Time:** 60 seconds - **Interactive Demo:** https://www.eachlabs.ai/ai-models/alibaba-wan-3-0-prime-reference-to-video ## Pricing - **Charge Type:** dynamic - **Estimate:** $0.068–$0.28 per second, depending on resolution - **Pricing Details:** 1080p: $0.28 - **Pricing Details:** 720p: $0.14 - **Pricing Details:** 480p: $0.068 ## Input Schema | Parameter | Type | Required | Default | Constraints | Description | |-----------|------|----------|---------|-------------|-------------| | prompt | string | Yes | - | - | Text description of the desired video. Reference assets can be mentioned by order, such as Image1, Video1, and Audio1. | | reference_images | array | No | - | - | Reference image URLs. Maximum 10 images. Supported formats: JPEG, JPG, PNG without alpha, BMP, WEBP. Each side must be between 240 and 8000 pixels, aspect ratio must be between 1:8 and 8:1, and each file must be no larger than 20 MB. | | reference_videos | array | No | - | - | Reference video URLs. Maximum 5 videos. Supported formats: MP4 and MOV. Total reference video duration must not exceed 15 seconds. | | reference_audios | array | No | - | - | Reference audio URLs. Maximum 5 audio clips. Supported formats: WAV and MP3. Total reference audio duration must not exceed 15 seconds. | | file | string | No | - | - | Reference file URL. Only one file can be provided. File and link cannot be used together. | | link | string | No | - | - | Public web page URL used as reference. Only one link can be provided. Link and file cannot be used together. | | resolution | string | No | 1080P | 480P, 720P, 1080P | Output video resolution. | | ratio | string | No | adaptive | adaptive, 16:9, 9:16, 1:1, 4:3, 3:4 | Output video aspect ratio. Adaptive lets the model choose a suitable aspect ratio based on the reference content. | | duration | string | No | 5 | auto, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 | Generated video duration. Select Auto to let the model determine the duration automatically. | | audio | boolean | No | true | - | Whether to generate an audio track for the output video. | | prompt_extend | boolean | No | true | - | Enable automatic prompt rewriting. Enabled by default. Disabling it may reduce latency but can also reduce generation quality. | | seed | integer | No | - | 0–2147483647 | Random seed for reproducibility. The same seed may produce similar, but not identical, results. | ## Example Request ```bash curl -X POST https://api.eachlabs.ai/v1/prediction/ \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "alibaba-wan-3-0-prime-reference-to-video", "input": { "prompt": "Start with @Image1. The woman stands in front of the ornate gold mirror, phone raised in her right hand, her left hand tucked into her jeans pocket. She looks at the camera and gives a small smile. She shifts her weight slightly and her hair moves.\n\nShe pulls her left hand out of her pocket, raises it beside her face, and snaps her fingers once.\n\nOn the snap, her outfit changes instantly. She is now wearing a fitted black evening gown, floor-length and elegant, with silver stiletto heels and a small silver clutch bag held in her left hand. She straightens her posture, lifts her chin, and turns slightly to one side to show the silhouette of the dress.\n\nShe raises her hand and snaps her fingers a second time.\n\nThe outfit changes instantly again. She now wears a sharply tailored orange blazer suit, matching orange trousers, black stiletto heels, and a structured black handbag on her shoulder. She adjusts the lapel of the blazer with her free hand and gives a confident look at the camera.\n\nShe raises her hand and snaps her fingers a third time.\n\nThe outfit changes instantly one final time. She now wears an oversized cherry-red knit sweater, loose black trousers, and black sneakers. No bag. She relaxes her shoulders, shifts her weight onto one hip, laughs, and gives a small wave at the mirror.\n\nThroughout the entire video, the following must stay completely unchanged: her face, her hair and hairstyle, her skin tone, the phone in her right hand, the ornate gold mirror frame, the cream panelled walls, the white armchair, the marble side table with the vase, the patterned rug, the wooden floor, the daylight coming from the right, and the camera position. The camera does not move, zoom or pan at any point.\n\nOnly the clothing, shoes and bag change. Each change happens in a single instant on the finger snap, with no fade, no dissolve, no morph and no transition effect. The old outfit does not blend into the new one; it is replaced in one frame.\n\nPhotorealistic, natural window light, subtle film grain, unretouched. Sound of three crisp finger snaps and quiet room tone." } }' ``` ## Output Schema Response returned by `GET /v1/prediction/{id}` when the job completes: ```json { "status": "success", "predictionID": "string", "output": "string (URL of generated video)", "metrics": { "predict_time": "number (seconds)" } } ``` ## Polling ```bash curl https://api.eachlabs.ai/v1/prediction/{PREDICTION_ID} \ -H "Authorization: Bearer YOUR_API_KEY" ``` | Status | Meaning | |--------|---------| | `processing` | Still running — poll again | | `success` | Done — read `output` | | `error` | Failed — read `message` / `details` | ## Webhook (alternative to polling) Pass `"webhook_url": "https://your.host/path"` in the create request. Eachlabs POSTs this payload when the job ends: ```json { "exec_id": "prediction-uuid", "status": "succeeded", "output": "https://...", "error": "" } ``` `status` is `"succeeded"` or `"failed"`. `exec_id` equals the `predictionID` from create. Return 2xx within 30 seconds. ## Errors Error body: `{ "status": "error", "message": "...", "details": "..." }` | Code | Meaning | |------|---------| | `400` | Invalid input | | `401` | Missing / invalid `Authorization` bearer token | | `404` | Unknown model or prediction id | | `429` | Rate limit — 100 creates / min, 10 concurrent per key | | `5xx` | Retry with backoff | ## Overview **Alibaba | Wan | 3.0 | Prime | Reference to Video Overview** Alibaba | Wan | 3.0 | Prime | Reference to Video is a multimodal AI video generation model that turns prompts plus reference assets into polished short clips with native audio. It sits within Alibaba’s Wan 3.0 video family, which is designed for *all‑in‑one* text-to-video, image-to-video, reference-to-video, and video editing workflows through a single API endpoint. The primary differentiator of Alibaba | Wan | 3.0 | Prime | Reference to Video is its omni-reference capability: you can drive a single 30‑second shot from mixed inputs including images, video clips, audio files, documents, and even web pages, while keeping subjects, style, and scene layout consistent across the entire clip. On each::labs, this variant focuses specifically on reference-to-video, making it ideal for users who want to transform existing creative assets and data into high-quality video drafts. ## Usage Notes - API Base URL: `https://api.eachlabs.ai/v1` - Authentication: send `Authorization: Bearer YOUR_API_KEY`. Generate a key from the Eachlabs dashboard at https://www.eachlabs.ai/dashboard/api-keys. - File-typed parameters (`*_url`, `image_url`, `video_url`, `audio_url`, etc.) accept publicly-reachable HTTPS URLs only. Upload your asset first (GCS / S3 / your CDN) and pass the resulting URL. Data-URIs and localhost URLs are rejected. - For structured parameters (arrays / objects) send real JSON values, not stringified payloads. - Monetary values are reported in USD; per-token / per-megapixel rates may be billed in micro-cents internally. - Prefer `webhook_url` over polling for long-running predictions — see the Webhook Callback section.