# MiniMax | H3 | Max | Reference to Video H3 Max Reference-to-Video creates short AI video clips from prompts and reference images, preserving subject cues for character-driven creative scenes. ## API Information - **Model Slug:** minimax-h3-max-reference-to-video - **Branded URL:** https://www.eachlabs.ai/fal/minimax-h3-max/minimax-h3-max-reference-to-video - **Provider:** fal - **Category:** Reference to Video - **Output Type:** video - **Status:** active - **Base Cost:** 768P video costs $0.08 per second, so a 5 second clip is $0.40. The first 4096 reference tokens are free, so up to 4 reference images at 1024x1024 add no extra charge. - **Estimated Processing Time:** 300 seconds - **Interactive Demo:** https://www.eachlabs.ai/ai-models/minimax-h3-max-reference-to-video ## Pricing - **Charge Type:** dynamic - **Estimated Price (default example):** $0.8000 - **Pricing Details:** 768P video costs $0.08 per second, so a 5 second clip is $0.40. The first 4096 reference tokens are free, so up to 4 reference images at 1024x1024 add no extra charge. ### Pricing Rules | Condition | Pricing | | --- | --- | | resolution eq_i "480P" AND output.reference_image_tokens <= "4096" | 480P video costs $0.05 per second, so a 5 second clip is $0.25. The first 4096 reference tokens are free, so up to 4 reference images at 1024x1024 add no extra charge. | | resolution eq_i "480P" AND output.reference_image_tokens > "4096" | 480P video costs $0.05 per second, plus $0.02 per 1000 reference tokens above the free 4096. A 1024x1024 reference image counts as 1024 tokens, a 2048x2048 image as 4096 tokens. | | resolution eq_i "1080P" AND output.reference_image_tokens <= "4096" | 1080P video costs $0.16 per second, so a 5 second clip is $0.80. The first 4096 reference tokens are free. | | resolution eq_i "1080P" AND output.reference_image_tokens > "4096" | 1080P video costs $0.16 per second, plus $0.02 per 1000 reference tokens above the free 4096. | | output.reference_image_tokens <= "4096" | 768P video costs $0.08 per second, so a 5 second clip is $0.40. The first 4096 reference tokens are free, so up to 4 reference images at 1024x1024 add no extra charge. | | output.reference_image_tokens > "4096" | 768P video costs $0.08 per second, plus $0.02 per 1000 reference tokens above the free 4096. A 1024x1024 reference image counts as 1024 tokens, a 2048x2048 image as 4096 tokens. | ## Input Schema | Parameter | Type | Required | Default | Constraints | Description | |-----------|------|----------|---------|-------------|-------------| | prompt | string | Yes | - | - | Text prompt for video generation. Refer to reference assets as Image 1, Video 1, Audio 1, and so on. | | reference_image_urls | array | No | - | - | Subject or style reference image URLs. Reference images as Image 1, Image 2, and so on in the prompt. | | reference_video_urls | array | No | - | - | Motion or visual reference video URLs, 2-15 seconds each and 15 seconds combined. Reference videos as Video 1, Video 2, and so on in the prompt. | | reference_audio_urls | array | No | - | - | Reference audio URLs, 2-15 seconds each and 15 seconds combined. Audio cannot be the only reference input; provide at least one reference image or video with it. | | duration | integer | No | 5 | 5–15 | Output video duration in seconds. Default: 5. | | resolution | string | No | 768P | 480P, 768P, 1080P | Native generation resolution of the video. Default: 768P. | | aspect_ratio | string | No | adaptive | adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | Aspect ratio of the generated video. Default: adaptive. | | prompt_expansion_mode | string | Yes | balanced | disabled, balanced, quality | How much effort to spend rewriting the prompt before generation. Default: balanced. | | seed | integer | No | - | - | Optional random seed. A random seed is selected when omitted. | | enable_safety_checker | boolean | No | true | - | Enable fal's safety checker. Default: true. | ## Example Request ```bash curl -X POST https://api.eachlabs.ai/v1/prediction/ \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "minimax-h3-max-reference-to-video", "input": { "prompt": "Use Image 1 as the locked character and set reference. Preserve the young woman exactly as she appears — her face shape, green eyes, full natural brows, dewy skin with visible texture and the small beauty mark, slicked-back dark bun with loose face-framing strands, gold hoop earrings, thin gold necklace, and white off-shoulder textured knit top. Preserve the glossy black mascara tube and its wand exactly. Keep the woman, the product, and the beige living-room set — matte wall, slim olive tree in a pot, framed print, candle on a wooden ledge, cream sofa — identical and consistent in every shot. Use Image 1 for storyboard order and pacing, following the twelve panels left to right, top to bottom.\n\nRender in high-quality 4K, vertical 3:4 live-action UGC beauty commercial with clean commercial production value: intimate, effortless, and quietly premium. Use soft diffused daylight from camera left, a warm neutral palette (cream, sand, soft gold), and an unhurried morning-routine atmosphere throughout. Shoot handheld with subtle natural sway and a shallow depth of field, real skin texture with no beauty smoothing, and no visible on-screen text. Follow the storyboard beat by beat with natural camera movement and seamless cuts on action, never a slideshow.\n\nMotion beats in order: she lifts the closed tube beside her cheek and tilts her head to camera; she lowers her gaze and blinks, bare lashes visible; extreme close-up as she blinks slowly twice; her hand rotates the closed tube as light slides across it; both hands twist the cap off in one smooth motion; the hand lifts the wand and turns the bristles to camera; she sweeps the wand upward through her upper lashes and pauses; she brushes twice more from a slightly different angle; she finishes a final upward stroke and lowers the wand; extreme close-up as her eye opens wider and blinks, lashes lifted and separated; she turns slightly and breaks into a soft smile; she tilts the tube toward camera and smiles with a small shoulder shrug.\n\nAudio: a soft modern commercial soundtrack throughout — warm minimal pop with a gentle four-on-the-floor pulse, muted plucked synth, light finger snaps, and airy pads. Understated and confident, never energetic or club-like. The track builds subtly through the application beats and resolves on a warm sustained chord at the final smile. Add quiet ambient product sound design under the music: a soft click as the cap twists off, a delicate brush sweep on the lashes. No voiceover, no lyrics, no dialogue.\n\nKeep her face fully visible and undistorted in every shot with natural eye anatomy and symmetrical features. The applicator hand always stays to the side at ear level and never crosses or covers the face. Render hands with correct anatomy and five fingers, especially in the macro cap-twisting and wand shots. Show the face in medium close-up, close-up, or extreme close-up only; use hands-only macro framing for the three product shots, with no face in those frames. ", "prompt_expansion_mode": "quality" } }' ``` ## Output Schema Response returned by `GET /v1/prediction/{id}` when the job completes: ```json { "status": "success", "predictionID": "string", "output": "string (URL of generated video)", "metrics": { "predict_time": "number (seconds)" } } ``` ## Polling ```bash curl https://api.eachlabs.ai/v1/prediction/{PREDICTION_ID} \ -H "Authorization: Bearer YOUR_API_KEY" ``` | Status | Meaning | |--------|---------| | `processing` | Still running — poll again | | `success` | Done — read `output` | | `error` | Failed — read `message` / `details` | ## Webhook (alternative to polling) Pass `"webhook_url": "https://your.host/path"` in the create request. Eachlabs POSTs this payload when the job ends: ```json { "exec_id": "prediction-uuid", "status": "succeeded", "output": "https://...", "error": "" } ``` `status` is `"succeeded"` or `"failed"`. `exec_id` equals the `predictionID` from create. Return 2xx within 30 seconds. ## Errors Error body: `{ "status": "error", "message": "...", "details": "..." }` | Code | Meaning | |------|---------| | `400` | Invalid input | | `401` | Missing / invalid `Authorization` bearer token | | `404` | Unknown model or prediction id | | `429` | Rate limit — 100 creates / min, 10 concurrent per key | | `5xx` | Retry with backoff | ## Usage Notes - API Base URL: `https://api.eachlabs.ai/v1` - Authentication: send `Authorization: Bearer YOUR_API_KEY`. Generate a key from the Eachlabs dashboard at https://www.eachlabs.ai/dashboard/api-keys. - File-typed parameters (`*_url`, `image_url`, `video_url`, `audio_url`, etc.) accept publicly-reachable HTTPS URLs only. Upload your asset first (GCS / S3 / your CDN) and pass the resulting URL. Data-URIs and localhost URLs are rejected. - For structured parameters (arrays / objects) send real JSON values, not stringified payloads. - Monetary values are reported in USD; per-token / per-megapixel rates may be billed in micro-cents internally. - Prefer `webhook_url` over polling for long-running predictions — see the Webhook Callback section.