# MiniMax H3 | Reference-to-Video MiniMax H3 Reference-to-Video creates videos from image, video, and audio references, keeping characters and visual style consistent across shots. ## API Information - **Model Slug:** minimax-h3-reference-to-video - **Branded URL:** https://www.eachlabs.ai/minimax/minimax-h3/minimax-h3-reference-to-video - **Provider:** Minimax - **Category:** Reference to Video - **Output Type:** video - **Status:** active - **Version:** 0.0.1 - **Base Cost:** MiniMax H3 2K reference-to-video billing: output and input video seconds at $0.13/s; the first 5 reference images and audio are free. - **Estimated Processing Time:** 400 seconds - **Last Updated:** 2026-08-06 - **Interactive Demo:** https://www.eachlabs.ai/ai-models/minimax-h3-reference-to-video ## Pricing - **Charge Type:** dynamic - **Pricing Details:** MiniMax H3 2K reference-to-video billing: output and input video seconds at $0.13/s; the first 5 reference images and audio are free. ### Pricing Rules | Condition | Pricing | | --- | --- | | resolution == "768P" AND output.task.usage.input_image_count > "5" | MiniMax H3 768P reference-to-video billing: output and input video seconds at $0.09/s, plus $0.04 for each reference image after the first 5; audio is free. | | resolution == "768P" | MiniMax H3 768P reference-to-video billing: output and input video seconds at $0.09/s; the first 5 reference images and audio are free. | | output.task.usage.input_image_count > "5" | MiniMax H3 2K reference-to-video billing: output and input video seconds at $0.13/s, plus $0.04 for each reference image after the first 5; audio is free. | | Rule 4 | MiniMax H3 2K reference-to-video billing: output and input video seconds at $0.13/s; the first 5 reference images and audio are free. | ## Input Schema No input parameters documented. ## Example Request ```bash curl -X POST https://api.eachlabs.ai/v1/prediction/ \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "minimax-h3-reference-to-video", "input": {} }' ``` ## Output Schema Response returned by `GET /v1/prediction/{id}` when the job completes: ```json { "status": "success", "predictionID": "string", "output": "string (URL of generated video)", "metrics": { "predict_time": "number (seconds)" } } ``` ## Polling ```bash curl https://api.eachlabs.ai/v1/prediction/{PREDICTION_ID} \ -H "Authorization: Bearer YOUR_API_KEY" ``` | Status | Meaning | |--------|---------| | `processing` | Still running — poll again | | `success` | Done — read `output` | | `error` | Failed — read `message` / `details` | ## Webhook (alternative to polling) Pass `"webhook_url": "https://your.host/path"` in the create request. Eachlabs POSTs this payload when the job ends: ```json { "exec_id": "prediction-uuid", "status": "succeeded", "output": "https://...", "error": "" } ``` `status` is `"succeeded"` or `"failed"`. `exec_id` equals the `predictionID` from create. Return 2xx within 30 seconds. ## Errors Error body: `{ "status": "error", "message": "...", "details": "..." }` | Code | Meaning | |------|---------| | `400` | Invalid input | | `401` | Missing / invalid `Authorization` bearer token | | `404` | Unknown model or prediction id | | `429` | Rate limit — 100 creates / min, 10 concurrent per key | | `5xx` | Retry with backoff | ## Overview **MiniMax H3 | Reference-to-Video Overview** MiniMax H3 | Reference-to-Video is a multimodal video generation model from MiniMax’s H3 family that turns reference assets into short, polished video clips. It is designed for creators who need consistent characters, visual style, motion, and pacing across shots, without manually animating every frame. The primary differentiator is its ability to combine image, video, and audio references in one generation flow while producing native 2K output with 24fps playback and integrated audio generation. That makes MiniMax H3 | Reference-to-Video especially useful for reference-driven scene creation, visual continuity work, and short-form cinematic content. Public documentation and third-party summaries describe H3 as a general-purpose multimodal model that supports text, image, video, and audio inputs, with reference-to-video workflows as one of its core uses. It is commonly positioned for short, high-fidelity clips rather than long-form editing. ## Usage Notes - API Base URL: `https://api.eachlabs.ai/v1` - Authentication: send `Authorization: Bearer YOUR_API_KEY`. Generate a key from the Eachlabs dashboard at https://www.eachlabs.ai/dashboard/api-keys. - File-typed parameters (`*_url`, `image_url`, `video_url`, `audio_url`, etc.) accept publicly-reachable HTTPS URLs only. Upload your asset first (GCS / S3 / your CDN) and pass the resulting URL. Data-URIs and localhost URLs are rejected. - For structured parameters (arrays / objects) send real JSON values, not stringified payloads. - Monetary values are reported in USD; per-token / per-megapixel rates may be billed in micro-cents internally. - Prefer `webhook_url` over polling for long-running predictions — see the Webhook Callback section.