MiniMax H3 API
MiniMax H3 Image-to-Video turns still images into 2K video clips with native audio, using first and last frames to control how each shot moves.
- Runtime (p50)
- 5m
- Estimated price
- $0.13 / unit
Overview
MiniMax H3 | Image-to-Video Overview
MiniMax H3 | Image-to-Video is a multimodal video generation model from Minimax that turns still images into high-fidelity 2K video clips with native stereo audio in a single pass. It belongs to the MiniMax H3 (Hailuo 3.0) family, a frontier-level text-to-video and image-to-video system designed for cinematic short-form content. Its primary differentiator is native 2K output at 24 fps with synchronized dialogue, sound effects, and ambience, while letting you control motion and camera behavior using first and last frames. On each::labs, MiniMax H3 | Image-to-Video focuses on workflows where creators provide one or two key images plus a brief prompt to define how the scene evolves over 5–15 seconds. This makes it ideal for story beats, product shots, and character moments that need precise visual continuity from still to video.
Capabilities
Capabilities
- Turn single or paired still images into smooth 2K video clips with coherent motion between first and last frames.
- Generate native stereo audio alongside video, including dialogue-like voice, sound effects, ambience, and simple music cues.
- Support text-to-video and image-to-video workflows in the same MiniMax H3 | Image-to-Video API, using prompts plus references to drive story and style.
- Honor fixed cinematic aspect ratios (21:9, 16:9, 4:3, 1:1, 3:4, 9:16) and provide an adaptive mode that selects a suitable framing when dimensions are omitted.
- Produce 5–15 second clips at 24 fps for short-form stories, ads, teasers, and social content.
- Use “first-and-last-frame” conditioning so uploaded images define the visual start and end of a shot while the model animates the in-between.
- Ingest mixed references (up to 9 images, 3 videos, and 3 audio files) to guide character design, environments, motion style, and soundscape.
- Handle long-form prompts (up to ~7,000 characters) to encode detailed briefs, camera directions, and narrative beats in one generation.
Use cases
Use Cases for MiniMax H3 | Image-to-Video
For content creators, MiniMax H3 | Image-to-Video can animate character portraits or key art into short cinematic moments, using first-and-last-frame control to keep faces and poses on-model while the camera moves around them. A creator might prompt: “From this character illustration, 10-second slow push-in, subtle breathing motion, quiet room tone.” Marketers can turn static product renders into polished promo clips by leveraging native 2K resolution and stereo audio, for example: “Camera orbit around this sneaker render, 8 seconds, clean white background, soft electronic beat.” Designers can test motion concepts by feeding UI mockups or environment art and asking H3 to animate transitions over exactly 5–12 seconds at a chosen ratio. Developers integrating the MiniMax H3 | Image-to-Video API in pipelines can combine text prompts with mixed references to auto-generate video variations for campaigns or personalization flows.
Tips & tricks
Tips and Tricks
For MiniMax H3 | Image-to-Video, detailed prompts paired with strong reference images give the most predictable motion. Describe camera moves (“slow dolly in,” “handheld pan”), subject actions (“character turns to the window”), and pacing (“gentle, unhurried motion over 10 seconds”) in explicit terms. Use one image as the first frame for the opening composition and optionally a second image as the last frame to lock the final pose; H3 will interpolate between them while preserving aspect ratio. When you add mixed references, prioritize a few high-quality assets instead of many weak ones to keep the visual direction clear. Start with mid-range durations (8–12 seconds) before pushing to the extremes, and refine prompts iteratively to control subtle details like lighting shifts and facial expressions. Example prompts:
“Cinematic close-up of a young woman from the starting portrait, slowly turning to face camera as the evening city lights bloom behind her, soft bokeh, 10-second shot, subtle orchestral swell.”
“Product hero shot: from the initial packshot, the camera orbits the smartwatch in a clean studio, glossy reflections, minimal ambient soundtrack, 8 seconds, 16:9.”
“Fantasy landscape from the concept art image, clouds drifting, camera glides forward over the valley, gentle wind ambience and distant birds, 12-second wide 21:9 frame.”Technical spec
Technical Specifications
- Official model name: MiniMax-H3 (multimodal video generation model).
- Output resolution: Native 2K; common documented presets include 2560×1440 (16:9), 2944×1248 (21:9), 1920×1440 (4:3), 1440×1440 (1:1), 1440×1920 (3:4), and 1440×2560 (9:16).
- Frame rate: 24 fps fixed for cinematic and broadcast-style motion.
- Clip duration: Integer durations in roughly the 5–15 second range; several docs cite 4–15 seconds as the official API window.
- Audio: Native stereo audio generated in the same pass as video, modeling voice, music, ambience, and sound effects.
- Aspect ratios: Fixed set of ratios (21:9, 16:9, 4:3, 1:1, 3:4, 9:16), plus an adaptive mode when width/height are not specified.
- Inputs: Text prompts, image start/end frames (0–2 images), and optional mixed references (up to 9 images, 3 videos, 3 audio clips; total up to 12 files).
- Prompt length: Up to about 7,000 characters in documented guides.
Things to be aware of
Things to Be Aware Of
MiniMax H3 | Image-to-Video operates within fixed duration and aspect-ratio bands, so you cannot request arbitrary lengths or non-standard resolutions; plan your storyboard around 5–15 second beats and the documented ratios. Very complex briefs with many mixed references may yield less predictable motion, especially when image, video, and audio cues conflict. Because H3 generates both video and audio, outputs are relatively heavy files, which can impact storage, bandwidth, and preview performance in production environments. Users sometimes under-specify camera behavior and pacing, leading to generic or overly static shots; adding explicit motion verbs and timing in the prompt usually improves results. Finally, region-specific policies or platform wrappers may expose only a subset of documented settings, so verify available options in your each::labs integration.
Key considerations
Key Considerations
MiniMax H3 | Image-to-Video works best when you treat your input images as keyframes: they anchor composition, character pose, and style while the model handles motion and transition in between. For reliable results, plan within the 5–15 second range and choose an aspect ratio that matches your target channel, such as 9:16 for vertical social feeds or 21:9 for cinematic trailers. Because H3 generates native stereo audio along with video, it is particularly valuable when you need coherent sound design without a separate audio model. If you only need visuals, you can still use the MiniMax H3 | Image-to-Video API and ignore audio tracks at export, but keep runtime and resolution in mind for compute and storage budgets.
Limitations
Limitations
MiniMax H3 | Image-to-Video is tuned for short clips, not long-form video; current public specs point to a practical ceiling around 15 seconds per generation. It only supports a fixed set of aspect ratios and a 24 fps frame rate, which may not fit every broadcast or product requirement. While the model handles native audio, it does not guarantee fully controllable dialogue scripts or music structures in the same way specialized audio models do. Input limits on reference files, sizes, and prompt length mean extremely large projects must be broken into multiple shots and stitched externally. As with most generative video systems, fine-grained physical accuracy and complex multi-character choreography can be challenging, requiring iteration and careful prompt design.
Related models
4 modelsAbout MiniMax H3 API
What is MiniMax H3 Image-to-Video?
MiniMax H3 Image-to-Video animates a still image into a 5 to 15 second video in 2K resolution. The output follows the aspect ratio of your image, so the composition you designed stays intact, while a text prompt directs the motion, camera work, and audio. Native stereo sound is included in every generation.



