MiniMax H3 API

Video·minimax-h3·by Minimax

MiniMax H3 Reference-to-Video creates videos from image, video, and audio references, keeping characters and visual style consistent across shots.

Runtime (p50)
7m
Estimated price
Usage-based
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "minimax-h3-reference-to-video",
    "version": "0.0.1",
    "input": {
        "ratio": "16:9",
        "prompt": "Use Image 1 as the locked character reference. Preserve the chubby, adorable little girl with dark hair in two small buns tied with pink flower ties, rosy cheeks, big expressive brown eyes, and her soft pink floral cardigan, and the fawn pug with its wrinkly face, dark mask, curled tail, and collar. Keep both the girl and the pug identical and consistent in every shot. Use Image 1 for storyboard order and pacing.\nRender in high-quality 4K, 16:9 warm Pixar-style 3D animation with cinematic heartwarming production value: cozy, tender, and full of gentle emotion. Use soft golden-hour lighting, rich warm tones (peach, cream, gold), and a whimsical storybook atmosphere throughout. Follow the storyboard beat by beat, with natural camera movement and seamless transitions, never a slideshow. Show the face only in close-up or extreme close-up. In wide shots, use back view, rear three-quarter view, or empty environment shots (the cottage, the garden, the sunset); never show a distant frontal fac",
        "duration": 10,
        "reference_image_urls": [
            "https://cdn-us.eachlabs.ai/defaults/15ada5aed89a4abfa0f373a4afeebe24.png"
        ]
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    MiniMax H3 | Reference-to-Video Overview

    MiniMax H3 | Reference-to-Video is a multimodal video generation model from MiniMax’s H3 family that turns reference assets into short, polished video clips. It is designed for creators who need consistent characters, visual style, motion, and pacing across shots, without manually animating every frame. The primary differentiator is its ability to combine image, video, and audio references in one generation flow while producing native 2K output with 24fps playback and integrated audio generation. That makes MiniMax H3 | Reference-to-Video especially useful for reference-driven scene creation, visual continuity work, and short-form cinematic content.

    Public documentation and third-party summaries describe H3 as a general-purpose multimodal model that supports text, image, video, and audio inputs, with reference-to-video workflows as one of its core uses. It is commonly positioned for short, high-fidelity clips rather than long-form editing.

  • Capabilities

    Capabilities

    • Generates short clips from image, video, audio, and text references in a single multimodal workflow.
    • Maintains character and visual-style consistency across shots better than prompt-only generation.
    • Produces native 2K output for sharper detail than lower-resolution reference-to-video systems.
    • Renders at 24fps, which supports a more cinematic and broadcast-like look.
    • Includes native stereo audio generation, with voice, ambience, sound effects, or music.
    • Supports multiple aspect ratios, including vertical, square, and widescreen formats.
    • Accepts multiple references, including up to 9 images, 3 video clips, and 3 audio tracks.
    • Can follow start and end frame guidance when precise scene entry or exit is needed.
  • Use cases

    Use Cases for MiniMax H3 | Reference-to-Video

    Creators can use MiniMax H3 | Reference-to-Video to turn a character concept image into a short cinematic shot with matching mood and motion. A useful prompt is: “Keep the character from the reference image, add a slow camera orbit, nighttime city lighting, and restrained facial movement.”

    Marketers can build short product spots from a hero image and a brand reference clip, preserving color and style while adding camera movement and audio. Example: “Animate the product from the reference photo with a gentle zoom, glossy reflections, clean studio audio, and premium commercial pacing.”

    Designers can prototype motion for interfaces, packaging reveals, or environment concepts using a single visual anchor. Example: “Use the reference frame as the base, animate a smooth parallax move, keep the palette minimal, and match the original composition.”

    Developers can integrate the MiniMax H3 | Reference-to-Video API into content pipelines that need repeatable short-form generation from structured prompts and reference files. Example: “Generate a 9-second scene from the uploaded image set, preserve wardrobe details, and keep the camera motion subtle.”

  • Tips & tricks

    Tips and Tricks

    For MiniMax H3 | Reference-to-Video, write prompts in layers: subject, motion, camera, style, and audio. Keep the reference assets aligned with the intended scene so the model does not have to reconcile conflicting visual cues. If you need exact composition, use start or end frames where supported, and describe motion in simple terms such as “slow push-in,” “left-to-right pan,” or “subtle handheld movement.” When using Minimax image-to-video workflows, one strong image often performs better than several weak references.

    Example prompts:

    • “A calm product reveal on a reflective table, slow camera push-in, soft studio lighting, clean shadows, premium commercial style.”
    • “Keep the same character from the reference image, walking through a rainy neon street, subtle head turn, cinematic depth of field.”
    • “Match the reference video’s color palette and pacing, but change the setting to a modern gallery with gentle ambient sound.”

    For best results, specify duration only when your storyboard needs it, and leave room for the model to maintain motion coherence.

  • Technical spec

    Technical Specifications

    • Output resolution: Native 2K, with some references describing a 768p layer used internally in the stack.
    • Frame rate: 24fps.
    • Duration: Typically 5–15 seconds per generation, with some documentation mentioning extension paths to around 30 seconds.
    • Aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus adaptive framing in some implementations.
    • Inputs: Text prompts, images, videos, and audio references; some docs mention zero, one, or two start/end frames.
    • Reference limits: Up to 9 images, up to 3 video clips, and up to 3 audio tracks, with a maximum of 12 files combined in reference workflows.
    • File limits: Reported caps include video up to 50MB, images up to 30MB, and audio up to 15MB.
    • Output: Video with native stereo audio, including voice, sound effects, ambience, or music in a single pass.
  • Things to be aware of

    Things to Be Aware Of

    MiniMax H3 | Reference-to-Video performs best when the references agree with one another. Strongly mismatched images, videos, or audio tracks can make the output less stable. The model is also optimized for short clips, so trying to force long story beats into one generation can weaken continuity. Users commonly overstuff prompts with style words instead of clear motion direction, which makes scene control harder. File size, reference count, and duration limits also matter, so large asset sets may need preprocessing before upload.

  • Key considerations

    Key Considerations

    MiniMax H3 | Reference-to-Video is best when continuity matters more than long runtime. It works well for short branded scenes, character-driven clips, and motion studies where a reference image or clip anchors the visual identity. It is less suited to long narrative sequences because the documented generation window is short. The MiniMax H3 | Reference-to-Video API is also reference-heavy, so users should prepare clean source assets with matching style, lighting, and framing. Mixed references can be powerful, but conflicting inputs may reduce consistency. For many teams, the tradeoff is clear: stronger multimodal control and native audio in exchange for tighter duration and asset constraints.

  • Limitations

    Limitations

    MiniMax H3 | Reference-to-Video is not a long-form video editor, and it is not built for extended scenes or multi-minute timelines. Documented outputs are short, usually 5–15 seconds, with strict duration and file constraints. Public sources describe native 2K and 24fps output, but production behavior may vary by implementation and account settings. Some reference modes accept many assets, yet the model can still struggle when inputs conflict or when motion is overly complex.

Related models

4 models
* FAQ

About MiniMax H3 API

01 / 05

What is MiniMax H3 Reference-to-Video?

MiniMax H3 Reference-to-Video generates video conditioned on media you supply: images, video clips, and audio in a single request. The model reads them together to keep a subject's identity, follow a clip's motion or camera style, and match a soundtrack or voice. Output runs 5 to 15 seconds in 2K with native audio.