MiniMax H3 API

Video·minimax-h3·by Minimax

MiniMax H3 Text-to-Video turns written prompts into 5 to 15 second 2K clips with native sound, from cinematic shots to ads and stylized motion design.

Runtime (p50)
5m
Estimated price
$0.13 / unit
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "minimax-h3-text-to-video",
    "version": "0.0.1",
    "input": {
        "ratio": "16:9",
        "prompt": "A photorealistic cinematic 5-second video of a professional chef cooking in a warm, busy restaurant kitchen. The chef, wearing a clean white chef's jacket, stands at the stove tossing vegetables in a hot pan as flames flare up and steam rises. Quick, skilled hand movements, sizzling ingredients, droplets of oil catching the light. Rich warm lighting, shallow depth of field, realistic textures on the food and stainless steel surfaces, subtle handheld camera movement, authentic documentary-style food film aesthetic, ultra-detailed, high quality.",
        "duration": 5
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    MiniMax H3 | Text to Video Overview

    MiniMax H3 | Text to Video is a multimodal generative model from Minimax designed to turn written prompts and visual or audio references into short, cinematic video clips with native stereo sound. It focuses on high-fidelity 2K video at a film-standard 24 fps, making it well suited for marketing assets, social content, and design explorations that need polished motion in seconds rather than minutes. As part of the Minimax H3 family (also known as Hailuo 3.0), it unifies understanding of text, images, video, and sound to produce 5–15 second clips that already include dialogue, sound effects, and ambience in a single pass. Integrated through each::labs, MiniMax H3 | Text to Video gives creators and developers a fast path from concept to production-ready motion without separate audio pipelines.

  • Capabilities

    Capabilities

    • Text-to-video generation of 5–15 second clips at native 2K resolution and 24 fps from detailed written prompts.
    • Image-to-video and reference-guided video, using start/end frames or multiple images to anchor composition and style.
    • Omni-reference mode that combines up to 9 images, 3 videos, and 3 audio files into a single coherent output.
    • Native stereo audio generation in the same pass, including dialogue-like voices, sound effects, ambience, and music.
    • Fixed high-quality aspect ratios from ultrawide 21:9 to vertical 9:16, plus adaptive framing where MiniMax H3 chooses the best ratio.
    • Long-form prompt understanding with support for descriptions up to thousands of characters for complex scenes and narrative beats.
    • Start/end frame control, allowing zero, one, or two images to define opening and closing shots while preserving input ratio.
    • Multimodal context understanding that unifies text, image, video, and audio inputs into a single generative process.
  • Use cases

    Use Cases for MiniMax H3 | Text to Video

    Content creators can use MiniMax H3 | Text to Video to quickly produce short cinematic sequences for channels or portfolios by combining rich prompts with native 2K output and adaptive aspect ratios. For example: “10-second 16:9 travel vlog intro, drone-style shot over mountains at sunrise, warm grading, gentle acoustic soundtrack.” Marketers can build product teasers and ad cutdowns that ship with sound already embedded, leveraging the model’s native stereo audio and 9:16 vertical framing for mobile campaigns.

    Designers and motion artists can explore typographic animations or abstract motion graphics using start/end frames and omni-reference mode to lock keyframes while allowing creative transitions. Developers integrating the MiniMax H3 | Text to Video API through each::labs can automate generation of short explainers or UI demos from structured text descriptions and style references, keeping durations within 5–15 seconds for consistent performance.

  • Tips & tricks

    Tips and Tricks

    To get strong results from MiniMax H3 | Text to Video, describe camera movement, lighting, and mood directly in the prompt, and specify your desired aspect ratio and duration up front. The model responds well to structured prompts that separate scene description, motion, and audio cues, especially when paired with a few high-quality reference images or short video clips. Avoid overly long or ambiguous prompts; instead, use clear phrases like “slow dolly shot,” “dynamic typography,” or “soft ambient soundtrack.” When using omni-reference mode, keep references consistent in style so the model does not have to reconcile conflicting visual directions.

    Example prompts:

    “16:9, 10-second cinematic shot of a cyberpunk city at night, neon reflections on wet streets, slow camera pan, subtle synth ambience and distant traffic.”

    “9:16 product teaser for a smartwatch, macro close-ups of rotating watch, glossy reflections, high-contrast studio lighting, energetic electronic music.”

    “21:9 title sequence with bold kinetic typography, abstract shapes morphing to the brand logo, smooth camera motion, modern orchestral trailer-style audio.”

  • Technical spec

    Technical Specifications

    • Provider / Family: Minimax, MiniMax H3 (Hailuo 3.0) multimodal video generation model.
    • Output resolution: Native 2K-class video; documented 1440p layer and fixed 2K aspect ratio presets.
    • Frame rate: Film-standard 24 frames per second.
    • Clip duration: Integer durations from about 5 to 15 seconds per generation; some integrations note 4–15 seconds.
    • Aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus adaptive framing where the model chooses.
    • Inputs: Text prompt, optional start/end images, image/video/audio references; up to 12 reference files in omni-reference mode.
    • Audio: Native stereo audio jointly modeled with video (voices, effects, music) in one pass.
    • Prompt length: Up to roughly 7,000 characters in documented APIs.
    • Typical latency: Implementations report interactive-generation times suitable for creative workflows; exact seconds depend on provider and hardware.
  • Things to be aware of

    Things to Be Aware Of

    MiniMax H3 | Text to Video is tuned for short clips, so attempts to tell long, multi-act stories in a single generation often result in rushed or visually inconsistent pacing. The model only supports specific 2K aspect ratios and a fixed frame rate, which means you cannot request arbitrary resolutions or frame rates directly. Complex omni-reference setups with many mixed-style images or conflicting audio can confuse scene composition or sound design, leading to less coherent outputs. As a newly released model, some implementation details—like precise latency or fine-grained control of audio content—may vary across providers, so you should validate behavior in your each::labs environment before relying on it for production-critical workflows.

  • Key considerations

    Key Considerations

    MiniMax H3 | Text to Video is optimized for short-form motion, so it works best when your idea fits into 5–15 seconds with a clear visual narrative. Because the model enforces a fixed set of 2K aspect ratios, you should choose the ratio that matches your target platform—such as 16:9 for widescreen or 9:16 for vertical social content—before generation. Native audio is generated along with the video, reducing the need for separate sound design, but you may still want post-production to fine tune timing and mix. For cost-performance, use concise prompts and minimal references first, then add more images, clips, or audio only when necessary to steer style or continuity.

  • Limitations

    Limitations

    MiniMax H3 | Text to Video currently supports clips up to about 15 seconds, so it cannot natively generate long-form videos or multi-minute narratives in one pass. Output is constrained to 2K-class resolutions and a fixed 24 fps cadence, limiting direct control over resolution tiers and motion blur characteristics. While audio is generated jointly, users have limited fine-tuning over specific voice identity, lyrics, or detailed sound design compared with dedicated audio tools. Finally, documentation does not yet expose full architecture parameters or training data details, so absolute performance benchmarks and dataset-level guarantees are not publicly confirmed.

Related models

4 models
* FAQ

About MiniMax H3 API

01 / 05

What is MiniMax H3 Text-to-Video?

MiniMax H3 Text-to-Video is a video generation model that turns a written prompt into a finished 5 to 15 second clip in 2K resolution. It follows detailed directions for camera movement, pacing, on-screen text, and sound design, which makes it a fit for ads, brand films, game visuals, and short-form social content.