MiniMax | H3 | Max | Reference to Video

Video·minimax-h3·by Minimax

H3 Max Reference-to-Video creates short AI video clips from prompts and reference images, preserving subject cues for character-driven creative scenes.

Runtime (p50)
5m
Estimated price
From $0.05
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "minimax-h3-max-reference-to-video",
    "version": "0.0.1",
    "input": {
        "prompt": "Use Image 1 as the locked character and set reference. Preserve the young woman exactly as she appears — her face shape, green eyes, full natural brows, dewy skin with visible texture and the small beauty mark, slicked-back dark bun with loose face-framing strands, gold hoop earrings, thin gold necklace, and white off-shoulder textured knit top. Preserve the glossy black mascara tube and its wand exactly. Keep the woman, the product, and the beige living-room set — matte wall, slim olive tree in a pot, framed print, candle on a wooden ledge, cream sofa — identical and consistent in every shot. Use Image 1 for storyboard order and pacing, following the twelve panels left to right, top to bottom.\n\nRender in high-quality 4K, vertical 3:4 live-action UGC beauty commercial with clean commercial production value: intimate, effortless, and quietly premium. Use soft diffused daylight from camera left, a warm neutral palette (cream, sand, soft gold), and an unhurried morning-routine atmosphere throughout. Shoot handheld with subtle natural sway and a shallow depth of field, real skin texture with no beauty smoothing, and no visible on-screen text. Follow the storyboard beat by beat with natural camera movement and seamless cuts on action, never a slideshow.\n\nMotion beats in order: she lifts the closed tube beside her cheek and tilts her head to camera; she lowers her gaze and blinks, bare lashes visible; extreme close-up as she blinks slowly twice; her hand rotates the closed tube as light slides across it; both hands twist the cap off in one smooth motion; the hand lifts the wand and turns the bristles to camera; she sweeps the wand upward through her upper lashes and pauses; she brushes twice more from a slightly different angle; she finishes a final upward stroke and lowers the wand; extreme close-up as her eye opens wider and blinks, lashes lifted and separated; she turns slightly and breaks into a soft smile; she tilts the tube toward camera and smiles with a small shoulder shrug.\n\nAudio: a soft modern commercial soundtrack throughout — warm minimal pop with a gentle four-on-the-floor pulse, muted plucked synth, light finger snaps, and airy pads. Understated and confident, never energetic or club-like. The track builds subtly through the application beats and resolves on a warm sustained chord at the final smile. Add quiet ambient product sound design under the music: a soft click as the cap twists off, a delicate brush sweep on the lashes. No voiceover, no lyrics, no dialogue.\n\nKeep her face fully visible and undistorted in every shot with natural eye anatomy and symmetrical features. The applicator hand always stays to the side at ear level and never crosses or covers the face. Render hands with correct anatomy and five fingers, especially in the macro cap-twisting and wand shots. Show the face in medium close-up, close-up, or extreme close-up only; use hands-only macro framing for the three product shots, with no face in those frames. ",
        "duration": 10,
        "resolution": "768P",
        "aspect_ratio": "3:4",
        "reference_image_urls": [
            "https://cdn-us.eachlabs.ai/defaults/533ec6794ee34df785d1ae44d1be6fad.png"
        ],
        "enable_safety_checker": true,
        "prompt_expansion_mode": "quality"
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/

Related models

4 models
* FAQ

About MiniMax | H3 | Max | Reference to Video

01 / 03

What is H3 Max Reference-to-Video?

H3 Max Reference-to-Video is a MiniMax video generation model for creating short clips from a text prompt and reference images. It can preserve visual subject cues from the images while the prompt describes motion, setting, and scene direction. The model returns a generated video file.