all dispatches
Aug 5, 20268 min read

MiniMax H3:Speaks Your Visual Language

One good clip is not hard any more. Write a decent prompt, generate a few times, pick the best one. That problem is solved and has been for a while. Try making the second clip. Same character, same wardrobe, same face, a different camera angle. Suddenly her jacket is a slightly different green. Her jaw is narrower. The light comes from the other side of the room and nobody can say why. You regenerate. Now the jacket is right and the hair is wrong. Twenty minutes later you're choosing between t

MiniMax H3:Speaks Your Visual Language

One good clip is not hard any more. Write a decent prompt, generate a few times, pick the best one. That problem is solved and has been for a while.

Try making the second clip.

Same character, same wardrobe, same face, a different camera angle. Suddenly her jacket is a slightly different green. Her jaw is narrower. The light comes from the other side of the room and nobody can say why. You regenerate. Now the jacket is right and the hair is wrong. Twenty minutes later you're choosing between two takes that are each wrong in a different way, and the sequence you were building has quietly become impossible.

That's the wall. Not quality, consistency. And it's the reason so much generative video ends up as one-shot novelty rather than something you can build a scene, a campaign or a product feature on.

MiniMax H3 is interesting because it attacks that wall directly, and it does so by changing what you hand the model.

0:00
/0:15

What MiniMax H3 Actually Is

An open-weight multimodal generation model from MiniMax, also known as Hailuo 3.0. It produces clips of roughly five to fifteen seconds at native 2K, twenty-four frames per second, with stereo audio generated in the same pass as the picture.

Those numbers are worth sitting with for a second, because two of them are unusual.

2K, natively. Not upscaled from something smaller. The documented presets go up to 2560×1440 for widescreen and 2944×1248 for ultrawide, with matching sizes for square and vertical. That's a resolution you can actually cut into a deliverable rather than a proof of concept you'll have to regenerate later at a size that works.

Twenty-four frames per second. Film cadence, not a compromise between film and web. Motion reads the way an audience expects motion to read, which sounds like a small thing until you put generated footage next to shot footage and the difference announces itself.

But the resolution isn't the reason to pay attention. The inputs are.

0:00
/0:15

Show It, Don't Just Describe It

Most video models give you one channel: text. You describe the character, you describe the camera, you describe the mood, and the model interprets. Interpretation is exactly where consistency dies, because the model interprets slightly differently every time you ask.

MiniMax H3 accepts references. Up to nine images, up to three video clips, up to three audio tracks — twelve files in total in a single request, alongside a prompt that can run to roughly seven thousand characters.

Read that list again and think about what each slot is for. Images lock identity: this face, this wardrobe, this palette. Video clips lock behaviour: move the camera like this, pace the edit like this. Audio locks sound: this voice, this ambience, this musical register.

You're no longer describing a character and hoping the model builds the same one twice. You're showing it the character. The difference between telling and showing is the difference between a clip and a sequence.

There's a real constraint attached, and it's worth knowing before you load up twelve files: the references have to agree with each other. Mixed styles, conflicting lighting, a video reference whose energy fights your image reference — the model has to reconcile those, and reconciliation costs you the consistency you came for. Three strong references beat nine mediocre ones every time.

Every Generation Ships With Audio

Stereo audio comes out of MiniMax H3 by default, in the same pass as the frames: voice, sound effects, ambience, music. There's no separate audio step, no attaching a track afterward and nudging it into alignment.

Where this earns its place is the boring middle of production. A product spot needs room tone. An explainer needs a voice. A social cut needs something under it that isn't silence. Each of those is normally a small errand — find the file, licence it, drop it in, adjust. Individually trivial, collectively the reason a two-hour job takes a day.

Be honest about the ceiling, though. Generated audio here is strong on ambience and effects and weaker on precision. You don't get fine control over exact voice identity, specific lyrics, or the kind of detailed sound design a mixer would sign off on. Treat it as a good scratch track that sometimes survives to final, not as a replacement for a sound designer.

0:00
/0:10

Three Models, Three Ways to Anchor a Shot

There are three MiniMax H3 video models on each::labs, and they aren't interchangeable. Each one anchors a different part of the shot.

Text to Video anchors nothing but the description. Prompt, aspect ratio, duration. This is the exploration mode — where you find out what an idea looks like before you commit assets to it. It's also the one that rewards long prompts most, since text is the only steering you have.

Image to Video anchors the frame, and optionally both ends of it. You can pass a first frame, a last frame, or both, and the model animates the path between them. That second slot changes the character of the tool completely: with a first and last frame you're not requesting motion, you're specifying a start and a destination and letting the model solve the middle.

The exploded-view product animation is the obvious use — assembled burger in, separated burger out, model handles the float. But it applies anywhere a client has approved an opening and closing composition and you need both to appear exactly as approved. Output follows your image's aspect ratio, so the composition you designed stays the composition you get.

Reference to Video anchors identity. This is where the nine-image, three-video, three-audio ceiling lives, and it's the model you reach for when the same character has to appear in shot after shot without drifting. Character work, episodic content, anything with a mascot or a recurring face.

The pattern is simple once you see it: text for exploring, frames for composition, references for continuity. Most teams end up using more than one on the same project.

Why Open Weights Matter Here

MiniMax H3 shipped with open weights, and that matters more for this kind of model than it might seem.

If you're building a product feature on generated video, you're taking on a dependency. Closed models change underneath you — a new version behaves differently, a parameter gets deprecated, what you tuned against last quarter renders differently this quarter. With open weights there's a version that stays the version. Anyone who has had to explain why last month's output can't be reproduced knows that isn't a technical detail. It's the difference between a feature you can promise and one you can only demo.

0:00
/0:05

Where People Actually Use MiniMax H3

Brand and campaign teams producing across formats, where 2K output and the full ratio set from 21:9 down to 9:16 means one concept covers cinema, web and vertical without a regeneration pass for each.

Studios and agencies doing character work, where reference-driven identity locking is the whole reason the model is on the shortlist. A mascot that stays the same mascot across eight shots beats any individual shot being marginally prettier.

Product teams turning packshots into motion, using first-and-last-frame control so the hero composition arrives approved and departs approved. Creators shipping short-form volume, where fifteen seconds with sound attached is a finished post rather than the start of an editing session.

And teams building AI-native apps, where open weights and a documented reference schema make it something you build on rather than something you rent.

Getting Better Results Out of MiniMax H3

Prompt in layers. Subject, motion, camera, style, audio. The model handles prompts up to around seven thousand characters, but structure matters more than volume — a prompt organised by layer beats a longer one that mixes everything together.

Say the audio out loud. Name the ambience, the register, whether there's music. Audio gets generated whether you specify it or not, so specifying is free and guessing isn't.

Fewer, better references. The temptation with a nine-image ceiling is to use nine. Resist it. Aligned references in consistent style and lighting produce far better continuity than a mixed set the model has to average.

Pick your ratio first. Ratios are fixed presets, frame rate is locked at 24fps. You can't crop your way out of a wrong choice without losing composition.

Start mid-range. Eight to twelve seconds is the comfortable zone for validating motion. Push to the edges once the look is settled.

Using MiniMax H3 on each::labs

All three models run through one endpoint on each::labs. Switching between them is a change of model string and input shape:

bash

curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "minimax-h3-text-to-video",
    "version": "0.0.1",
    "input": {
      "ratio": "16:9",
      "prompt": "A chef tosses vegetables in a hot pan as flames flare, warm kitchen light, shallow depth of field, handheld camera. Audio: sizzling, extractor hum, distant service noise.",
      "duration": 5
    },
    "webhook_url": ""
  }' \
  https://api.eachlabs.ai/v1/prediction/

Point model at minimax-h3-image-to-video and pass first_frame_url, optionally with last_frame_url. Point it at minimax-h3-reference-to-video and pass reference_image_urls. Same auth, same call, same response handling.

The response is a prediction ID rather than a video, which you poll or take by webhook. Design for that from the start — generation latency varies, and any integration built on a synchronous assumption breaks the first time it meets real traffic.

The Honest Limitations

I don't want this to read like a brochure with the rough edges sanded off, so here's the other side.

Fifteen seconds is the ceiling. Anything longer is a multi-shot workflow you assemble yourself. That's a real constraint on narrative work, and the reference system is what makes it survivable — consistency across shots is the only reason multi-shot assembly is viable at all here.

The duration window isn't consistently documented. Some sources describe five to fifteen seconds, others four to fifteen. Small discrepancy, but if you're building automated pipelines with hard duration values, validate the actual accepted range rather than trusting a spec sheet.

No arbitrary resolutions or frame rates. Fixed 2K presets, fixed 24fps. If your delivery spec calls for something else, you're conforming in post.

It isn't fast. Median runtime sits around five minutes for text and image inputs, closer to seven for reference-driven generation. Fine for a queue, wrong for anything a user is waiting on. Build the waiting into your UX.

Conflicting references degrade output. This is the failure mode people hit hardest, precisely because the reference system is the feature. Mismatched inputs don't get creatively reconciled — they get averaged, and averaged is worse than either input alone.

Asset limits are real, and output is heavy. Video references cap around 50MB, images around 30MB, audio around 15MB, twelve files total — large sets need preprocessing first. On the way out, 2K with audio is not a light payload, and storage, bandwidth and preview performance all feel it at volume.

Independent benchmarks are thin. Full architecture and training details aren't public, and most published performance numbers trace back to the lab. Read them as a starting point, and test against your own material before you commit a pipeline to them.

Wrapping Up

The headline numbers on MiniMax H3 — 2K, 24fps, native audio — are the reason it gets on a shortlist. They're not the reason it stays there.

What makes it worth building on is that it takes references seriously. Nine images, three clips, three audio tracks, and a real attempt at holding identity steady across shots. That's the difference between a model that produces impressive clips and one you can plan a sequence with. Add open weights, and it becomes something you can commit to rather than something you're hoping doesn't change.

Shot one was never the hard part.

You can try MiniMax H3 on each::labs.

Frequently Asked Questions

What makes MiniMax H3 different from prompt-only video models?

The inputs. Text-only generation means the model reinterprets your description on every call, which is why characters drift between shots. MiniMax H3 accepts images, video and audio as references in the same request, so identity, motion style and sound are shown rather than described. If you only ever need one clip, this barely matters. If you need six that belong together, it's the whole game.

Which of the three MiniMax H3 models should I use?

Ask what needs to stay fixed. Nothing yet, use Text to Video and explore. A specific opening or closing composition, use Image to Video with first and last frames. A character or style that has to survive across multiple shots, use Reference to Video. Plenty of projects use two of the three at different stages.

Does MiniMax H3 audio replace a sound designer?

No, and it's worth being clear about that. Every generation includes stereo audio covering voice, effects, ambience and music, which removes a whole category of small errands from short-form work. What you don't get is precise control over voice identity, specific lyrics or detailed mix decisions. It's a strong first pass. On anything where sound carries the piece, budget for a real one.