MiniMax H3 Explained: 2K Video, Native Audio, and What It Really Costs

MiniMax H3 Explained: 2K Video, Native Audio, and What It Really Costs
You finish a personality test, you get four letters and a description that feels uncomfortably accurate, and then… you screenshot it. That has been the end of the road for years. The result is a static card, and static cards do not travel well.
What changed is that turning a character description into a short film stopped requiring a production budget. A model called MiniMax H3 will now take a paragraph about a personality type and return a 15-second, 2K clip with synchronized audio, for roughly the price of a bus fare. This article explains what the model actually is, does the pricing math honestly, and then shows what it looks like applied to something personal.
First, the Naming Confusion
You will see three names for the same thing, and it causes real confusion:
- MiniMax H3 — the model name, used in research and API contexts.
- Hailuo 3.0 (or Hailuo 03) — the consumer product brand MiniMax ships it under.
- H3 — shorthand, used everywhere.
They are the same model. If a comparison article pits "MiniMax H3 vs Hailuo 3," it is comparing a model to itself and you can safely skip it.
What the Model Actually Does
The headline capability is that H3 is multimodal on the input side. Most video models accept a text prompt and maybe one image. H3 accepts text, images, video, and audio in the same request, holds them in one context, and generates the result in a single pass — video and sound together, not sound stitched on afterwards.
| Specification | MiniMax H3 |
|---|---|
| Resolution | Native 2K at 24fps (768p and 4K tiers available) |
| Duration | 4–15 seconds per clip |
| Audio | Native stereo, generated with the video |
| Image inputs | Up to 9 reference images |
| Video inputs | Up to 3 clips (motion and camera transfer) |
| Audio inputs | Up to 3 tracks (pacing, ambience, voice) |
| Frame control | First frame, optional last frame |
| Aspect ratios | 16:9, 9:16, 1:1, 4:3, 3:4 |
| Prompt limit | 2,000 characters |
| Licensing | Open weights |
Two consequences follow from that table, and they are the reason H3 got attention.
Audio is directable. Because sound is generated in the same pass as the picture, you can art-direct it the way you art-direct a shot. "Rain on a window, no music, she says 'I'll wait' quietly" is a valid instruction, and the delivery lands on the visual beat instead of drifting against it. Models that generate silent video and add audio later cannot do this reliably, and the mismatch is the thing your eye catches even when you cannot name it.
References are constraints, not suggestions. Nine image slots plus three video slots plus three audio slots means you can specify identity, motion, and rhythm separately. Use a photo or illustration for who the subject is, a video clip for how the camera moves, an audio track for pacing. Each input governs a different axis instead of all of them fighting over the same one.
The Pricing Math, Done Honestly
This is where H3 made its case, so it deserves actual numbers rather than adjectives.
| Model | Approx. cost per second | Resolution | Native audio |
|---|---|---|---|
| MiniMax H3 | ~$0.13 | 2K | Yes |
| Veo 3.1 Standard | ~$0.40 | 720p / 1080p | Yes |
| Veo 3.1 Fast | ~$0.12–0.14 | 720p / 1080p | Yes |
Against Veo 3.1 Standard, that is roughly a 3× difference — for output at a higher resolution. That is the comparison the launch coverage led with, and it holds up.
The honest caveat: Veo 3.1 Fast is priced within a cent or two of H3, and on some third-party hosts it comes in cheaper. So "H3 is a third of the price" is true against the standard tier and misleading against the fast tier. What survives either comparison is the resolution and the reference stack — you are getting 2K and nine-image conditioning at fast-tier prices, and that is the actual argument. Compare like for like on what actually comes out the other end, rather than on headline rates, and the gap reopens.
On the consumer side, access is credit-based rather than per-second: plans run from around $7.50/month for 1,000 credits up to about $49/month for 11,000, with generation starting at 50 credits per second and commercial usage rights included at every tier. Current plans and credit costs are worth checking directly, since credit pricing on these platforms moves faster than articles about it.
What Open Weights Change
MiniMax released H3 as an open-weight model. This is not a small footnote.
Closed video models are rented. You send a prompt, you get a clip, and the terms can change beneath you — pricing, content policy, availability, all outside your control. Open weights mean the model can be run through community tooling, integrated into local pipelines, and kept working after any particular company loses interest in it. For anyone building something that needs to still exist in three years, that difference outweighs a few points of benchmark quality.
It also means the prompting knowledge you build is durable. Learning to write for a hosted API that gets deprecated is wasted effort. Learning to write for a model whose weights are public is not.
Turning a Personality Result Into a Short Film
Here is the part that connects to why you are on this site.
A personality type is a character description. It has traits, a way of moving through the world, a specific emotional register. That is exactly the input a video model wants — far better than the vague "cinematic shot of a woman, beautiful, 4K" prompts that produce nothing memorable.
The structure that works is six parts in order: subject, environment, action, camera, audio, ending state.
Template 1 — The devoted, steady type
A woman in her late twenties in a soft grey knit sweater, sitting at a kitchen
table by a rain-streaked window in early evening. She is writing something by
hand, pauses, and looks up without hurry. Camera holds still at a medium
three-quarter angle, shallow depth of field, warm practical lighting.
Audio: rain against glass, a kettle settling in the background, no music,
pen on paper. She ends looking toward the window, calm, half-smiling.
Template 2 — The guarded, observant type
A person standing at the back of a crowded night bus, city lights sliding
across the window behind them, wearing a dark coat with the collar up. They
watch the reflections rather than the people. Camera is handheld, slightly
unstable, framed at a low medium shot. Audio: engine hum, muffled conversation
too indistinct to follow, no music. They end turning their face away from
camera as the bus takes a corner.
Template 3 — The warm, expressive type
A young man in a bright yellow raincoat crossing a wet plaza at midday,
laughing at something off-camera and gesturing broadly with one hand.
Camera tracks alongside him at walking pace, eye level, bright and saturated.
Audio: rain, footsteps in shallow water, his laugh, distant traffic.
He ends mid-stride looking back over his shoulder toward camera.
Notice the pattern across all three. Each one is a single shot with a single emotional beat. None tries to compress a whole narrative arc into fifteen seconds. The temperament is expressed through how the camera behaves and what you hear, not through the subject explaining themselves.
If you have an image already — an avatar, an illustration, a character card — feed it as the opening frame through the image-to-video mode and describe only the motion and the sound. Starting from nothing but a written description? Text-to-video takes the same six-part structure directly.
And if you would rather start from something that already works than from a blank field, the official prompt library pairs finished videos with the complete prompts that produced them. Reading four or five before writing your own is the fastest way to internalize the structure — you stop guessing at what level of detail the model actually responds to.
The Limits Worth Planning Around
Every video model has a shape. Knowing H3's in advance saves you the credits you would otherwise spend discovering it.
Fifteen seconds is a hard ceiling. H3 makes shots, not films. Anything longer means generating several clips and cutting them together. Plan around it: write your lighting and wardrobe description as a fixed block of text and reuse it verbatim across every generation. Clips built that way cut together far more cleanly than they have any right to.
One shot per generation. Ask for a cut between two locations inside a single clip and you will usually get an incoherent blend. Plan around it: storyboard first, generate each shot separately. This is how commercial work gets made regardless — nobody shoots a three-location sequence in one take either.
Text rendering is good, not perfect. On-screen text and brand marks are handled noticeably better than earlier models managed, but anything past a few words still needs checking. Plan around it: keep in-frame copy to short phrases and add longer text in an editor, where you control the typography completely.
Faces drift under fast motion. The nine-image reference stack helps a great deal; it does not fully solve it. Plan around it: slower, more deliberate motion holds identity better — which, conveniently, is exactly what the reflective character clips in this article call for. When you do need speed, frame wide enough that the face is not carrying the shot.
None of these are dealbreakers once you know them. They are the difference between treating H3 as a slot machine and treating it as a camera.
Who This Is Actually For
If you need broadcast-grade photorealism and have the budget for it, the premium closed models still have an edge in specific scenarios and you should test both. If you are making social clips, character pieces, product shots, or anything where a fifteen-second window is the format rather than a limitation, the value argument is straightforward: 2K output, real audio, nine-image conditioning, and open weights, at roughly what other models charge for 720p. For most people reading this, that is not a close call.
The fastest way to know whether it fits your work is to run one prompt from this article through the MiniMax H3 video generator and see whether what comes back looks like the person you had in mind. That takes about two minutes, which is less time than you have already spent reading about it.