Skip to content
MiniMax H3·Hailuo 3.0·AI Video Generator·AI Video Models·

What Is MiniMax H3? Specs, Modes & First Look (2026)

MiniMax H3 (Hailuo 3.0) explained: 15-second video with native stereo audio, 12-file multimodal input, instruction editing — and open weights with a catch.

Pixo Team·9 min read
What Is MiniMax H3? Specs, Modes & First Look (2026)

On July 31, 2026, MiniMax shipped the most ambitious model in its Hailuo line — and quietly changed what "video model" means. MiniMax H3 (the community already calls it Hailuo 3.0) doesn't just turn prompts into clips. It reads text, images, video and audio as one unified context, generates 15-second scenes with stereo sound baked in — and then lets you edit the result with a plain-language instruction instead of re-rolling the dice.

We build Pixo, a multi-model AI video platform, which means every time a frontier model drops we do two things: read the official documentation line by line, and test it against the models we already run. This first look is based on MiniMax's official video-generation docs and the H3 model card — including details that haven't made it into most English coverage yet, like the exact input limits, the three generation modes, and the licence restriction that decides whether you can run this model at all.

Here's what H3 is, what's genuinely new, and what it means if you make videos for a living.

MiniMax H3 at a Glance

SpecMiniMax H3
ReleaseJuly 31, 2026 (weights followed August 3)
Architecture33B-parameter dense, single-stream omni-modal transformer
Output length4–15 seconds, 24 fps
Resolution768p default; up to 2K
Aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, 9:16
AudioNative stereo at 32 kHz — dialogue, ambience, SFX in the same pass
Multimodal inputUp to 12 files total: ≤9 images + ≤3 video clips + ≤3 audio tracks
EditingInstruction-based edits on existing footage
LanguagesSpeech in 11 languages
AccessHailuo AI app, MiniMax API (model ID MiniMax-H3), open weights

The Big Idea: One Context, Every Modality

Most AI video tools are pipelines of special-purpose models: one for text-to-video, another for lip-sync, another for upscaling, something else for sound. MiniMax's pitch with H3 is to collapse that pipeline into a single "omni-modal" system — architecturally, a 33B-parameter dense transformer that reads every modality in one stream rather than bolting encoders onto a video backbone. In practice that means it stops treating image, video and sound generation as separate tasks and instead reads all of your materials — a script, a face reference, a motion clip, a music track — as one brief, then produces a finished audiovisual scene from it.

Concretely, you can hand H3 up to 12 files in one prompt: as many as 9 images, 3 video clips (15 seconds total), and 3 audio tracks. Each reference gets a declared role — this image locks the character's face, this video is the motion reference, this track sets the pacing. The model fuses them rather than collaging them. If you've ever tried to keep a character consistent while also matching a motion reference and syncing to a beat across separate tools, you'll understand why this matters.

Sound is not an afterthought

Every H3 generation ships with native stereo audio at 32 kHz — dialogue with lip-sync, room tone, effects — created in the same pass as the pixels. Only a handful of models do true native audio, and H3 pushes further on the language side: its speech covers 11 languages — Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian and Spanish. It can also clone and migrate voice timbre from a reference track, which turns "localize this ad into three markets" from a production project into a prompt.

Editing is the headline feature

Here's the part that separates H3 from almost every clip generator we track: it takes edit instructions against existing footage. MiniMax's own examples are wonderfully mundane — "replace the cat in the video with a dog" — and escalate to things like swapping a glowing convenience-store sign, changing a can of soda into a branded cola, and rewriting the closing line of dialogue, all in one instruction, all while everything else in the shot stays put. You can change backgrounds and relight scenes, replace a line of dialogue and have the performance subtly adjust to match, or migrate a character's voice to a cloned timbre.

For working creators, this attacks the single most expensive habit in AI video: re-rolling. When a shot is 90% right, you don't want a new shot — you want the other 10%.

The Three Generation Modes

The docs define three distinct ways to drive H3, and knowing which one you're in changes how you should prompt (we wrote a full MiniMax H3 prompt guide based on the official formula):

1. Reference mode. The full 12-file multimodal input described above. You label every asset's role — face lock, motion reference, pacing track — and describe the scene.

2. Image-to-video / first-last-frame mode. Give it one or two images. With two, H3 treats them as the first and last frames and fills in the motion, lighting and sound between them — notably, it will not invent camera cuts in this mode. Aspect ratio follows your input image.

3. Pure text-to-video. No references at all; the model builds subject, scene and action from your description. H3 rewards specific, visual, camera-aware writing over vibes.

Open Weights — With a Territorial Catch

This is the detail most coverage has skipped, and it's the one that decides whether H3 is usable for you.

On August 3, MiniMax published H3's weights to Hugging Face — a genuinely open release of a frontier video model, and the strategic opposite of ByteDance keeping Seedance closed. But it landed under the MiniMax H3 Community Licence, not Apache or MIT, and the licence text defines its "Applicable Territory" as worldwide excluding the European Union, the United Kingdom, the Republic of Korea and the United States of America.

Read plainly: if you're in the US, the EU, the UK or South Korea, the community licence does not grant you rights to run these weights locally. Commercial use elsewhere also requires separate written authorisation above $20M in yearly revenue, plus visible "MiniMax H3" attribution in your product.

For most of our readers, that makes the hosted API — or a platform that licenses access — the practical route, not a local deployment. It's an unusual restriction, and worth knowing before you budget GPU time for a model you may not be licensed to run.

How H3 Fits the Current Model Landscape

The honest answer: H3 doesn't dominate every axis, and it picked a crowded day to launch. Kling 3.0 and Veo 3.1 still output higher raw resolution. And on July 31 — the very same day H3 shipped — ByteDance launched Seedance 2.5 globally on Dreamina, which pushes generation to 30-second single takes with a long-video mode reaching three minutes, and adds its own timestamped edit modes. The generate-then-edit idea clearly isn't H3's alone anymore; we've compared the two flagships in detail in MiniMax H3 vs Seedance 2.5.

What H3 still distinctly owns is the voice and language half of editing: rewriting a line of dialogue and having the performance adjust, cloning a voice from a reference track, and speech across 11 languages in the same system that generated the picture. If your work involves brands, dialogue or multiple markets — rather than one-off spectacle shots — that combination is worth more than another notch of resolution.

One model still isn't a movie, though. A 15-second generator becomes genuinely useful when it slots into a storyboard next to other models — Seedance for a physics-heavy action beat, H3 for the dialogue scene you'll want to revise later. That per-shot, multi-model workflow is exactly what we've built Pixo's agent around, and it's why we're preparing to bring H3 into Pixo's storyboard alongside Seedance 2.0, Kling 3.0 and Veo 3.1.

Access and Pricing

H3 is live in the consumer Hailuo AI app and through MiniMax's open-platform API under the model ID MiniMax-H3. List pricing is $0.13 per second of 2K video — about $0.78 for a 6-second clip with audio — with a cheaper 768p tier. That sits noticeably above the legacy Hailuo 2.3 rates, which tells you where MiniMax thinks this model belongs in the market. And the API-first strategy is clearly working: within the first week, H3 surfaced across a wave of third-party creative platforms. Raw access to this model is spreading fast, and won't be anyone's moat for long.

If you'd rather work at the level of finished videos than raw clips, Pixo runs the leading video models — Seedance, Kling, Veo, Hailuo and more — behind one agent-driven storyboard workflow, and MiniMax H3 is on its way into that lineup. Sign up now — new users get 200 free credits on sign-up — and you'll be ready to direct H3 the moment it lands.

Frequently Asked Questions

What is MiniMax H3?

MiniMax H3 (also called Hailuo 3.0) is an omni-modal video model released by MiniMax on July 31, 2026. It treats text, images, video and audio as one unified context and generates 4–15 second videos at up to 2K with native stereo audio in a single pass.

Is MiniMax H3 the same as Hailuo 3.0?

Yes. MiniMax is the company, Hailuo is its consumer video brand, and H3 is the third-generation model. The API model ID is MiniMax-H3, and much of the community refers to it as Hailuo 3.0 — they're the same model.

Does MiniMax H3 generate audio?

Yes — every H3 output includes native stereo audio at 32 kHz generated in the same pass as the video: dialogue with lip-sync, ambient sound and effects. Its speech covers 11 languages, including Chinese, English, Japanese, Korean, Arabic, French, German, Italian, Portuguese, Russian and Spanish.

Can MiniMax H3 edit videos?

Yes, and this is its most distinctive capability. H3 accepts plain-language edit instructions against existing footage — swap a character or object, change the background or lighting, rewrite a line of dialogue, or migrate a voice — while keeping unedited regions stable.

Is MiniMax H3 open source?

MiniMax published the 33B-parameter weights to Hugging Face on August 3, 2026, but under a community licence rather than a standard open-source one. The licence excludes the United States, European Union, United Kingdom and South Korea from its Applicable Territory, so creators in those regions should use the hosted API instead of self-deploying.

Specs in this article were read from MiniMax's official video-generation documentation, the MiniMax-H3 Hugging Face model card and the published licence text on August 5, 2026, and cross-checked against the models we run in production on Pixo. Pricing and licence terms change — verify against the primary sources before you budget.

Ready to Revolutionize your workflow?

Join thousands of creators using Pixo to turn their stories into visual reality.

Sign Up Now

No credit card required • Free 200 credits