MiniMax H3·Hailuo 3.0·AI Video Generator·AI Video Models·

What Is MiniMax H3? Specs, Modes & First Look (2026)

MiniMax H3 (Hailuo 3.0) explained: 15-second multi-shot video with native stereo audio, 12-file input, and instruction-based editing. Full specs inside.

What Is MiniMax H3? Specs, Modes & First Look (2026)

On July 31, 2026, MiniMax shipped the most ambitious model in its Hailuo line — and quietly changed what "video model" means. MiniMax H3 (the community already calls it Hailuo 3.0) doesn't just turn prompts into clips. It reads text, images, video and audio as one unified context, generates 15-second multi-shot scenes with stereo sound baked in — and then lets you edit the result with a plain-language instruction instead of re-rolling the dice.

We build Pixo, a multi-model AI video platform, which means every time a frontier model drops we do two things: read the official documentation line by line, and test it against the models we already run. This first look is based on MiniMax's official H3 usage manual and launch materials — including details that haven't made it into most English coverage yet, like the exact input limits, the three generation modes, and how the prompt format actually works.

Here's what H3 is, what's genuinely new, and what it means if you make videos for a living.

MiniMax H3 at a Glance

SpecMiniMax H3
ReleaseJuly 31, 2026 (previewed at WAIC 2026)
Output length4–15 seconds, 24 fps
Resolution768p mode (upgradable to 1440p); 1440p "2K" mode officially recommended — up to 2976×1248 at 21:9
Aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, 9:16
AudioNative stereo on every output — dialogue, ambience, SFX in the same pass
Multimodal inputUp to 12 files: 9 images + 3 video clips + 3 audio tracks
EditingInstruction-based edits on existing footage (characters, backgrounds, dialogue, voices)
LanguagesMultilingual prompts; TTS covers 11 languages precisely, ~40 via derivation
Prompt limitUp to 7,000 characters
AccessHailuo AI app, MiniMax API (model ID MiniMax-H3)

The Big Idea: One Context, Every Modality

Most AI video tools are pipelines of special-purpose models: one for text-to-video, another for lip-sync, another for upscaling, something else for sound. MiniMax's pitch with H3 is to collapse that pipeline into a single "omni-modal" system. In the official manual's words, H3 stops treating image, video and sound generation as separate tasks and instead reads all of your materials — a script, a face reference, a motion clip, a music track — as one brief, then produces a finished audiovisual scene from it.

Concretely, that means you can hand H3 up to 12 files in one prompt: as many as 9 images, 3 video clips (15 seconds total), and 3 audio tracks. Each reference gets a declared role — this image locks the character's face, this video is the motion reference, this track sets the pacing. The model fuses them rather than collaging them. If you've ever tried to keep a character consistent while also matching a motion reference and syncing to a beat across separate tools, you'll understand why this matters.

Sound is not an afterthought

Every H3 generation ships with native stereo audio — dialogue with lip-sync, room tone, effects — created in the same pass as the pixels. Only a handful of models do true native audio (Seedance 2.0, Veo 3.1, Wan 2.6 are the notable others), and H3 pushes further on the language side: its text-to-speech covers 11 languages precisely — Chinese, English, Japanese, Korean, French, German, Spanish among them — with around 40 more reachable through derivation. It can also clone and migrate voice timbre from a reference track, which turns "localize this ad into three markets" from a production project into a prompt.

Editing is the headline feature

Here's the part that separates H3 from almost every clip generator we track: it takes edit instructions against existing footage. The manual's own examples are wonderfully mundane — "replace the cat in the video with a dog" — and escalate to things like swapping a glowing convenience-store sign, changing a can of soda into a branded cola, and rewriting the closing line of dialogue, all in one instruction, all while everything else in the shot stays put. You can change backgrounds and relight scenes, replace a line of dialogue and have the performance subtly adjust to match, or migrate a character's voice to a cloned timbre. You can try this today in Pixo's Playground — upload your footage as a video reference and write the instruction against it.

For working creators, this attacks the single most expensive habit in AI video: re-rolling. When a shot is 90% right, you don't want a new shot — you want the other 10%.

Pixo

Create AI videos with Pixo

Turn any idea into a publish-worthy video. One sentence is all it takes.

The Three Generation Modes

The manual defines three distinct ways to drive H3, and knowing which one you're in changes how you should prompt (we wrote a full MiniMax H3 prompt guide based on the official formula):

1. Omni-reference mode. The full 12-file multimodal input described above. You label every asset's role — face lock, motion reference, pacing track — and describe the scene.

2. First/last-frame mode. Give it one or two images. With two, H3 treats them as the first and last frames and fills in the motion, lighting and sound between them — notably, it will not invent camera cuts in this mode. Aspect ratio follows your input image.

3. Pure text-to-video. No references at all; the model builds subject, scene and action from your description. Prompts can run to 7,000 characters, and H3 rewards specific, visual, camera-aware writing over vibes.

How H3 Fits the Current Model Landscape

The honest answer: H3 doesn't dominate every axis, and it picked a crowded day to launch. Kling 3.0 and Veo 3.1 still output higher raw resolution (4K). And on July 31 — the very same day H3 shipped — ByteDance released Seedance 2.5, which pushes single-pass generation to as long as 180 seconds and adds its own instruction-based edit modes. The generate-then-edit idea clearly isn't H3's alone anymore; we've compared the two flagships in detail in MiniMax H3 vs Seedance 2.5.

What H3 still distinctly owns is the voice and language half of editing: rewriting a line of dialogue and having the performance adjust, cloning a voice from a reference track, and TTS across 11 languages in the same system that generated the picture. If your work involves brands, dialogue or multiple markets — rather than one-off spectacle shots — that combination is worth more than another notch of resolution.

One model still isn't a movie, though. A 15-second generator becomes genuinely useful when it slots into a storyboard next to other models — Seedance for a physics-heavy action beat, H3 for the dialogue scene you'll want to revise later. That per-shot, multi-model workflow is exactly what we've built Pixo's agent around. H3 is now live in Pixo's Playground for single-shot generation, and the storyboard runs Seedance 2.0, Kling 3.0 and Veo 3.1 today — use the Hailuo hub as your starting point.

Access and Pricing

H3 is live in the consumer Hailuo AI app and through MiniMax's open-platform API under the model ID MiniMax-H3. Launch API pricing works out to roughly $0.78 for a 6-second 2K clip with audio — noticeably above the legacy Hailuo 2.3 rates ($0.28 at 768p), which tells you where MiniMax thinks this model sits in the market. MiniMax has also said open weights are coming. And the API-first strategy is clearly working: within the first week, H3 has already surfaced across a wave of third-party creative platforms — raw access to this model is spreading fast, and won't be anyone's moat for long.

If you'd rather work at the level of finished videos than raw clips, Pixo runs the leading video models — Seedance, Kling, Veo, Hailuo and more — behind one agent-driven storyboard workflow. MiniMax H3 is now live in Pixo's Playground — generate 4–15 second shots at 768p or 2K, with image, video and audio references. Sign up now — new users get 200 free credits on sign-up, and plans are currently up to 55% off.

FAQ

What is MiniMax H3?

MiniMax H3 (also called Hailuo 3.0) is an omni-modal video model released by MiniMax on July 31, 2026. It treats text, images, video and audio as one unified context and generates 4–15 second multi-shot videos at up to 2K (1440p) with native stereo audio in a single pass.

Is MiniMax H3 the same as Hailuo 3.0?

Yes. MiniMax is the company, Hailuo is its consumer video brand, and H3 is the third-generation model. The API model ID is MiniMax-H3, and much of the community refers to it as Hailuo 3.0 — they're the same model.

What resolution and length can MiniMax H3 output?

H3 generates 4–15 second clips at 24 fps in six aspect ratios from 21:9 to 9:16. It runs in a 768p mode (results upgradable to 1440p) and a 1440p mode that MiniMax officially recommends — at 21:9 that's up to 2976×1248.

Does MiniMax H3 generate audio?

Yes — every H3 output includes native stereo audio generated in the same pass as the video: dialogue with lip-sync, ambient sound and effects. Its text-to-speech covers 11 languages precisely, with roughly 40 more available through derivation.

Can MiniMax H3 edit videos?

Yes, and this is its most distinctive capability. H3 accepts plain-language edit instructions against existing footage — swap a character or object, change the background or lighting, rewrite a line of dialogue, or migrate a voice — while keeping unedited regions stable.

How can I try MiniMax H3?

H3 is live in the Hailuo AI app, via the MiniMax open-platform API (model ID MiniMax-H3), and is already appearing on a number of third-party creative platforms. API pricing at launch works out to roughly $0.78 for a 6-second 2K clip with audio.

From idea to finished video.
In one conversation.

Start Creating

No credit card required • Free 200 credits