Skip to content
AI Video With Sound·Native Audio·MiniMax H3·AI Video Generator·

AI Video Generator With Audio: MiniMax H3 Native Sound

Most AI video generators are silent. MiniMax H3 generates dialogue, ambience and effects with the picture — stereo, lip-synced. How it works on Pixo.

Pixo Team·7 min read
AI Video Generator With Audio: MiniMax H3 Native Sound

There's a moment every AI video workflow hits: the clip looks great, and it's completely silent. So you go find a music bed, a TTS tool for the voiceover, something for the sound effects — and then you discover the hard part, which is that a voice recorded over a video never quite belongs to it. The room tone is wrong, the mouth doesn't move, the footsteps land nowhere. Sound design becomes a second production.

That's the gap MiniMax's H3 closed when it shipped on July 31: it doesn't add audio to video, it generates them together. We run Pixo, a multi-model video platform where H3 is live in the Playground, and of everything in H3's spec sheet, this is the capability that changes daily work most — so here's exactly how it behaves and how to prompt it.

What "Native Audio" Actually Means

Every H3 output ships with stereo sound generated in the same pass as the picture. Three layers come out of one prompt:

  • Dialogue. Write the line verbatim and a character speaks it — with lip-sync, because the mouth movement and the audio come from the same generation. Rewrite the line and the performance follows.
  • Ambience. The scene sounds like the scene: wind over a wetland, the hum of a kitchen, traffic behind glass. You don't request it so much as fail to exclude it.
  • Effects. Contact sounds land where the physics happen — a lid clicks when it opens, wingbeats hit when the bird takes off.

The difference from dubbing isn't subtle. Layered post-audio is about the video; generated audio is of it. When a character turns away mid-sentence, the voice turns with them.

Who Has Sound, Who Doesn't

Audio has quietly become the dividing line between model generations:

ModelNative audio
MiniMax H3✅ Stereo on every output — dialogue, ambience, effects
Seedance 2.5✅ Yes
Veo 3.1✅ Yes
Wan 2.6✅ Yes — synced dialogue and effects
Kling 3.0❌ Silent output

What H3 adds on top of "has audio" is the voice stack: text-to-speech across multiple languages, and voice cloning from a reference track — upload a voice, and dialogue generates in that timbre. (That capability earns its own deep-dive; here we'll stay on the everyday case.)

Four Rules for Prompting Sound

1. Dialogue must be written out. H3 doesn't improvise lines. A tired voice says: "We open in five minutes." generates that sentence — nothing else. Vague requests ("she says something reassuring") produce vague results.

2. Decide who's on screen. A visible speaker gets lip-sync; an off-screen voice reads as narration. Both are one line of prompt — the difference is whether you put the speaker in frame.

3. Kill the music unless you want it. H3 will happily score your clip. If you're adding licensed music in post — which most brand work is — end the prompt with non-narrative music: N/A and keep the track clean.

4. Audio references ride along, never alone. You can upload an audio track (≤15MB) to set pacing or voice character, but it must accompany an image or video reference — audio can't be the sole input.

Three Copy-Paste Prompts

1. On-screen dialogue with lip-sync (16:9)

The showcase's reference files

Character reference
Character reference
Scene reference
Scene reference

A manual showcase in the same register

A cozy late-night diner, warm tungsten light, light rain on the window. A woman in her
thirties sits at the counter, hands around a coffee mug. She looks up at the camera and
says: "You always order the same thing." She smiles slightly after the line.
Quiet diner ambience — rain, a distant radio, cutlery. Non-narrative music: N/A.

Why it works: one speaker, one written line, one reaction beat. The ambience list gives the sound stage without competing with the dialogue.

2. Narrated product moment (9:16)

The showcase's reference files

Voice timbre sample

A manual showcase in the same register

Vertical video. A pour-over coffee setup on a wooden counter, morning light. Hot water
spirals from a kettle into the filter; steam rises; the camera pushes in slowly.
A calm male voice says: "Thirty grams. Three minutes. No shortcuts."
Sound of water pouring and gentle steam. Non-narrative music: N/A.

Why it works: off-screen narration means no lip-sync risk at all, and the physical sounds (pour, steam) do the atmosphere work a music bed usually fakes.

3. The reference-driven performance (three videos in, one song out)

The showcase's reference files

Camera-move reference (dolly zoom)
Subject footage
Song & performance reference

The manual showcase this prompt adapts

Video 1 provides the camera move: a Hitchcock dolly zoom. Video 2 is the subject: a
woman having coffee at a street cafe. Video 3 provides the song and the singing
performance. Make the woman in Video 2 sing — voice, phrasing and performance from
Video 3 — while the camera executes Video 1's dolly zoom on her.

Why it works: three references, three jobs — camera, subject, song — and the audio is generated as a performance, not laid over as a track. This is the manual's own demonstration that sound direction is just another reference role.

Try It

MiniMax H3 is live in Pixo's Playground: select Video → MiniMax H3, write a prompt with the line you want spoken, and generate a 4–15 second shot at 768p or 2K — sound included, export watermark-free. For the full model picture, start with what MiniMax H3 is; for how its audio stack compares to the other flagship released the same day, see MiniMax H3 vs Seedance 2.5.

FAQ

Can AI video generators create sound?

Some can, most can't. MiniMax H3, Seedance, Veo 3.1 and Wan 2.6 generate audio natively; Kling 3.0 outputs silent video. H3 goes furthest: every output ships with stereo audio — dialogue, ambience and effects — generated in the same pass as the picture, not layered on afterward.

Does MiniMax H3 lip-sync dialogue?

Yes. Write the exact line into your prompt and H3 generates the voice with the picture, with the character's mouth movements matching the words. Its text-to-speech spans multiple languages, so dialogue isn't limited to English.

How do I make an AI video with a voiceover?

Write the voiceover text verbatim in your prompt — for example: A calm male voice says: "Good tools disappear into the work." H3 generates the speech with the video. If the speaker is on screen, the lip-sync follows; if off-screen, it reads as narration.

Can I turn the background music off?

Yes — end your prompt with "non-narrative music: N/A" and H3 keeps the track clean: dialogue, ambience and effects only. That's the mode to use when you'll add licensed brand music in post.

Can I use my own audio as a reference?

Yes. Upload an audio track (up to 15MB) alongside an image or video reference — audio can't be the only input — and use it to set pacing or voice character. H3 also supports voice cloning from a reference track.

How do I try it?

MiniMax H3 is live in Pixo's Playground — generate 4–15 second shots at 768p or 2K. New users get 200 free credits on sign-up, and plans are currently up to 55% off.


MiniMax H3 is now live in Pixo's Playground — generate 4–15 second shots at 768p or 2K, with image, video and audio references. Sign up now — new users get 200 free credits on sign-up, and plans are currently up to 55% off. Putting products on screen? See e-commerce product videos with H3.

Ready to Revolutionize your workflow?

Join thousands of creators using Pixo to turn their stories into visual reality.

Sign Up Now

No credit card required • Free 200 credits