Skip to content
MiniMax H3·Prompt Guide·AI Video Prompts·Hailuo 3.0·

MiniMax H3 Prompt Guide: Official Formula & Modes

How to write MiniMax H3 prompts that work: the official three-part formula, @file references, all three generation modes, and the six mistakes to avoid.

Pixo Team·9 min read
MiniMax H3 Prompt Guide: Official Formula & Modes

Most of what's been written about prompting MiniMax H3 is guesswork — the model is barely a week old. This guide isn't guesswork. MiniMax published a detailed usage manual alongside H3's July 31 launch, and buried in it is something rare: an explicit, engineer-written prompt formula, a taxonomy of file-reference roles, and a frank list of the mistakes that most often ruin generations. Very little of it has surfaced in English until now.

We build Pixo, a multi-model AI video platform, and reading model manuals the day they drop is part of the job — our agent has to write production-grade prompts for every model we run. Here's the H3 prompting system, translated and organized for working creators, cross-checked against MiniMax's official video-generation docs. (New to the model itself? Start with our H3 first look.)

The Official Formula

Straight from the manual:

Full prompt = reference material notes + core idea + scene-by-scene description

Three blocks, in that order. Skip the first if you uploaded nothing. Most failed H3 prompts break exactly one of these blocks — usually by mashing everything into a single undifferentiated paragraph, which is the manual's #1 listed mistake.

Block 1: Reference material notes

Every file you upload gets a number in upload order, cited inline as @image1, @video2, @audio3. And every file needs a declared role. The role taxonomy is worth memorizing, because it's effectively H3's API surface in natural language:

  • Character reference — lock a face or figure
  • Object reference — lock a product or prop
  • Scene reference — lock a location
  • Keyframe — lock a frame (say explicitly whether it's the first or last)
  • Voice reference — lock a voice timbre
  • Storyboard — generate shots following a story panel
  • Style reference — match an image's look
  • Composition reference — match framing and layout
  • Audio reuse — use the track directly as the video's audio (or reuse part of it)
  • Motion reference — lock an action from a video
  • Camera reference — lock a camera move
  • Video edit — the video to be modified

The manual's example, translated: "@image1 is the character reference (lock this woman's face), @video1 is the motion reference (use the sword-dance moves in it), @audio1 is the mood reference (classical guqin score). Have this woman perform the video's sword dance in a cherry-blossom courtyard."

Remember the hard ceiling from the model card: 12 files total — at most 9 images, 3 video clips and 3 audio tracks, with video and audio each capped at 15 seconds in total.

Block 2: Core idea

Four elements, explicitly named in the manual: subject (who or what), place (where), event (doing what), and genre/style (live-action, animation, cinematic, commercial, documentary — or a named aesthetic like cyberpunk or neon). Plus, optionally, camera behavior: H3 cuts between shots by default, so if you want one continuous take, say so. For camera moves, the manual is unusually specific — write concrete moves like "truck left + pan right" rather than vague requests like "orbit the scene." For cuts, you can specify the style: hard cut, fade, beat-synced, rapid montage.

Block 3: Scene-by-scene description

Describe what happens over time, segment by segment — "0–3s: …, 3–7s: …" — connecting elements to files with @-references wherever they're linked. This is also where dialogue lives, and the rule is absolute: write the exact line. "She says something moving" produces mush; "She says: 'You came. The blade has waited long enough.'" produces the line, lip-synced.

Rules That Save Generations

Four smaller rules from the manual punch far above their weight:

Write what the camera sees, not what it means. H3 rewards concrete, visual, direct description and stumbles on metaphor. "Rain-slick neon reflecting off her leather jacket" beats "an atmosphere of urban melancholy" every time.

On-screen text must be quoted exactly. If you want text in the frame, write it: "The phone screen shows the title 'AI Video Creation' with a button reading 'Start Now'."

Kill unwanted music explicitly. No background music? End the prompt with non_diegetic_music: N/A. And never contradict yourself — requesting a soundtrack in one line and banning background music in another is a listed failure mode.

Name the shot when you cut. When specifying a cut, state the new shot size and which established subject it holds — that's what keeps faces consistent across cuts.

Prompting Each of the Three Modes

Reference mode (up to 9 images + 3 videos + 3 audio): the full formula applies — the reference notes block is mandatory, and unlabeled files are wasted files.

Image-to-video / first-last-frame: upload one image and say whether it's the opening or closing frame; upload two and H3 fills in the motion, lighting and sound between them. Critically, it will not invent camera cuts in this mode — the manual's example: "@image1 is the first frame: a woman holding a sword under a cherry tree. Take her from ready stance through the full sword dance, flowing naturally, no cuts."

Pure text-to-video: no files, and a generous ceiling — the manual puts the prompt limit at 7,000 characters. The floor for a usable result: subject appearance + scene details + action + style. Prompts shorter than that are a documented failure mode, not a style choice.

Copy-Paste Starting Points

Three prompts built on the official formula (adapt freely):

[References] @image1 is the character reference (lock her face). @image2 is the scene reference.
[Core idea] A woman in a rain-soaked neon alley in Tokyo at night; cinematic live-action; one continuous take, slow push-in.
[Process] 0–4s: she walks toward camera, rain on her jacket, distant traffic hum. 4–8s: she stops, looks up; she says: "It was never about the money." Ambient rain continues. non_diegetic_music: N/A
[References] @image1 is the object reference (lock the perfume bottle, label included). @audio1 is the audio reuse track.
[Core idea] Luxury product film, 21:9, studio-to-sunset lighting shift; beat-synced cuts.
[Process] 0–3s: bottle rises from black marble, side light. 3–6s: cut on beat to golden-hour terrace, bottle identical. 6–8s: macro on the label, on-screen text reads "NOCTURNE".
[Edit instruction, against an existing video] Replace the cat in the video with a golden retriever. Keep the camera move, lighting and background unchanged.

The Six Mistakes the Manual Warns About

MistakeFix
One undifferentiated paragraphSplit into the three formula blocks
Files uploaded, roles unstatedAdd "@image1 is the X reference" for every file
Requesting music and banning background musicDelete one — or scope them to different scenes
Wanting one take but writing "Shot 1 / Shot 2"Keep one continuous narrative paragraph, no shot structure
Wanting facial consistency without a reference imageUpload one, labeled "character reference (lock face)"
Prompt too short with no filesCover subject appearance + scene + action + style, minimum

The Honest Shortcut

The manual's closing advice — hand prompt-writing to a professional partner — is MiniMax's own admission of the pattern we see across every frontier model: prompts have quietly become specifications, with reference manifests, timing blocks and formatting rules. Knowing the formula is what separates usable output from mush, but writing spec-grade prompts for every shot of a multi-scene video is real work.

That's precisely the job we built Pixo's agent for: describe the video you want, and the agent writes the script, the storyboard, and the per-shot prompts in each model's native dialect — the formula above included. Different models want different dialects, which is the whole point of our H3 vs Seedance 2.5 comparison. MiniMax H3 is coming to Pixo's storyboard alongside Seedance, Kling and Veo. Sign up now — new users get 200 free credits on sign-up — and you'll be ready to put H3 to work the moment it lands.

Frequently Asked Questions

What is the official prompt formula for MiniMax H3?

MiniMax's usage manual defines it as: full prompt = reference material notes + core idea + scene-by-scene description. First declare what each uploaded file is for, then state subject, place, event and style, then describe the action over time — with timestamps if you need them.

How do I reference uploaded files in a MiniMax H3 prompt?

Number them in upload order and cite them inline — @image1, @video2, @audio3. Every file needs a declared role: face lock, object lock, scene reference, motion reference, voice reference, pacing track and so on. Files without a stated purpose are the single most common cause of ignored references.

How do I stop MiniMax H3 from adding background music?

State it explicitly at the end of the prompt: "non_diegetic_music: N/A" (or the plain-language equivalent). Negative wishes that aren't written down tend to be ignored — and never request a soundtrack in one line while banning background music in another.

How do I keep a character's face consistent in MiniMax H3?

Upload a character reference image and label it explicitly as a face lock — for example "@image1 is the character reference (lock this woman's face)". Wanting facial consistency without uploading a labelled reference is one of the manual's listed common mistakes.

What are MiniMax H3's three generation modes?

Reference-to-video (up to 12 files — 9 images, 3 videos, 3 audio — each with a declared role), image-to-video using one or two keyframes (H3 fills the motion between them and won't invent cuts), and pure text-to-video.

This guide translates and reorganizes MiniMax's own H3 usage manual, cross-checked against the official video-generation documentation and the published model card on August 5, 2026. The prompt patterns are the ones our agent uses in production; model behaviour changes with updates, so treat any single example as a starting point rather than a guarantee.

Ready to Revolutionize your workflow?

Join thousands of creators using Pixo to turn their stories into visual reality.

Sign Up Now

No credit card required • Free 200 credits