MiniMax H3 reads text, images, video, and audio as one context. These MiniMax prompts show how to direct all of it: timed beats, sound, camera moves, and the details you want held in place.
The MiniMax prompts here read more like production paperwork than a sentence. Six things do the heavy lifting.
Don't just attach files, assign jobs. "Image 1 sets the location, Image 2 is the lead, Audio 1 is her voice." MiniMax H3 treats mixed inputs as one context, so naming each reference's duty is what turns a pile of assets into direction.
Write the clip as a timeline: "0 to 5s she walks the aisle, 5 to 10s she lifts the bottle into the light, 10 to 15s she sets it down." Timed beats are how a MiniMax prompt gets shot order instead of one averaged motion.
Name the capture, not just the scene. 16mm grain, restrained palette, hard overhead light, handheld tremor, delayed autofocus. These are the details that separate footage that feels filmed from output that feels rendered.
Because MiniMax H3 generates stereo audio in the same pass as picture, sound gets directed rather than accepted. Name instruments, name specific effects, say where a cue lands, and say what stays silent. "Add some music" wastes the capability.
State what holds and what's banned: hold the wardrobe, no subtitles, no watermark, no soft dissolves. Naming locked elements is what stops a character or product detail from drifting halfway through the clip.
Push in, pull out, truck left, tilt up, tracking shot. One clear move per beat reads better than three stacked into a single instruction, and it keeps the subject where you put it.
A prompt written for a single-input model leaves most of MiniMax H3 unused. Four habits change the output the most.
The whole point of an omni-modal model is that you can say "borrow the camera movement from Video 1, put the character from Image 2 in it, match the delivery in Audio 1." That sentence is the prompt. Other models need you to pick one input type and accept the rest.
No separate audio stage means the line a character speaks, the room tone under it, and the one music cue that lands on the title are all prompt text. MiniMax prompts that skip this get generic ambience by default.
Instruction-based editing changes the one element you name and leaves approved lighting, framing, and performance intact. "Replace the background with a night street, hold everything else" beats regenerating a shot you already liked.
MiniMax H3 accepts detailed briefs and keeps the nuance. Where a terse prompt is a virtue on some models, here the specific, structured version is usually the one that survives to the render.
Six kinds of MiniMax prompt, each weighted differently. Start from the one closest to the shot you need.
No source frame, so the prompt carries the whole composition: subject, action, setting, light, and camera. One scene with one primary motion outperforms a prompt that tries to be a screenplay.
With a first or last frame supplied, appearance and framing are already fixed. These image to video prompts spend their words on motion, the transition between frames, and what should stay exactly where it is.
Stills, clips, and audio mixed together, each with a stated role: identity from one, motion rhythm from another, voice from a third. The prompts that work restate traits in words instead of trusting the file to carry them.
Several shots inside one generation, blocked into timed segments with named transitions such as a hard cut, whip pan, or match cut on a shape. Left unnamed, cutting drifts toward generic smoothness.
Character-driven clips where the line, the delivery, and the timing all live in the prompt. Voice references plus written traits keep the same character sounding and looking consistent across shots.
Swap a subject, replace a background, relight day to night, change a spoken line. These prompts show how to name the target and the locks so one change doesn't cost you the rest of the take.
MiniMax H3 is MiniMax's latest video model. What makes MiniMax prompts different is that the model reads text, images, video, and audio as a single context, and generates synchronized stereo audio in the same pass as the picture. One prompt can direct camera, performance, dialogue, and sound at once, then change a single element without disturbing the rest. Every prompt in this collection is shown in full, so you can see which choices did the work instead of guessing at them. Updated as new MiniMax H3 examples surface from the community.
Copy any MiniMax prompt and use it on Hailuo AI, or through API providers that host MiniMax H3. The prompt text itself is portable. It's plain structured description rather than a proprietary parameter syntax, so it carries over wherever the model is available.
Longer than you would write for most image models. MiniMax H3 holds onto detail, so a brief that assigns reference roles, blocks out timing, and specifies sound usually beats a single descriptive sentence. Length by itself isn't the goal though. Every line should do a job: padding with adjectives adds nothing, while an unwritten audio track or an unstated lock leaves the model guessing.
Yes, and it's a reasonable starting point, since none of these models need a special parameter syntax. What a ported prompt usually lacks is sound direction and reference roles, because the model it was written for couldn't act on either. Add the audio you want, give each attached file a job, and state your locks.
Yes, and that's one of the model's stronger uses. Instruction-based editing changes only what you name: swap an object, replace a background, relight the scene, or change a line of dialogue, while approved framing, lighting, and performance stay put. The editing prompts in this collection show how to phrase the target and the locks together.
Completely free to browse and copy, no signup required. Generating with MiniMax H3 itself happens on Hailuo AI or through a provider that hosts the model.

































