Character consistency is less about finding one magic prompt and more about controlling a chain of references, descriptions, and shot-level changes. This is a repeatable way to do that in MiniMax H3.
Disclosure: I work with PixMind.
Why Character Consistency Is the Hard Problem in AI Video
The "different person every shot" failure mode is structural, not a tuning issue. A video model that samples each frame from a text prompt has no memory of the face it generated two seconds ago, let alone the face it generated in the previous clip. Each frame is a fresh draw from a distribution, so features drift. Cheekbones sharpen between cuts. Eye color shifts half a shade. A mole on the left cheek in shot one migrates to the right cheek in shot two. By the time you cut three shots together, the audience reads three different people.
This matters more for story work than for any other format. A landscape shot of a city skyline does not need continuity. A product spin of a perfume bottle does not either. But the moment a character carries the narrative, the audience tracks that face frame by frame. Even small drift reads as a continuity error, and large drift breaks the fiction entirely. Social series, recurring spokespeople, branded characters, multi-shot ads, and short films all hit the same wall.
Re-describing the character in the prompt does not solve it. A paragraph that says "woman in her thirties, short black hair, green eyes, freckles, denim jacket" leaves every visual detail up to the model's interpretation on each render. Two generations from the same prompt produce two different women who both match the description. The prompt is a specification, not a lock. Without a visual anchor, the model will keep inventing.
The fix is to stop specifying the character in words and start supplying it as a reference. That is the entire premise of MiniMax H3's reference system, and the rest of this guide is about how to use it well.
MiniMax H3 complete model guide
How MiniMax H3 Solves Character Consistency
MiniMax H3 tackles the consistency problem with two complementary mechanisms: a multimodal reference system that lets you attach the character as an input, and native multi-shot consistency that keeps the subject coherent across shots of a scene. Together they shift character identity out of the prompt and into the reference layer, where the model can read it instead of imagining it.
Reference files are the identity layer
The core idea is simple. Instead of describing the character, you show it. MiniMax H3 accepts up to 12 multimodal reference files in a single request, drawn from images, videos, audio, and text. When you supply reference images of a character, the model holds the appearance consistent across angles and shots without re-describing it. The reference file is the lock. The prompt just directs the action.
This works because a reference image removes interpretation. There is exactly one face in the reference, not a distribution of faces that match a description. The model's job changes from "invent a person matching these words" to "use this exact person in this new shot". That is a far easier and more stable task.
The budget caps inside the 12-file limit are well documented: up to 9 images, 3 videos, and 3 audio files, combined with the text prompt. The exact split is yours to allocate, which is where most of the craft lives. We cover allocation in the next section.
Native multi-shot consistency
MiniMax H3 treats multi-shot consistency as a first-class capability, meaning the same subject stays coherent across multiple shots of a scene rather than only within a single clip. In practical terms, a character who walks into a room in shot one and sits down in shot three still looks like the same person, because the reference file travels with every shot.
This is what separates a true multi-shot model from a single-shot model that happens to render multiple clips. A single-shot model can hold a face together for five seconds but cannot guarantee the face in clip two matches clip one. A multi-shot model can, because the identity is pinned by the reference, not by the previous frame's residual statistics.
How to allocate the 12-file reference budget
The most common mistake with multimodal references is feeding conflicting inputs. Two character images with different identities, or a motion reference video whose wardrobe contradicts the character reference image, will produce flicker and drift. Each of the 12 slots should have exactly one job.
A reliable allocation pattern for a character-driven sequence:
- Character identity: Two to three images of the same character from different angles (front, three-quarter, profile) if you have them, or one clean front view if you do not.
- Outfit and props: One to two images locking the wardrobe, accessories, or product the character carries, kept separate from the identity reference so changes to costume do not leak into the face.
- Setting and environment: One to two images of the location, lighting, or style, so the character is rendered into a consistent world across shots.
- Motion or performance: One short reference video (optional) that demonstrates the pacing, camera move, or body language you want.
- Audio (optional): One reference audio clip for voice or scoring, kept separate from picture references.
This leaves room in the budget for iteration. You do not need to fill all 12 slots on every request. A single subject-reference image is enough for many shots, and adding references only helps when each one adds a clear, non-conflicting constraint.
MiniMax H3 multimodal input deep dive
MiniMax H3 in Action: Real Character Consistency Examples
Before the method, watch what subject reference actually produces. The video below walks through MiniMax H3's single-image setup, where one clean photo locks a character across new angles, lighting, and shots.
MiniMax H3 shipped subject reference as a first-class input for exactly this workflow, which is the capability the tutorial above relies on.
MiniMax H3 Subject Reference announcement
The Master-Character Method, Step by Step
The most reliable workflow for a recurring character is what community guides call the master-character method: design a master character, generate a clean front-view reference image, then feed that reference on every shot. It works because it gives MiniMax H3 one canonical source of truth for the face, and reuses it as a fixed input rather than a fresh prompt.
Step 1: Design the master character
Start by defining the character in writing before you render anything. Note the fixed attributes that must never change across the project: age range, ethnicity, hair style and color, eye color, build, distinguishing marks, default wardrobe. This document is the character bible. It exists so that when you revise shots weeks apart, you still agree with your earlier self on what the character looks like.
This is also where you decide what is invariant and what is variable. The face is invariant. The wardrobe may be variable across scenes. The haircut may be variable across a time jump. Mark each attribute as locked or flexible, and keep the locked attributes out of the prompt text on subsequent shots. Locked attributes belong in the reference, not in the prose.
Step 2: Generate a clean front-view reference image
Use MiniMax H3 (or any image tool you prefer) to generate a single, clean, front-view portrait of the master character. The goal is a well-lit, head-and-shoulders or head-to-waist image where the face is clearly visible, the expression is neutral, and there are no occlusions (no sunglasses, no hand in front of the face, no harsh shadows across the features).
Three rules make a strong reference image. Light the face evenly so the model can read the geometry. Keep the camera at eye level and front-facing so there is no perspective distortion to interpret. Use a plain or simple background so nothing competes with the character. A reference image is a measurement tool first and an aesthetic object second. You can always render more artful compositions later, but the reference itself should be the clearest possible statement of the face.
Generate three to five variations of this reference before you commit. Faces drift across variations even at the same settings, so pick the one that best matches your character bible. This chosen image becomes the canonical reference for every shot in the project.
Step 3: Feed the reference on every subsequent shot
Once you have the canonical reference, every video shot in the sequence starts the same way: attach the reference image as a subject-reference input, then write a prompt that describes only what is new to that shot (the action, the camera, the setting, the lighting). Do not re-describe the character in the prompt. Re-describing invites the model to reinterpret, which is the exact failure mode the reference is supposed to prevent.
Across a multi-shot sequence, the reference image is the constant and the prompt is the variable. That asymmetry is what holds the character together. The face comes from the file, the staging comes from the words, and the two layers do not compete.
Subject Reference: When One Photo Is All You Have
Not every project starts with a designed master character. Sometimes the input is a single photo: a real person for a spokesperson spot, a product shot for an ad, an existing illustration for a brand mascot. MiniMax H3's subject-reference style input handles this case directly. You supply one image of the subject, and the model renders that subject across new angles, lighting, and contexts.
Creators working with single-image subject references report character consistency in the range of 95 percent and above when the source image is clean. That figure is a community observation rather than a benchmark, and it assumes you follow a few rules. The source image should be high resolution, evenly lit, and unambiguous about the subject. The subject should fill a meaningful portion of the frame. The face should not be occluded, blurred, or shot at an extreme angle.
Test three to five variations before committing
A single reference image produces a distribution of outputs, not a single deterministic face. Generate three to five test clips from the same reference and review them side by side. If the subject holds across all five, the reference is strong enough to build on. If it drifts, either the source image is weak or the prompt is asking for something that conflicts with the reference (heavy stylization, extreme age change, conflicting wardrobe).
This test-before-committing step is the single biggest lever for one-photo workflows. It catches weak references early, when reworking is cheap, instead of late, when you have already built a sequence around a reference that does not lock.
Use one photo, or design a master character?
The choice between a one-photo workflow and the master-character method comes down to source material. If you already have a photo of a real person or an existing character design, subject reference is the right starting point. If you are inventing a character from scratch, the master-character method gives you more control because you design the reference deliberately rather than inheriting the constraints of an existing image.
Both workflows end in the same place: a canonical reference image attached to every shot, and a prompt that describes only the staging.
Tuning Reference Influence: The 65 to 75 Percent Sweet Spot
Most subject-reference implementations expose a control that sets how strongly the reference should drive the output. The naming differs by surface ("influence", "strength", "adherence"), but the mechanics are the same: a low value lets the model improvise around the reference, and a high value forces the output to match the reference closely. Community group tips for MiniMax H3-style subject reference consistently land this setting in the 65 to 75 percent range, and that band is the right starting point for character consistency work.
Under-adherence: the reference is ignored
When influence is too low, the model treats the reference as a suggestion. The face in the output resembles the reference but does not match it, and the resemblance weakens as the clip runs. The failure mode reads as "same kind of person" rather than "the same person". This is the wrong failure for a recurring character, where the whole point is exact identity.
Over-adherence: the output looks frozen
When influence is too high, the model rigidly copies the reference instead of reposing it for the new shot. The face locks into the exact expression and head angle of the source image, motion becomes stiff, and the character looks pasted onto the scene rather than inhabiting it. Consistency is high but the performance is dead.
Finding the sweet spot
The 65 to 75 percent band is where the reference holds the identity while the prompt controls the performance. Start at 70 percent for a clean front-view reference. Move up if the face drifts during a clip. Move down if the motion looks stiff or the character cannot turn their head. Treat the setting as a per-shot dial, not a global constant, because the right value depends on how much the shot asks the character to move and turn.
Two cases warrant a move outside the band. Fast-moving action shots where the character turns away from camera may need a slightly higher value to preserve identity through the motion. Stylized shots where you want the character rendered in a different visual style may need a slightly lower value to let the style through. In both cases, change one variable at a time so you can attribute the result.
Building a Multi-Shot Sequence with a Recurring Character
A multi-shot sequence is where character consistency earns its keep. Each shot is a single MiniMax H3 generation with the canonical reference attached, and the sequence is held together by the reference, not by luck. Planning the sequence as a shot list before you render is what separates a coherent piece from a pile of clips.
Build a shot list
Before any generation, write the sequence as a table with one row per shot. For each shot, capture the shot number, the shot size (wide, medium, close), the camera move, the action, the setting, and the reference or references attached. The reference column is the one that enforces consistency: the same character reference appears on every row, and shot-specific references (a product, a location, a wardrobe change) appear only where relevant.
A minimal shot list for a three-shot character sequence might look like this:
- Shot 1 (wide): Character enters the kitchen, crosses to the counter, morning light. Reference: character front view, kitchen setting.
- Shot 2 (medium): Character at the counter, opens a box, reacts. Reference: character front view, product shot.
- Shot 3 (close): Character's face, subtle smile. Reference: character front view, three-quarter angle reference if available.
Every row carries the character reference. Only the supporting references change. That structure is what makes the cut hold.
Storyboard before you render
Academic work on training-free "video storyboarding" for multi-shot consistent characters underpins this direction, and the practical version is straightforward. Sketch or describe each shot as a single image first, generate those stills, and approve them as a sequence before spending render budget on video. Stills are cheaper than video, they let you check continuity at a glance, and they become first-frame references when you animate the shots later.
The discipline here is to review the sequence as a sequence, not as independent images. Lay the approved stills out side by side and ask one question: does the character read as the same person across all of them? If yes, move to video. If no, fix the reference before you spend video budget on a sequence that will not cut together.
Render shots in order, in the same session
Render the shots in story order within a single session if you can. Models and settings drift over time, and rendering shots days apart introduces variance that breaks continuity even when the reference is constant. A same-session render with the same reference and the same settings is the most controlled path to a consistent sequence.
Worked Examples: Character Consistency in Practice
Three worked examples show how the method adapts to different formats. Each one starts from the same foundation: a canonical character reference attached to every shot, with prompts that describe only the staging.
Example 1: A three-shot product ad with a recurring spokesperson
A skincare brand needs a three-shot social ad featuring one spokesperson holding and reacting to a jar of cream. The character reference is a clean front-view portrait of the spokesperson. The product reference is a separate still of the jar, included only on the shots where the product is on screen.
- Shot 1 (medium): Spokesperson walks into frame, product in hand. References: character front view, product still.
- Shot 2 (close): Spokesperson unscrews the lid. References: character front view, product still.
- Shot 3 (medium-close): Spokesperson speaks to camera. Reference: character front view only.
Because the character reference is identical across all three shots and the product reference only appears where the product is visible, the cut holds the spokesperson's identity while letting the product appear and disappear cleanly. At the verified MiniMax H3 rates ($0.13 per second at 2K, $0.09 per second at 768P), three six-second 2K shots cost roughly $2.34, which makes this format cheap to iterate on.
Example 2: A recurring character across a Reels series
A creator wants the same animated host across a weekly series of short Reels. The master-character method applies directly. The first session is spent designing and locking the master character into a canonical front-view reference. Each weekly episode then starts from that reference, with a per-episode prompt describing the topic, setting, and action.
The reference image never changes week to week, which is the entire point. The audience recognizes the host across episodes because the host is literally the same face, supplied as the same file. Wardrobe and setting can vary by episode through supporting references, but identity stays locked. For a series running dozens of episodes, this is the difference between a recognizable character and a parade of lookalikes.
Example 3: A multi-shot story scene
A short narrative scene requires one character across five shots: entering a room, sitting at a desk, reacting to a phone call, standing, and leaving. The character reference is attached to every shot. A location reference (the same room from a consistent angle) is attached to the wide shots. A shot list is built before any video is rendered, and stills are approved as a sequence before animation.
The tricky shot here is the reaction shot, where the character's expression has to change without the identity changing. The fix is to keep the character reference at the standard influence, and describe the new expression in the prompt while keeping every other invariant in the reference. The reference holds the face, the prompt supplies the emotion.
Common Failures and How to Fix Them
Character consistency work fails in predictable ways. Most failures trace back to the reference, the prompt, or the interaction between them.
Low-quality reference image
A blurry, dark, or extreme-angle reference cannot lock identity because the model cannot read the face clearly. The output drifts because the model has to invent the details the reference hides. The fix is to regenerate the reference with even lighting, a front-facing camera at eye level, and a neutral expression. The reference is a measurement tool. Treat it like one.
Conflicting references
Two character images with different faces, or a character reference whose wardrobe conflicts with an outfit reference, force the model to arbitrate. Arbitration shows up as flicker and drift. The fix is to audit the reference set before rendering and make sure each attribute is defined by exactly one reference. If you need a wardrobe change, change the outfit reference, not the identity reference.
Re-describing the character in the prompt
Writing "woman with short black hair and green eyes" in the prompt when a reference image already defines her invites reinterpretation. The model reads both inputs and tries to satisfy both, which can pull the face away from the reference. The fix is to remove identity description from the prompt entirely once a reference is attached, and let the reference do its job.
Wrong aspect ratio or framing
A reference shot in 9:16 used on a 16:9 output can distort the face or crop out distinguishing features. The fix is to generate references at the aspect ratio you intend to deliver, or to use head-and-shoulders framing that survives cropping across ratios.
Motion that breaks identity
Fast spins, occlusion (a hand passing in front of the face), and extreme head turns can cause the model to lose the face mid-clip and reconstruct a slightly different one when the face reappears. The fix is to plan shots so the face stays at least partially visible through the motion, and to keep extreme turns for moments where the face is not the focus. If a shot must break identity through motion, cut around the break in the edit rather than trying to hold the face through it.
MiniMax H3 prompt templates for character work
Cost of a Multi-Shot Sequence
MiniMax H3's pricing makes multi-shot character work practical to iterate on. The verified rates are $0.13 per second at native 2K and $0.09 per second at 768P. A six-shot sequence of six-second 2K clips costs roughly $4.68; the same sequence at 768P costs roughly $3.24. That is cheap enough to render variations and pick the best take, which is the right way to approach character consistency work.
A useful budget pattern is to iterate at 768P and finalize at 2K. Use 768P for the test renders that check whether a reference locks and whether a sequence cuts together. Once the references and the shot list are stable, render the final sequence at native 2K. This keeps the cost of iteration low and reserves the higher fidelity, higher cost renders for output you will actually ship.
The per-second math scales linearly with clip length, so keep each shot as short as the story allows. A character consistency test does not need a fifteen-second clip. Five or six seconds is usually enough to judge whether the face holds, and at 2K that costs less than a dollar per shot.
Free credits guide for MiniMax H3
MiniMax H3 Character Consistency FAQ
How many references do I need for a consistent character?
One clean front-view image is enough to lock identity for many shots. Three to five images from different angles (front, three-quarter, profile) give the model more to work with and improve consistency across head turns. Allocate the rest of the 12-file budget to outfit, setting, motion, and audio only when each adds a non-conflicting constraint.
Can I use a real photo as the character reference?
Yes. A real photo works as a subject-reference style input. The photo should be high resolution, evenly lit, and unobstructed, with the subject filling a meaningful portion of the frame. Generate three to five test clips from the same photo and check that the subject holds across all of them before building a sequence. Ensure you hold the rights to any real person's likeness before publishing.
Does subject reference work for non-human subjects like products or mascots?
Yes. The mechanism is the same. Supply a clean reference image of the product, mascot, or object and the model carries its appearance across shots. This is especially useful for branded characters and product ads where the same object has to look identical across a sequence. The same rules apply: one clear reference, no conflicting inputs, and an influence setting in the 65 to 75 percent range to start.
How much does a multi-shot character sequence cost?
At the verified rates of $0.13 per second at 2K and $0.09 per second at 768P, a six-shot sequence of six-second clips costs roughly $4.68 at 2K or $3.24 at 768P. Iterating at 768P and finalizing at 2K keeps total cost down while reserving high-fidelity renders for the cuts you plan to ship.
Are there free options for MiniMax H3?
The fastest no-code path is the hosted MiniMax H3 tool, which runs the same reference system through a UI without setup. Free credits and starter offers rotate, so check the current credits guide for what is available when you produce. Free credits guide for MiniMax H3
Does this work for a recurring character across separate videos, not just one sequence?
Yes. The master-character method is designed for exactly that case. Lock the canonical reference once, then attach it to the first shot of every new video in the series. The character will read as the same person across separate videos released weeks or months apart because the face is supplied as the same file each time.
What influence setting should I start with?
Start at 70 percent with a clean front-view reference. Move up if the face drifts during a clip, move down if the motion looks stiff. Treat the setting as a per-shot dial rather than a global constant, because the right value depends on how much the shot asks the character to move and turn.
Conclusion: Make the Reference Do the Work
Character consistency stops being a gamble the moment you stop describing the character in prose and start supplying it as a reference file. MiniMax H3's 12-file multimodal budget and native multi-shot consistency are built for exactly this workflow, and the master-character method gives you a repeatable path: design the master character, generate a clean front-view reference, and attach that reference to every shot. Test three to five variations before committing, set influence in the 65 to 75 percent band, and let the prompt carry only the staging. The face comes from the file, every time.
The next step is to put the method on a real project. Pick one character, build the canonical reference, run a three-shot sequence at 768P to test the cut, and finalize at native 2K once the references lock.
PixMind MiniMax H3 video tool Viral video guide for MiniMax H3 Free credits guide for MiniMax H3
Specs, pricing, and capabilities in this guide were verified on 2026-08-01 against the MiniMax official blog, the MiniMax platform documentation, OpenRouter, Vercel AI Gateway, and EvoLink. Community workflow details (the master-character method, the 65 to 75 percent influence band, and the roughly 95 percent single-image consistency observation) reflect creator and industry practice and are not formal benchmarks. Model cards and rate cards change quickly; always confirm the live values in your route before producing.
If you apply this workflow, start with one representative asset, record the settings that matter, and only then scale it across a larger batch. That makes the result easier to compare, debug, and reuse.
Originally published by the PixMind Editorial Team
https://www.pixmind.io/posts/minimax-h3-character-consistency


Top comments (0)