Most MiniMax H3 tutorials stop the moment text-to-video or image-to-video produces a usable clip.
That is enough to prove the model runs, but it misses the part that makes H3 interesting: reference-to-video (R2V). With R2V, images can define identity and style, video can define motion or camera work, and audio can define a voice—all within one generation context.
In practice, adding more files isn't the hard part. The real challenge is telling H3 exactly what each reference controls, what must remain unchanged, and what may be transformed.
This guide focuses on four practical skills:
| Skill | What it gives you |
|---|---|
| Mixed-reference R2V | Reuse identity, style, motion, camera work, and audio |
| Instruction-based video editing | Change one part of a clip while preserving the rest |
| Voice-timbre transfer | Generate new dialogue using a consented voice reference |
| Agent-written prompts | Convert a rough brief into H3's structured prompt format |
Before you start
This article assumes that MiniMax H3 already runs in ComfyUI. The official ComfyUI integration requires version 0.30.0 or later and provides T2V, I2V, and R2V templates under Template Library → Video.
Two details are easy to overlook:
- H3-Base generates at 768p. MiniMax's separate H3-Regenerate-2K stage produces 2K output, but that module is not currently part of the open-weight release.
- R2V uses a different checkpoint from T2V and I2V. In ComfyUI, use the
ref2vaweights for reference-driven generation, not thefl2vaweights.
1. Think in reference roles, not reference files
H3-Base-Ref2VA accepts a mixed context of:
- Up to 9 images
- Up to 3 video clips
- Up to 3 audio clips
- Up to 12 files in total
Each video or audio clip must be between 2 and 15 seconds, and the total duration for each media type cannot exceed 15 seconds.
Those limits are generous, but filling every slot rarely improves the result: a short target clip can't express twelve competing ideas clearly, so start with the smallest reference set that describes the shot.
For example:
-
<Picture 1>: character identity -
<Picture 2>: costume and color palette -
<Video 1>: body movement and camera motion -
<Audio 1>: voice timbre
Then state those jobs directly in the prompt:
<Picture 1> defines the character's face and hairstyle.
<Picture 2> defines the wardrobe and blue-silver color palette.
<Video 1> defines the walking motion and slow left-to-right tracking shot.
<Audio 1> is the voice-timbre reference for the character.
The order matters. ComfyUI identifies references by the order in which they are connected, so <Picture 1> in the prompt must really be the first connected image.
A reliable R2V workflow
- Decide what the final 4–15 second clip should accomplish.
- Assign one explicit role to each reference.
- Remove references that do not contribute to that result.
- Describe the action chronologically.
- Specify camera movement, dialogue, ambience, effects, and music.
- Check every label against the order of the connected inputs.
References carry reusable traits. The prompt decides what those traits do over time.
2. Edit video by separating preservation from change
R2V can also behave like an instruction-based video editor. Give it a source video, identify what must remain stable, and describe the local change.
Typical edits include:
- Replacing a character or object
- Changing the background
- Relighting a scene
- Transferring a visual style
- Adding a localized effect
- Preserving or replacing the original audio
A prompt like "make this warmer" rarely gets useful results. What works better is a small edit specification with an explicit preservation contract.
Here is a relighting example:
subject_definitions:
<Video 1> is the source video. It shows a product on a table in daylight.
summary:
[video editing] Relight <Video 1> as a warm evening interior.
retention_analysis:
<Video 1> (composition, camera movement, product position): fully_preserved.
Only the lighting condition changes.
detailed_description:
The target video keeps the composition and timing of <Video 1>.
[Shot 1] Replace the daylight with warm evening light entering from the left.
Soft amber highlights move across the product. The camera, product, table,
background geometry, and duration remain unchanged.
overall_soundscape:
Reuse the original quiet room tone without modification.
non_diegetic_music:
N/A
This structure reduces a common failure mode: the model "helpfully" redesigns the entire shot when only one attribute should change.
Use fully_preserved for composition, motion, timing, or audio that must remain intact. Use an explicit transfer or replacement instruction only for the attributes that should change. If a section is irrelevant, write N/A instead of leaving the intent ambiguous.
3. Treat voice cloning as timbre transfer—not signal copying
With an audio reference, H3 can generate new speech that follows a speaker's vocal characteristics while producing new audio for the target scene.
A minimal prompt might look like this:
<Subject 1> is the character shown in <Picture 1>.
<Audio 1> is the consented voice-timbre reference for <Subject 1>.
Use <Audio 1> only as a timbre reference; do not reproduce its original words.
<Subject 1> turns toward the camera and says:
<d>[Japanese] こんにちは、今日は新商品をご紹介します。</d>
For cleaner transfer, use a short recording with:
- One speaker
- Little or no music
- Minimal room echo
- No overlapping dialogue
- A speaking style close to the desired output
Consent is part of the workflow
Only use a voice when you have the speaker's informed consent or another clear legal right to use it. Do not imitate a real person deceptively, and disclose synthetic media where appropriate. A technically successful clone can still violate privacy, publicity, copyright, platform, or employment rules.
Check this before the audio ever enters the workflow, not as a final QA step tacked on afterward.
4. Let an agent write the structured prompt
MiniMax includes an official h3-prompt-writing skill in the H3 repository. It's a Markdown-based instruction package, portable to any coding agent or harness that can read a SKILL.md file and its local references.
Install it with:
npx skills add https://github.com/MiniMax-AI/MiniMax-H3 --skill h3-prompt-writing
The skill first selects the input mode:
| Mode | Purpose |
|---|---|
| T2VA | Generate an audiovisual timeline from text |
| I2VA | Continue forward from a first frame |
| FL2VA | Connect supplied first and last frames |
| L2VA | Build toward a supplied last frame |
| Ref2VA | Generate from mixed image, video, and audio references |
For the base modes, it produces three sections in this order:
integrated_multimodal_description
overall_soundscape
non_diegetic_music
For full-reference Ref2VA, it uses six:
subject_definitions
summary
retention_analysis
detailed_description
overall_soundscape
non_diegetic_music
Handing this to an agent works well because the task is mechanical: preserve exact field names, resolve labels, arrange events chronologically, keep the sound synced to the visual timeline.
A practical request to the agent can be short:
Use the MiniMax H3 prompt-writing skill.
Mode: Ref2VA
Duration: 8 seconds
Picture 1: character identity
Video 1: camera motion only
Audio 1: consented voice-timbre reference
Goal: The character walks into a small studio, stops at the desk,
looks at the camera, and says “Build the first version today.”
Preserve the character identity and the reference camera motion.
Generate quiet room ambience. No music.
The resulting prompt should still be reviewed. Check that every reference exists, every preservation rule is intentional, the dialogue fits the duration, and the generated structure has not introduced details you did not ask for.
Common failure modes
1. A reference has no declared job
If the prompt does not say whether a video controls motion, style, composition, or all three, the model has to guess.
Fix: assign a role to every input.
2. The edit and preservation rules conflict
“Keep the original lighting” and “change the scene to sunset” cannot both be true.
Fix: separate preserved attributes from transferred or replaced attributes.
3. Too much story for the duration
A 6-second clip cannot reliably contain an establishing shot, three actions, a costume change, a camera orbit, and two lines of dialogue.
Fix: choose one visual beat, or increase the duration within H3's 15-second limit.
4. Voice audio contains music or multiple speakers
The reference no longer represents one clean voice identity.
Fix: prepare a rights-cleared, single-speaker sample with minimal background sound.
5. The open-weight license is treated like a permissive OSS license
It is not. The MiniMax H3 Community License currently excludes the European Union, the United Kingdom, the Republic of Korea, and the United States from its applicable territory. It also restricts use and display of the model's outputs outside that territory, requires separate authorization for certain commercial products above the stated revenue threshold, and includes labeling obligations.
Fix: read the current license before downloading the weights, deploying a service, or publishing an output. If your use crosses jurisdictions, get qualified legal advice. Because the restriction covers display of outputs, I have intentionally not embedded an H3-generated video in this cross-post.
Pre-generation checklist
- [ ] The target duration and aspect ratio are defined.
- [ ] Every image, video, and audio input has one explicit role.
- [ ] Reference labels match ComfyUI's connection order.
- [ ] Preserved and changed attributes do not conflict.
- [ ] The action fits within 4–15 seconds.
- [ ] Camera, dialogue, ambience, effects, and music are specified.
- [ ] Voice and visual references are consented or rights-cleared.
- [ ] The deployment and publication plan complies with the current license.
Final takeaway
The skill in advanced H3 prompting has little to do with cinematic paragraph length. It comes down to a clear contract between the references and the target clip:
- What does each reference control?
- What must remain unchanged?
- What should change, and when?
- What should the audience hear at each moment?
Once those decisions are explicit, R2V becomes much more predictable. And an agent can handle much of the repetitive prompt structuring for you.
Official references
- MiniMax H3 repository and model documentation
- MiniMax H3 workflows in ComfyUI
- Official H3 prompt-writing skill
- MiniMax H3 Community License
This article was edited with AI assistance.
*Originally published in Japanese on EdgeHUB.
Top comments (0)