MiniMax shipped H3 on July 31, 2026, and it belongs in the "rethink the pipeline" bucket rather than the "another text-to-video model" bucket. The interesting part is not the demo reel. It is that H3 treats text-to-video, image-to-video, first/last-frame generation, reference-based generation, motion transfer, audiovisual editing, and audio-conditioned creation as different instructions to one multimodal generation problem.
Here is what I found worth knowing before you wire it into anything.
The shape of it
H3 accepts a context made of text, images, video, and audio, and returns synchronized video plus stereo sound. Hosted output runs 4 to 15 seconds at 24 FPS with 32 kHz stereo audio, and MiniMax markets 2K by default. Under the hood the hosted 2K path is a two-step thing: H3-Base produces a 768p short-side result, then H3-Regenerate-2K reconstructs at 2K using that result plus the original multimodal context.
The contrast with the usual assembly line is the point:
typical: prompt -> silent video -> TTS -> SFX -> music -> sync pass
H3: multimodal context -> one joint video + audio latent prediction
MiniMax states the H3-Omni-Transformer predicts video and audio latents jointly, so dialogue, ambience, music, and effects are generated in context rather than stitched afterward. Dialogue is reported stable across 11 languages: English, Chinese, Japanese, Korean, French, German, Spanish, Portuguese, Italian, Russian, Arabic.
Spec sheet
| Item | Value |
|---|---|
| Developer | MiniMax |
| Category | General-purpose omni-modal video generation |
| Inputs | Text, image, video, audio |
| Output | Video + native stereo audio |
| Duration | 4–15 seconds |
| Base resolution | 768p short side (H3-Base) |
| Hosted resolution | Up to 2K via H3-Regenerate-2K, offered by default |
| Frame rate | 24 FPS |
| Audio | 32 kHz stereo |
| Stable dialogue languages | 11 |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus others |
| Reference limits | 9 images, 3 videos, 3 audio clips, 12 mixed files max |
| Core model | 33B dense H3-Omni-Transformer |
| Position encoding | 3D Multimodal RoPE |
| Open weights | H3-Base-FL2VA, H3-Base-Ref2VA |
| Precision | BF16 |
| License | MiniMax H3 Community License |
The 33B number describes the H3-Omni-Transformer, not every component in the hosted pipeline.
Pipeline: Context-IR → Base → Regenerate-2K
H3-Context-IR -> structured Context Intermediate Representation
H3-Base -> 768p video + audio
H3-Regenerate-2K -> 2K reconstruction using base output + original context
The last stage is the one I care about. A conventional upscaler interpolates pixels and cannot invent back detail that was never resolved, which is why small lettering and fine object edges usually turn to mush. H3-Regenerate-2K re-injects the original prompt and references while rebuilding the scene at higher resolution, so the model gets a second chance to resolve semantics rather than smooth them.
Architecture notes
Contextual Omni Representation is the design idea: language describes the relationship between the context, the target, and each modality. H3-Context-IR is the hosted preprocessing and orchestration layer that parses instructions, associates modalities, reasons over time, and serializes the result for H3-Base.
H3-Encoder lives inside H3-Base. It uses pretrained Qwen3-VL-32B weights and passes hidden states from layer 50 into the H3-Omni-Transformer.
H3-VAE is a family name covering both latent spaces. The open-weight docs split it:
-
H3-VisualVAE: temporally causal encoding, 16× spatial compression, 4× temporal compression, 24 latent channels, followed by 1 × 2 × 2 patchification. -
H3-AudioVAE: 32 kHz stereo compressed to latent tokens at a 40 Hz temporal rate.
H3-Omni-Transformer is a 33B dense single-stream Transformer. Roughly 13B parameters sit in AdaLN-related branches, and their modulation outputs can be precomputed and cached, so those branches do not need to stay resident in an inference-only deployment. Modality-specific parts are confined to input/output layers and AdaLN branches; the central Transformer runs on one packed sequence. 3D Multimodal RoPE encodes temporal and spatial position across (t, h, w).
What "open weight" gets you
Open weight, not fully open source. H3-Base-FL2VA and H3-Base-Ref2VA are downloadable, along with the processor, tokenizer, text encoder, Transformer, visual VAE, and audio VAE. What is not in the initial release: H3-Context-IR and H3-Regenerate-2K. Local generation targets a 768p short side, so reproducing the official 2K workflow means calling hosted stages.
Reference handling in Ref2VA mode is generous but constrained: up to 9 images, 3 video clips, and 3 audio clips, capped at 12 mixed files, with reference video and audio each 2 to 15 seconds and per-modality duration limits. Audio cannot be the only reference input. The design intent is that H3 infers relationships between assets rather than just receiving a bag of files: preserve a character, copy motion, transfer style, keep or replace audio, use a voice as timbre reference, or edit an existing scene via natural language.
Licensing is the MiniMax H3 Community License, not Apache 2.0 or MIT. Read the territorial, attribution, safety, and commercial terms before shipping.
Benchmark position
MiniMax's launch material leans on demonstrations rather than a VBench-style matrix, so Artificial Analysis's blind Video Arena is the more useful neutral read. Snapshot from September 2, 2026:
| Task | H3 rank | H3 Elo | Leader | Leader Elo |
|---|---|---|---|---|
| Text-to-Video with Audio | #4 | ≈1,228 | Wan 3.0 | 1,242 |
| Image-to-Video with Audio | #3 | 1,185 | H3 Max (fal post-train) | 1,202 |
| Video Editing with Audio | #2 | 1,129 | Wan 3.0 | 1,190 |
These move. What stands out is that H3 sits near proprietary leaders while being one of the few open-weight entries in that company. The pitch is the bundle (quality, native audiovisual generation, multimodal references, downloadable base weights), not a single top score.
Pricing
Rechecked against MiniMax's official page on September 24, 2026. Verify before you budget.
| Usage | List price |
|---|---|
| H3 generation, 768P | $0.08/second |
| H3 generation, 2K | $0.13/second |
| 768P → 2K regeneration output | $0.05/second |
| Reference audio, standard generation | Free |
| Reference images, standard generation | First 5 free, then $0.04/image |
| Reference images, regeneration | First 5 free, then $0.025/image |
| Reference video, regeneration | $0.05/second of original 768P input |
| H3-Context-IR input | $0.90/M tokens |
| H3-Context-IR output | $3.60/M tokens |
At list, 10 seconds lands near $0.80 at 768P and $1.30 at 2K; 15 seconds near $1.20 and $1.95. Those figures exclude billable references and Context-IR tokens. If you would rather not manage separate keys per vendor, a unified multi-model gateway such as CometAPI also fronts H3 under model ID minimax-h3 at an advertised starting rate of $0.064/second, though third-party headline rates depend on route, resolution, and billing rules and should not be read as a universal unit price.
Versus Wan3.0, Seedance 2.5, Vidu Q3
| Dimension | MiniMax H3 | Wan3.0 | Seedance 2.5 | Vidu Q3 |
|---|---|---|---|---|
| Max single generation | 15s | 30s | 30s | 16s |
| Resolution | 2K hosted; 768p open-weight base | Up to 1080P | Route-dependent | Up to 1080P |
| Native audio+video | Yes | Yes | Yes | Yes |
| Reference inputs | Text, image, video, audio | Image, video, audio, documents, webpages | Images, videos, audio | Image/reference workflows |
| Reference capacity | 12 mixed files | Up to 20 materials | Up to 50 materials | Not stated |
| Differentiator | Open-weight base plus 2K regeneration | 30s and broad inputs | 30s storytelling, large reference sets | Short narrative and dialogue control |
| Fits best | Customizable multimodal pipelines | Long all-in-one production | Reference-heavy commercial work | Short drama and dialogue |
No universal winner. Pick H3 when open weights, multimodal conditioning, stereo audio, editing, and 2K finishing outweigh clip length. Pick Wan 3.0 for 30s and 1080P. Pick Seedance 2.5 for huge reference sets and identity consistency. Pick Vidu Q3 for dialogue timing and camera control. Run the same internal prompt set across all of them; public leaderboards do not always expose comparable versions, and swapping in an older or turbo variant invalidates the comparison.
Known limits
- Duration: 15s ceiling against 30s competitors.
- Open-weight scope: only H3-Base downloads; Context-IR and 2K regeneration stay hosted.
- Hardware: a 33B dense video model is expensive locally even with the AdaLN caching trick.
- Cost surface: generation, regeneration, reference media, and Context-IR bill separately.
- License: Community License needs a legal pass.
- Arena volatility: rankings drift and never replace workload-specific tests.
Access paths
- MiniMax hosted products, including Hailuo and MiniMax Design.
- The first-party MiniMax Open Platform for H3-2K, H3-Context-IR, and H3-Regenerate-2K.
- Downloading H3-Base checkpoints from the official repo for local serving.
Notes from a real evaluation
Split local and hosted deliberately. Use local H3-Base for exploration, prompt iteration, reference testing, and anything with data-locality constraints. Push the keepers to hosted H3-Context-IR for richer instruction parsing and H3-Regenerate-2K for delivery resolution. The right split depends on GPU capacity, turnaround, governance, and whether the hosted quality gain justifies transfer, latency, and the extra line items.
Build the eval set from real briefs, not pretty prompts, and score instruction adherence, subject and brand consistency, temporal stability, text rendering, camera-motion accuracy, audio-video sync, dialogue quality, regeneration fidelity, latency, failure rate, and cost per accepted clip. Compare across models with identical assets and acceptance criteria.
Prepare references with one job each: one image for identity, one clip for motion, one audio sample for voice or atmosphere. Strip low-quality or contradictory assets, trim clips to the relevant action, and state each asset's relationship to the target explicitly in the instruction. Starting with the smallest sufficient set makes it obvious which reference helped and which introduced ambiguity.
Cost modeling should include reference videos, extra images, Context-IR tokens, regeneration, retries, rejected generations, local GPUs, storage and transfer, moderation, integration engineering, and human review. The number that matters is cost per approved deliverable at the required resolution.
And when reading "2K by default," keep the layers straight. It describes the hosted commercial experience; 768p describes the downloadable H3-Base stage; H3-Regenerate-2K rebuilds the final output from base plus original context. Test them as separate workflow stages and check whether regeneration restores text, branding, faces, and fine detail, or just adds pixels.
If you need much longer single-pass clips, fully local 2K finishing, a permissive license, minimal GPU spend, or plain text-to-video, H3's generalized architecture may cost more operationally than it returns. A short proof of concept on representative prompts is the cheapest way to find out.
Originally published at cometapi.com
Top comments (0)