DEV Community

Cover image for MiniMax H3 Notes: Omni-Modal Video, a 33B Backbone, and 2K In-Context Regeneration
Dylan Foster
Dylan Foster

Posted on Originally published at cometapi.com

MiniMax H3 Notes: Omni-Modal Video, a 33B Backbone, and 2K In-Context Regeneration

MiniMax shipped H3 on July 31, 2026, and it belongs in the "rethink the pipeline" bucket rather than the "another text-to-video model" bucket. The interesting part is not the demo reel. It is that H3 treats text-to-video, image-to-video, first/last-frame generation, reference-based generation, motion transfer, audiovisual editing, and audio-conditioned creation as different instructions to one multimodal generation problem.

Here is what I found worth knowing before you wire it into anything.

The shape of it

H3 accepts a context made of text, images, video, and audio, and returns synchronized video plus stereo sound. Hosted output runs 4 to 15 seconds at 24 FPS with 32 kHz stereo audio, and MiniMax markets 2K by default. Under the hood the hosted 2K path is a two-step thing: H3-Base produces a 768p short-side result, then H3-Regenerate-2K reconstructs at 2K using that result plus the original multimodal context.

The contrast with the usual assembly line is the point:

typical: prompt -> silent video -> TTS -> SFX -> music -> sync pass
H3:      multimodal context -> one joint video + audio latent prediction
Enter fullscreen mode Exit fullscreen mode

MiniMax states the H3-Omni-Transformer predicts video and audio latents jointly, so dialogue, ambience, music, and effects are generated in context rather than stitched afterward. Dialogue is reported stable across 11 languages: English, Chinese, Japanese, Korean, French, German, Spanish, Portuguese, Italian, Russian, Arabic.

Spec sheet

Item Value
Developer MiniMax
Category General-purpose omni-modal video generation
Inputs Text, image, video, audio
Output Video + native stereo audio
Duration 4–15 seconds
Base resolution 768p short side (H3-Base)
Hosted resolution Up to 2K via H3-Regenerate-2K, offered by default
Frame rate 24 FPS
Audio 32 kHz stereo
Stable dialogue languages 11
Aspect ratios 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus others
Reference limits 9 images, 3 videos, 3 audio clips, 12 mixed files max
Core model 33B dense H3-Omni-Transformer
Position encoding 3D Multimodal RoPE
Open weights H3-Base-FL2VA, H3-Base-Ref2VA
Precision BF16
License MiniMax H3 Community License

The 33B number describes the H3-Omni-Transformer, not every component in the hosted pipeline.

Pipeline: Context-IR → Base → Regenerate-2K

H3-Context-IR  -> structured Context Intermediate Representation
H3-Base        -> 768p video + audio
H3-Regenerate-2K -> 2K reconstruction using base output + original context
Enter fullscreen mode Exit fullscreen mode

The last stage is the one I care about. A conventional upscaler interpolates pixels and cannot invent back detail that was never resolved, which is why small lettering and fine object edges usually turn to mush. H3-Regenerate-2K re-injects the original prompt and references while rebuilding the scene at higher resolution, so the model gets a second chance to resolve semantics rather than smooth them.

Architecture notes

Contextual Omni Representation is the design idea: language describes the relationship between the context, the target, and each modality. H3-Context-IR is the hosted preprocessing and orchestration layer that parses instructions, associates modalities, reasons over time, and serializes the result for H3-Base.

H3-Encoder lives inside H3-Base. It uses pretrained Qwen3-VL-32B weights and passes hidden states from layer 50 into the H3-Omni-Transformer.

H3-VAE is a family name covering both latent spaces. The open-weight docs split it:

  • H3-VisualVAE: temporally causal encoding, 16× spatial compression, 4× temporal compression, 24 latent channels, followed by 1 × 2 × 2 patchification.
  • H3-AudioVAE: 32 kHz stereo compressed to latent tokens at a 40 Hz temporal rate.

H3-Omni-Transformer is a 33B dense single-stream Transformer. Roughly 13B parameters sit in AdaLN-related branches, and their modulation outputs can be precomputed and cached, so those branches do not need to stay resident in an inference-only deployment. Modality-specific parts are confined to input/output layers and AdaLN branches; the central Transformer runs on one packed sequence. 3D Multimodal RoPE encodes temporal and spatial position across (t, h, w).

What "open weight" gets you

Open weight, not fully open source. H3-Base-FL2VA and H3-Base-Ref2VA are downloadable, along with the processor, tokenizer, text encoder, Transformer, visual VAE, and audio VAE. What is not in the initial release: H3-Context-IR and H3-Regenerate-2K. Local generation targets a 768p short side, so reproducing the official 2K workflow means calling hosted stages.

Reference handling in Ref2VA mode is generous but constrained: up to 9 images, 3 video clips, and 3 audio clips, capped at 12 mixed files, with reference video and audio each 2 to 15 seconds and per-modality duration limits. Audio cannot be the only reference input. The design intent is that H3 infers relationships between assets rather than just receiving a bag of files: preserve a character, copy motion, transfer style, keep or replace audio, use a voice as timbre reference, or edit an existing scene via natural language.

Licensing is the MiniMax H3 Community License, not Apache 2.0 or MIT. Read the territorial, attribution, safety, and commercial terms before shipping.

Benchmark position

MiniMax's launch material leans on demonstrations rather than a VBench-style matrix, so Artificial Analysis's blind Video Arena is the more useful neutral read. Snapshot from September 2, 2026:

Task H3 rank H3 Elo Leader Leader Elo
Text-to-Video with Audio #4 ≈1,228 Wan 3.0 1,242
Image-to-Video with Audio #3 1,185 H3 Max (fal post-train) 1,202
Video Editing with Audio #2 1,129 Wan 3.0 1,190

These move. What stands out is that H3 sits near proprietary leaders while being one of the few open-weight entries in that company. The pitch is the bundle (quality, native audiovisual generation, multimodal references, downloadable base weights), not a single top score.

Pricing

Rechecked against MiniMax's official page on September 24, 2026. Verify before you budget.

Usage List price
H3 generation, 768P $0.08/second
H3 generation, 2K $0.13/second
768P → 2K regeneration output $0.05/second
Reference audio, standard generation Free
Reference images, standard generation First 5 free, then $0.04/image
Reference images, regeneration First 5 free, then $0.025/image
Reference video, regeneration $0.05/second of original 768P input
H3-Context-IR input $0.90/M tokens
H3-Context-IR output $3.60/M tokens

At list, 10 seconds lands near $0.80 at 768P and $1.30 at 2K; 15 seconds near $1.20 and $1.95. Those figures exclude billable references and Context-IR tokens. If you would rather not manage separate keys per vendor, a unified multi-model gateway such as CometAPI also fronts H3 under model ID minimax-h3 at an advertised starting rate of $0.064/second, though third-party headline rates depend on route, resolution, and billing rules and should not be read as a universal unit price.

Versus Wan3.0, Seedance 2.5, Vidu Q3

Dimension MiniMax H3 Wan3.0 Seedance 2.5 Vidu Q3
Max single generation 15s 30s 30s 16s
Resolution 2K hosted; 768p open-weight base Up to 1080P Route-dependent Up to 1080P
Native audio+video Yes Yes Yes Yes
Reference inputs Text, image, video, audio Image, video, audio, documents, webpages Images, videos, audio Image/reference workflows
Reference capacity 12 mixed files Up to 20 materials Up to 50 materials Not stated
Differentiator Open-weight base plus 2K regeneration 30s and broad inputs 30s storytelling, large reference sets Short narrative and dialogue control
Fits best Customizable multimodal pipelines Long all-in-one production Reference-heavy commercial work Short drama and dialogue

No universal winner. Pick H3 when open weights, multimodal conditioning, stereo audio, editing, and 2K finishing outweigh clip length. Pick Wan 3.0 for 30s and 1080P. Pick Seedance 2.5 for huge reference sets and identity consistency. Pick Vidu Q3 for dialogue timing and camera control. Run the same internal prompt set across all of them; public leaderboards do not always expose comparable versions, and swapping in an older or turbo variant invalidates the comparison.

Known limits

  • Duration: 15s ceiling against 30s competitors.
  • Open-weight scope: only H3-Base downloads; Context-IR and 2K regeneration stay hosted.
  • Hardware: a 33B dense video model is expensive locally even with the AdaLN caching trick.
  • Cost surface: generation, regeneration, reference media, and Context-IR bill separately.
  • License: Community License needs a legal pass.
  • Arena volatility: rankings drift and never replace workload-specific tests.

Access paths

  1. MiniMax hosted products, including Hailuo and MiniMax Design.
  2. The first-party MiniMax Open Platform for H3-2K, H3-Context-IR, and H3-Regenerate-2K.
  3. Downloading H3-Base checkpoints from the official repo for local serving.

Notes from a real evaluation

Split local and hosted deliberately. Use local H3-Base for exploration, prompt iteration, reference testing, and anything with data-locality constraints. Push the keepers to hosted H3-Context-IR for richer instruction parsing and H3-Regenerate-2K for delivery resolution. The right split depends on GPU capacity, turnaround, governance, and whether the hosted quality gain justifies transfer, latency, and the extra line items.

Build the eval set from real briefs, not pretty prompts, and score instruction adherence, subject and brand consistency, temporal stability, text rendering, camera-motion accuracy, audio-video sync, dialogue quality, regeneration fidelity, latency, failure rate, and cost per accepted clip. Compare across models with identical assets and acceptance criteria.

Prepare references with one job each: one image for identity, one clip for motion, one audio sample for voice or atmosphere. Strip low-quality or contradictory assets, trim clips to the relevant action, and state each asset's relationship to the target explicitly in the instruction. Starting with the smallest sufficient set makes it obvious which reference helped and which introduced ambiguity.

Cost modeling should include reference videos, extra images, Context-IR tokens, regeneration, retries, rejected generations, local GPUs, storage and transfer, moderation, integration engineering, and human review. The number that matters is cost per approved deliverable at the required resolution.

And when reading "2K by default," keep the layers straight. It describes the hosted commercial experience; 768p describes the downloadable H3-Base stage; H3-Regenerate-2K rebuilds the final output from base plus original context. Test them as separate workflow stages and check whether regeneration restores text, branding, faces, and fine detail, or just adds pixels.

If you need much longer single-pass clips, fully local 2K finishing, a permissive license, minimal GPU spend, or plain text-to-video, H3's generalized architecture may cost more operationally than it returns. A short proof of concept on representative prompts is the cheapest way to find out.


Originally published at cometapi.com

Top comments (0)