I’d evaluate MiniMax H3 Max primarily as a latency and throughput option. The independent quality scores favor it over H3, but the gap is modest. The more substantial claim is fal’s reported generation speed: a five-second 768p clip in roughly three seconds or less on its optimized infrastructure.
H3 Max comes from fal Research’s post-training of MiniMax H3 open weights. Its targets are prompt adherence, visual aesthetics, and inference efficiency. Public materials do not establish a larger Transformer or a new parameter scale behind the “Max” name.
That distinction matters when choosing an endpoint. H3 Max targets fast audiovisual generation. Standard H3 offers the broader system: richer multimodal conditioning, video editing, open H3-Base checkpoints, and a path to 2K output.
All benchmark and pricing figures below refer to the September 18, 2026 snapshot.
Start with the workflow you need
Before comparing Elo or price, I’d check whether both models support the actual job.
| Requirement | H3 Max | Standard H3 |
|---|---|---|
| Fast iterative generation | Primary optimization target | Broader generation system |
| Text-to-video and image-to-video with audio | Supported | Supported |
| First/last-frame and reference workflows | Supported | Broader multimodal conditioning |
| Video editing | No ranking in the comparison below | Included in the broader workflow |
| Maximum documented output path | 1080p refinement from 768p | H3-Regenerate-2K |
| Open-weight deployment | Do not infer availability from H3 ancestry | H3-Base checkpoints available |
For short-form production with frequent retries, I’d put H3 Max on the evaluation shortlist. For editing, local validation of open weights, or 2K delivery, I’d start with standard H3.
H3 Max’s exposed capabilities
| Property | Documented behavior |
|---|---|
| Foundation | MiniMax H3 |
| Post-training developer | fal Research |
| Generation types | Text-to-video, image-to-video, first-to-last-frame, reference-to-video |
| Duration | 5–15 seconds |
| Native resolution | 480p / 768p |
| Higher-resolution option | 1080p latent refinement from native 768p |
| Frame rate | 24 FPS |
| Audio | Synchronized stereo audio |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
The text-to-video endpoint also exposes prompt-expansion modes including disabled, balanced, and quality. Those modes add another variable to latency and prompt interpretation, so I’d keep them fixed during comparisons.
Workflow selection matters too. Text-to-video fits language-defined scenes; image-to-video anchors the result to a supplied frame. First-to-last-frame control is useful when both endpoint states matter, while reference-to-video supports consistency with visual references.
What fal changed, and what comes from H3
The underlying H3 system jointly generates audio and video. Sound is part of the generation architecture, which supports coordinated dialogue, ambience, sound effects, and motion.
MiniMax describes H3-Omni-Transformer as a 33-billion-parameter dense single-stream Transformer, with approximately 13B parameters in AdaLN-related branches. H3-Encoder uses pretrained Qwen3-VL-32B weights, and separate visual and audio VAEs encode their respective modalities. The Transformer jointly predicts video and audio latents.
The complete H3 workflow has three components:
- H3-Context-IR interprets free-form multimodal instructions.
- H3-Base performs core audiovisual generation at 768p.
- H3-Regenerate-2K uses the original context and lower-resolution result to regenerate output at higher resolution.
fal’s contribution sits on top of that foundation. It added post-training data and evaluation loops targeting instruction fidelity and visual quality, while developing the inference stack alongside the model.
That joint development is central to the speed claim. Lowering sampling steps or precision can reduce latency while hurting output quality. fal says it retained candidate optimizations only when the resulting checkpoints continued to perform well in its preference evaluations.
I read this as a change to H3’s quality, speed, and cost trade-off. It does not establish that every capability of the complete H3 system carries over to H3 Max.
Open weights cover only part of standard H3
MiniMax released the H3-Base FL2VA and Ref2VA checkpoints under the H3 Community License, allowing local validation of core 768p generation.
The complete production workflow remains partly hosted. H3-Context-IR and H3-Regenerate-2K are hosted components, so reproducing the official end-to-end 2K workflow requires MiniMax APIs.
That is a useful distinction for deployment planning: access to H3-Base weights does not provide the entire hosted system.
Separate the speed claim from the quality evidence
fal reports a five-second 768p generation in under three seconds on its optimized infrastructure, with roughly 35× the throughput of the official MiniMax H3 endpoint in its comparison.
That is a provider-specific result. I would measure the exact endpoint I intended to deploy before using that number in a product latency budget.
Observed latency also includes queueing, input uploads, network transfer, prompt expansion, safety processing, and provider routing. Duration and resolution affect the workload as well. A different backend can expose the same model family without reproducing fal’s timing.
Independent preference scores
The Artificial Analysis snapshot gives H3 Max the higher measured score in both directly comparable audio-video generation categories.
| Category, September 18, 2026 | H3 Max Elo | H3 Elo | Difference |
|---|---|---|---|
| Text-to-video with audio | 1227 ±9 | 1220 ±8 | +7 |
| Image-to-video with audio | 1195 ±10 | 1181 ±8 | +14 |
| Video editing with audio | Not ranked in this comparison | 1132 ±6 | — |
The text-to-video sample counts were 5,689 for H3 Max and 8,602 for H3. Image-to-video used 5,569 for H3 Max and 6,949 for H3.
These are dated blind human-preference Elo measurements, with 95% confidence intervals. The intervals overlap. I’d treat the results as directional evidence for H3 Max, then test whether that preference survives on my own prompt distribution.
Exact prompts, generation settings, and evaluation procedures should be checked against the Artificial Analysis methodology. A small aggregate lead is not enough to predict which model will handle a particular camera instruction, identity constraint, or scene transition better.
Provider evaluations answer a different question
fal reports first place in overall preference, prompt understanding, and aesthetics in its head-to-head evaluation against twelve video models.
That helps explain the post-training targets, but it remains provider-run evidence. Its public cost-versus-quality chart does not disclose every prompt, hardware detail, or serving parameter needed for independent reproduction.
The evidence supports a stronger speed claim than a universal quality claim. I’d expect to validate prompt adherence carefully, especially for prompts combining actions, camera directions, temporal constraints, style instructions, and scene transitions.
Resolution labels hide different processing paths
H3 Max’s documented 1080p option uses latent refinement from native 768p. Standard H3 reaches 2K through H3-Regenerate-2K, which reuses both the original multimodal instructions and the 768p result.
The latter is an in-context regeneration stage that can reconstruct details. It is distinct from a conventional super-resolution pass and from H3 Max’s documented refinement path.
My evaluation sequence would be:
- 480p: inexpensive drafts and concept screening.
- 768p: normal quality evaluation and many web deliveries.
- H3 Max 1080p refinement: higher-resolution delivery while prioritizing fast production.
- H3 2K regeneration: workloads where detail and the broader H3 workflow justify the added cost and latency.
I’d review final artifacts at the intended display size. A resolution label alone cannot tell me whether motion, texture, or identity remains acceptable.
Price the accepted output
The public rates checked on September 18, 2026 were:
| Provider and output path | H3 Max | H3 |
|---|---|---|
| fal 480p | $0.025/sec, promotional | — |
| fal 768p | $0.04/sec, promotional | — |
| fal 1080p refinement | $0.08/sec, promotional | — |
| MiniMax official 768p | — | $0.08/sec |
| MiniMax official 2K | — | $0.13/sec |
| MiniMax 768p → 2K regeneration | — | $0.05/sec |
fal labels its rates as promotional launch pricing and states that the discount ends September 30, 2026. Those figures need rechecking before forecasting a production workload.
For a unified multi-model API comparison, CometAPI lists H3 Max with model ID minimax-h3-max and a starting price of $0.064 per second; standard H3 also starts at $0.064 per second there.
The metric I care about is cost per accepted clip. That includes retries, rejected outputs, refinement or regeneration, storage, transfer, review time, and downstream editing.
Faster generation can shorten iteration cycles. Stronger adherence can reduce rejected outputs. Conversely, H3’s editing or multimodal controls may avoid extra tools and rework. The listed rate captures only one part of that calculation.
Wire up the asynchronous job lifecycle
The unified API described above exposes generation through /v1/videos. Its documented workflow is submission, polling, then content retrieval:
POST /v1/videos
GET /v1/videos/{task_id}
GET /v1/videos/{task_id}/content
Create an API token and store it securely. Submit a generation request using model minimax-h3-max, a prompt, duration, and output size. For image-to-video, add the image URL according to the live schema.
Read the returned task ID and poll with an interval until the task becomes completed or failed. For a completed task, retrieve the content, save the MP4, and inspect the result.
I’d verify the live request schema before implementing the payload. Routing, exposed parameters, and prices can change independently of the underlying model. The endpoint lifecycle alone does not establish exact request field names or every supported option.
Build the comparison around production failures
For a fair H3 Max versus H3 test, I’d hold these inputs and policies constant:
- Prompt and reference inputs.
- Duration, aspect ratio, and resolution.
- Audio settings and prompt-expansion configuration.
- Retry policy and acceptance criteria.
Then I’d score the outputs on prompt adherence, motion coherence, visual artifacts, audio quality and synchronization, and identity consistency. Reviewer acceptance rate matters more to the application than an isolated attractive frame.
Operational measurements should include task-failure rate, retry rate, p50 and p95 end-to-end latency, and cost per accepted clip. Segment results by generation workflow, duration, resolution, aspect ratio, and prompt complexity. An aggregate win can conceal a regression in the scenario that produces most of your workload.
My default decision would be H3 Max when iteration speed, throughput, and prompt adherence dominate. I’d choose standard H3 when the job depends on video editing, deeper multimodal conditioning, open H3-Base deployment, or 2K regeneration.
The switching criterion would be concrete: better acceptance-adjusted cost and latency on representative jobs, with no unacceptable regression in the workflows the product depends on.
Originally published at cometapi.com
Top comments (0)