Both models take text or images in and hand you video with synchronized audio. Both do reference-driven work. If you compare them on a feature grid you will conclude they overlap almost entirely, and you will pick based on price per second, which is the wrong axis.
The useful split is about where your production cost concentrates. One model makes a rejected take cheap. The other makes a finished scene coherent. Those are different problems and they rarely live in the same job.
The decision rule I actually use
- If the bottleneck is drafts — discarded takes, alternate camera moves, hook variants, prompt iteration — route to MiniMax H3 Max.
- If the bottleneck is continuity — a scene longer than ~15 seconds, a large reference pack, timestamp-level edits — route to Seedance 2.5.
Everything below is justification for that rule.
Provenance matters here more than usual
MiniMax H3 Max
H3 Max is not "the official MiniMax H3 endpoint with a faster tier." fal Research started from the MiniMax H3 open weights and applied post-training aimed at prompt adherence and aesthetics, while co-designing the serving stack for throughput. That last part is the whole product: a 5-second 768p clip in under three seconds of inference on fal's infrastructure (roughly 2.8–3s reported).
The route-level spec: minimax-h3-max, 5–15 seconds, 24 FPS, synchronized stereo audio, six aspect ratios (21:9, 16:9, 4:3, 1:1, 3:4, 9:16), 480P/768P output, up to 12 total references with documented caps of 3 video and 3 audio.
Latency matters even when the model isn't the final renderer. Every rejected take is part of the cost of arriving at the accepted clip. During pre-production you can evaluate more hypotheses per hour, and that's usually the metric that matters.
Seedance 2.5
ByteDance Seed's audio-video model, officially launched July 31, 2026. The launch materials push three things: single-pass clips up to 30 seconds, large multimodal reference packs, and timestamp-level editing.
The 30-second ceiling is not a bigger number, it's a different topology. Setup, development, transitions, and payoff stay inside one generation instead of forcing a mid-scene seam. ByteDance describes it as organizing multiple connected shots in that window, plus multi-round extension when the story continues.
Reference capacity at model level is up to 50 in a single pass: 30 images, 10 video clips, 10 audio clips, with modes for motion, clay/3D blocking, characters, props, scene style, and audio. On the current CometAPI route it's exposed through the async Video API as seedance-2-5, 4–30 seconds, 480p/720p.
Side-by-side, route-level
| MiniMax H3 Max | Seedance 2.5 | |
|---|---|---|
| Model ID | minimax-h3-max |
seedance-2-5 |
| Origin | MiniMax H3 open weights + fal post-training | ByteDance Seed |
| Duration (route) | 5–15 s | 4–30 s |
| Resolution (route) | 480P / 768P | 480p / 720p |
| Frame rate | 24 FPS | — |
| Audio | Synchronized stereo | Joint audio-video generation |
| Inputs | Text, image, reference media, first/last frame | Text, image; broader model-level image/video/audio reference workflows |
| Reference cap | 12 total (≤3 video, ≤3 audio documented) | 50 at model level (30 image + 10 video + 10 audio) |
| Editing / extension | Not the differentiator | Timestamp editing, green-screen, perspective, reference-based edits; multi-round extension |
| Headline price | $0.064/s | $0.0824/s |
That resolution row is deliberately narrow. fal exposes a 1080P H3 Max option, described as a latent refinement of a 768P render — but the CometAPI H3 Max docs list 480P and 768P. Seedance 2.5's route documents 480p and 720p even where other providers expose more tiers. "The model supports 1080p" is a meaningless sentence unless it names the endpoint.
Where the differences actually bite
Latency and the iteration loop
H3 Max has the better public number and there is no current Seedance 2.5 latency figure measured under the same hardware and queue conditions, so any "N× faster" claim would be made up. What is defensible: H3 Max is the natural fit for interactive prompt exploration, ad variants, storyboarding, and batch jobs where most outputs get thrown away.
Duration and seam count
5–15s versus 4–30s. For a ten-second product hook this is a non-issue. For a 25–30 second brand film it changes the production topology: H3 Max forces multiple generations and an edit point; Seedance 2.5 absorbs the scene in one pass. Fewer generation boundaries means fewer places where character appearance, lighting, camera language, or audio drift.
References and editing
12 inputs versus 50. The gap shows up when a shot is constrained by a lead character, a secondary character, a packshot, a location, a camera-motion reference, a soundtrack, a voice, and styling. H3 Max handles compact packs fine. Seedance 2.5 is architecturally built for the heavy case — and it treats editing as first-class rather than generate-once, which makes it closer to an iterative production system.
Audio
Both compose audio alongside picture, so neither forces a separate "generate silent, then align" stage for most workloads. That does not make their audio behavior identical — dialogue, effects, ambience, musical structure, controllability, and provider toggles all need per-prompt testing.
Benchmarks are asymmetric
H3 Max has the stronger point-in-time public results, including preference-style evaluations. Those tests do not measure 30-second continuity, 50-reference workflows, or timestamp editing under matched conditions. There is no clean public table where every Seedance 2.5 capability is scored against H3 Max at the same duration, resolution, prompt set, references, audio conditions, and provider latency. A short blind-preference benchmark disadvantages a continuity model; a long-scene-only test disadvantages a speed model. The closest controlled comparison available is fal's six-scenario head-to-head across text-to-video, image-to-video, and reference-to-video — provider-run, not a leaderboard.
Pricing: per generated second is the wrong unit
| H3 Max | Seedance 2.5 | |
|---|---|---|
| Headline starting rate | $0.064/s | $0.0824/s |
| 480p documented route rate | verify live route | $0.103/s |
| 720p documented route rate | — | $0.231/s |
| Billing basis | per generated second | resolution-dependent |
These are route prices, not intrinsic model prices, and they can move independently of any model release.
What you care about is cost per accepted clip. If a team burns eight attempts to approve one ten-second shot, H3 Max's unit cost and turnaround compound across seven dead outputs. If Seedance 2.5 returns one coherent 30-second shot that would otherwise be two or three clips plus an edit session, the higher per-second rate can still win on total effort.
Routing table
| Workflow | Fit | Reason |
|---|---|---|
| 5-second social hook | H3 Max | Latency and cost per draft |
| A/B testing 10 ad concepts | H3 Max | Throughput beats a 30s ceiling |
| 15-second product teaser | Depends | H3 Max for drafts; Seedance if references/editing are heavy |
| 20–30-second brand story | Seedance 2.5 | One generation, fewer seams |
| Large character/product reference pack | Seedance 2.5 | 50 model-level references |
| Reference-heavy shot needing edits | Seedance 2.5 | Editing and extension are design centers |
| Storyboard / camera exploration | H3 Max | Faster reject-and-refine loop |
| High-volume short-video API | H3 Max | Unit economics + throughput |
Calling both: one async pattern
I run both through a single unified endpoint — CometAPI in my case — because the job architecture is identical regardless of which model handles the task. Submit a multipart POST to /v1/videos, keep the returned task ID, poll, fetch the result. Only the model ID and parameter values change.
Five-second 768P H3 Max:
curl https://api.cometapi.com/v1/videos \
-H "Authorization: Bearer $COMETAPI_KEY" \
--form-string 'model=minimax-h3-max' \
--form-string 'prompt=A premium product shot, slow dolly-in, soft daylight, no text.' \
--form-string 'seconds=5' \
--form-string 'size=1360x768'
Twenty-second 720p Seedance 2.5:
curl https://api.cometapi.com/v1/videos \
-H "Authorization: Bearer $COMETAPI_KEY" \
--form-string 'model=seedance-2-5' \
--form-string 'prompt=A continuous cinematic tracking shot through a rain-lit night market.' \
--form-string 'seconds=20' \
--form-string 'size=1280x720'
Because the shape is the same, routing is a policy decision rather than an integration project: send short drafts and high-volume variants to H3 Max, send anything over 15 seconds or carrying a large reference pack to Seedance 2.5.
The hybrid loop, concretely
H3 Max belongs early, where you're exploring prompts, shot structures, camera moves, and hooks. Once the concept stabilizes, Seedance 2.5 takes the longer, reference-heavy, editing-intensive version. Model selection becomes a router rule, not a rewrite.
Read the current model page and API reference before hard-coding resolutions or reference limits into a service. Video endpoints move fast, and provider-specific capabilities do not always land on every aggregator route the same day.
Originally published at cometapi.com
Top comments (0)