ByteDance's Seedance 2.5 generates 30-second video in one pass with 30 image, 10 video, 10 audio references. Follows Seedance 2.0 from February 2026.
ByteDance's Seed team released Seedance 2.5 in late July 2026, enabling a single generation pass to produce a 30-second video clip — tripling Seedance 2.0's typical single-shot output.
Key facts
- Seedance 2.5 generates 30-second video in one pass
- Supports 30 image, 10 video, 10 audio references
- Multi-round extension chains longer narratives
- Seedance 2.0 debuted Feb 16, 2026 on Jimeng AI
- Timestamp-controlled editing for precise modifications
The update extends the model's reach beyond raw duration. According to the ByteDance Seedance 2.5 release, the new version supports up to 30 image references, 10 video references, and 10 audio references in a single prompt — a substantial jump in multi-modal conditioning. The company positions this as a long-narrative tool, not just a longer clip generator.
Key Takeaways
- ByteDance's Seedance 2.5 generates 30-second video in one pass with 30 image, 10 video, 10 audio references.
- Follows Seedance 2.0 from February 2026.
How the 30-second generation changes workflow
Seedance 2.0, which debuted February 16, 2026 on ByteDance's Jimeng AI platform, typically produced shorter clips requiring multiple passes for extended scenes. Seedance 2.5 collapses that pipeline: one generation call yields 30 seconds of coherent footage. For production teams, this cuts the number of API round-trips needed for a single shot, reducing both latency and the risk of style drift between segments.
The multi-round extension feature chains successive generations, allowing narratives to stretch well beyond the base 30 seconds. Combined with timestamp-controlled editing — which lets creators specify precise moments in the timeline for targeted modifications — the model moves closer to a directable video editor than a pure text-to-video generator.
Multi-modal references: 30 images, 10 videos, 10 audio clips
Seedance 2.5 accepts a dense reference stack. A creator can feed it 30 images for character consistency, 10 video clips for motion reference, and 10 audio tracks for dialogue or ambient sound. This is markedly higher than most competing models, which typically cap at a handful of images and a single audio input.
The company did not disclose training compute, parameter count, or inference cost per generation. Pricing and API availability details also remain unspecified in the release, though the model is expected to reach ByteDance's Jimeng AI platform and the Volcano Engine enterprise API in the coming weeks.
What to watch
Watch for Seedance 2.5's arrival on the Volcano Engine API and Jimeng AI, where pricing per generation will reveal whether the 30-second capability is cost-competitive with multi-pass alternatives. Also track benchmark comparisons against Flux 3, which beat Seedance 2.0 in BFL tests in late July 2026.
Source: pandaily.com
[Updated 01 Aug via scmp_tech]
The competitive landscape is heating up: MiniMax has launched H3, a rival video generation model with open weights and a price of 0.8 yuan per second, directly challenging ByteDance's closed-source approach [per SCMP]. H3 supports 15-second 2K native dual-channel audio-video generation and currently ranks first in video editing on benchmark platform Artificial Analysis, though it trails Google's Gemini Omni Flash in text-to-video tasks [per Pandaily].
Originally published on gentic.news


Top comments (0)