DEV Community

Cover image for FLUX 3: What the Video API Gives You, What It Costs, Where It Still Fails
Mason Reed
Mason Reed

Posted on Originally published at cometapi.com

FLUX 3: What the Video API Gives You, What It Costs, Where It Still Fails

FLUX 3 is Black Forest Labs' unified multimodal foundation model, and the architectural bet is the interesting part: one shared representation trained across images, video and audio instead of three separate models. BFL's framing is that images, video and sound are different observations of the same world. A moving car has appearance, motion, sound, spatial structure, material behaviour and causal behaviour, and learning those signals together gives the network more constraints than learning frames in isolation.

That is why BFL describes FLUX 3 as a step toward a reality or world model rather than just another video generator. The claim is broader than the shipped product. As of September 2026 the video endpoint is live, while FLUX 3 Image and the open-weight FLUX 3 Dev model are separate rollout items. Here is what you can call today, what it costs, and where it breaks.

What has shipped

Capability Status
FLUX 3 Video generation Generally available
Native video audio Generally available
Text-to-video Generally available
Image-to-video / keyframes Generally available
Video continuation Generally available
Richer image/video/audio reference combinations Planned expansion
FLUX 3 Image generation and editing Upcoming
FLUX 3 Dev open weights Planned
Action-prediction deployment Research / commercial partner track

If you need stills from this family, FLUX.2 is still the dedicated option (FLUX.2 Max among them). The public FLUX 3 product is video and audio, full stop. Do not read the family roadmap as a shipped image endpoint.

Specs

Specification Value
Developer Black Forest Labs
Model family FLUX 3
Model type Unified multimodal foundation / generative video model
Public video release August 4, 2026
Text-to-video Supported
Image-to-video / keyframes Supported
Video continuation Supported
Max T2V / I2V duration 20 seconds
Continuation duration 5–15 seconds
Frame rate 24 fps
Resolution tiers HD, FHD, QHD/2K, UHD/4K
FHD 16:9 output 1920 × 1088
Image references Up to 10 for I2V / keyframe workflows
Native audio Yes
Multilingual dialogue + lip sync Yes
Multi-shot generation Yes
Draft Mode Yes
Research foundation Self-Flow
Core training direction Joint image, video and audio representation learning
Public still-image endpoint Not yet generally released
Open-weight FLUX 3 Dev Planned

Aspect ratios cover 21:9, 2:1, 16:9, 4:3, 1:1, 3:4 and 9:16. Resolution bands are defined by total pixels per frame rather than aspect ratio, and Draft Mode is HD only. Note the duration split: text-to-video and image-to-video accept 5–20 seconds, continuation accepts a 5–15 second generation window.

Pricing, and a working request

BFL bills by generated duration.

Mode BFL list price
T2V / I2V Draft HD $0.06/s
T2V / I2V HD $0.17/s
T2V / I2V FHD $0.29/s
T2V / I2V QHD (2K) $0.40/s
T2V / I2V UHD (4K) $0.80/s
Continuation Draft $0.12/s
Continuation HD $0.41/s
Continuation FHD $0.53/s
Continuation QHD (2K) $0.65/s
Continuation UHD (4K) $0.95/s

A 10-second HD render lands around $1.70 at list; the same length at FHD is about $2.90. Continuation is substantially more expensive per second than ordinary generation at every tier, so if your product extends clips often, model that as its own cost line instead of folding it into an average.

The request shape is a standard form POST to an OpenAI-style /v1/videos endpoint:

curl --request POST "https://api.cometapi.com/v1/videos" \
  --header "Authorization: Bearer $COMETAPI_KEY" \
  --form-string "model=flux-3" \
  --form-string "prompt=A paper boat glides across a still pond in soft morning light" \
  --form-string "seconds=5" \
  --form-string "size=1280x720"
Enter fullscreen mode Exit fullscreen mode

The workflow is asynchronous. The initial call returns a task ID, you poll until the job completes, then you fetch the resulting video. Do not hold an HTTP request open for the full render; push generation through a background job or task queue and let the client poll your own status endpoint. I route this through CometAPI, which lists the model as flux-3 (marked Available, platform release date August 13, 2026) at $0.136/s for 720p and $0.232/s for 1080p, roughly 20% under BFL's listed $0.17/s and $0.29/s. That is about $1.36 and $2.32 for a 10-second run. One caveat worth flagging: some older descriptive copy on that site still reflects the July Early Access period, so check the live model card and pricing fields before you commit to an integration.

Under the hood: Self-Flow

The research foundation is Self-Flow, BFL's self-supervised flow-matching approach. Standard generative pipelines lean on separately pretrained representation models or auxiliary supervision. Self-Flow is built to learn useful representations directly from the generative objective.

The mechanism to know about is Dual-Timestep Scheduling: different groups of tokens get different noise levels during training. That creates information asymmetry, where some parts of the input carry more usable signal than others, which forces the network to build representations that help infer what is missing. In practice FLUX 3 is pushed toward learning relationships between objects, frames, motion, sound and temporal events rather than memorizing what a single frame should look like.

Self-Flow method architecture

Self-Flow architecture, image source: Black Forest Labs

BFL also ran joint multimodal training experiments on a FLUX.2-derived backbone with millions of videos and hundreds of millions of images. Those experiments support the broader hypothesis that representation learning and generation can share a system.

Why video eats the compute

Video carries information a still cannot. A single frame shows a ball above the floor; a video shows it falling, bouncing, deforming on impact, making a sound and changing direction. BFL disclosed that video prediction accounts for more than 95% of FLUX 3's training compute. Audio is comparatively cheap because it is a much lower-dimensional signal tightly coupled to visible events. That allocation is a decent explanation for why video shipped first.

Capability notes

Keyframes and references. The current API accepts up to ten image references in image-to-video workflows, so you can animate an initial image, drive toward an end frame, or place multiple images as timed keyframes instead of letting the model improvise the whole sequence.

Continuation. Feed an existing video plus its audio and FLUX 3 generates what happens next, preserving movement, camera behaviour, dialogue and sound across the seam. If you have ever generated a second clip and concatenated the files, you know how visible that join usually is.

Audio. Sound is produced at generation time, not in a separate speech or SFX pass. Dialogue, effects and ambience are all in scope, with multilingual dialogue and lip sync.

Multi-shot. A single prompt can carry multiple scenes or camera angles. Most earlier models are strongest on one continuous shot; this one is explicitly built for multi-shot sequences in one generation.

Draft Mode. You can generate a cheaper preview and then enhance the selected draft while keeping the composition and motion you picked. Per current BFL pricing a standard T2V/I2V draft is $0.06/s against $0.17/s for regular HD. For prompt-heavy iteration this is the feature that matters most.

Benchmarks: vendor versus third party

There are two useful evidence sets now.

BFL's August 4 evaluation of the released model reports a text-to-video ELO of 1135 in an internal all-vs-all human-preference test, the highest score in the group BFL tested. Image-to-video tied the strongest tested competitor and beat the rest. This supersedes the preliminary July win rates, though it is still a vendor-run protocol.

Megaton's v-benchmark v2 is the independent check, run in August 2026.

Dimension Megaton score
Megaton Index 76.51 (rank #3 at evaluation time)
Prompt Adherence 87.07
Scene Consistency 94.76
Physics 70.63
Human Fidelity 73.70
Object & Product Fidelity 83.82
Causal & Semantic Coherence 80.91
Text Fidelity 82.41
Cinematography 71.43
Taste & Art Direction 65.00

The profile tells you more than the index. Scene consistency is the standout, and prompt adherence, object fidelity, causal coherence and text fidelity all clear 80. Physics, human fidelity and art direction are the weak spots, which is the useful takeaway: better temporal representation does not solve synthetic motion or taste. The two rankings diverging (first internally, third externally) is not a contradiction. Different prompt distributions, scoring dimensions, evaluator pools and aggregation rules will do that.

Against Seedance 2.5, MiniMax H3 and Wan 3.0

Dimension FLUX 3 Seedance 2.5 MiniMax H3 Wan 3.0
Developer Black Forest Labs ByteDance Seed MiniMax Alibaba
Max advertised clip 20 s 30 s 15 s 30 s
Resolution HD / FHD / QHD (2K) / UHD (4K) API tier dependent Up to 2K Up to 1080p
Native audio Yes Yes Yes, stereo Yes
Text-to-video Yes Yes Yes Yes
Image / reference input Yes Yes Yes Yes
Keyframe / reference control Strong timed keyframes Strong multimodal reference control Multimodal conditioning Multimodal reference control
Continuation / editing Continuation available Editing-oriented workflow Omni generation focus Editing and reference workflows
Multi-shot Yes Yes Yes Yes
Distinctive strength Reality-model architecture, scene consistency, Self-Flow Long-form storytelling, reference control 2K audiovisual generation, omni-modal context Long clips, lower entry cost
Listed entry price / s $0.136 $0.0824 $0.064 $0.040

Treat the price row as a rough comparison only. These APIs bill against different resolutions, durations, quality tiers and rules, so the cheapest per-second number is not the cheapest equivalent-quality output.

Seedance 2.5 wins on duration outright (30 s against 20 s) and leans hard into reference-driven generation and editing. Megaton currently ranks it above FLUX 3 overall, so expect real competition at the top of this market. MiniMax H3 is the spec argument: up to 2K with native stereo audio, but a 15 s ceiling against FLUX 3's 20 s, and no equivalent of Draft Mode for cheap iteration. Wan 3.0 goes for broad multimodal inputs, audiovisual output, editing and 30 s clips at a lower listed entry price; FLUX 3 counters with scene consistency, the Self-Flow foundation and the action-prediction roadmap.

Where it holds up, and where it does not

Strengths I would build on:

  • Scene and temporal consistency, and it is the top Megaton category at 94.76.
  • Native audiovisual generation in one pass, no stitching two models together.
  • Prompt adherence at 87.07 independently, so it is not just visual polish.
  • Keyframe control with up to ten image references.
  • Draft-to-final iteration, which addresses a real cost problem in video work: most generations get thrown away.
  • A deliberately wide style range rather than one house look.

Limitations to plan around:

  • The full multimodal roadmap has not shipped. Do not assume FLUX 3 Image or Dev weights from the family announcement.
  • 20 seconds is no longer industry-leading.
  • Physics is not solved. 70.63 against 94.76 scene consistency is a wide gap.
  • Human fidelity at 73.70 still fails on hard anatomy and realistic motion.
  • Continuation pricing is steep relative to ordinary generation.
  • BFL's headline competitive benchmark is internal. Run your own prompts, shot types, languages and reference assets before you commit.

The part that is not video

Robot control signals are temporal predictions conditioned on visual observations, so BFL uses the FLUX 3 video representation as a base for action prediction. Its work with mimic robotics produced FLUX-mimic, where an action decoder reads representations from the video model instead of training a separate perception stack. BFL reports the FLUX-mimic backbone running in under 80 ms on a single RTX 5090, with the full robot system closing a roughly 101 ms reaction loop. Audi has been discussed as an industrial evaluation partner.

None of that makes the public video API a robotics API. It does explain why BFL spent heavily on multimodal video representation rather than treating generation as an isolated media task.

FAQ

Is FLUX 3 available now? Yes. FLUX 3 Video went generally available through BFL on August 4, 2026, and is listed as Available under the model ID flux-3. The still-image release and FLUX 3 Dev open weights remain separate roadmap items.

Does it generate audio? Yes, synchronized natively: dialogue, ambient sound and effects, with multilingual dialogue and lip sync.

How long can clips be? 5–20 seconds for T2V and I2V. Continuation generates 5–15 seconds of new content.

What resolutions? HD (up to 1 MP per frame), FHD (up to 2 MP), QHD/2K (up to 4 MP) and UHD/4K (up to 8 MP). 16:9 FHD is 1920 × 1088. Bands are total pixels per frame, not aspect ratio, and Draft Mode is HD only.

Is it open source? Not currently. FLUX 3 Dev has been announced as an open-weight multimodal variant with no public release date.

Can I use it for still images? Architecturally it is trained across images, video and audio, and BFL has demonstrated image generation and editing research, but the FLUX 3 Image product is still upcoming. For production stills, use an existing FLUX image model.

Is it a world model? BFL positions it as a reality-model foundation because the shared representations extend from audiovisual prediction toward physical action prediction. Calling it a general-purpose world model goes beyond the available evidence. A multimodal generative foundation model designed around world-modelling objectives is the accurate description.

My take

FLUX 3 is less an upgrade to FLUX.2 than a different product line with a longer thesis behind it: train one shared multimodal representation, then reuse it for video, sound, images and eventually actions. Self-Flow is the research basis and the released video model is the first production proof.

What you can ship against today is 20-second clips at 24 fps, HD through UHD/4K, synchronized audio, multilingual dialogue, keyframes, continuation, multi-shot and Draft Mode. It is not the cheapest per second, not the longest, and not top of the independent leaderboard. Where it earns its place is scene consistency, prompt adherence, object persistence and controllable keyframes, plus native sound in the same pass. Pick it for the workload that needs those, and test against your own prompts rather than anyone's ranking.


Originally published at cometapi.com

Top comments (0)