Black Forest Labs' FLUX 3 raises the question enterprises can't price yet: can one model generate 20-second AI video with audio and also serve as a robotics backbone? The Freiburg-based AI lab has launched FLUX 3, its first public video generation model, expanding the FLUX family beyond images into audio-video generation and action prediction, according to VentureBeat.
The headline claim is aggressive. FLUX 3 can generate images or combined video and audio clips up to 20 seconds from a single prompt, while using the same underlying architecture as the basis for robotic vision and actions.
BFL is not pitching this as three models hidden behind one product page. The company says FLUX 3 is jointly trained across images, video and audio, with the architecture extended toward action prediction. Its term for the broader bet is visual intelligence: models "that can perceive, predict, and act across physical and digital environments."
Can FLUX 3 make one architecture do the work of several?
BFL's central argument is that creative generation, simulation, computer use and robotics should not be treated as separate AI markets. The company wants buyers to see them as different outputs from a shared model family.
That claim builds on Self-Flow, BFL's technique for aligning multimodal understanding and generation inside one architecture. In its technical blog, BFL says FLUX 3 learned from video, images and audio at the same time, rather than stitching together isolated systems.
"You can't cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds."
That line from Robin Rombach, BFL co-founder and CEO, is the cleanest version of the pitch. If a model learns motion and sound together, BFL argues, it should produce more plausible video and give robotics teams a better starting point for predicting what happens next.
The company says FLUX 3 targets creative tooling, media, design, e-commerce and physical AI. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea and Picsart, according to VentureBeat.
How limited is FLUX 3 early access right now?
The launch comes through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and the upcoming FLUX 3 Dev.
For now, the usable rollout is narrow. FLUX 3 Video, with optional native audio generation, and FLUX 3 Action are entering a gated Early Access program. Anyone can apply, but BFL must approve access. There is no public access yet through BFL's API or partner APIs.
FLUX 3 Image is expected in the coming weeks, followed by general availability. That leaves enterprise buyers with demos, claims and early partner testing, but not enough commercial detail to model deployment.
The missing pieces are material:
- Pricing: BFL has not announced prices.
- Service levels: No production SLA has been published.
- Benchmarks: Full evaluation methodology, sample sizes and rater counts are not available.
- Image metrics: BFL has not published image-model benchmarks.
- Weights: FLUX 3 is not launching with downloadable weights or an open source license.
The access question also sits inside a wider fight over who controls advanced AI systems and data access. XOOMAR has tracked that tension in Big Tech Blocks Digital Services Act Data Access in EU Test and the public-facing backlash covered in Avoiding AI Workshops Turn Libraries Into Big Tech Revolt. FLUX 3's immediate issue is narrower: enterprises can't evaluate cost, latency or deployment risk until BFL opens more of the stack.
BFL has published early preference results, but they come with a major caveat. The company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google's Gemini Omni Flash in 52%.
Those tests used 10-second, 720p text-to-video clips with audio. BFL labeled the chart a "preliminary evaluation of an early FLUX 3 candidate," meaning the numbers do not directly measure the model now entering early access.
Can 20-second FLUX 3 video matter without pricing or resolution?
The most concrete part of the launch is FLUX 3 Video. BFL says it can generate clips up to 20 seconds with native audio in one generation.
That duration is one of the launch's strongest claims. But the resolution ceiling is not stated. BFL's published evaluations ran at 720p, which makes the 20-second figure harder to compare against rivals that disclose resolution and price.
| Model | Max single-generation duration | Max resolution | Key constraint | 10-second 720p price |
|---|---|---|---|---|
| FLUX 3 Video | 20 seconds | Not stated, evaluations at 720p | Early access, no public pricing or SLA | Not announced |
| HappyHorse 1.1 | 15 seconds | 1080p | No 4K, closed weights | Not published |
| Gemini Omni Flash | 10 seconds, 3s minimum | 720p at 24 FPS | Preview, uploaded-video editing unavailable in EEA, Switzerland and UK | $1.00 |
| Veo 3.1 Fast | Per-second billing | 4K | Preview | $1.00 |
For creative teams, the bigger issue is not whether one prompt can produce one impressive clip. It is whether characters, products, lighting and motion hold together across multiple shots.
BFL says FLUX 3 supports text-to-video, image-to-video, video-to-video, video-audio continuation, keyframe-to-video, multilingual dialogue, typography generation and agentic chaining of clips into longer multi-shot sequences. It also says visual references can help keep characters consistent across scenes.
That is where the product will be judged. HappyHorse 1.1 is pushing reference-based identity control. Gemini Omni Flash is already generally available through Google's Gemini API at $0.10 per second of generated 720p video, or $1.00 for a 10-second clip. BFL has a longer single-generation claim, but Google has public API access and pricing.
A regional wrinkle may help BFL in Europe. VentureBeat reports that editing uploaded video is unavailable to Omni Flash users in the European Economic Area, Switzerland and the United Kingdom, though editing video generated by Omni itself is allowed.
Will FLUX 3 Dev and FLUX-mimic prove the robotics claim?
The biggest delay is FLUX 3 Dev. BFL is not releasing downloadable weights at launch, even though open-weight FLUX releases helped drive developer adoption.
That matters because FLUX 3 Dev is described as more than another image model. BFL calls it "open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction." The company has not yet shared the license, parameter count, quantizations or hardware requirements.
The robotics test case is FLUX-mimic, developed with Swiss firm Mimic Robotics. It combines the FLUX 3 video backbone with Mimic's work in robot learning and dexterous manipulation.
BFL and Mimic Robotics say the model can be fine-tuned for some manipulation tasks with as little as 30 minutes of robot data, compared with prior approaches that required 30 or more hours, depending on task difficulty.
That is the claim to watch. Public API access, pricing, full benchmarks, FLUX 3 Dev licensing and real production examples will decide whether FLUX 3 becomes a unified enterprise platform or remains an impressive gated demo with unanswered economics.
The Bottom Line
- FLUX 3 pushes generative AI beyond still images into combined video, audio, and potential robotics use cases.
- Enterprises may need to rethink whether multimodal models can replace separate tools for media generation and simulation.
- The launch highlights growing competition to build AI systems that can perceive, predict, and act across digital and physical environments.
Originally published on XOOMAR. For more news and analysis, visit XOOMAR.
Top comments (0)