<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Leila Rowe</title>
    <description>The latest articles on DEV Community by Leila Rowe (@leilarowe).</description>
    <link>https://dev.to/leilarowe</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4136837%2F0e76c59b-722d-4765-a409-77e58811f701.png</url>
      <title>DEV Community: Leila Rowe</title>
      <link>https://dev.to/leilarowe</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/leilarowe"/>
    <language>en</language>
    <item>
      <title>FLUX 3 Review 2026: Black Forest Labs' Multimodal Video Model, Priced Per Second</title>
      <dc:creator>Leila Rowe</dc:creator>
      <pubDate>Wed, 07 Oct 2026 15:49:49 +0000</pubDate>
      <link>https://dev.to/leilarowe/flux-3-review-2026-black-forest-labs-multimodal-video-model-priced-per-second-p76</link>
      <guid>https://dev.to/leilarowe/flux-3-review-2026-black-forest-labs-multimodal-video-model-priced-per-second-p76</guid>
      <description>&lt;p&gt;FLUX 3 is Black Forest Labs’ first multimodal foundation model — video with native audio, images, and action prediction for robotics — built on a new Self-Flow architecture and sold per second of output. I went through the model page, the research post and the pricing sheet to explain what FLUX 3 does, what it costs at each resolution, and where it fits next to the rest of the FLUX&amp;nbsp;family.&lt;/p&gt;

&lt;p&gt;Black Forest Labs made its name with FLUX.1 and FLUX.2, image models that became the default open-weight choice for a large slice of the industry. FLUX 3, announced on 23 July 2026, is a different kind of release: one model that “jointly learns from images, videos, and audio within a unified architecture”, with video as the first shipping modality, image “soon”, and an action-prediction head aimed at robots. The claim is that each modality makes the others better. I spent a week with the FLUX 3 model page, the “Real World Models” research write-up, the pricing calculator and the early community output to work out how much of that is real&amp;nbsp;today.&lt;/p&gt;

&lt;p&gt;Because FLUX 3 is sold only through BFL’s own per-second API and licensing tiers, I kept a second inference platform that hosts FLUX and other media models on per-image and per-GPU-second billing, &lt;a href="https://synexa.ai/?utm_source=devto&amp;amp;utm_medium=ugc&amp;amp;utm_campaign=leilarowe&amp;amp;utm_content=m11-intro" rel="noopener noreferrer"&gt;Synexa&lt;/a&gt;, open in the next tab for comparison — more at the&amp;nbsp;end.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpd94rtkarfuikt6zrq2m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpd94rtkarfuikt6zrq2m.png" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Black Forest Labs pitches FLUX 3 as “one multimodal model” — video with native audio now, images soon, action prediction for robotics.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What FLUX 3&amp;nbsp;is
&lt;/h2&gt;

&lt;p&gt;FLUX 3 is a diffusion-transformer foundation model trained jointly across images, video and audio. Black Forest Labs’ argument is that no single modality describes the world: images capture spatial structure, video captures time and physics, audio captures causal relationships that vision misses, and language links them. Train on all of them together and you get a model whose generations are “truer to life in every kind of style”. FLUX 3 is also BFL’s first step toward “real-world visual intelligence” — models that perceive, predict and act — hence the robotics&amp;nbsp;head.&lt;/p&gt;

&lt;h2&gt;
  
  
  FLUX 3 Video: what it can&amp;nbsp;do
&lt;/h2&gt;

&lt;p&gt;Video is the modality you can use today, and the capability list is&amp;nbsp;long:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Text to video, image to video, keyframes&lt;/strong&gt; — start frame, end frame, or multiple key frames in&amp;nbsp;order.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Up to 20 seconds in one generation&lt;/strong&gt;, with multiple shots and scenes in one&amp;nbsp;take.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Native audio, optional&lt;/strong&gt; — multilingual speech with accents and strong lip-sync, effects and ambience generated with the&amp;nbsp;frames.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Video editing&lt;/strong&gt; — edit an existing clip with a text prompt; everything the prompt does not mention stays as&amp;nbsp;shot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Video continuation&lt;/strong&gt; — extend a clip from its final frame with momentum and framing carried forward, and agentic chaining to build longer&amp;nbsp;stories.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Typography&lt;/strong&gt; that reads as native to the scene, and a stylistic range BFL describes as “far beyond conventional cinematic output”.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Resolutions from HD (720p) to UHD&amp;nbsp;(4K).&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The early community reaction on the model page is enthusiastic in a specific way: people praise FLUX 3’s dialogue from vague prompts (“ranting about ai”), multishot cooking tutorials, first-person POV shots and 20-second continuous action sequences. Partners quoted at launch include Nous Research (whose Hermes agent chains FLUX 3 shots into long-form pieces), Picsart, Burda Media, Magnific and&amp;nbsp;Envato.&lt;/p&gt;

&lt;h2&gt;
  
  
  Draft mode
&lt;/h2&gt;

&lt;p&gt;FLUX 3 ships with a &lt;strong&gt;Draft&lt;/strong&gt; variant: a fast, cheap preview of your prompt. When a draft is right you send it back and FLUX 3 renders the same video at full quality (“Draft Enhance”). Draft is HD-only, and the enhance step is billed as a regular full-quality generation. For iterative work this is the feature that keeps FLUX 3 affordable.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhnqxdjkt2rp79h7843zk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhnqxdjkt2rp79h7843zk.png" width="800" height="986"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The 23 July 2026 write-up introduces Self-Flow, the alignment approach under FLUX 3, with a Fréchet-distance comparison against flow matching.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-Flow: the architecture claim
&lt;/h2&gt;

&lt;p&gt;FLUX 3 builds on Self-Flow, BFL’s approach to aligning multimodal generation and understanding in one architecture. The research post shows Self-Flow beating standard flow matching on generation error (Fréchet distance) per modality and on manipulation-task success after fine-tuning — the latter being the robotics story. Whether that translates into better video than single-modality competitors is something the community is still working out, but the direction — a model that learns physics from video and causality from audio — is coherent.&lt;/p&gt;

&lt;h2&gt;
  
  
  FLUX 3 pricing: per second, no subscription
&lt;/h2&gt;

&lt;p&gt;BFL sells FLUX 3 with no subscriptions and no seat fees; you pay per second of video output, rounded up to the whole second, with audio included at no extra charge. Rates depend on modality, variant and resolution band:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Text / Image → Video (FLUX 3&amp;nbsp;Video)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;HD: &lt;strong&gt;$0.17/s&lt;/strong&gt; · FHD: &lt;strong&gt;$0.29/s&lt;/strong&gt; · QHD (2K): &lt;strong&gt;$0.40/s&lt;/strong&gt; · UHD (4K):&amp;nbsp;&lt;strong&gt;$0.80/s&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Video → Video (continuing, extending, editing, restyling)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;HD: &lt;strong&gt;$0.41/s&lt;/strong&gt; · FHD: &lt;strong&gt;$0.53/s&lt;/strong&gt; · QHD: &lt;strong&gt;$0.65/s&lt;/strong&gt; · UHD:&amp;nbsp;&lt;strong&gt;$0.95/s&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;FLUX 3 Video Draft&lt;/strong&gt;: &lt;strong&gt;$0.06/s&lt;/strong&gt;, HD&amp;nbsp;only.&lt;/p&gt;

&lt;p&gt;So a 5-second HD clip is $0.85, a 20-second 4K one-take is $16, and extending it in 4K costs $19 for another 20 seconds. Resolution bands are defined by megapixels per frame (HD ≤1.0 MP, FHD ≤2.0 MP, QHD ≤4.0 MP, UHD ≤8.0 MP) regardless of aspect ratio. Enterprise pricing adds volume discounts, SLAs and dedicated support, and separate self-hosting licences (Builder, Platform, Professional, Enterprise, Synthetic Data) cover running FLUX weights on your own infrastructure — though FLUX 3 itself was API-only at&amp;nbsp;launch.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftgt7u52qox5kcxxw4gdy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftgt7u52qox5kcxxw4gdy.png" width="800" height="1160"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Per-second rates by resolution band, a calculator (5 s at HD = $0.85), and enterprise and open-weights licensing below.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What FLUX 3 gets&amp;nbsp;right
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Audio and video from one model&lt;/strong&gt;, with real lip-sync and multilingual speech.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;20-second multi-shot generations&lt;/strong&gt; and continuation for longer&amp;nbsp;pieces.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompt-based editing&lt;/strong&gt; that leaves untouched frames&amp;nbsp;alone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Draft mode&lt;/strong&gt; for cheap iteration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transparent pricing&lt;/strong&gt; — a public per-second rate card with a calculator, no plan to&amp;nbsp;buy.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where FLUX 3 falls&amp;nbsp;short
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;4K is expensive.&lt;/strong&gt; $0.80/s for generation and $0.95/s for video-to-video adds up fast on long&amp;nbsp;clips.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Image modality “soon”.&lt;/strong&gt; The unified model’s still-image side was not available at launch; FLUX.2 remains the image workhorse.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API-only, closed weights&lt;/strong&gt; for FLUX 3, in contrast to the open FLUX.1/FLUX.2 lineage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rounding.&lt;/strong&gt; Partial seconds round up, so short clips carry a hidden&amp;nbsp;premium.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Robotics claims are early.&lt;/strong&gt; Action prediction is a research direction with a “learn more” link, not a product you can buy&amp;nbsp;today.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  FLUX 3 vs the alternatives
&lt;/h2&gt;

&lt;p&gt;Against Seedance 2.5, FLUX 3 trades longer single-pass clips (30 s) for a broader stylistic range and a public per-second price. Against Sora-class models, community consensus at launch was that FLUX 3 is “somewhere between Seedance 2 and Sora” with a distinctive look. Against the rest of the FLUX family, FLUX 3 is the video and audio model; FLUX.2 is still where images live — and for images and open-weight FLUX, a per-image host is cheaper than BFL’s own&amp;nbsp;API.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict: is FLUX 3 worth it in&amp;nbsp;2026?
&lt;/h2&gt;

&lt;p&gt;For anyone producing short video with dialogue, product and brand motion, or agent-driven storytelling, FLUX 3 is worth a serious look: the audio-in-the-model design and 20-second multishot output are real advantages, and Draft mode keeps exploration cheap. Budget carefully above HD. Read the model page and run the official per-second calculator before committing to&amp;nbsp;4K.&lt;/p&gt;

&lt;h2&gt;
  
  
  The alternative worth keeping next to it:&amp;nbsp;Synexa
&lt;/h2&gt;

&lt;p&gt;Two FLUX 3 realities pushed me to keep a second platform open: the still-image side is not shipped, so images still come from FLUX.1/FLUX.2, and BFL’s per-second API is the only way to run FLUX 3. Synexa hosts the open FLUX models and other media models on per-output and per-GPU-second billing.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;FLUX images at a fraction of the price.&lt;/strong&gt; FLUX.1 [dev] is $0.0125 per image, FLUX.1 [schnell] $0.0015, FLUX.1 [pro] $0.02 — Synexa’s own comparison table shows 50–60% below the providers it&amp;nbsp;lists.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Video and 3D on the same account&lt;/strong&gt; — Wan 2.1 video at $0.20 per clip, Hunyuan 3D at $0.025 per model, Stable Diffusion XL at $0.002 per&amp;nbsp;image.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Raw GPU by the second, scaling to zero.&lt;/strong&gt; H100 at $2.99/hour ($0.00083/s), A100 80GB at $2.49/hour, RTX 4090 at $0.69/hour — run your own FLUX fine-tunes or LoRAs without a self-hosting licence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Billing on model output&lt;/strong&gt;, so nothing is charged for a failed&amp;nbsp;run.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Generate the video on FLUX 3. Generate the stills, the 3D assets and anything on open weights through&amp;nbsp;&lt;a href="https://synexa.ai/?utm_source=devto&amp;amp;utm_medium=ugc&amp;amp;utm_campaign=leilarowe&amp;amp;utm_content=m11-outro" rel="noopener noreferrer"&gt;Synexa&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F1024%2F0%2AfHRBzvfF922ekPzi" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F1024%2F0%2AfHRBzvfF922ekPzi" width="1024" height="1820"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Photo by Daniel Olah on&amp;nbsp;Unsplash&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>api</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
