Guides

FLUX 3: Video With Audio Now, Open Weights Coming in 2026

FLUX 3 launched July 23, 2026: 20-second video with native audio, early access only. Open weights are promised later this year. What is real and what is not.

By Aditya Bawankule11 min read

FLUX 3 is Black Forest Labs' first multimodal model: one network trained jointly on images, video and audio, extended to predicting robot actions. It was announced on July 23, 2026, it generates video up to 20 seconds long with synchronized audio in a single pass, and as of today you cannot buy it, download it, or call it from an API. Video is in application-gated early access. Image is unreleased. Weights do not exist publicly.

That gap between what the model does and what you can actually touch is the whole story of this launch, and it is where most of the coverage goes wrong. Below is what shipped, what the numbers do and do not say, and what is still an open question, sourced to Black Forest Labs' own posts rather than to a reseller's landing page.

What launched on July 23, 2026

Black Forest Labs, the Freiburg and San Francisco lab behind Stable Diffusion and the FLUX image models, announced FLUX 3 as a multimodal frontier model rather than an image model with extras bolted on. Note the naming: it is FLUX 3, with a space and no dot, not FLUX.3. The previous families were FLUX.1 and FLUX.2.

The lab's framing is that image generation, video generation, world models and robotics are all the same underlying capability, which it calls "visual intelligence": models that perceive, predict and act across physical and digital environments. Co-founder and CEO Robin Rombach put the argument for joint training bluntly in the launch press release:

You can't cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds.

The technical version of the same claim, from the launch post: each modality is a lossy projection of one reality, and training on all of them at once means the constraints cross-check each other. The sound has to match the impact, the motion has to obey the mass, the future has to follow from the past.

The FLUX 3 lineup is split by modality, not by speed tier

This is the structural break from FLUX.1 and FLUX.2, and it changes how you should read every announcement that follows. Those families were segmented by how much compute you wanted to spend: pro, dev, flex, schnell, klein. Same job, different price and latency.

FLUX 3 is segmented by what modality you are asking for. There is no pro tier and no schnell equivalent announced. There are four named products, at four different stages of not-quite-existing:

NameWhat it doesStatus as of August 1, 2026
FLUX 3 VideoVideo with native audio, plus editingEarly access, application only
FLUX 3 ImageImage synthesis and editingNot released. "In the following weeks"
FLUX 3 Action and FLUX-mimicAction prediction for roboticsSelected research and commercial partners
FLUX 3 DevOpen-weight multimodal backboneAnnounced only. "Later this year"

All four are described as capabilities of the same underlying multimodal flow matching model, which is why the lineup reads as one model with four doors rather than four models. Video and image are both slated to arrive "through APIs and private weight access", which is the enterprise arrangement, not a public download.

FLUX 3 Video: 20 seconds with audio in one pass

The headline number is 20 seconds of video with audio in a single generation. That matters because most current video models cap out around 8 to 10 seconds, and because the audio is generated jointly rather than dubbed on afterwards: speech is synchronized to lip movement, and sound effects land on the physical events that cause them.

The documented capabilities, all of which come with native audio:

  • Text to video, the standard path.
  • Image to video, either animating from a starting frame or using images as visual references.
  • Video to video from a reference clip, carrying elements such as a character into a new scene.
  • Generative video and audio continuation from an input clip.
  • Keyframe to video for controlled transitions between defined moments.
  • Multilingual dialogue, and multi-language text rendering inside the frame.
  • Agentic chaining of individual clips into multi-shot sequences, which is how you get past 20 seconds. Visual references keep characters consistent across scenes that run several minutes.

Black Forest Labs calls out three areas as unusually strong even mid-development: human facial expressions, associating sounds with physical events, and multilingual work. Style range is a stated design goal too, from candid camcorder footage through animation to cinematics, alongside strong typography and animated design, which is the historical FLUX strength carried into motion.

Aspect ratio support is described as "a broad range" extending beyond conventional cinematic output. If you need an exact list of supported ratios or resolutions, there is no published spec sheet yet, and anyone giving you one is guessing.

The FLUX 3 benchmarks, and what they do not cover

The published evaluations are human-preference win rates on 10-second, 720p, text-to-video clips with audio, from a model the lab itself describes as still in development. Read them with that scope attached.

FLUX 3 preferred overWin rate
Luma Ray 3.293%
Runway Gen-4.577%
Grok Imagine Videoup to 69%
Kling v3 Pro60%
Seedance 2.052%
Gemini Omni Flash52%

Four things worth holding onto before you repeat these numbers.

  1. 52% is a tie, not a win. Against Seedance 2.0 and Gemini Omni Flash, FLUX 3 is at statistical parity. The two clear results are Luma Ray 3.2 and Runway Gen-4.5.
  2. The tests ran at 10 seconds while the marketing runs at 20. Nobody has published a like-for-like comparison at the length FLUX 3 is being sold on, and holding quality across 20 seconds is harder than across 10.
  3. No sample sizes and no methodology are published. Black Forest Labs says so itself: full benchmark results and methodology will come "alongside broader availability". Until then these are vendor-reported preliminaries, and the lab flags them as such.
  4. The comparison set has holes. Neither Sora nor Veo appears in it. And there are zero published image benchmarks, on a model whose image product has not shipped. Any FLUX 3 image quality claim you read today is extrapolation.

For what a currently purchasable video stack looks like across Veo, Seedance and Grok, we keep a running comparison in AI video models in 2026, and the prompt craft that actually transfers between them is in our Veo 3 prompt guide.

Is FLUX 3 open source? The FLUX 3 open weights question

No. Not today, and not in the sense people usually mean.

FLUX 3 shipped on July 23, 2026 with no downloadable weights, no open-source license, and no public inference code. There is no FLUX 3 repository on Black Forest Labs' Hugging Face org, which still tops out at the FLUX.2 klein family. Early access is granted per application, and the launch post describes video and image as arriving through "APIs and private weight access", which is a commercial arrangement with a lab, not an open release.

What is promised is FLUX 3 Dev: an open-weight multimodal backbone covering video, audio, image and action prediction. The press release commits to "faster and open-weight versions of FLUX 3 later this year". As of today there is no release date, no parameter count, and no published license.

That last one is the question worth being careful about, because Black Forest Labs has used two very different licenses for its open releases:

  • Apache 2.0, genuinely open, commercial use included: FLUX.1 [schnell], and the FLUX.2 [klein] distilled models.
  • The FLUX non-commercial license, weights you can download and research with but not sell output from without a separate commercial license: FLUX.1 [dev] and FLUX.2 [dev].

FLUX 3 Dev carries the "dev" name, and every prior model wearing that name was non-commercial. So the informed guess is non-commercial. It is a guess. Treat any page that states FLUX 3 Dev is Apache 2.0 as fabricated, because Black Forest Labs has not published a license.

Worth noting how much the release posture changed. FLUX.2 launched on November 25, 2025 with the 32B FLUX.2 [dev] weights on Hugging Face the same day. FLUX 3 launched with nothing downloadable at all. The lab still describes itself as open-core and still points at the open-weight roadmap, but the frontier model now goes out closed first and open later, which is a different bargain than the one FLUX.1 and FLUX.2 offered.

How to get FLUX 3 access, and why there is no FLUX 3 API yet

There is exactly one route: apply for early access through Black Forest Labs. There is no self-serve endpoint, no published model ID, and no pricing.

We checked the usual places on August 1, 2026:

  • fal has a FLUX 3 page describing text to video and image to video up to 20 seconds with native audio. It is marked coming soon. No endpoint, no rate.
  • Krea has a FLUX 3 model page. Also coming soon, with FLUX.2 as the model you can actually run there today.
  • Replicate lists no FLUX 3 model of any kind. The newest Black Forest Labs entries there are the FLUX.2 family.
  • Hugging Face has no FLUX 3 weights, quantized or otherwise.

Named partners already testing it are Canva, Burda, Magnific (formerly Freepik), Krea and Picsart, plus Audi through the robotics collaboration below. That is the shape of this launch: platform partners first, general availability "to follow" with no date attached.

The practical consequence, and the thing to take away if you take away one thing: any FLUX 3 price you see quoted today is invented. Not estimated, not leaked. Invented. No public rate card exists, from Black Forest Labs or from any reseller, and the credit-cost tables circulating on aggregator sites are filled in from FLUX.2 numbers or from nothing.

We do not serve FLUX 3 in Dream Pixel Forge, and neither does any other self-serve tool, because there is nothing to serve yet. When the API opens we will evaluate it against the video models we already run, on the same clips and the same brief, and add it if it earns a slot. If you want that comparison the day it happens, the fastest way to be in the loop is to have an account: our video studio already runs Veo 3.1, Seedance 2.0 and Grok Imagine Video on one credit balance, and new engines show up in the same model picker.

FLUX-mimic: the same backbone driving robots

The launch had a second half that got much less attention. Black Forest Labs and mimic robotics built FLUX-mimic, a video-action model that puts a lightweight action decoder on top of intermediate features from FLUX 3's video prediction path. The thesis is that a model good enough to generate convincing video has already learned contact, weight and cause and effect, so you can decode actions out of that representation instead of training a separate robotics foundation model.

The reported results: the action decoder outperforms previous vision-language-action models even with the FLUX backbone completely frozen, a setting where those models fail outright, and reaches state-of-the-art success rates when the backbone is fine-tuned alongside the decoder. Fine-tuning a new manipulation task takes as little as 30 minutes of robot data depending on difficulty, against 30 or more hours for prior approaches.

On latency, the FLUX-mimic backbone runs input to world representation in under 80ms on a single NVIDIA RTX 5090, and the full robot system reacts in 101ms. That figure applies to the robot policy variant, which is deliberately small; it says nothing about what the video or image models cost to run. Audi's Production Lab has been testing and deploying the result on soft-body manipulation tasks, the kind of flexible-material work that conventional automation has never handled economically.

Architecture in brief: Self-Flow, and why video is 95% of the bill

FLUX 3 is a scaled-up application of Self-Flow, a Black Forest Labs method published in March 2026 (Chefer and Esser, arXiv 2603.06507). The short version: standard flow matching models lean on frozen external encoders like CLIP or DINOv2 for semantics, which caps how far scaling helps. Self-Flow folds representation learning into the generative objective itself, so the same features serve generation and understanding. The lab reports the two improving together, better generation quality and better decodability for downstream control tasks.

Two numbers from the training mix are worth knowing because they explain the whole design:

  • Video is over 95% of total training compute. Video prediction is the expensive, hard part, and it is where the world model comes from.
  • Audio is under 0.5% of the tokens in a 720p video with audio. It is low-dimensional and cheap once video understanding exists, which is why native synchronized audio comes almost as a byproduct rather than as a bolted-on second model.

Training data is described as tens of millions of hours of general video plus hundreds of thousands of hours focused on human and robot manipulation. When action prediction was added to the curriculum, human ratings on text-to-video and image-to-video fell by up to 10% and recovered fully after 3,500 steps, which is the evidence for the claim that one backbone carries both jobs.

The parameter count is undisclosed. So is the full technical report, which the lab says is coming.

What to run while FLUX 3 is gated

Nothing about FLUX 3 changes what you can ship this week. If the job is a product clip, an ad, or a UGC-style hook, the models that are actually purchasable today handle it, and the workflow that gets the best results out of any of them is the same one FLUX 3 will inherit: get the still right first, cheaply, then animate the frame you approved.

That order matters more than the engine. A generated still costs a fraction of a video clip, and fixing composition, product accuracy and branding in a still is faster than re-rolling 8 seconds of motion. Our FLUX prompt guide covers the still side for the current FLUX.2 models, and best AI image generators covers which model to reach for per job.

Try it now

Free to try, no account needed

Reference images

Select up to 8 images to guide the result.

0/8
Example output for Freeform Image Generator
Example output

Your generated image will replace this example.

The short version

  • FLUX 3 is real, announced July 23, 2026, and genuinely new: one model across image, video, audio and action.
  • FLUX 3 Video is in application-gated early access. FLUX 3 Image has not shipped. FLUX 3 Dev has no date and no license.
  • The 20-second-with-audio capability is the differentiator. The benchmarks behind it were run at 10 seconds, without published methodology, and two of the six results are ties.
  • There are no open weights, no public API, no model IDs and no prices. Anything you read quoting one of those is fabricated.

We will re-test and update this post when the API opens.

Tools for this guide

Frequently asked questions

Does Flux AI do video?

As of FLUX 3, yes. FLUX.1 and FLUX.2 were image-only models. FLUX 3, announced on July 23, 2026, is a multimodal model trained jointly on images, video and audio, and FLUX 3 Video generates clips up to 20 seconds long with synchronized native audio in a single pass. It is not something you can sign up for yet: FLUX 3 Video is in early access and Black Forest Labs grants it by application. If you want Flux for images today, that is still FLUX.2, which is widely available through APIs and open weights.

Is FLUX 3 open source?

No. FLUX 3 shipped with no downloadable weights, no open-source license and no public inference code, and there is no FLUX 3 repository on Black Forest Labs' Hugging Face organization. The lab has promised FLUX 3 Dev, an open-weight multimodal backbone, later in 2026, but has published no date, no parameter count and no license for it. Its previous open releases split between Apache 2.0 for FLUX.1 schnell and FLUX.2 klein, and a non-commercial FLUX license for FLUX.1 dev and FLUX.2 dev, so a non-commercial license is the informed guess for FLUX 3 Dev. That is a guess, not a fact: any page stating FLUX 3 Dev is Apache 2.0 is making it up.

What is the FLUX 3 release date?

FLUX 3 was announced on July 23, 2026, and FLUX 3 Video went into application-gated early access the same day, alongside FLUX 3 Action for selected robotics partners. FLUX 3 Image was described as arriving in early access in the following weeks and had not shipped as of August 1, 2026. FLUX 3 Dev, the open-weight backbone, is promised later in 2026 with no date attached. General availability is stated as to follow, also with no date.

Is there a FLUX 3 API, and what does FLUX 3 cost?

There is no public FLUX 3 API. Access is by early-access application to Black Forest Labs only, with no self-serve endpoint and no published model IDs. fal and Krea both have FLUX 3 pages marked coming soon, and Replicate lists no FLUX 3 model at all. Because no rate card exists from Black Forest Labs or any reseller, every FLUX 3 price circulating today is fabricated rather than estimated. Black Forest Labs has said full benchmark results and methodology will publish alongside broader availability.

FLUX 3 vs FLUX.2: what actually changed?

Two things. First, modality: FLUX.2 generates and edits images, while FLUX 3 is trained jointly across images, video and audio and extends to robot action prediction. Second, how the lineup is split. FLUX.1 and FLUX.2 were segmented by speed and cost, with pro, flex, dev, schnell and klein tiers doing the same job at different prices. FLUX 3 is segmented by modality: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and FLUX 3 Dev, with no announced speed tiers. Availability changed too. FLUX.2 launched on November 25, 2025 with 32B open weights on Hugging Face the same day; FLUX 3 launched with none.

Can I run FLUX 3 locally?

Not today. No FLUX 3 weights have been published in any form, quantized or otherwise, so there is nothing to load into ComfyUI, diffusers or any local runner. That will only change if and when FLUX 3 Dev ships. When it does, expect video-scale hardware requirements rather than image-scale ones: Black Forest Labs reports that over 95% of FLUX 3's training compute went to video. The one published local-hardware figure, a backbone running in under 80ms on a single RTX 5090, is for the FLUX-mimic robot policy variant and does not describe the video or image models.

Can I use FLUX 3 in Dream Pixel Forge?

Not yet, and no self-serve tool can serve it while access is application-gated. We run Veo 3.1, Seedance 2.0 and Grok Imagine Video for video today, on one credit balance with per-second pricing shown before you spend. When the FLUX 3 API opens we will evaluate it on the same briefs against those engines and add it to the model picker if it earns a slot. Creating a free account is the simplest way to get it the day it lands, since new engines appear in the same studio you already use.