Black Forest Labs expands beyond images with a multimodal video model
Black Forest Labs has introduced Flux 3, a multimodal foundation model that the company says learns from images, video and audio together and can generate video clips with native sound for the first time. The release marks a notable step for the German AI company, which is positioning the new system as part of a broader push toward models that can better represent how events unfold in the real world.
According to the company’s description, Flux 3 can produce videos up to 20 seconds long with synchronized audio and supports a range of generation modes, including text-to-video, image-to-video and video-to-video. It also includes tools for keyframe-based transitions, multilingual dialogue and linking individual clips into longer multi-shot sequences.
That feature set places Flux 3 in a crowded and fast-moving part of the AI market, where model makers are racing to improve visual coherence, physical realism and scene continuity. The addition of native audio is especially significant because many video systems still rely on separate workflows for soundtrack generation, sound design or speech synthesis. A model that can generate both moving images and sound in one pass could simplify production pipelines if the quality holds up in practice.
Why Black Forest Labs is emphasizing multimodal training
The company frames Flux 3 as more than a video generator. Its argument is that images, video and audio each capture different parts of reality, and that combining them during training can help a model build a richer representation of events. Images provide spatial detail, video adds temporal change, and audio can capture cause-and-effect cues that are invisible in a single frame.
That logic aligns with a wider industry movement toward so-called world models, systems meant to reason across multiple forms of sensory information rather than treat each medium in isolation. In Black Forest Labs’ formulation, multimodal training is a step toward what it calls “real-world visual intelligence,” meaning models that can perceive, predict and act across physical and digital environments.
In practical terms, the company says Flux 3 is particularly strong at human facial expressions and at matching sounds to visible physical events. Those are both difficult problems in generative video. Faces tend to reveal artifacts quickly, and audio-video synchronization breaks immersion when timing is even slightly off. If those claims prove out beyond internal testing, they would address two of the most visible weaknesses in current AI video tools.

The benchmark claims come with an important caveat
Black Forest Labs also shared early comparison results for 10-second, 720p clips. In those company-reported evaluations, Flux 3 was preferred over Luma Ray 3.2 in 93% of comparisons, over Runway Gen-4.5 in 77%, and over Grok Imagine Video in 69%. The margins were much narrower against stronger rivals: 60% over Kling v3 Pro, 59% over Happy Horse v1, 57% over Happy Horse 1.1, and 52% over both Seedance 2.0 and Gemini Omni Flash.
Those numbers are enough to make the release noteworthy, but not enough to settle the competitive picture. The source material explicitly notes that the results are preliminary and that no independent tests were available at the time of publication. That matters because internal preference studies can be useful indicators, yet they do not provide the same level of confidence as third-party benchmarking or broader user validation in production settings.
So the launch should be read as a serious product move, but not as definitive proof that Flux 3 has overtaken the field. The stronger claim supported by the available information is narrower: Black Forest Labs has entered the top-tier video-model conversation with a system that appears competitive on company-reported tests and differentiated by integrated audio generation.
More than media generation: a robotics angle
Alongside Flux 3, Black Forest Labs introduced Flux-mimic, described as a video action model aimed at robotics applications. The company says the system is already being tested at Audi. That detail broadens the strategic significance of the launch.
Much of the recent public focus in generative AI has centered on creative tooling, entertainment and advertising workflows. By pairing a media-generation model with a robotics-oriented action model, Black Forest Labs is signaling that it sees multimodal systems as useful not only for content creation, but also for embodied or industrial applications where visual understanding and action prediction matter.

The connection between the two products is conceptual as much as commercial. A model trained to relate motion, appearance and sound may be better suited to understanding physical processes than a model trained on still images alone. Whether that translates into robust robotics performance remains to be seen, but the company’s framing suggests it wants to compete in a category where perception and control begin to overlap.
What the launch changes now
For creators and developers, the clearest immediate impact is option expansion. Native-audio video generation up to 20 seconds, multilingual dialogue support and clip-chaining all point toward workflows that are more usable for short-form storytelling, prototyping and possibly ad production. The more steps a single model can absorb, the less manual stitching is needed between video generation, voice, sound effects and editing tools.
For the market, the release adds pressure on rivals to improve multimodal integration rather than treat audio as an afterthought. It also reinforces the idea that the next competitive frontier in video AI is not just prettier frames, but systems that can maintain consistency across motion, timing, speech and physical events.
At the same time, caution is warranted. The supplied evidence does not yet establish how Flux 3 performs under independent evaluation, how well it generalizes across prompts, or whether its audio quality remains stable across longer or more complex scenes. Those questions will determine whether the model becomes a durable industry contender or simply another short-lived benchmark headline.
The bigger picture
Even with those caveats, Flux 3 stands out because it reflects where the market is heading. Video generation is no longer just about silent clips assembled from text. The more ambitious goal is multimodal synthesis that behaves coherently enough to support narrative, interaction and eventually physical-world reasoning.
Black Forest Labs is arguing that the path forward runs through joint training on images, video and audio. With Flux 3, it has attached that thesis to a product with features that matter today and research ambitions that point beyond media alone. The result is not yet a verified leap over the field, but it is a meaningful launch in one of AI’s most competitive arenas.
This article is based on reporting by The Decoder. Read the original article.
Originally published on the-decoder.com








