Black Forest Labs has released FLUX 3, a multimodal foundation model that the company says jointly learns from images, video, and audio to build a single representation of the physical world. Announced on July 23, 2026, and now available in early access, FLUX 3 represents a significant architectural departure from the model-by-modality approach that has dominated generative AI, collapsing image creation, video generation with native sound, and even robot control into one unified system.

The launch arrives at a moment of intensifying competition in the generative media space, where players including Runway, OpenAI's Sora, and Google's Veo have been racing to produce higher-quality and more capable video models. FLUX 3 enters this contest with a fundamentally different premise: that separate models for each media type are an artificial limitation. Its progress is being monitored as part of our ongoing AI industry coverage of the generative media landscape.

A Unified Model of Reality

In a blog post accompanying the release, Black Forest Labs laid out the theoretical motivation for training a single model across multiple modalities. The company argued that no single modality — whether image, video, or audio — provides a complete description of reality. Each is a projection of the same underlying physical world, captured by different sensors, each losing information in the process.

Images capture spatial structures at a specific point in time. Videos restore the dimension of time and reveal temporal dynamics and physical laws. Audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect. Language links these perceptions to goals, abstractions, and instructions, the company wrote.

By learning from all of these simultaneously, the model's mutual constraints provide more information than any single modality alone. The sound has to match the impact, the motion has to obey the mass, the future has to follow from the past. The modalities stop being separate and start being evidence about one underlying reality.

FLUX 3 is the company's first model built entirely on this principle, which Black Forest Labs calls Self-Flow — an approach for efficiently aligning multimodal generation and understanding within the same underlying architecture.

Video Generation With Native Audio

The most immediately striking capability of FLUX 3 is its ability to generate videos up to 20 seconds in length with synchronized native audio in a single generation pass. This distinguishes it from most existing video models, which produce silent footage that must be paired with separately generated audio.

According to the company, FLUX 3's video output is particularly strong in capturing human facial expressions, associating sounds with physical events, and supporting multilingual capabilities. These capabilities can be combined to create sequences lasting several minutes, where visual references help ensure that characters remain consistent across all scenes.

In preliminary human evaluation conducted during training, FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93% of comparisons. Against newer competitors, it was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, and Seedance 2.0 and Gemini Omni Flash in 52%.

The company cautioned that these results are preliminary and that further improvements are expected during the early access phase.

Images and Text Rendering

Beyond video, FLUX 3 can synthesize and edit images across a wide variety of styles, aspect ratios, and resolutions. Black Forest Labs reported significant improvements over earlier FLUX versions in handling complex prompts and generating accurate text within images — a persistent challenge for generative image models.

An early access phase for FLUX 3 Image capabilities will open in the weeks following the video release, the company said.

A Bridge to Physical AI

Perhaps the most unconventional aspect of FLUX 3 is its extension into robotics. The model's world understanding extends to action prediction, the company explained, taking two routes to physical AI: integrating native action prediction directly into FLUX 3, and using the pretrained video backbone as a foundation for specialized action models fine-tuned with limited task-specific data.

Black Forest Labs partnered with mimic robotics, which was among the first companies to gain early access to FLUX 3. Together they developed FLUX-mimic, a video-action model combining the FLUX 3 backbone with mimic's expertise in robot learning for dexterous manipulation. According to the company, the system is already being tested on real production tasks at Audi.

This represents a notable convergence: the same underlying technology used to generate entertainment content is being applied to control physical robots in manufacturing environments. The company's thesis is that physical AI and content creation run on the same foundation, because both require a model that understands how objects hold together, how things move, and how physical events produce sound.

Competition and Market Position

The launch positions Black Forest Labs — the Berlin-based startup behind the widely used FLUX image generation models — as a more ambitious competitor in the generative AI space. While companies like Runway have focused narrowly on video and OpenAI on text and multimodal chatbots, Black Forest Labs is betting that a truly unified model across visual, auditory, and physical domains will prove more capable than specialized alternatives.

The company said it is already working on the next generation of models, with a goal of unifying perceptual, action, and language prediction in the same model. Over the coming weeks and months, it will roll out additional capabilities built from the same underlying multimodal flow matching architecture, each following an early access phase for feedback and safety testing.

FLUX 3 is available now in early access for video generation, with image and additional capabilities rolling out incrementally.

Stay Ahead of AI

FLUX 3's unified approach to images, video, audio, and robotics marks a notable shift in how generative AI models are being designed. For more on the generative media race and the latest developments across artificial intelligence, follow our breaking AI news coverage.

Read more AI news →