Creators and filmmakers
Generate cinematic clips with native audio, multilingual dialogue, and style diversity — from candid camcorder footage to animation. Chain clips into longer multi-shot sequences with consistent characters.
Black Forest Labs
Black Forest Labs' new multimodal foundation model jointly learns from images, videos, and audio within a unified architecture — because visual intelligence requires understanding how objects hold together, how things move, and how events sound.
FLUX 3 is built on Self-Flow — efficiently aligning multimodal generation and understanding within the same architecture — scaled across video, images, and audio at the same time.
Drag or click to upload up to 10 JPG, PNG, or WEBP reference images
0 / 5000
Sample Image
(The Created Image results will appear here)

FLUX 3 is built on Self-Flow — efficiently aligning multimodal generation and understanding within the same architecture — scaled across video, images, and audio at the same time.
Images, videos, and audio are projections of the same world captured by different sensors. FLUX 3 learns from all of them at once, so sound matches impact, motion obeys mass, and the future follows from the past.

Based on Self-Flow, BFL significantly scaled compute and data to train generation and understanding in the same underlying model — improving both quality and downstream task performance.
FLUX 3 can generate images and video with audio jointly from text prompts, or combine input references such as images and video clips to carry characters, styles, and context into new scenes.

FLUX 3 is a checkpoint on BFL's mission to develop real-world visual intelligence: models that perceive, predict, and act across physical and digital environments.
Black Forest Labs is rolling out FLUX 3 capabilities in phases — each built from the same underlying multimodal flow matching model, with early access for feedback and safety testing.
FLUX 3 is built for creators, marketers, designers, robotics teams, and anyone who wants an intelligent creative partner — not just a single-modality converter.
Generate cinematic clips with native audio, multilingual dialogue, and style diversity — from candid camcorder footage to animation. Chain clips into longer multi-shot sequences with consistent characters.
Produce campaign concepts, polished variants, and typography-heavy designs with agentic reasoning. Ground generations in references and iterate faster across image and video workflows.

Use FLUX 3's world understanding for action prediction. FLUX-mimic finetunes the video backbone for dexterous manipulation with limited task-specific data — bridging content creation and physical AI.
From interactive image and video editing to simulation, computer use, and physical AI — FLUX 3 opens a wide frontier. Video generation is available in early access today.
One unified model powers video, image, and action — each modality benefiting from shared world understanding trained across images, videos, and audio.
Text-to-video, image-to-video animation, video-to-video reference transfer, generative video-audio continuation, and keyframe-to-video control — all with synchronized native audio generation.
Strong at capturing human facial expressions, associating sounds with physical events, and generating multilingual dialogue — combining capabilities into sequences lasting several minutes.
Handle ranges from candid camcorder footage to animation and cinematics. Strong typography generation and animated designs across a broad range of visual styles and aspect ratios.
Chain individual clips into longer, multi-shot sequences. Visual references help ensure characters remain consistent across all scenes throughout extended narratives.
Synthesize and edit images in a wide variety of styles, aspect ratios, and resolutions. Improved complex prompt adherence and high-accuracy text rendering in multiple languages.
Native action prediction integrated into FLUX 3 directly, plus finetuned specialized models like FLUX-mimic — using the pretrained video backbone as a dynamics-aware foundation.
In preliminary 720p text-to-video evaluations, FLUX 3 was preferred over several leading models in head-to-head comparisons — with further improvements expected during early access.
All capabilities are built from the same underlying multimodal flow matching model — released in phases through early access.
Video and audio generation and editing with native audio, multi-shot chaining, and reference-based workflows. Available in early access.

Image synthesis and editing with improved prompt handling, style diversity, and multilingual text rendering. Early access opening in the following weeks.

Open-weight multimodal backbone for content creation and action prediction — for researchers and developers building on unified visual intelligence.

Video-action model combining the FLUX 3 backbone with robotics expertise for dexterous manipulation and production deployment.
Explore FLUX 3 from Black Forest Labs and start creating with our platform.
This page summarizes publicly available information about FLUX 3 from Black Forest Labs. We are not affiliated with BFL.