Black Forest Labs

FLUX 3 — Real World Models

Black Forest Labs' new multimodal foundation model jointly learns from images, videos, and audio within a unified architecture — because visual intelligence requires understanding how objects hold together, how things move, and how events sound.

FLUX 3 is built on Self-Flow — efficiently aligning multimodal generation and understanding within the same architecture — scaled across video, images, and audio at the same time.

Flux 3 AI

0/10

Drag or click to upload up to 10 JPG, PNG, or WEBP reference images

0 / 5000

Credits required:6030

Sample Image

(The Created Image results will appear here)

Sample result

FLUX 3 is built on Self-Flow — efficiently aligning multimodal generation and understanding within the same architecture — scaled across video, images, and audio at the same time.

Why FLUX 3 changes multimodal generation

One model, one underlying reality

Images, videos, and audio are projections of the same world captured by different sensors. FLUX 3 learns from all of them at once, so sound matches impact, motion obeys mass, and the future follows from the past.

Self-Flow architecture

Self-Flow architecture

Based on Self-Flow, BFL significantly scaled compute and data to train generation and understanding in the same underlying model — improving both quality and downstream task performance.

Mix modalities freely

FLUX 3 can generate images and video with audio jointly from text prompts, or combine input references such as images and video clips to carry characters, styles, and context into new scenes.

Toward visual intelligence

Toward visual intelligence

FLUX 3 is a checkpoint on BFL's mission to develop real-world visual intelligence: models that perceive, predict, and act across physical and digital environments.

Black Forest Labs is rolling out FLUX 3 capabilities in phases — each built from the same underlying multimodal flow matching model, with early access for feedback and safety testing.

FLUX 3 launch plan

FLUX 3 is built for creators, marketers, designers, robotics teams, and anyone who wants an intelligent creative partner — not just a single-modality converter.

Who is FLUX 3 for

Creators and filmmakers

Generate cinematic clips with native audio, multilingual dialogue, and style diversity — from candid camcorder footage to animation. Chain clips into longer multi-shot sequences with consistent characters.

Marketers and visual teams

Produce campaign concepts, polished variants, and typography-heavy designs with agentic reasoning. Ground generations in references and iterate faster across image and video workflows.

Physical AI and robotics

Physical AI and robotics

Use FLUX 3's world understanding for action prediction. FLUX-mimic finetunes the video backbone for dexterous manipulation with limited task-specific data — bridging content creation and physical AI.

The next backbone of visual intelligence

From interactive image and video editing to simulation, computer use, and physical AI — FLUX 3 opens a wide frontier. Video generation is available in early access today.

One unified model powers video, image, and action — each modality benefiting from shared world understanding trained across images, videos, and audio.

FLUX 3 capabilities

Video with native audio (up to 20s)

Text-to-video, image-to-video animation, video-to-video reference transfer, generative video-audio continuation, and keyframe-to-video control — all with synchronized native audio generation.

Multilingual dialogue and sound design

Strong at capturing human facial expressions, associating sounds with physical events, and generating multilingual dialogue — combining capabilities into sequences lasting several minutes.

Style diversity and typography

Handle ranges from candid camcorder footage to animation and cinematics. Strong typography generation and animated designs across a broad range of visual styles and aspect ratios.

Agentic multi-shot chaining

Chain individual clips into longer, multi-shot sequences. Visual references help ensure characters remain consistent across all scenes throughout extended narratives.

Image synthesis and editing

Synthesize and edit images in a wide variety of styles, aspect ratios, and resolutions. Improved complex prompt adherence and high-accuracy text rendering in multiple languages.

Action prediction

Native action prediction integrated into FLUX 3 directly, plus finetuned specialized models like FLUX-mimic — using the pretrained video backbone as a dynamics-aware foundation.

Early benchmark performance

In preliminary 720p text-to-video evaluations, FLUX 3 was preferred over several leading models in head-to-head comparisons — with further improvements expected during early access.

Learn more about multimodal AI

Explore FLUX 3 from Black Forest Labs and start creating with our platform.

Frequently asked questions about FLUX 3

This page summarizes publicly available information about FLUX 3 from Black Forest Labs. We are not affiliated with BFL.