freym
freym / blog / FLUX 3 and FLUX 3 Video: Black Forest Labs goes multimodal

FLUX 3 and FLUX 3 Video: Black Forest Labs goes multimodal

Published · by freym · sourced from official announcements

FLUX 3 is Black Forest Labs' first unified multimodal model, announced July 23, 2026, generating image, video, audio and action-prediction from one jointly trained architecture. FLUX 3 Video, fully released August 4, produces video with native audio in a single pass — clips up to 20 seconds, with dialogue in multiple languages.

What did Black Forest Labs actually release?

Black Forest Labs — the lab behind the FLUX image model family — released two things in quick succession, both tracked on our FLUX 3 and FLUX 3 Video pages:

Why is single-pass native audio a big deal?

Until mid-2026, most AI video pipelines generated silent clips and bolted sound on afterwards with a separate model — a workflow that makes lip sync, timing and sound design fragile. FLUX 3 Video generates picture and sound together in one pass. Per Leonardo's August 7 integration announcement, that extends to clips up to 20 seconds and keeps characters consistent across chained multi-shot sequences. Native audio arrived across the market almost simultaneously: ByteDance's Seedance 2.5 and MiniMax H3 shipped the same capability within the same two weeks, making single-pass audio the new baseline for serious video models.

Where is FLUX 3 available?

As of August 8, 2026: FLUX 3 image generation is available through Pika's API Club (announced August 5), and FLUX 3 Video is integrated in Leonardo (announced August 7) alongside BFL's own access. The model page tracks new availability as it is announced.

What should you watch next?

Two things. First, whether BFL extends the unified architecture beyond the four launch modalities — the July 23 announcement explicitly framed it as extensible. Second, how FLUX 3 Video's Draft mode pricing plays against Seedance 2.5's aggressive platform deals (Higgsfield launched Seedance with a free unlimited week). The single-pass-audio generation is now table stakes; the fight is moving to cost per usable second.