SOON
FLUX 3 · LAUNCHING SOON

Generate Cinematic Video with FLUX 3

Explore text-to-video, image-to-video, and reference-guided creation with synchronized audio, precise motion control, and generations up to 20 seconds.

FLUX 3 Video is currently in Early Access. This independent platform uses a third-party Preview Engine while direct FLUX 3 generation access is being prepared.

Flux 3 Video

Text to Video · Image to Video · Native Audio · Precise Control

What is FLUX 3?

FLUX 3 is Black Forest Labs' multimodal foundation model for image, video, audio, and action prediction.
In Early Access, FLUX 3 can generate video with native audio from text, images, and video references, with clips up to 20 seconds.

Official FLUX 3 Text-to-Video Showcase

These official FLUX 3 showcase clips are presented through media links published by Black Forest Labs.
Black Forest Labs has not published the original prompts for these clips; use the source link below each video to view the official showcase.

FLUX 3 Image to Video

FLUX 3 supports animation from a starting frame and image references for subject, character, and style guidance.
The first item uses an official FLUX 3 showcase preview published by Black Forest Labs; its original input image and prompt have not been published. The remaining items illustrate image-to-video workflows and are not presented as FLUX 3 outputs.
Original image
Golden hour, soft lighting, warm colors, saturated colors, wide shot, left-heavy composition. A weathered gondolier stands in a flat-bottomed boat, propelling it forward with a long wooden pole through the flooded ruins of Venice. The decaying buildings on either side are cloaked in creeping vines and marked by rusted metalwork, their once-proud facades now crumbling into the water. The camera moves slowly forward and tilts left, revealing behind him the majestic remnants of the city bathed in the amber glow of the setting sun. Silhouettes of collapsed archways and broken domes rise against the golden skyline, while the still water reflects the warm hues of the sky and surrounding structures.
Prompt

Golden hour, soft lighting, warm colors, saturated colors, wide shot, left-heavy composition. A weathered gondolier stands in a flat-bottomed boat, propelling it forward with a long wooden pole through the flooded ruins of Venice. The decaying buildings on either side are cloaked in creeping vines and marked by rusted metalwork, their once-proud facades now crumbling into the water. The camera moves slowly forward and tilts left, revealing behind him the majestic remnants of the city bathed in the amber glow of the setting sun. Silhouettes of collapsed archways and broken domes rise against the golden skyline, while the still water reflects the warm hues of the sky and surrounding structures.

Video
Original image
Make the changes happen instantly.
Prompt

Make the changes happen instantly.

Video

FLUX 3 Early Access Capabilities

FLUX 3 jointly learns from images, video, and audio in one multimodal architecture, supporting native synchronized sound, visual references, video continuation, keyframe transitions, and multilingual dialogue.

Up to 20-Second Clips

Create diverse video clips with audio up to 20 seconds in a single generation during Early Access.

Multimodal References

Guide generation with text, images, and reference video for animation, transformation, and continuation workflows.

Native Video and Audio

Generate visuals and synchronized sound together, including dialogue, ambience, and effects.

Video Continuation

Continue from input video and audio or carry central elements from a reference clip into a new context.

Keyframe Control

Define important visual moments and generate controlled transitions between them.

Multilingual Dialogue

Create scenes with spoken dialogue across multiple languages and a broad range of visual styles.

What Can You Create with FLUX 3 Video?

FLUX 3 Video brings prompts, images, reference footage, motion, and sound into one multimodal creative workflow. That foundation makes it relevant to more than isolated AI clips. Creators can plan complete audiovisual scenes for storytelling, marketing, design, and pre-production while directing how a subject looks, moves, speaks, and responds over time.

Cinematic Concepts and Previsualization

Use FLUX 3 Video to explore a scene before committing to a full production. Describe the location, subject, action, framing, lens behavior, camera movement, lighting, pacing, dialogue, and ambience in one brief. Directors and creative teams can test visual ideas, compare approaches, and communicate a clearer direction for storyboards, pitch films, and cinematic previsualization.

Product and Campaign Videos

Start with product images or other visual references to establish shape, materials, color, styling, and brand context. FLUX 3 Video can then place the central subject into a directed scene with motion and synchronized sound. This workflow can support product reveals, launch concepts, campaign treatments, ecommerce visuals, and early creative exploration for advertising teams.

Character-Driven Visual Stories

Reference-guided generation can help carry a character or central visual element from one scene into another. Combine that guidance with prompts for expression, action, environment, camera, and tone. FLUX 3 Video is designed for workflows where identity and story context matter across connected shots, including short narratives, branded characters, music concepts, and episodic ideas.

Multilingual Dialogue and Social Content

Build scenes in which spoken dialogue, facial expression, ambience, and physical sound are considered together. The announced multilingual capabilities of FLUX 3 Video can support creative concepts for different audiences and markets. Vertical scenes, short campaign ideas, dialogue-led posts, and localized variations can all begin from the same prompt-first audiovisual workflow.

Animation, Typography, and Visual Design

FLUX 3 Video is not limited to conventional cinematic realism. Its announced capabilities include broad style diversity, animated designs, and strong typography generation. Creators can explore stylized animation, graphic motion, title-driven sequences, candid footage, or polished cinematic treatments while using references and prompts to keep the visual language connected to the original brief.

Video Continuation and Multi-Shot Sequences

Continue from input video and audio, transform reference footage into a new context, or connect individual clips into a longer sequence. Visual references can help preserve central elements while each prompt changes the location, action, camera, or mood. This gives FLUX 3 Video a practical role in developing multi-shot stories instead of treating every generation as an unrelated clip.

How to explore Flux 3 Video

Build a Strong Video Brief in Four Steps

Start with a clear scene, add the right references, direct motion and sound, then iterate on the result.

1

Describe the Scene

Write a prompt that explains the subject, setting, camera direction, motion, mood, and how the scene should begin and end.

2

Add Visual References

Upload images, video, or audio so the model can follow your product, character, scene, motion style, and brand direction.

3

Direct Motion and Sound

Specify camera movement, action, pacing, dialogue, ambience, and effects so the scene has a clear audiovisual direction.

4

Generate and Refine

Generate the scene, review motion and sound, then refine one creative variable at a time.

FAQ

Flux 3 Video FAQ

Answers about FLUX 3 Early Access, multimodal references, and announced capabilities.

1

What is FLUX 3?

FLUX 3 is Black Forest Labs' multimodal foundation model for image, video, audio, and action prediction. It is currently available through an Early Access program.

2

What video capabilities has Black Forest Labs announced?

The official Early Access announcement includes text-to-video, image-to-video, video-to-video, video and audio continuation, keyframe transitions, native audio, and multilingual dialogue.

3

How long can FLUX 3 videos be?

Black Forest Labs says FLUX 3 can create videos with audio up to 20 seconds in a single generation during Early Access.

4

Does FLUX 3 Video generate native audio?

Yes. Black Forest Labs describes native audio as part of FLUX 3 Video generation rather than a separate soundtrack added afterward. Announced capabilities include synchronized dialogue, ambience, physical sound effects, and multilingual speech, allowing the visual and audio direction to be planned together in one scene.

5

Can FLUX 3 Video create longer multi-shot stories?

A single FLUX 3 Video generation can run for up to 20 seconds during Early Access. Black Forest Labs also describes agentic chaining of individual clips into longer multi-shot sequences. Visual references can help maintain a character or central element while prompts direct new locations, actions, camera choices, and story beats.

6

How do I get better preview results?

Describe the subject, action, setting, camera, lighting, pacing, and sound. Add a relevant image when visual consistency matters.

Explore Flux 3 Video Workflows

Explore FLUX 3 text-to-video, image-to-video, reference-guided video, and native-audio workflows.