FLUX 3The First Video Model from Black Forest Labs

FLUX 3 is the multimodal model from Black Forest Labs: text to video and image to video up to 20 seconds, with sound generated inside the same pass. Coming soon to fal.

What Makes FLUX 3 Different

Real World Understanding

Give It the Idea and It Knows the Rest

The whole prompt for this clip was “Visualization of ‘Annabel Lee’ by Edgar Allan Poe.” No shot list, no wardrobe notes, no lighting direction. FLUX 3 worked out the setting, the period, and the order of events on its own, because it already had the context. That is what real world understanding buys you: write the idea at the level you actually think about it, and let the model fill in the hundred details you would otherwise be spelling out.

Native Audio

Sound Generated With the Picture

FLUX 3 renders audio inside the same pass as the video, so an event arrives carrying the sound it makes. Every hammer blow here lands with its own ring off the anvil and the hiss of sparks leaving the hot blade, on the exact frame the metal is struck. Nothing needs lining up afterward and no second model gets called, because the soundtrack was never a separate render.

Keyframes to Video

Lock the Shot to Your Keyframes

Keyframes and storyboards are how a real production plans a shot, and a keyframes to video endpoint is coming to FLUX 3. Here four frames at 0, 5, 10, and 15 seconds set the whole arc, from a rock headland above a harbour to a colossus standing above the clouds, and the model returns one continuous take that hits every one of them. Pinning the shot this way is what keeps a long generation from drifting: you hand the model your creative guidelines instead of hoping one paragraph of text lands the same way twice.

Physical Cause and Effect

Chain Reactions That Follow Through

FLUX 3 is trained jointly across image, video, audio, and action, and it shows in shots where one event has to trigger the next. A red marble rolls in from the left, strikes only the first of six dominoes, and the run completes cleanly to the right while the marble comes to rest beside the piece it started. Motion carries mass and momentum from one object to the next instead of resetting between beats.

Camera and Materials

Real Camera Moves, Real Surfaces

A slow lateral dolly travels behind three dark tree trunks to reveal a lighthouse on a distant cliff. Each trunk crosses the foreground at its own speed, the lighthouse stays anchored where it belongs, and the storm keeps working through the whole move as waves break against the rock below and rain beads on the lens. Orbits, focus racks, and tracking shots hold their geometry the same way.

20 Second Clips

Long Shots and Time-Lapse That Hold Together

Clips run up to 20 seconds, and the one in this row is exactly that: a single 20 second generation with no cuts and no stitching. Kids roll snow into a snowman, it stands alone through a golden sunset and then a moonlit night, and the snow around it recedes to bare grass as the season turns. The snowman and the yard stay recognizable from the first frame to the last, and prompts can chain clips together when you need longer than 20 seconds.

Examples

See what FLUX 3 can create

Turn on audio to hear the native sound generation. Every clip below came back from a single request, video and audio together, with no post-production.

Text to video

Reverse-motion glass reassembly

"In a surreal but physically clean reverse-motion shot, scattered shards of a green glass bottle slide and leap from a concrete floor, reassembling into one intact bottle standing upright. Every shard contributes to the final bottle; dust lifts back into cracks. Static camera, reversed glass sounds, no hands."

Text to video

Airflow and hair in a stable close-up

"Profile close-up of a woman with long curly hair standing beside an open train window. A passing tunnel briefly blocks the light as airflow pushes individual curls backward; the hair settles naturally when the train slows. Stable face, realistic strand motion, rhythmic rail audio."

Text to video

A full facial reaction, held

"tight shot of an elderly man's face laughing hard at a joke, deep wrinkles creasing, eyes squeezing shut, then wiping a tear from the corner of his eye"

Text to video

Two hands playing independently

"a pianist's hands playing a fast passage across the keys, fingers crossing over each other, both hands visible and independent"

Text to video

Fire, steam, and night-market sound

"street food vendor in bangkok flipping noodles in a wok over roaring flames at night"

Text to video

Fur and wind at highway speed

"a golden retriever sticking its head out a car window on the highway, ears flapping in the wind"

Text to video

Handmade stop-motion look

"Handmade stop-motion scene of a tiny paper sailor raising a blue fabric sail on a cardboard boat. The cloth catches a gust, billows to the right, and pulls the boat forward across a rippled cellophane sea. Visible handcrafted texture, consistent puppet proportions, papery rustles and gentle wooden creaks."

Image to video

A still photo animated with sound

"Cat walks across the windowsill, tail swaying. It pauses to look outside, then continues and hops down. Soft paw steps and a gentle meow."

FAQ

Common questions about FLUX 3

What is FLUX 3?

FLUX 3 is a multimodal model from Black Forest Labs, and the first FLUX release to generate video. One model is trained jointly across image, video, audio, and action prediction, so it generates clips of up to 20 seconds with a synchronized soundtrack rendered in the same pass. Text to video and image to video are coming soon to fal.

Is FLUX 3 available yet?

Not yet. FLUX 3 is coming soon to fal. Leave your details below and our team will let you know as soon as it goes live.

What can FLUX 3 generate?

Text to video takes a written prompt and returns a clip with audio. Image to video pins your image as the first frame and animates forward from it, which keeps the look of your source frame while the prompt describes the motion, camera work, and sound. Both paths cover live-action realism, stylized and handmade looks such as stop-motion, macro and product shots, and time-lapse.

What durations, resolutions, and aspect ratios are supported?

Clips can be 5, 10, 15, or 20 seconds, or you can leave duration on auto and let the model choose. Output is 480p or 720p. Aspect ratios run from 21:9 through 16:9, 4:3, 1:1, 3:4, 9:16, and 9:21, with auto available as well. Audio generation is on by default and can be switched off.

How does the native audio work?

Audio is part of the same generation, not a second step, so sound effects land on the frame where the event happens. Naming the sound you want in the prompt gives the most reliable result, for example a single bell strike, papery rustles, rhythmic rail noise, or kitchen ambience. Quiet, low-motion scenes are the ones most likely to come back with little audible sound, so it is worth adding an explicit audio cue for those.

What does the grounding parameter do?

Grounding ties the generation to the real-world context implied by your prompt, which helps when a scene needs plausible physics, materials, and scale. It is enabled by default. Turning it off gives the model more freedom, which can suit surreal or heavily stylized shots.

What kinds of prompts work best with FLUX 3?

Shots built around one clear event are the most reliable, especially when motion travels from a main object to the next object in a visible sequence. Say what the camera does, name the material behavior you expect, and name the sound. Asking for several coordinated actions at once, such as a group of people crossing the same point or a full-body sequence where hands, feet, and spine all have to move together, is harder, so it helps to split those into separate shots.

Is FLUX 3 image generation available?

Black Forest Labs has announced image generation and editing for FLUX 3, with improvements to complex prompt following and multilingual text rendering, and an open-weight FLUX 3 Dev release planned. Video is the part coming to fal. In the meantime you can run FLUX.2 and the rest of the FLUX family for image generation and editing on fal.ai today.

Can I use FLUX 3 for commercial projects?

Yes. Content generated through the fal.ai API can be used in commercial projects. Check fal.ai's terms of service for full details on usage rights and licensing.

Get notified about FLUX 3

FLUX 3 is coming soon to fal. Leave your details and our team will let you know when it goes live.

Contact Sales