Now on fal

FLUX 3

The first video model from Black Forest Labs. Generate up to 20 seconds of video with native audio from text, a still image, first and last frames, keyframes, or an existing clip.

Your prompt opens in the playground, where you generate.

Made with FLUX 3

What Makes FLUX 3 Different

Real World Understanding

Give It the Idea and It Knows the Rest

The whole prompt for this clip was “Visualization of ‘Annabel Lee’ by Edgar Allan Poe.” No shot list, no wardrobe notes, no lighting direction. FLUX 3 worked out the setting, the period, and the order of events on its own, because it already had the context. That is what real world understanding buys you: write the idea at the level you actually think about it, and let the model fill in the hundred details you would otherwise be spelling out.

Native Audio

Sound Generated With the Picture

FLUX 3 renders audio inside the same pass as the video, so an event arrives carrying the sound it makes. Every hammer blow here lands with its own ring off the anvil and the hiss of sparks leaving the hot blade, on the exact frame the metal is struck. Nothing needs lining up afterward and no second model gets called, because the soundtrack was never a separate render.

Keyframes to Video

Lock the Shot to Your Keyframes

Keyframes and storyboards are how a real production plans a shot, and FLUX 3's keyframes to video endpoint pins up to 10 frames to exact positions. Here four frames at 0, 5, 10, and 15 seconds set the whole arc, from a rock headland above a harbour to a colossus standing above the clouds, and the model returns one continuous take that hits every one of them. Pinning the shot this way is what keeps a long generation from drifting: you hand the model your creative guidelines instead of hoping one paragraph of text lands the same way twice.

Physical Cause and Effect

Chain Reactions That Follow Through

FLUX 3 is trained jointly across image, video, audio, and action, and it shows in shots where one event has to trigger the next. A red marble rolls in from the left, strikes only the first of six dominoes, and the run completes cleanly to the right while the marble comes to rest beside the piece it started. Motion carries mass and momentum from one object to the next instead of resetting between beats.

Camera and Materials

Real Camera Moves, Real Surfaces

A slow lateral dolly travels behind three dark tree trunks to reveal a lighthouse on a distant cliff. Each trunk crosses the foreground at its own speed, the lighthouse stays anchored where it belongs, and the storm keeps working through the whole move as waves break against the rock below and rain beads on the lens. Orbits, focus racks, and tracking shots hold their geometry the same way.

20 Second Clips

Long Shots and Time-Lapse That Hold Together

Clips run up to 20 seconds, and the one in this row is exactly that: a single 20 second generation with no cuts and no stitching. Kids roll snow into a snowman, it stands alone through a golden sunset and then a moonlit night, and the snow around it recedes to bare grass as the season turns. The snowman and the yard stay recognizable from the first frame to the last, and prompts can chain clips together when you need longer than 20 seconds.

Endpoints

Five endpoints, full quality and fast drafts

Generate from text, a still image, first and last frames, keyframes, or an existing clip, at full quality or as a fast, low-cost draft. Draft Enhance re-renders the draft you pick at full quality with the same seed and motion.

Flux 3 Text to Video
Black Forest Labs logo
blackforestlabs/flux-3/text-to-video
text-to-video

FLUX 3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene.

stylizedtransformlipsync
FLUX 3 Image to Video
Black Forest Labs logo
blackforestlabs/flux-3/image-to-video
image-to-video

FLUX 3 is Black Forest Labs' frontier video model. This endpoint animates a single still image into video, extending one frame into coherent, natural motion.

stylizedtransformlipsync
Flux 3 First Last Frame to Video
Black Forest Labs logo
blackforestlabs/flux-3/first-last-frame-to-video
image-to-video

FLUX 3 is Black Forest Labs' frontier video model. This endpoint generates the video between a defined start and end frame, interpolating a smooth, coherent transition from the first image to the last.

stylizedtransformlipsync
Flux 3 Image to Video
Black Forest Labs logo
blackforestlabs/flux-3/keyframes-to-video
image-to-video

FLUX 3 is Black Forest Labs' frontier video model. This endpoint builds video from a sequence of keyframes, generating the motion between each anchor point for precise control over how a shot progresses.

stylizedtransformlipsync
Flux 3 Extend Video
Black Forest Labs logo
blackforestlabs/flux-3/extend-video
video-to-video

FLUX 3 is Black Forest Labs' frontier video model. This endpoint continues an existing clip beyond its final frame, generating additional footage that stays consistent with the original motion and scene.

stylizedtransformlipsync
Flux 3 Text To Video Draft
Black Forest Labs logo
blackforestlabs/flux-3/text-to-video/draft
text-to-video

FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews from a text prompt, with a reusable draft cache for full-quality enhancement.

stylizedtransformlipsync
Flux 3 Image To Video Draft
Black Forest Labs logo
blackforestlabs/flux-3/image-to-video/draft
image-to-video

FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews that animate a still image, with a reusable draft cache for full-quality enhancement.

stylizedtransformlipsync
Flux 3 First Last Frame to Video Draft
Black Forest Labs logo
blackforestlabs/flux-3/first-last-frame-to-video/draft
image-to-video

FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews between a start and an end frame, with a reusable draft cache for full-quality enhancement.

stylizedtransformlipsync
Flux 3 Keyframes To Video Draft
Black Forest Labs logo
blackforestlabs/flux-3/keyframes-to-video/draft
image-to-video

FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews pinned to your keyframe images, with a reusable draft cache for full-quality enhancement.

stylizedtransformlipsync
Flux 3 Extend Video Draft
Black Forest Labs logo
blackforestlabs/flux-3/extend-video/draft
video-to-video

FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews that continue an existing clip, with a reusable draft cache for full-quality enhancement.

stylizedtransformlipsync
Flux 3 Draft Enhance
Black Forest Labs logo
blackforestlabs/flux-3/draft-enhance
video-to-video

FLUX.3 is Black Forest Labs' frontier audio/video model. Re-render a previously generated draft at full quality — same seed, same motion, no re-planning.

stylizedtransformlipsync
Examples

See what FLUX 3 can create

Turn on audio to hear the native sound generation. Every clip below came back from a single request, video and audio together, with no post-production. Copy any prompt and try it in the playground.

Text to video

Reverse-motion glass reassembly

"In a surreal but physically clean reverse-motion shot, scattered shards of a green glass bottle slide and leap from a concrete floor, reassembling into one intact bottle standing upright. Every shard contributes to the final bottle; dust lifts back into cracks. Static camera, reversed glass sounds, no hands."

Text to video

Airflow and hair in a stable close-up

"Profile close-up of a woman with long curly hair standing beside an open train window. A passing tunnel briefly blocks the light as airflow pushes individual curls backward; the hair settles naturally when the train slows. Stable face, realistic strand motion, rhythmic rail audio."

Text to video

A full facial reaction, held

"tight shot of an elderly man's face laughing hard at a joke, deep wrinkles creasing, eyes squeezing shut, then wiping a tear from the corner of his eye"

Text to video

Two hands playing independently

"a pianist's hands playing a fast passage across the keys, fingers crossing over each other, both hands visible and independent"

Text to video

Fire, steam, and night-market sound

"street food vendor in bangkok flipping noodles in a wok over roaring flames at night"

Text to video

Fur and wind at highway speed

"a golden retriever sticking its head out a car window on the highway, ears flapping in the wind"

Text to video

Handmade stop-motion look

"Handmade stop-motion scene of a tiny paper sailor raising a blue fabric sail on a cardboard boat. The cloth catches a gust, billows to the right, and pulls the boat forward across a rippled cellophane sea. Visible handcrafted texture, consistent puppet proportions, papery rustles and gentle wooden creaks."

Image to video

A still photo animated with sound

"Cat walks across the windowsill, tail swaying. It pauses to look outside, then continues and hops down. Soft paw steps and a gentle meow."

API Documentation

How to access the FLUX 3 API

The fal client handles the request submit protocol. It submits the job, streams status updates, and returns the result when the generation is complete.

import { fal } from "@fal-ai/client";

const result = await fal.subscribe("blackforestlabs/flux-3/text-to-video", {
  input: {
    prompt: "A street food vendor flips noodles in a wok over roaring flames at night, sparks and steam rising, sizzling audio and night-market chatter",
    duration: 10,
    resolution: "1080p",
    generate_audio: true
  },
  logs: true,
  onQueueUpdate: (update) => {
    if (update.status === "IN_PROGRESS") {
      update.logs.map((log) => log.message).forEach(console.log);
    }
  },
});
console.log(result.data);
console.log(result.requestId);
FAQ

Common questions about FLUX 3

What is FLUX 3?

FLUX 3 is a multimodal model from Black Forest Labs, and the first FLUX release to generate video. One model is trained jointly across image, video, audio, and action prediction, so it generates clips of up to 20 seconds with a synchronized soundtrack rendered in the same pass. It is available now on fal across five core endpoints, each with a fast draft variant for cheap previews.

What can FLUX 3 generate?

FLUX 3 has five core video endpoints on fal. Text to video and image to video build a clip from a prompt or a starting frame. First-last frame and keyframes to video pin exact frames, up to 10, so a long shot hits the beats you set. Extend video continues an existing clip from its final frames. Every endpoint returns synchronized audio in the same pass, across live-action realism, stylized and handmade looks such as stop-motion, macro and product shots, and time-lapse, and each one has a fast draft variant for previewing a shot before you commit to a full render.

How does the draft workflow work?

Every core endpoint has a draft variant that renders a lower-quality preview quickly and cheaply, so you can judge a shot in seconds instead of waiting on a full render. Drafts take a batch count, so you can preview several variations of one prompt at once. When you find the take you want, Draft Enhance continues rendering that exact draft to full quality, keeping the same seed and motion, so the final clip matches the preview you picked. You can also try the text to video draft free, five generations a day for signed-in accounts.

How do I get started with the FLUX 3 API?

Install the fal.ai SDK (Python or JavaScript), grab an API key from your dashboard, and make your first request in a few lines of code. The API is serverless, so there are no GPUs to manage and no infrastructure to set up. Check the API documentation for every parameter.

What durations, resolutions, and aspect ratios are supported?

Clips can be any whole number of seconds from 5 to 20, or you can leave duration on auto and let the model choose. Output is 720p or 1080p at 24 FPS. Aspect ratios run from 21:9 through 16:9, 4:3, 1:1, 3:4, 9:16, and 9:21, with auto available as well. Audio generation is on by default and can be switched off.

How does the native audio work?

Audio is part of the same generation, not a second step, so sound effects land on the frame where the event happens. Naming the sound you want in the prompt gives the most reliable result, for example a single bell strike, papery rustles, rhythmic rail noise, or kitchen ambience. Quiet, low-motion scenes are the ones most likely to come back with little audible sound, so it is worth adding an explicit audio cue for those.

What does the grounding parameter do?

Grounding ties the generation to the real-world context implied by your prompt, which helps when a scene needs plausible physics, materials, and scale. It is enabled by default. Turning it off gives the model more freedom, which can suit surreal or heavily stylized shots.

What kinds of prompts work best with FLUX 3?

Shots built around one clear event are the most reliable, especially when motion travels from a main object to the next object in a visible sequence. Say what the camera does, name the material behavior you expect, and name the sound. Asking for several coordinated actions at once, such as a group of people crossing the same point or a full-body sequence where hands, feet, and spine all have to move together, is harder, so it helps to split those into separate shots.

Is FLUX 3 image generation available?

Black Forest Labs has announced image generation and editing for FLUX 3, with improvements to complex prompt following and multilingual text rendering, and an open-weight FLUX 3 Dev release planned. Video is the part on fal today. Alongside it you can run FLUX.2 and the rest of the FLUX family for image generation and editing on fal.ai.

Can I use FLUX 3 for commercial projects?

Yes. Content generated through the fal.ai API can be used in commercial projects. Check fal.ai's terms of service for full details on usage rights and licensing.

Get in touch about FLUX 3

Want help integrating FLUX 3 into your workflow? Leave your details and our team will reach out.

Contact Sales