Best Flux 3 Text to Video Examples + Prompting Guide

Explore all models

A few days of testing FLUX 3, Black Forest Labs' first video model, collected into 30 clips. Period documentaries, historical explainers, public domain trailers, music videos, and comedy, almost all from a single sentence.

last updated
7/30/2026
edited by
Bennett Heyn
read time
20 minutes
Best Flux 3 Text to Video Examples + Prompting Guide

FLUX 3 is Black Forest Labs' first video model, and it is coming soon to fal. I have spent the past few days testing it, and this article collects my favorite clips I kept coming back to.

Two things shape everything below. The first is length. FLUX 3 produces up to 20 seconds in a single pass, which is enough room for a sequence of shots rather than one continuous moment, and it changes what a prompt can ask for. The second is how much the model already knows. Almost every prompt in this article is one sentence long, with no shot list, no wardrobe notes, and no lighting direction.

Every clip has audio generated in the same pass as the picture, so turn your sound on before you start. Everything here is 720p text to video.

What FLUX 3 can do

FLUX 3 is not limited to one entry point. It runs text to video, image to video, and reference to video, with more modes on the way. Output comes back at 480p or 720p, at 5, 10, 15, or 20 seconds.

Pricing has not been announced yet, and neither has the launch date. What follows is what a few days with the model actually produced.

Historical documentaries are one of the strongest thing FLUX 3 does

This is the category that sold me. A period documentary is a hard ask because it is not one decision, it is a hundred of them: the film stock, the lens, the clothing, the way a crowd behaves when it knows a camera is present, and the emotional register of the moment. Get one wrong and the whole thing reads as a costume party.

FLUX 3 tends to get them all pointing in the same direction, and it captures the raw feel of an era rather than a surface impression of it.

A 1969 documentary about Woodstock

Prompt: a 1969 documentary about Woodstock

What impressed me most is how much FLUX 3 understood from such a simple prompt. It picked up the handheld camera work, the faded colors, the rough film texture, the crowded framing, and the candid feeling of footage captured in the moment.

It did not just make Woodstock look vintage. It made the whole clip feel like it came from that period. The camera, lighting, clothing, crowd behavior, and overall atmosphere all work together. It opens on a hillside packed back to the treeline, cuts to a guitarist in bell bottoms working in front of a stack of amps, and finds a young woman walking back through the blankets and smiling at the lens.

This is where FLUX 3 feels especially strong. You can describe an era, a format, and a subject in one sentence, and it turns that into a surprisingly complete visual style.

A 1987 local news report from the mall

Prompt: a 1987 local news report about teenagers hanging out at the mall

FLUX 3 picked up on nearly every part of the era. The soft video quality, fluorescent lighting, oversized clothing, crowded storefronts, awkward interviews, and slightly washed out colors all feel connected.

It did not simply add an old camera filter. The framing and the performances feel like footage recorded by a local news crew in the late 1980s. Look at how it is cut: a reporter stand up in the concourse with a stick microphone, b roll of teenagers around a food court pizza, a row of lit arcade cabinets, then back to the reporter while kids walk behind her. That is the structure of a local news package, and nothing in the prompt asked for it.

This is what makes simple prompts so interesting with FLUX 3. The model does not only visualize the subject, it interprets the media format around it.

A 1980s documentary about the Loch Ness monster

Prompt: a 1980s documentary about the loch ness monster

FLUX 3 is excellent at historical documentaries, and it builds this one the way these programmes were actually built. A wide of the loch under low cloud with a single boat on it, a researcher in an oilskin jacket hunched over an echo sounder, the paper trace and the green CRT beside him, and then a hump breaking the surface at a distance far enough away to stay deniable.

It even changes shape partway through. The later shots narrow into a squarer frame with black edges, the way a broadcast cuts away to older archive material.

A 1995 television documentary explaining the internet

Prompt: a 1995 television documentary explaining the internet

FLUX 3 captured the specific optimism and awkwardness of early internet coverage. The bulky computers, office cubicles, slow demonstrations, animated graphics, formal narration, and people carefully explaining email all felt accurate to the period.

It was not simply a modern video with older technology placed inside it. The presenter in the grey suit, the beige tower and CRT on the desk, the external modem with its row of LEDs, and the rotating wireframe globe with little computer icons hung on it are all period correct. The pacing matches how television used to introduce a new technology: show the object, show a person using it, then cut to the graphic that explains the concept.

FLUX 3 is extremely good at reconstructing not only an era, but how that era imagined the future.

A 1989 television documentary about the fall of the Berlin Wall

Prompt: a 1989 television documentary about the fall of the Berlin Wall

Here it captured the handheld news footage, the harsh camera lights, the crowded streets, the concrete dust, the emotional reunions, and the chaotic feeling of history unfolding in real time.

It did not just recreate the Berlin Wall. It recreated the experience of watching it come down through a late 1980s television broadcast: a night crowd pressed against a graffitied section under sodium light, a border guard beside a checkpoint gate as people push through, figures hauling each other up onto the top of the wall, then a hammer and chisel close on the concrete with the Brandenburg Gate visible down the line.

FLUX 3 feels especially strong when the event already has a recognizable visual language. A year, a format, and a subject can be enough to generate an entire historical atmosphere.

A 1969 television broadcast of the moon landing

Prompt: a 1969 television broadcast of the moon landing

It got the statement right on its own: "that's one small step for man, one giant leap for mankind." Nothing in the prompt asked for it.

What impressed me most is how much FLUX 3 understood from one sentence. It captured the harsh contrast, the blown out highlights, the unstable transmission quality, the slow astronaut movement, and the eerie emptiness of the lunar surface.

The most telling decision is one it did not make. It never cuts. The whole clip sits in a single locked off camera position beside the lander, with the sun flaring at the left edge and the frame bowed slightly at the corners like a signal arriving through a tube. Reaching for something cinematic would have meant cutting four times. This understood that the format is a fixed camera and a poor signal.

This is where FLUX 3 feels especially powerful. You can give it a historical event and a recording format, and it starts filling in the visual language around both.

Archival footage of the Wright brothers in 1903

Prompt: archival footage of the Wright brothers' first flight in 1903. No sound

FLUX 3 understood the simplicity of the moment. A fragile aircraft, an open field, a small group of witnesses, heavy clothing, rough wind, and a camera positioned far enough away to observe rather than dramatize. It even simulated the edge of a film frame, with a sprocket notch sitting in the corner of the picture.

The result did not feel like a cinematic reenactment. It felt like an improbable machine briefly lifting off the ground while someone happened to be recording.

It also honored the last two words of the prompt. The audio track it returned is effectively silent, averaging around -68 dB, while every other clip in this article came back with a full mix.

FLUX 3 is especially impressive when it turns a familiar historical photograph into a living moment.

What the dinosaurs saw

I asked for what the dinosaurs saw as the asteroid hit. This one is a visualization rather than a recording format, and it is the clearest example of FLUX 3 building an arc inside a single generation.

It opens calm, with hadrosaurs and a horned herd grazing along a river in low morning light and wading birds in the shallows. Then the horizon goes white. The next shots track alongside the herd as it stampedes through the trees, and it ends with the treeline bending all one way while dust swallows the frame and the animals disappear into it.

None of that is one shot held for twenty seconds. It is a beginning, a turn, and an end.

How it was made: Explainer Historical

These are a different kind of test. How something was built is a question with a real answer, so atmosphere alone will not carry the clip. The model has to know the process and put the steps in the right order.

How the Golden Gate Bridge was built

I asked for archival footage of how the Golden Gate Bridge was built, and this is the clip that made me stop and rewind. FLUX 3 did not just generate black and white construction footage. It generated the title cards to go with it, and it got the dates right.

It opens on a card reading JANUARY 5, 1933 over the anchorage works, which is the day construction actually started. Then it moves through labeled phases in the correct order: foundations and anchorages, then the towers, then the main cables being spun, then the suspended roadway deck, and a closing card reading CONSTRUCTION COMPLETED, APRIL 19, 1937, which matches the completion date weeks before the bridge opened to the public. Each lower third carries its own date range, and each range lines up with the phase it labels.

Legible on screen typography is hard for a video model. Legible typography that is also historically correct, and in sequence, is a different thing entirely.

How the Great Wall of China was made

Prompt: How the great wall of china was made

If the Golden Gate clip was a surprise, this is the single most impressive generation in this article. From that one line, FLUX 3 produced a structured explainer with a title card, a split screen comparison, a labeled stage for every construction method, and a closing argument.

It opens on HOW THE GREAT WALL WAS MADE over a three panel split screen labeled LOCAL GRAVEL, REEDS + SOIL, and LOCAL STONE, captioned as early regional walls from different places and different periods. Then it walks the actual methods in order: rammed earth tamped in thin layers inside wooden forms, the Ming era rebuilding on stone foundations, a rammed earth or rubble core, fired brick facing with lime mortar, and finally battlements, stairs, and watchtowers.

It closes on a card reading BUILT AND REBUILT ACROSS MANY PERIODS, NOT ONE PROJECT, NOT UNIFORMLY BRICK. That card is the part I did not expect. The most common misconception about the Great Wall is that it is one continuous brick wall, and the model chose to end by correcting it. It is not only illustrating the subject, it has a point of view about what you probably have wrong.

How the pyramids were built

Prompt: How pyramids are made? Visualize it.

No title cards on this one, just the logistics, in order. A limestone quarry where workers cut blocks free with mallets and chisels, a finished block lashed onto a wooden sledge, the sledge loaded onto a barge on the river with the pyramid already rising behind it, hauling teams leaning into ropes on a ramp, a close shot of a block being levered into place, and a wide of the construction ramp running up the face.

Quarry, sledge, river, ramp, placement. That chain is the mainstream account of how the stone actually moved, and it arrived in the right sequence from a five word question.

Cinematic trailers from short public domain prompts

If the model already knows a book, you do not have to describe the book. These four prompts come close to just naming the work and getting out of the way. If you use non public domain titles, you may hit an error.

Frankenstein by Mary Shelley

Prompt: Visualize Frankenstein by Mary Shelley as a movie.

I generated an entire movie trailer with FLUX 3 from one simple prompt.

From that single sentence it created a complete cinematic concept with characters, environments, camera movement, atmosphere, pacing, and a clear sense of story. It felt less like generating an isolated clip and more like watching an early version of a film take shape.

What sold me is that it adapted the novel rather than the famous 1931 film. The creature wakes in a candlelit stone laboratory, Victor recoils from what he has made, and then the trailer follows the book: the creature outside a snowbound cottage watching a lit window, and an ending out on the Arctic ice with a dog sled crossing the frame. It signs off on a card reading FRANKENSTEIN; Or, The MODERN PROMETHEUS, and beneath it, based on the 1818 novel by Mary Shelley. Right subtitle, right year, right ending.

The pace of progress in generative media is incredible. We can already turn rough ideas into images, videos, music, voices, and cinematic sequences in minutes, yet these tools are still in their earliest stages of adoption.

Most people have not seriously experimented with AI generated media yet. That creates a huge opportunity for anyone willing to learn how the models work, test different creative workflows, and develop a real instinct for prompting and directing them.

Now is the time to explore. Try different models, study what each one does best, and push beyond the obvious prompts. The tools will continue improving, but knowing how to use them creatively will be the real advantage.

Sherlock Holmes

Same idea, and this time barely more than a name.

Every piece of the iconography arrives unasked. The wood panelled Baker Street study with a chemistry bench and glassware on it, the fire going, a magnifier held over a handwritten letter, the pipe, and then a foggy gaslit street with a hansom cab and wet cobbles. The model knows what this character's world looks like, so the prompt does not have to carry it.

The Odyssey by Homer

Prompt: Visualization of 'The Odyssey' by Homer.

720p, not 70mm.

No shot list, no wardrobe notes, no lighting direction, no storyboard. FLUX 3 worked out the setting, the period, and the order of events on its own, because it already had the context.

It opens in the Cyclops cave, where the crew drive a burning stake into Polyphemus. Then a card reading THE ODYSSEY, and under it, traditionally attributed to Homer. Then the ship: a galley under a square sail rowing past a rocky headland, Odysseus lashed to the mast, figures watching from the rocks.

Two of the most famous episodes in the poem, in the order Homer put them, plus a credit line that hedges the authorship the way a scholar would.

That is what real world understanding buys you. Write the idea at the level you actually think about it, and let the model fill in the hundred details you would otherwise be spelling out.

Annabel Lee by Edgar Allan Poe

Prompt: Visualization of 'Annabel Lee' by Edgar Allan Poe.

FLUX 3 is a really smart model. You can just tell it what you want and it understands.

It staged the poem's real sequence: a castle above a sea cliff, two lovers in period dress down on the shingle, the storm that takes her, her highborn kinsmen carrying her through a graveyard to a sepulchre, and the lover lying down beside her tomb under a low moon. That is the poem, in order, from one line naming it.

Modern television and film examples

Away from history, the thing worth watching for is coverage: whether the model cuts between angles the way an editor would, or just hands you one pretty shot.

A modern day spy thriller

Prompt: recreate a modern-day spy thriller in full cinematic quality

This is insane. FLUX 3 exceeded my expectations here.

It established multiple cinematic shots as starting frames and made them flow together seamlessly. It turned a one line idea into a complete sequence about a woman taking a valuable item, trying to protect it, and fleeing from the people chasing her.

The run of shots: a rain soaked street full of umbrellas, a gloved handoff of a small black case, a fast walk across a hotel lobby, a masked man waiting in a parking garage, a sprint between parked cars, and an escape on a motorcycle into wet night traffic.

What impressed me most was how much story it created without being given a detailed script. The locations, pacing, camera movement, tension, and continuity all felt intentional. It did not just generate a collection of cool shots. It made something that actually felt like a movie trailer.

An episode of television

Prompt: generate an episode of television

I have to give it to the model, I am intrigued.

A detective in a belted trench coat lets herself into a dim apartment with rain running down the windows, searches it, crouches under an overturned desk, and comes up with a Polaroid of a man lit by her flashlight. Then it cuts outside to the man from the photograph, out in a rain slicked alley under teal light.

Who is he, and why was his picture in that apartment? The pacing and the editing between shots are the best part of this one, and it does all of it in ten seconds.

A Ghibli style anime clip

FLUX 3 handles hand drawn styles too, and it did a ghibli style clip really well.

A girl in a straw hat stands on a green hillside above terraced rice paddies with seed heads blowing past her, and an enormous translucent antlered spirit rises over the hill behind her. It ends on her face beside the creature's amber eye.

The painted backgrounds, the towering cumulus, the wind moving through grass, and a benign forest spirit are all load bearing parts of that tradition, and it assembled them without a reference image.

Einstein explaining probabilities

Prompt: Einstein explaining probabilities to a professor

It made a really good scene out of a simple prompt. The setting is a wood panelled study at night, green shaded desk lamps, papers spread across the desk.

What matters here is the coverage. It opens on a two shot, moves to an over the shoulder, cuts to a close single on Einstein mid explanation, cuts to a matching single on the professor listening, then returns to the wide. That is standard dialogue grammar with eyelines that match, and none of it was specified.

A gym interview

Prompt: Man with a big right arm and a small left arm gets interviewed in a gym explaining his training routine.

A funny television style clip: a man on a bench in a gym with a microphone held into frame from off camera, explaining his training routine, with one enormous right arm and a normal left one.

The joke works because the model keeps it consistent. The asymmetry holds across a wide shot, an over the shoulder, a close insert on the arm, and a reverse on his face. Continuity on an absurd detail is harder than continuity on a normal one.

A man eating soup with a fork

This one felt like it could be part of a television show. A man eats soup with a fork and slowly registers that the woman across the table is judging him for it.

A couple at a restaurant table under warm bulbs, brick wall behind, other diners filling the background. There is an insert on the fork going into the bowl, then singles on his face getting defensive, then back to the two shot. Nothing really happens, and it is still legible as a scene, which is the whole point.

A baby ranting about AI

Prompt: baby ranting about AI

An irreverent clip of a baby getting worked up and delivering a little monologue about AI, from a four word prompt.

A baby in denim overalls in a high chair in a bright kitchen, gesturing with both hands. The camera stays locked off for the whole clip, which is the right instinct for a rant, and the performance is all in the hands.

Nature documentary about shopping carts returning to the wild

Prompt: A nature documentary about shopping carts returning to the wild.

My favorite of the comedy prompts, because it takes the format completely seriously.

It opens on a cart corral outside a supermarket at dusk, then a trail flattened through a misty meadow like a game trail, a lone cart standing in wildflowers at dawn, a herd fording a woodland creek, and a wide of a dozen of them grazing in a clearing in golden light.

Every one of those is a real wildlife documentary shot. The joke lands because the model committed to the grammar instead of the gag.

Music videos

FLUX 3 is not a music model, and it still turned out three watchable music videos from one line each. I was impressed by the model's ability to compose music and sync it to music videos that felt realistic.

A pop punk ballad about the girl next door

Prompt: make a catchy pop punk ballad about a boy and his next door girl neighbor

It built the video the genre would have built. A kid in a band tee with an electric guitar playing at his bedroom window at dusk, the girl visible at her window across the yard, suburban clapboard houses with the lights coming on, inserts on his hands moving up the fretboard, and the two of them meeting at the fence and then running across the lawn.

A modern R&B hit set in the Bronx

Prompt: make a modern day R&B hit that takes place in the Bronx in New York

The locations do the work here, and they are specific. Under the elevated line with bodegas running along the street, inside a bodega beside the drink fridges, a night shot with a black sedan on wet asphalt below the tracks, and a rooftop at blue hour with water towers and tenement roofs behind him.

A 2000s boy band video

Prompt: make a music video for an 2000s boy band

Honestly a bit catchy.

The visual references are exact: five guys walking a glowing white tunnel, a rooftop at sunset with a skyline behind them, a formation dance in a mirrored corridor, and a rain soaked stage number in backlight, with coordinated wardrobe holding across all of it.

Historical people meeting modern conveniences

A lighter category, and three variations on one joke: put people from one era in front of an object from another and let them work it out. What makes them work is that FLUX 3 plays the discovery straight instead of winking at it.

Medieval peasants discover a bounce house

Prompt: Medieval peasants discovering an inflatable bounce house

A muddy village of thatched roofs and wattle fences with cabbages in the garden, and a red and yellow inflatable castle sitting in the middle of it. The peasants approach the way you would approach a large animal: prod the column, lean on the vinyl, then commit, until a group of them is inside it jumping.

The detail I like is that the inflatable is shaped like a castle, which drops a plastic keep into a village that has probably never seen a real one.

Cavemen encounter a revolving door

Prompt: Cavemen encountering a revolving door

I thought this one was pretty funny. Two cavemen in furs with clubs outside a glass office tower. They press their hands against the glass, get carried around inside the drum, take a swing at it with a club, and end up standing in a marble lobby having accidentally succeeded.

It looks like it has the basis to be a good commercial.

Medieval peasants encounter an escalator

Prompt: Medieval peasants encountering an escalator

A torch lit stone market hall with straw on the floor and produce stalls, and a modern escalator running up out of the middle of it. The peasants peer at the moving steps, and there is a close insert of a leather boot landing on the treadplate before a group rides up holding the handrail and grinning.

The insert is the giveaway. Nobody asked for a cutaway to the boot, but that is the shot the scene needs.

What I would tell you about prompting FLUX 3

Four things I took away from the week.

Name the format, not just the subject. "A 1987 local news report about teenagers hanging out at the mall" returns a news package. Teenagers at a 1980s mall returns a shot. The format carries the camera, the lighting, and the edit along with it, so it is the highest leverage word in the prompt.

Write at the level you actually think. The strongest results here came from prompts that name a work or an idea and then stop. Every detail you add is a detail the model is no longer choosing, and on subjects it knows well its choices are better informed than mine.

Let the length do the storytelling. Twenty seconds is enough for a beginning, a turn, and an end. The dinosaur clip and the spy thriller both use it that way, and neither prompt asked for a sequence.

Trust it with real facts. The dates on the Golden Gate title cards, the construction stages of the Great Wall, the subtitle and publication year of Frankenstein, and the order of episodes in The Odyssey were all correct without being supplied. When a subject has a real answer, the model often knows it.

FLUX 3 FAQ

Is FLUX 3 available on fal yet?

Not yet. FLUX 3 is coming soon to fal, and pricing has not been announced.

How long can a FLUX 3 clip be?

5, 10, 15, or 20 seconds in a single generation. Most of the clips in this article are the full 20 seconds, which is what makes multi shot sequences possible in one pass.

What resolutions does FLUX 3 support?

480p and 720p. Everything in this article is 720p.

Does FLUX 3 generate audio?

Yes, in the same pass as the picture. Every clip here came back with sound except the Wright brothers flight, where the prompt asked for none and the model returned a silent track.

What can you do with FLUX 3 besides text to video?

FLUX 3 is more than a single entry point. Alongside text to video it handles image to video and reference to video, with more modes on the way.

What kind of prompt works best with FLUX 3?

Name the format alongside the subject. A year, a medium, and a subject in one sentence gave the strongest results in my testing, because the format carries the camera, the lighting, and the edit with it.

Try FLUX 3 when it lands

Every clip in this article came out of a single sentence and a single model, which is a genuinely new thing to have on hand. FLUX 3 is coming soon to fal, and the useful thing to do between now and then is build the instinct: pick formats you know well, name them precisely, and find out how much the model already has. If you want to see some of my generations, I've been posting clips to my twitter.

In the meantime, it is worth comparing notes with the rest of the field. We keep running lists of the best AI video generators and the best text to video APIs, plus a walkthrough on how to generate videos with AI if you are starting from scratch.

about the author
Bennett Heyn
Bennett is a growth marketer at fal leading SEO & AEO. He is interested in creative engineering and testing new AI models

Related articles