Seedance 2.5 Prompting Guide + Real Examples

Explore all models

Prompt structures for 30-second shots, reference control, native audio, physical continuity, and cleaner motion in Seedance 2.5. All 10 examples include the exact prompt and the generated video.

last updated
8/7/2026
edited by
Ilker
read time
15 minutes
Seedance 2.5 Prompting Guide + Real Examples

Most Seedance 2.5 prompting advice is still stuck at "describe the subject, camera, and lighting." Everyone knows that part. It works for a short mood clip, then falls apart when a shot has several actions, objects changing hands, dialogue, references, or a camera that needs to hit a specific mark.

For anything beyond a simple shot, I write the prompt like a short production note. It says what happens, when it happens, what each reference controls, what stays unchanged, and where the shot ends.

All 10 videos below were generated on fal with the prompt printed directly under the result. I did not cut them, retime them, or add audio afterward.

A 30-second shot needs a timeline

The 30-second setting only changes the available duration. It does not add more events to the prompt. Give it one small movement and the extra time usually turns into waiting, repetition, or slow motion.

The bodega prompt uses six blocks. The messenger enters, gets a drink, pays, and leaves. Each block picks up the physical state left by the one before it. The red helmet stays in the left hand. Once the bottle comes out of the cooler, it stays in the right hand. The camera has its own route through the store too.

Generated using Seedance 2.5 on fal.

Prompt

30-second continuous single take inside a small New York City bodega on a rainy morning, all action at natural real-time speed. The same bike messenger wears a yellow rain jacket and carries one red bicycle helmet in the left hand throughout. 0-5 seconds: the door bell rings as the messenger enters, closes the glass door with the right hand, and shakes rain from the shoulders without dropping the helmet. 5-10 seconds: the camera follows from behind at chest height as the messenger walks to the drink cooler, opens the cooler with the right hand, removes one clear bottle of seltzer, and closes the door. 10-16 seconds: the messenger turns toward the counter and walks around one stationary customer without changing hands; red helmet remains in the left hand, bottle remains in the right. 16-22 seconds: the messenger sets only the bottle on the counter, taps a black phone once on the card reader, waits for one confirmation beep, then picks up the same bottle with the right hand. 22-27 seconds: the clerk gives a small nod; the messenger turns back toward the entrance while the camera backs up and keeps a medium full-body frame. 27-30 seconds: the messenger opens the door with the right forearm, exits into the rain, and the door closes behind. The camera stops inside the store. Keep the same messenger, jacket, helmet, bottle, phone, clerk, counter, cooler, and store layout from first frame to last. No cuts, no slow motion, no repeated entrance, no duplicated bottle or helmet, no object teleportation. Audio: rain outside, door bell, cooler hum, footsteps, bottle on counter, one card-reader beep, quiet store room tone, no music.

I do not put timestamps in every prompt. A simple action can stay simple. Timing blocks become useful when several events share one shot and the order matters. They give Seedance fewer chances to skip the middle or rush to the ending.

Write the cause before the reaction

"The customer looks surprised" leaves too much open. The reaction may arrive before its cause, or everyone in the scene may react at the same time.

In this diner shot, the tray touches the cup handle first. The cup tips after contact. The customer hears the impact, looks down, and moves the sketchbook before the spill reaches it. The other diners react at slightly different times.

Generated using Seedance 2.5 on fal.

Prompt

15-second continuous single take inside a busy Brooklyn neighborhood diner in the morning, natural real-time speed. Eye-level medium-wide camera at the end of the counter. 0-4 seconds: a server places a full ceramic coffee cup beside an open sketchbook and walks away. 4-8 seconds: a busboy passes behind the seated customer; the edge of his tray lightly catches the cup handle. The cup tips only after contact, strikes the counter, and coffee begins spreading toward the sketchbook. 8-11 seconds: the customer hears the impact, looks down, then lifts the sketchbook just before the coffee reaches it. Nearby diners turn toward the sound at slightly different delays. 11-15 seconds: the server returns with a towel and stops the spill. The camera begins a slow 30-centimeter push-in only after the cup tips. Keep the same cup, sketchbook, server, customer, counter layout, and clothing throughout. Coffee follows the counter surface and never moves uphill. No cuts, no slow motion, no repeated action, no duplicated props, no music. Audio: ordinary diner room tone, dishes, the ceramic impact, liquid spill, chair movement.

Write cause and effect as separate actions: contact, resulting movement, sound, then reaction. I do not leave the connection between them for the model to invent.

Hidden objects still need continuity

An object does not stop existing when a column, car, or person blocks it. If the prompt only describes the entrance and exit, the covered part of the motion can become a reset point.

On the Chicago platform, the woman and her suitcase disappear behind the column for three seconds. The prompt repeats the identity, clothing, suitcase, direction, and speed that need to come back on the other side.

Generated using Seedance 2.5 on fal.

Prompt

11-second continuous locked wide shot on a Chicago elevated train platform in overcast afternoon light, natural real-time speed. A woman in a bright green wool coat pulls one small red rolling suitcase from frame left toward frame right. 0–3.2 seconds: she walks steadily beside the yellow platform line, right hand holding the extended suitcase handle, suitcase rolling one step behind her. 3.2–5.1 seconds: she passes completely behind one thick concrete support column; both woman and suitcase are fully hidden for just under two seconds. 5.1–9.5 seconds: the same woman emerges from the opposite side of the column with the same face, hair, green coat, black boots, red suitcase, extended handle, walking speed, and direction. 9.5–10.8 seconds: she continues toward the right edge and the shot ends mid-stride as she is about to exit frame. The camera never moves. The column remains fixed. No person or suitcase appears on both sides of the column at once. No identity change, clothing change, color change, duplication, jump cut, morph, slow motion, or disappearing luggage. Audio: light wind, distant city traffic, suitcase wheels on concrete, a far train announcement, no music.

For occlusion, write how long the subject stays hidden and what must remain unchanged when it returns. "The same woman emerges" is weaker than repeating the few details you cannot afford to lose.

Give the camera a position inside the frame

"The camera follows the player" describes movement, but says nothing about composition. I also specify where the player stays in the frame, what event starts the pan, and whether the camera can cross the action axis.

In the basketball shot, the player in red stays in the left third until the pass leaves his hands. Only then does the camera pan with the ball. The basket remains on the same side of the frame.

Generated using Seedance 2.5 on fal.

Prompt

14-second continuous sideline tracking shot in a public high school basketball gym in Indiana, natural real-time speed. One player in a plain red jersey owns the ball at the start; one teammate in a plain white jersey waits near the right side of the lane. 0-5 seconds: the red-jersey player dribbles with the right hand from frame left toward the free-throw line. The camera moves parallel and keeps the red player in the left third of frame, with the hoop visible on the right. 5-8 seconds: the red player plants the left foot and makes one chest pass across frame. The camera does not pan until the ball has fully left both hands. 8-11 seconds: the camera pans right with the airborne ball; the white-jersey player catches it with both hands near the right block and takes one step toward the hoop. 11-14 seconds: the white player makes a right-handed layup, lands on both feet, and the ball drops through the net. Keep the basket on the same side of frame and never cross the court axis. Preserve both players, jersey colors, ball ownership, direction, court markings, and light. No cuts, no slow motion, no extra ball, no duplicated players, no impossible handoff. Audio: sneaker squeaks, two dribbles, pass impact, backboard and net, small gym crowd, no music.

The word "dynamic" is close to useless here. I put the subject in a real part of the frame and tie each camera move to an event the model can see.

Give each reference one job

Uploading several files does not tell the model how to divide responsibility between them. I assign a narrow job to each reference and say which parts of that file should not transfer.

Here, @Image1 owns the product shape and materials. @Image2 owns the kitchen layout and light. The studio background from the product image should not enter the kitchen, and the empty counter in the kitchen image should not redesign the product.

Portable espresso maker reference

@Image1 controls the product design.

Los Angeles kitchen reference

@Image2 controls the kitchen and light.

Generated using Seedance 2.5 on fal.

Prompt

14-second continuous product demonstration in the Los Angeles kitchen from @Image2. @Image1 controls only the exact portable espresso maker: preserve its short cylindrical proportions, matte cobalt-blue shell, black rubber grip ring, circular copper button, and clear lower chamber. Do not copy @Image1's studio background. @Image2 controls only the kitchen layout, oak countertop, white cup, beige towel, plants, window light, and warm daylight. Do not add the studio surface from @Image1. Start on a wide frame matching @Image2 with the product standing to the left of the white cup. 0-4 seconds: the camera makes a slow, level push toward the product while it remains still. 4-7 seconds: one natural right hand enters from frame right and presses the copper button once. 7-11 seconds: dark espresso begins flowing into the clear lower chamber; the liquid level rises naturally while the product body stays rigid and unchanged. 11-14 seconds: the hand withdraws and the camera shifts slightly right to place the product and cup side by side in the final frame. Keep the cup, towel, plants, countertop, product geometry, button position, and lighting consistent. No logo, no text, no extra machine, no extra hands, no cuts, no slow motion. Audio: quiet apartment room tone, soft button click, gentle brewing sound, distant Los Angeles traffic, no music.

Break physical movement into contacts

A result such as "jumps over the cone" leaves most of the movement unspecified. Approach, contact, transfer of force, and recovery give the model a path it can follow.

The skateboard prompt names the tail hitting concrete, the front foot sliding, the wheels landing, and the knees absorbing the impact. It also says how the move ends.

Generated using Seedance 2.5 on fal.

Prompt

12-second continuous low side-tracking shot at the Venice Beach skatepark in late afternoon, natural real-time speed. A skater wearing a faded red T-shirt rides a black skateboard toward one yellow traffic cone. 0-3 seconds: he pushes once with his right foot, places it back on the board, and centers his weight. 3-6 seconds: he bends both knees while approaching the cone; the camera moves parallel and keeps his full body centered. 6-8 seconds: the rear foot snaps the tail against the concrete, the board rises, the front foot slides forward, and both rider and board clear the cone together. 8-10 seconds: the front wheels contact first, then the rear wheels; his knees compress from the landing and his arms correct his balance. 10-12 seconds: he straightens and rides away without another trick. Preserve the same person, shirt, board, cone, direction of travel, and sunlight. Believable wheel rotation, board contact, gravity, and body weight. No cuts, no slow motion, no floating board, no duplicated limbs, no repeated jump. Audio: skateboard wheels on concrete, tail pop, landing impact, distant beach ambience, no music.

I use the same structure for sports, fights, dance, and product handling. The exact limb matters, as does where the weight goes next. I also tell the model how the body or object settles instead of stopping at the most dramatic frame.

I keep coming back to this reference pattern:

@Image1 controls only [identity, product design, clothing, or another invariant].
Do not copy [pose, background, lighting, camera angle, or text] from @Image1.

I keep each job narrow because reference conflicts get messy fast.

Do not make an image and a video reference do the same job

A video reference can carry camera movement, performance, or timing. An image is usually better for the look of a product or character. Those are different jobs.

For the car shot, @Image1 owns the vehicle design. @Video1 contributes only the low parallel tracking motion from the skateboard clip. The prompt explicitly rejects the skater, cone, clothing, and location from the motion reference.

Silver sports car reference

@Image1 controls the car design. The skateboard clip above is @Video1 and controls camera motion only.

Generated using Seedance 2.5 on fal.

Prompt

16-second continuous automotive tracking shot. @Image1 controls only the exact fictional silver sports car: preserve its body shape, silver paint, front light signature, black roof, wheel-spoke design, proportions, vents, and ride height. Do not copy @Image1's sunny San Francisco waterfront or parked composition. @Video1 controls only the low side-tracking camera height, parallel motion, subject framing, and real-time movement rhythm. Do not copy the skater, skateboard, cone, clothing, beach, or concrete setting from @Video1. Place the car on a rain-wet downtown Seattle avenue at blue hour. 0-4 seconds: the car waits at a red traffic light while the camera holds a low front-side angle. 4-11 seconds: the light turns green and the car accelerates smoothly; the camera tracks parallel at door height while keeping the full car centered and sharp. Wheels rotate at the correct speed, suspension settles under acceleration, reflections move across the exact silver body, and water sprays backward from the tires. 11-16 seconds: the car eases into the right lane and maintains speed while the camera falls half a car length behind into a rear three-quarter view. Preserve one car and the same design in every frame. No cuts, no redesign, no extra spoiler, no logo, no copied human subject, no skateboard, no cone, no slow motion. Audio: wet tire noise, restrained electric motor whine, distant traffic, rain on road, no music.

"Use @Video1 as a reference" is too vague. I name the motion I want, then list the subjects, objects, and setting that should stay behind.

Write dialogue as a performance

I block dialogue the same way I block movement. If a prompt gives the actor a line but says nothing about what happens before, during, and after it, the model may pile several actions on top of the speech.

The UGC clip has two short lines. The actor stays silent while turning the lid and waiting for the click. The leak test gets its own time block.

Generated using Seedance 2.5 on fal.

Prompt

12-second vertical 9:16 handheld phone video in a sunlit Austin apartment kitchen, natural real-time performance, one continuous take. A woman in her late twenties stands at the counter holding a small plain insulated travel cup with no logo. 0-3 seconds: she looks into the phone camera, briefly raises the cup, and says, "I bought this for the commute, but I use it at home every day." 3-6 seconds: she stops speaking, looks down, turns the lid one quarter turn with both hands, and waits until it clicks. 6-9 seconds: she tilts the sealed cup sideways over the sink for two seconds; no liquid leaks. 9-12 seconds: she returns the cup upright, looks back at the camera, and says, "That is the whole reason." Her mouth moves only during her own lines. Keep the same face, hair, shirt, cup, lid, kitchen, hand count, and daylight throughout. Subtle natural phone-camera shake, ordinary skin texture, no beauty filter, no cuts, no zoom, no slow motion, no extra products, no text, no logo, no music. Audio: clean lip-synced speech, quiet apartment room tone, soft lid click, distant traffic.

I put the spoken line in straight quotes and give it a start and end time. With more than one person, I also say who is speaking and who keeps their mouth closed.

Describe where fluid motion ends

"Syrup pours over the pancakes" starts an action. It does not say when the pour stops, where the syrup collects, or what should remain in the final frame.

The macro prompt follows the syrup through contact, flow, pooling, and separation. After the pitcher leaves, the last three seconds are just the existing motion settling. Nothing new starts.

Generated using Seedance 2.5 on fal.

Prompt

15-second continuous 21:9 macro food shot at a roadside diner breakfast counter, natural real-time speed. A stack of three plain buttermilk pancakes sits on one white ceramic plate with a square of butter centered on top and six blueberries around the base. 0-4 seconds: a small glass syrup pitcher enters from frame upper right and tilts above the butter; the camera begins a slow ten-degree clockwise arc around the plate. 4-9 seconds: one continuous amber stream lands on the butter, divides around it, and runs down the pancake edges under gravity. Syrup pools on the plate and does not flow uphill. 9-12 seconds: the pitcher tilts upright; the stream narrows, stretches, then breaks cleanly before the pitcher leaves frame. 12-15 seconds: the butter softens slightly and slides only a few millimeters before stopping; the syrup continues settling into one pool. Preserve the same three pancakes, butter, six blueberries, white plate, counter, and warm window light. No cuts, no slow motion, no extra fruit, no changing food count, no floating liquid, no text, no logo. Audio: quiet diner room tone, soft glass movement, syrup landing on food, distant dishes, no music.

For liquid, cloth, hair, smoke, or particles, name the direction of force and the state left behind after the motion stops. "Realistic physics" does not contain that information.

Use the final frame to continue a shot

When I need a clean continuation, I extract the final frame of the first clip and use it as the image reference for the next generation. That gives the second prompt a real starting state instead of asking the model to remember one.

@Image1 below is the actual final frame from the pancake video. The syrup pour is already over. The continuation starts from this frame and introduces the fork without replaying the pitcher action.

Final frame from the pancake clip

@Image1 is the exact final frame of the previous clip.

Generated using Seedance 2.5 on fal.

Prompt

Use @Image1 as the exact first frame and continue forward from that moment. Preserve the same stack of three buttermilk pancakes, softened square of butter, six blueberries, white ceramic plate, amber syrup pool, counter, warm diner light, 21:9 framing, macro lens, and camera position. Do not replay or recreate the syrup pour that happened before this frame. 0-4 seconds: hold the same composition while the syrup pool makes only small natural settling movements. 4-8 seconds: one clean stainless-steel fork enters from frame right with no hand visible, presses into the front edge of the top pancake, and separates one bite-sized piece. 8-10 seconds: the fork lifts the same piece upward; one thin syrup strand stretches from the piece to the stack. 10-12 seconds: the syrup strand narrows and breaks, and the fork exits toward frame right with the piece. The remaining stack stays on the plate with one small missing bite at the front. No new pour, no pitcher, no hand, no extra fork, no extra pancakes, no changing blueberry count, no cut, no slow motion, no text, no music. Audio is quiet diner room tone with a soft fork contact and sticky syrup separation.

I do not retell the first clip in the continuation prompt. I only describe what has already finished, what stays fixed, and the first new action.

The prompt structure I use

I do not fill out this whole template every time. A two-second action does not need a timeline or a long continuity list. I keep the sections that solve a real problem in the shot and delete the rest.

FORMAT
[Duration], [aspect ratio], [single take or cuts], [real-time or specified speed]

REFERENCE ROLES
@Image1 controls only [identity, product, wardrobe, environment, or another invariant].
Do not copy [pose, background, lighting, text, or camera angle] from @Image1.

@Video1 controls only [camera path, performance, motion rhythm, or timing].
Do not copy [subject, wardrobe, objects, or setting] from @Video1.

STARTING STATE
[Character positions, held objects, camera position, environment state]

TIMELINE
0-X seconds: [First action]
X-Y seconds: [Second action beginning from the first action's result]
Y-Z seconds: [Final action and settling state]

CAMERA
[Camera path, screen-space position, when movement starts and stops]

CONTINUITY
[Identity, object, direction, clothing, geometry, and state invariants]

AUDIO
[Dialogue, room tone, contact sounds, music or silence]

ENDING STATE
[Exact position of the character, objects, and camera in the final frame]

CONSTRAINTS
[Cuts, slow motion, repetition, extra objects, text, logos, or unwanted transfers]

Using Seedance 2.5 through the API

The text-to-video endpoint supports durations from 4 to 30 seconds. You can also set duration to auto. It supports 480p and 720p output, and native audio is enabled by default.

import { fal } from "@fal-ai/client";

const result = await fal.subscribe(
  "bytedance/seedance-2.5/text-to-video",
  {
    input: {
      prompt,
      duration: "30",
      aspect_ratio: "16:9",
      resolution: "720p",
      generate_audio: true,
    },
  }
);

console.log(result.data.video.url);

The reference-to-video endpoint accepts up to 30 images, 10 videos, and 10 audio files. The combined file count cannot exceed 50. References are numbered by upload order as @Image1, @Video1, and @Audio1.

import { fal } from "@fal-ai/client";

const result = await fal.subscribe(
  "bytedance/seedance-2.5/reference-to-video",
  {
    input: {
      prompt,
      image_urls: [productImageUrl, environmentImageUrl],
      video_urls: [cameraMotionVideoUrl],
      duration: "16",
      aspect_ratio: "16:9",
      resolution: "720p",
      generate_audio: true,
    },
  }
);

console.log(result.data.video.url);

Each video or audio reference must be between 1.8 and 30.2 seconds. The total reference duration within either modality cannot exceed 30.2 seconds. I check the upload order before sending the prompt because the right description attached to the wrong reference number is still wrong.

Change one thing when a generation misses

If 24 seconds of a 30-second result already work, I do not rewrite the whole prompt. I rewrite the physical state where the bad action begins and the state where it should end. If the camera is right, I leave the camera line alone. If the product is right, I keep the product reference and its invariant wording.

Changing the camera, character, location, and timing in one pass makes the next result hard to read. One controlled change tells you whether the edit worked.

Run Seedance 2.5 on fal

Use the text-to-video endpoint for a shot built from a prompt. Use reference-to-video when images, video, or audio need to control part of the result.

about the author
Ilker
Ilker is a creative engineer at fal-ai who specializes in training loras, testing models, pushing the boundaries of what's possible with AI video and image generation.

Related articles