If you're an advanced user looking to get the most out of Seedance 2.5, this documentation might be helpful for you. In this document, we cover how to use Seedance 2.5 and all the prompting techniques and capabilities of this new model.
Seedance 2.5 supports both text-only video generation and creation with image, video, and audio references. It can also edit existing videos. Start by stating what you want to generate, then add reference materials, event progression, visual treatment, and audio as needed.
1. Basic Prompting Techniques
1.1 The Core Prompt Formula
Prompts can flexibly combine the following elements:
💡 Subject + Action or Event + Scene and Environment (Optional) + Visual Style (Optional) + Camera Movement/Cut (Optional) + Audio (Optional)
- Subject + Action or Event: State who or what is doing what. This is the foundation of the video.
- Scene and Environment: Describe the location, time, weather, spatial relationships, and background state.
- Visual Style: Describe lighting, color, materials, image texture, or the overall mood.
- Camera Movement/Cut: Describe shot size, camera angle, camera movement, the focus subject, and shot transitions.
- Audio: Describe dialogue, voice characteristics, ambience, sound effects, and music.
Basic Template
<Subject> performs <primary action or event> in <scene and environment>.
The visuals feature <visual style>.
Use <shot size, camera angle, camera movement, or cuts>.
Audio includes <dialogue, ambience, sound effects, or music>.
Example
A ceramic artist finishes a pale blue cup in a studio at dawn, lifts it from the wheel, and places it in the center of a wooden shelf.
Soft morning light enters through the window. The wet clay has a delicate sheen, and the workbench remains tidy.
Begin with a medium shot of the wheel-throwing process, slowly push in toward the cup's surface texture, then cut to a frontal view of the shelf.
Retain the low hum of the pottery wheel, the friction of clay, and subtle indoor ambience.
You may omit any components you do not need. Generation parameters do not need to be included in the prompt; configurable parameters are set on the generation page or through the API.
1.2 When Using Reference Materials, Prepare Them and Define Their Roles
Material Quantity and Selection
Seedance 2.5 can combine up to 50 reference materials. Each material type follows the input limits below. The recommended ranges are intended to improve generation stability; they are not capability limits.
| Material Type | Input Limit | Recommended Range |
|---|---|---|
| Images | Up to 30 images, each no larger than 4K | Prefer 1-8 distinct subjects across subject-reference images |
| Videos | Up to 10 videos, with a combined duration of no more than 30 seconds | Prefer 1-5 distinct subjects and 5-10 seconds per subject video |
| Audio | Up to 10 audio clips, with a combined duration of no more than 30 seconds | Keep only dialogue, voice characteristics, ambience, or music directly relevant to the task |
| Video Editing | A source video may be used together with reference images | Prefer a source video under 20 seconds and 1-5 reference images |
You may still try ranges above these recommendations: 9-12 subjects in subject images, 6-10 subjects in subject audio/video, or 6-8 reference images for video editing. Stability may decrease as the number of materials grows. If more than five subjects also require multiple views, place different views in separate images; independent view images are usually more stable than combining several views into one collage.
Define Each Material's Role
After uploading reference materials, specify exactly what each one contributes. Add exclusions when people, backgrounds, or compositions in a material could be unintentionally carried into the output. Material mappings must be written in the prompt: do not rely only on text labels inside images, and do not make the model infer which person, prop, or scene each material represents.
Reference Role Template
@Image 1 defines <subject>'s <appearance, clothing, structure, or material>.
@Video 1 defines <motion, camera movement, or pacing>.
@Audio 1 defines <character or sound type>'s <voice, dialogue, ambience, or music>.
<Subject> completes <primary action or event> in <scene>.
The visuals feature <visual style>, with <camera treatment>.
Example
@Image 1 defines the ceramic artist's facial features, hairstyle, and dark green apron. Do not use the image background.
@Image 2 defines the wooden workbench, window placement, and morning light of the pottery studio. Do not use the people in the image.
@Video 1 defines the pacing of throwing clay with both hands, lifting the cup, and placing it down. Do not use the person's identity, clothing, or scene from the video.
The ceramic artist finishes a pale blue cup in the pottery studio at dawn, lifts it from the wheel, and places it in the center of a wooden shelf.
Begin with a medium shot of the wheel-throwing process, then slowly push in toward the cup's surface texture. Retain the sound of the wheel, the friction of clay, and indoor ambience.
If several images show different views of the same person or product, state this explicitly:
@Image 1 defines the front view of the same folding desk lamp.
@Image 2 defines the left-side structure of the same folding desk lamp.
@Image 3 defines the right-side structure of the same folding desk lamp.
@Image 4 defines the rear structure of the same folding desk lamp.
All four images define one folding desk lamp. The output must contain only one lamp throughout.
When a reference video already defines the motion, camera movement, and sequence accurately, state only which attributes to inherit; there is no need to restate every action. Repeating the motion description may conflict with the reference itself. A blockout video mainly provides motion and spatial structure, so the prompt must still define the intended subjects, scene, action, and visual style.
1.3 Special Syntax for Audio and Text
Prompts can be written entirely in natural language. When you need to distinguish music, sound effects, dialogue, and subtitles more explicitly, use the following syntax:
| Content | Syntax | Example |
|---|---|---|
| Music | () | (Soft, rhythmic piano music plays in the background) |
| Sound Effects | <> | <A bell rings in the distance> |
| Dialogue | {} | {Hello, welcome back.} |
| Subtitles | 【】 | 【Chapter One: Departure】 |
Dialogue Language Reinforcement
When dialogue is not in Chinese, specify the language before the line:
The girl says softly in Japanese: {もう大丈夫です}
If the dialogue text is in English but the model speaks it in Chinese, or if you need a specific regional variety, reinforce the language before the line. Use this formula:
💡 Dialogue Language + Regional Variety or Accent + Delivery Style + Speaker +
{Dialogue}
Dialogue language: American English. The girl says in natural, conversational American English: {I thought you weren't coming.}
Dialogue language: authentic Los Angeles English. The young man says in natural Los Angeles vernacular: {No way, you actually made it.}
2. New Capabilities in Dreamina Seedance 2.5
2.1 Multi-Reference Creation: Tell the Model Which Materials to Use in Each Scene
Seedance 2.5 supports up to 50 reference materials. When many materials are provided, the goal is not to put every reference into one sentence, but to define the relationships among characters, props, scenes, actions, and audio.
Use the following order:
💡 Define Each Material's Role → Map Subjects → Group by Type → Create Subject Profiles → Select References by Scene
Step 1: Name and Map Each Subject Individually
Bind each person, product, and prop to its reference material separately:
<Character A> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing.
<Character B> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing.
<Prop A> corresponds to @Image 3. Use only the structure, material, and color.
<Scene A> references @Image 4. Use only the spatial layout, architecture, and lighting. Do not use the people in the image.
Do not write, "@Images 1 through 4 define four characters respectively." That wording does not state which image corresponds to which character.
Step 2: Group Materials by Type
[Characters]
<Conservator> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing.
<Registrar> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing.
<Exhibition Installer> corresponds to @Image 3. Use only the appearance, hairstyle, and clothing.
<Guide> corresponds to @Image 4. Use only the appearance, hairstyle, and clothing.
Do not interchange the four characters' appearances, clothing, actions, positions, or dialogue.
[Props]
<Sample Case> corresponds to @Image 5 and belongs only to <Conservator>.
<Record Board> corresponds to @Image 6 and belongs only to <Registrar>.
[Scenes]
<Conservation Lab> references @Image 7. Use only the space, materials, and lighting.
<Gallery> references @Image 8. Use only the space, materials, and lighting.
[Motion and Audio]
@Video 1 defines the motion of <Conservator> opening <Sample Case>. Do not use the person or scene from the video.
@Audio 1 defines <Guide>'s voice and specified dialogue.
Step 3: Create a Centralized Profile for Important Subjects
When the same character uses several references across multiple scenes, add a subject profile:
[Subject Profile: Conservator]
Appearance and clothing: @Image 1.
Fixed prop: <Sample Case> from @Image 5.
Locations: <Conservation Lab> and <Gallery>.
Motion references: the case-opening motion from @Video 1 and the sample-placement motion from @Video 2.
Do not use: other characters' clothing. Do not give this character <Record Board> or guide equipment.
Step 4: Select References by Scene
Scene 1 | Inspection in the Conservation Lab
Use: <Conservator>, <Sample Case>, <Conservation Lab>, and the case-opening motion from @Video 1.
Event: <Conservator> opens <Sample Case> at the workbench and inspects the sample inside.
End state: <Conservator> remains on the inner side of the workbench. <Sample Case> stays beside the conservator's right hand, which is on the left side of the frame.
Scene 2 | Registration in the Gallery
Use: <Registrar>, <Record Board>, and <Gallery>.
Event: <Registrar> checks the number on <Record Board> beside the display case.
End state: <Registrar> still holds <Record Board> with both hands. No other character enters the display-case area.
The goal of multi-reference creation is to help the model select the correct materials for the current scene, not to make every material appear at the same time.
2.2 30-Second Videos: Organize Events with Stages and End States
When a video contains several events, divide the story into consecutive stages. Give each stage only one primary state change, and state what should be directly visible at the end of that stage.
Long-Video Template
[Generation Goal]
Generate a <video type>. The central subject is <subject>, and the primary event is <story summary>.
[Stage 1]
Initial state: <initial state of characters, props, and scene>.
Primary event: <one primary action or event>.
End state: <character positions, prop ownership, or visible scene state>.
[Stage 2]
Continue from the previous stage: <state that must remain unchanged>.
Primary event: <one primary action or event>.
End state: <observable state>.
[Stage 3]
Primary event: <closing event>.
End state: <final visible state>.
[Maintain Consistency]
Keep <character identity, number of characters, clothing, prop ownership, spatial direction, and audio relationships> consistent.
Example
[Generation Goal]
Generate an instructional video showing a flower shop's order-packing process. <Florist> and <Store Assistant> arrange, wrap, and hand off a bouquet together.
[Stage 1]
Initial state: <Florist> stands behind the workbench. Loose flower stems, scissors, and wrapping paper lie on the tabletop.
Primary event: <Florist> arranges the stems and trims them to length.
End state: <Florist> holds the bouquet in the left hand, and the scissors are back on the right side of the workbench.
[Stage 2]
Continue from the previous stage: both characters retain the same identities and clothing, and <Florist> still holds the bouquet.
Primary event: <Store Assistant> unfolds the wrapping paper. <Florist> places the bouquet inside and ties it with a green ribbon.
End state: the wrapped bouquet lies flat in the center of the workbench, with the ribbon bow facing the camera.
[Stage 3]
Primary event: <Store Assistant> picks up the bouquet and places it on the pickup shelf.
End state: the bouquet is centered on the pickup shelf, and both characters stand behind the workbench inspecting the finished order.
[Maintain Consistency]
Keep <Florist> and <Store Assistant>'s identities, clothing, workbench orientation, scissors position, and bouquet ownership consistent.
Timestamps and Pacing Control
For ordinary narratives, use stages by default. Use one-second precision only when you need to control a critical handoff, entrance or exit, transition, or explicit beat. Use time ranges to allocate pacing, exact time points for a single key event, and relative timing to describe a delay between events.
0-5 seconds: Show an empty wooden display table. A hand places a white ceramic plate on it. End state: the hand has left the frame, and only the white plate remains in the center of the table.
5-10 seconds: Remove the white plate, then place a clear glass on the table. End state: only the clear glass remains in the center of the table.
10-15 seconds: Remove the clear glass, then place a green ceramic vase on the table. End state: only the green vase remains in the center of the table.
| Pattern | Example |
|---|---|
| Time range | 0-3 seconds... 3-7 seconds... 7-12 seconds... |
| Exact time point | At 5 seconds, the camera whip-pans rapidly to the left and completes the transition. |
| Relative timing | Three seconds after the character presses the button, the room lights gradually turn off. |
Time ranges should be consecutive and non-overlapping. They represent an event's time budget, not a precise edit point, so actions may occur slightly before or after a boundary. Too little content in a range gives the model more freedom; too much can cause excessive cutting or omitted events. Do not use timestamps to demand frequencies such as "complete three actions in one second."
2.3 Parameter Rules for Editing, First/Last-Frame, and Extension Tasks
Video editing, first-frame or first-and-last-frame generation, and video extension automatically lock some generation parameters based on the input materials.
| Task Type | Aspect Ratio | Duration |
|---|---|---|
| Video editing | Automatically preserves the input video's aspect ratio; cannot be set separately | Automatically preserves approximately the input video's duration; cannot be set separately. Input-frame processing may introduce a difference of up to approximately 0.3 seconds |
| First-frame or first-and-last-frame generation | Automatically uses the first image's aspect ratio. The first and last images should use the same aspect ratio to avoid stretching the last frame | Can be set |
| Video extension | Automatically preserves the input video's aspect ratio; cannot be set separately | Can be set |
Parameters automatically locked for these tasks cannot be specified separately on the generation page or through the API. Other parameters depend on the options currently available.
2.4 Video Editing: Define the Master Video, Edit Scope, and Content to Preserve
When editing an existing video, first define the source video as the sole editing master. Then specify the edit target, edit scope, target material, and content to preserve. The output automatically preserves the input video's aspect ratio and approximately preserves its duration; neither can be set separately. Input-frame processing may introduce a difference of up to approximately 0.3 seconds, usually because of transition-frame handling, while the overall content and event order remain substantially unchanged.
General Editing Pattern
Template:
[Edit Goal]
Edit @Video 1. Within <the entire video or a specific time range>, <add, remove, replace, or adjust> <visual object, region, or audio category>.
[Source Video Role]
@Video 1 is the sole editing master. It defines <characters, scene, actions, composition, camera movement, occlusion relationships, audio, and event order>.
[Target Material Role]
@Image 1 or @Audio 1 defines <specified attributes of the target object or sound>.
[Edit Scope]
Modify only <object, region, time range, or audio category>.
[Content to Preserve]
Keep <visual content, motion, audio, and timing relationships that must not change> from @Video 1.
Example:
[Edit Goal]
Edit @Video 1. Only from 4-7 seconds, change the cool blue light on the right wall to warm orange light.
[Source Video Role]
@Video 1 is the sole editing master. It defines the character, room layout, actions, composition, camera movement, audio, and event order.
[Edit Scope]
Change only the light color on the right wall and the area it illuminates. Allow the character's skin tone to respond naturally to the environmental light.
[Content to Preserve]
Keep the character's identity, clothing, expression, position, motion, room structure, camera movement, dialogue, and ambience from @Video 1.
Subject Replacement
Template:
[Edit Goal]
Edit @Video 1. Change only <original object> to <target object>.
[Source Video Role]
@Video 1 is the sole editing master. It defines the original scene, camera position, camera movement, motion path, occlusion relationships, and event order.
[Target Reference Role]
@Image 1 defines <target object>'s <appearance, structure, or material>. Do not use <irrelevant background, people, or composition>.
[Edit Scope]
Modify only <specific object and area>. The entire video contains <number> target object(s). Do not modify <content to preserve>.
[Timeline Inheritance]
<Target object> inherits every appearance, motion, occlusion, and exit of <original object>, including timing, duration, path, and speed changes.
Except for the object or area explicitly modified above, keep all other people, props, scene content, camera movements, cuts, and event order from @Video 1 unchanged.
Example:
[Edit Goal]
Edit @Video 1. Replace only the yellow folding desk lamp with the white folding desk lamp in @Image 1.
[Source Video Role]
@Video 1 is the sole editing master. It defines the desk, books, hand movements, camera position, camera movement, occlusion relationships, and event order.
[Target Reference Role]
@Image 1 defines only the white folding desk lamp's appearance, structure, and material. Do not use the image's background, composition, or other objects.
[Edit Scope]
Keep exactly one white folding desk lamp throughout the video. Replace only the original yellow folding desk lamp. Do not modify the books, desk, hands, or background.
[Timeline Inheritance]
The white folding desk lamp inherits every appearance, lamp-arm rotation, hand occlusion, and exit of the original yellow folding desk lamp, including timing, path, and speed changes.
Except for the object or area explicitly modified above, keep all other people, props, scene content, camera movements, cuts, and event order from @Video 1 unchanged.
Background Replacement
Template:
[Edit Goal]
Edit @Video 1. Replace only <original background area> with <target environment> from @Image 1.
[Source Video Role]
@Video 1 is the sole editing master. It defines the people, foreground objects, actions, composition, camera movement, and event order.
[Target Reference Role]
@Image 1 defines only <target environment>'s spatial layout, materials, depth of field, ambient color, and lighting direction. Do not use the people or foreground objects in the image.
[Edit Scope]
Modify only <background outside the subject's silhouette>. Do not modify <subject identity, facial features, hairstyle, clothing, expression, position, size, or motion>.
[Timeline Inheritance]
Keep the character actions and occlusion relationships from @Video 1. Except for the object or area explicitly modified above, keep all other people, props, scene content, camera movements, cuts, and event order from @Video 1 unchanged.
Example:
@Video 1 is the sole editing master. It defines the people, actions, composition, camera treatment, and event order.
@Image 1 provides only the spatial layout, depth of field, ambient color, and lighting direction of a daylit glass greenhouse. Do not use the people in the image.
Replace only the light gray background outside the person's silhouette in @Video 1 with the daylit glass greenhouse from @Image 1.
Keep the person's identity, facial features, hairstyle, clothing, expression, position, size, and arm-raising motion from @Video 1.
Except for the object or area explicitly modified above, keep all other people, props, scene content, camera movements, cuts, and event order from @Video 1 unchanged.
Audio Editing
Dialogue, language, voice, background music, and sound effects can be edited separately. State the speaker or sound category, the intended change, and which other sounds must remain unchanged.
Edit @Video 1. Remove only the original background music. Keep the character dialogue, lip sync, ambience, and action sound effects; preserve the visuals, camera treatment, and editing rhythm from @Video 1.
Edit @Video 1. Change <Presenter>'s spoken language to natural American English while preserving the dialogue content and speaking times. Keep all other character voices, background music, ambience, and visuals from @Video 1.
2.5 Video Extension: Align the Boundary Frame Before Describing New Content
Video extension creates content beyond the boundary of a source video. For a forward extension after the original video, the extension's first frame should continue from the source video's last frame. For a backward extension before the original video, the extension's last frame should connect to the source video's first frame. In addition to the boundary frame, inspect whether the characters, props, background, and events in the extended segment are correct.
Forward Extension (After the Original Video)
First describe the continuous state of the source video's last frame, then describe what happens afterward.
Basic Template:
@Video 1 is the source video to extend forward.
Extend @Video 1 forward. The first frame of the extended segment directly continues from the last frame of @Video 1. Maintain continuity in <subject pose and orientation>, <prop position>, <background and spatial relationships>, <camera position and composition>, <lighting>, and <motion direction>.
Then, <describe the new action, event, camera treatment, or audio to add>.
Throughout the extension, maintain continuity in <character identity and clothing>, <key props>, <background layout>, and <axis of action>.
Keep each subject as the same continuous instance throughout: do not duplicate or split it, and keep the person's appearance or the object's number of parts stable.
Example:
@Video 1 is the source video to extend forward.
Extend @Video 1 forward. The first frame of the extended segment directly continues from the last frame of @Video 1. Maintain the same locked-off medium shot, the orange paper airplane's position and orientation, the classroom-window background, the afternoon lighting, and its movement toward the right side of the frame.
Then, the orange paper airplane continues gliding toward the right and exits the frame while the white curtain beside the window sways slightly. Keep the camera and classroom background in the state established by the source video's last frame.
With Additional Reference Materials, Template:
@Image 1 defines <Character A>'s facial features.
@Image 2 defines <Character A>'s clothing.
@Image 3 defines <key prop>'s structure and material.
@Video 1 is the source video to extend forward.
Extend @Video 1 forward. The first frame of the extended segment directly continues from the last frame of @Video 1. Maintain continuity in <boundary-frame state>.
Then, <Character A uses the key prop to complete a new action or event>.
Throughout the extension, maintain continuity in <character identity and clothing>, <key prop>, <background layout>, and <axis of action>.
Keep each subject as the same continuous instance throughout: do not duplicate or split it, and keep the person's appearance or the object's number of parts stable.
Example:
@Image 1 defines <Gardener>'s facial features.
@Image 2 defines <Gardener>'s light green work apron.
@Image 3 defines <Wicker Flower Basket>'s structure and material.
@Video 1 is the source video to extend forward.
Extend @Video 1 forward. The first frame of the extended segment directly continues from the last frame of @Video 1. Maintain the greenhouse workbench, <Gardener>'s position, and <Wicker Flower Basket>'s position.
Then, <Gardener> lifts <Wicker Flower Basket> with both hands and places it on the middle shelf of the wooden rack behind them.
Throughout the extension, maintain continuity in <Gardener>'s face, apron, greenhouse layout, and camera direction.
Backward Extension (Before the Original Video)
First describe what happens before the source video begins, then define the source video's first frame as the explicit end state of the extended segment. Writing only "then connect to the source video" may introduce later characters or effects too early, or cause the image to change again after reaching the target state.
Basic Template:
@Video 1 is the source video to extend backward.
Extend @Video 1 backward. Before the source video begins, <describe the preceding action, event, camera treatment, or audio>.
The last frame of the extended segment naturally connects to the first frame of @Video 1: <subject pose and orientation>, <prop position>, and <background and spatial relationships>. Match the <camera position and composition>, <lighting>, and <motion direction> of @Video 1's first frame.
Throughout the extension, maintain continuity in <character identity and clothing>, <key props>, <background layout>, and <axis of action>.
Keep each subject as the same continuous instance throughout: do not duplicate or split it, and keep the person's appearance or the object's number of parts stable.
Example:
@Video 1 is the source video to extend backward.
Extend @Video 1 backward. Before the source video begins, show an empty establishing shot of the same glass greenhouse. Morning mist drifts slowly near the floor, the overhead shade rises gradually, and no people are present yet.
The last frame of the extended segment naturally connects to the first frame of @Video 1. Match the greenhouse's central aisle, planting tables on both sides, glass frame, soft morning light, and locked-off wide composition. At the end, the shade is fully raised, the aisle is empty, and the leaves still sway slightly.
With Additional Reference Materials, Template:
@Image 1 defines <Character A>'s facial features.
@Image 2 defines <Character A>'s clothing.
@Image 3 defines <key prop>'s structure and material.
@Video 1 is the source video to extend backward.
Extend @Video 1 backward. Before the source video begins, <Character A completes a preceding action or event>.
The last frame of the extended segment naturally connects to the first frame of @Video 1: <Character A's pose and orientation>, <key prop's position and state>, and <other characters' positions>. Match the <background and spatial relationships>, <camera position and composition>, <lighting>, and <motion direction> of @Video 1's first frame.
Keep each subject as the same continuous instance throughout: do not duplicate or split it, and keep the person's appearance or the object's number of parts stable.
<Materials that should appear only after the source video begins> must not appear early in the backward extension.
Example:
@Image 1 defines <Curator>'s facial features.
@Image 2 defines <Curator>'s dark blue work jacket.
@Image 3 defines <Wooden Display Case>'s structure and material.
@Image 4 defines the gray workwear of two <Exhibition Assistants>.
@Image 5 defines <Exhibition Preparation Room>'s space and lighting.
@Video 1 is the source video to extend backward.
Extend @Video 1 backward. Before the source video begins, <Curator> walks to the workbench, picks up the closed <Wooden Display Case>, and opens its lid.
The last frame of the extended segment naturally connects to the first frame of @Video 1. <Curator> stands in the center of the frame, holding the open <Wooden Display Case> with both hands. The two <Exhibition Assistants> stand behind the curator, one on each side. Match the vertical frontal medium shot, workbench position, preparation-room background, and morning light from the left established by @Video 1's first frame.
Boundary frames should connect naturally at a visual level; this does not mean they will be pixel-identical. During review, inspect both sides of the boundary and the complete extended segment.
3. Advanced Techniques
3.1 Keyframes, Storyboards, and Blockout References
First and Last Frames with Additional References
In multimodal reference mode, you can state in the first line that @Image 1 is the first frame and @Image 2 is the last frame. There is no need to switch to a separate first/last-frame mode. The system locks the output aspect ratio to the first image, while duration is set on the generation page or through the API. The first and last images should use the same aspect ratio; mismatched ratios may stretch the last frame. Additional images can still define characters, props, scenes, and materials.
Basic Template:
@Image 1 is the first frame. It defines the opening composition, subject position, pose, prop state, scene, and camera direction.
@Image 2 is the last frame. It defines the ending composition, subject position, pose, prop state, scene, and camera direction.
@Image 3 defines <Subject A>'s <appearance, clothing, structure, or material>. Do not change the first-frame composition defined by @Image 1 or the last-frame composition defined by @Image 2.
@Image 4 defines <specified attributes> of <Subject B, prop, or scene>. Do not change the first-frame composition defined by @Image 1 or the last-frame composition defined by @Image 2.
<Describe one continuous action or event>.
The video begins naturally from the first frame defined by @Image 1 and reaches the last frame defined by @Image 2 after the continuous action.
Between the first and last frames, maintain continuity in <character identity, prop structure and ownership, scene layout, and camera direction>.
Example:
@Image 1 is the first frame. It defines the opening composition, character positions, poses, tabletop prop states, perfume-workshop scene, and camera direction.
@Image 2 is the last frame. It defines the ending composition, character positions, poses, tabletop prop states, perfume-workshop scene, and camera direction.
@Image 3 defines <Perfumer>'s face, hairstyle, and dark green apron. Do not change the first-frame composition defined by @Image 1 or the last-frame composition defined by @Image 2.
@Image 4 defines <Glass Perfume Bottle>'s shape, material, and label position. Do not change the first-frame composition defined by @Image 1 or the last-frame composition defined by @Image 2.
Starting from the first-frame pose, <Perfumer> picks up a dropper and <Glass Perfume Bottle>, drips amber fragrance oil into the bottle, swirls it gently, closes the stopper, places the finished bottle in the center of the table, and naturally reaches the last frame defined by @Image 2.
Between the first and last frames, maintain continuity in <Perfumer>'s identity and clothing, bottle count and structure, wooden-table layout, warm side lighting, and camera direction.
Describe each anchor image separately. Do not combine them into a sentence such as "@Images 1 and 2 are the first and last frames." The first and last images should use the same aspect ratio. Other references should supplement only their specified attributes and must not replace the first- or last-frame composition.
Multi-Keyframe Sequence Control
When separate images define different stages of a process, begin with "Use @Image 1 through @Image N as keyframes in this order," then describe the key state represented by each image. Independent keyframe images are usually easier to align than several frames combined into one grid. They control stage order and key states; they do not reproduce every frame exactly.
Basic Template:
Use @Image 1 through @Image N as keyframes in this order.
@Image 1 is the first frame. It defines <opening composition, subject position, pose, prop state, and camera direction>.
@Image 2 defines the second keyframe: <visible end state of Stage 1>.
@Image 3 defines the third keyframe: <visible end state of Stage 2>.
@Image N is the last frame. It defines <ending composition, subject position, pose, prop state, and camera direction>.
The video passes through the states defined by @Image 1, @Image 2, @Image 3, and @Image N in order, using continuous action to transition naturally between stages.
Maintain continuity in <subject identity, prop structure and ownership, scene layout, lighting, and axis of action> throughout.
Example:
Use @Image 1 through @Image 4 as keyframes in this order.
@Image 1 is the first frame. It shows an orange paper airplane resting on the left side of a classroom desk, pointed toward the right, in a locked-off medium shot.
@Image 2 defines the second keyframe: one hand lifts the same orange paper airplane from the desk without changing its direction.
@Image 3 defines the third keyframe: the same orange paper airplane passes the window while the curtain moves slightly to the right.
@Image 4 is the last frame. It shows the same orange paper airplane resting on the middle shelf of the bookcase on the right, still pointed toward the right.
The video passes through the states defined by @Image 1, @Image 2, @Image 3, and @Image 4 in order. Keep flight direction and speed continuous between stages.
Maintain the paper airplane's orange material, size, and folds, as well as the classroom layout, afternoon side lighting, and camera axis.
Storyboard Grids
A storyboard grid communicates the overall story, shot order, and approximate compositions. It is not intended for strict reproduction of every detail in every panel. Prefer no more than 15 panels, use clean line art or simple diagrams, and minimize text labels. State the reading order, then describe each panel's subject action, shot size or camera movement, final visual style, and audio.
Basic Template:
@Image 1 provides an <N-panel storyboard grid> for shot order and approximate composition. Read it <left to right, top to bottom>. Do not use the grid's <line-art style, text labels, or placeholder characters>.
@Image 2 defines <Subject A>'s <appearance and clothing>.
@Image 3 defines <key prop or scene>'s <structure, material, or lighting>.
Shot 1: <shot size, subject action, and scene state>.
Shot 2: <shot size, subject action, camera movement, or transition>.
...
Shot N: <closing action and final visible state>.
The final video uses <visual style>. Audio includes <dialogue, ambience, action sound effects, or music>.
Example:
@Image 1 provides a four-panel pottery-making storyboard for shot order and approximate composition. Read it left to right, top to bottom. Do not use the storyboard's line-art style or text labels.
@Image 2 defines <Ceramic Artist>'s face, short hair, and dark gray apron.
@Image 3 defines <Blue-Glazed Cup>'s proportions, glaze color, and curved handle.
Shot 1: a wide shot establishes a quiet pottery studio with <Ceramic Artist> seated at the wheel.
Shot 2: a side medium shot shows both hands shaping the rotating clay as the cup body takes form.
Shot 3: a close-up shows fingers refining the rim and handle joint while slip moves slowly over the fingertips.
Shot 4: a medium close-up shows the fired <Blue-Glazed Cup> placed on a wooden shelf as <Ceramic Artist> withdraws both hands.
Use a realistic documentary look. Retain the wheel's rotation, wet-clay friction, and studio ambience.
Blockout References and Rendering
Blockout references fall into two categories. Coarse blockouts mainly provide temporal information such as action, paths, blocking, camera movement, cuts, lighting, and sound. Fine blockouts already contain complete structures and are mainly used for re-rendering materials, colors, characters, scenes, and visual style. First determine whether the blockout controls a motion skeleton or a complete model, then use the corresponding prompt structure.
| Type | Best For | Material Requirements | Prompt Focus |
|---|---|---|---|
| Coarse blockout | Simple geometry that previews action, paths, blocking, camera movement, or cuts | Clear relationships between shapes and a complete action sequence; character, prop, and scene images may be added | Map every blockout subject and state which temporal and spatial information to inherit |
| Fine blockout | Complete modeling that needs new characters, materials, colors, scenes, or style | Complete, clean model; avoid path lines, coordinate axes, and camera frustums | Preserve structure, action, and camera treatment while defining the attributes to re-render |
Coarse Blockouts
Use a coarse blockout to lock action paths, motion direction, blocking, entrances and exits, camera paths, cut points, lighting changes, and sound rhythm. Map each geometric object separately to its final subject or prop. Additional images can define the appearance of characters, props, and scenes.
| Blockout Information | What to State in the Prompt |
|---|---|
| Path | Action trajectory, motion direction, subject blocking, and entrance/exit order |
| Camera movement | Camera position, path, direction, and speed changes |
| Lighting | Light direction, brightness changes, and when those changes occur |
| Cuts | Cut positions and the subject/composition before and after each cut |
| Audio | Whether to inherit dialogue, music, ambience, or action sound effects |
Prefer simple geometry with clear relationships. Arms, wings, and other appendages should be used only when the action sequence is complete; otherwise they may cause stiff motion or structural misinterpretation.
Basic Template:
@Video 1 is a coarse blockout reference. It provides only <motion paths, subject blocking, camera position, camera movement, cuts, lighting changes, sound rhythm, or spatial relationships>. Do not use its blockout appearance, materials, or scene.
<Blockout Subject A> in @Video 1 corresponds to <Subject A>.
<Blockout Subject B or geometric prop> in @Video 1 corresponds to <Subject B or key prop>.
@Image 1 defines <Subject A>'s <appearance, clothing, or structure>.
@Image 2 defines <specified attributes> of <Subject B, key prop, or scene>.
<Subject> completes <primary action or event> in <scene>.
Keep <motion path, blocking, camera movement, cuts, lighting, or sound rhythm> from @Video 1.
The final video uses <characters, scene, materials, and visual style>. Audio includes <dialogue, ambience, or action sound effects>.
Example:
@Video 1 is a coarse blockout reference. It provides only the character's walking path, cart direction, locked-off camera, one push-in, and two cuts. Do not use its gray geometry or empty scene.
The tall cylinder in @Video 1 corresponds to <Guide>.
The rectangular block in @Video 1 corresponds to <Mobile Display Cart>.
@Image 1 defines <Guide>'s face, blue uniform, and name badge.
@Image 2 defines <Mobile Display Cart>'s white metal frame and clear cover.
@Image 3 defines the technology gallery's curved walls, gray floor, and overhead strip lights.
<Guide> pushes <Mobile Display Cart> along the curved wall, stops in front of the central display, and opens the clear cover.
Keep the walking path, subject blocking, push-in direction, and cut points from @Video 1.
Use a bright, realistic museum-documentary style. Retain footsteps, wheel sounds, and gallery ambience.
Fine Blockouts
A fine blockout already contains complete character, prop, or scene structures. Use it to change materials, colors, character appearance, scene, or overall visual style. Keep the blockout clean: remove path lines, coordinate axes, controllers, camera frustums, and other production markers.
Basic Template:
@Video 1 is a fine blockout reference. Preserve <subject structure, action, spatial layout, camera position, camera movement, and cuts>. Do not use its original gray materials or empty background.
@Image 1 defines <subject>'s <character appearance, material, color, or surface details>.
@Image 2 defines <scene>'s <space, materials, lighting, or visual style>.
Re-render <subject> from @Video 1 as <final subject>, and re-render the scene as <final scene>.
Keep <structure, action, camera treatment, and spatial relationships> from @Video 1. Use <materials, colors, and style>. Audio includes <ambience, sound effects, or music>.
Example:
@Video 1 is a fine blockout reference. Preserve the kinetic sculpture's complete structure, three-ring rotation relationship, pedestal position, orbiting camera movement, and cuts. Do not use the gray materials or empty background.
@Image 1 defines the outer ring's brushed-brass material.
@Image 2 defines the inner blades' translucent blue-glass material.
@Image 3 defines a contemporary gallery with white curved walls, a dark gray floor, and soft overhead lighting.
Re-render the ring structure from @Video 1 as a kinetic sculpture made of brass and blue glass, and re-render the scene as a contemporary art gallery.
Keep the structure, rotation rhythm, orbiting camera movement, and cuts from @Video 1. Retain the sculpture's low mechanical rotation sound and quiet interior ambience.
3.2 One-Click Video
One-click video is designed to organize multiple images, or images plus a style-reference video, into a complete video with consistent pacing and visual packaging. State each material's role, image order, amount of motion, editing rhythm, visual treatment, and audio. Do not write only "turn these materials into a video."
💡 Material Roles → Image Order → Motion Amount → Editing Style → Visual Treatment → Audio
Basic Template:
[Material Roles]
@Image 1 is used for <character, product, scene, or opening image>.
@Image 2 is used for <character, product, scene, or process image>.
@Image 3 is used for <character, product, scene, or ending image>.
@Video 1 is used only for <editing rhythm, transitions, subtitle treatment, or music style>. Do not use its character identities or scene (optional).
[Arrangement]
Show the images in <upload order, a specified order, or a model-selected thematic order>.
<State the character, product, location, and event relationships that must remain consistent>.
[Image Motion]
Apply <subtle live motion, parallax, push-in/pull-out, lateral movement, or local action> to each image.
Keep <subject appearance, product structure, text, or background relationships> stable.
[Final Style]
Use <editing rhythm, transition style, subtitle or graphic treatment, and color style>.
[Audio]
Include <dialogue, ambience, sound effects, or music>.
Example:
[Material Roles]
@Image 1 is used for the night-market entrance and opening environment.
@Image 2 is used for <Traveler> walking along the street.
@Image 3 is used for the lantern stall and craft details.
@Image 4 is used for three friends eating together.
@Image 5 is used for the riverside night view and reflections.
@Image 6 is used for the final group photo by the bridge.
@Video 1 is used only for light editing rhythm, hand-drawn stickers, and transition style. Do not use its character identities or locations.
[Arrangement]
Show @Image 1 through @Image 6 in order to form a complete sequence: arrival, street exploration, dinner, riverside walk, and group photo.
Keep the three friends' appearances and clothing consistent. Do not mix their identities.
[Image Motion]
Use slow push-ins and subtle parallax for environment images. Add only natural blinking, head turns, glass-raising, and slight clothing movement to character images.
Keep stall structure, table position, and bridge railing stable.
[Final Style]
Use an upbeat travel-video rhythm. Connect scenes with natural occlusion and similar colors. Keep hand-drawn stickers at the frame edges.
[Audio]
Retain night-market chatter, light dish sounds, and riverside wind, with upbeat but unobtrusive instrumental music.
If image order matters, state the exact sequence. If the model may arrange the materials freely, say that it may organize them by theme. When several characters or products appear, continue to name and bind each one separately.
3.3 Seamless Video Transitions
A seamless video transition generates continuous bridge content between two videos. First identify the before-transition and after-transition videos, then describe the trigger action, camera movement, visual transformation, arrival state, and audio transition.
💡 Before Video → After Video → Trigger Action → Camera Movement → Visual Transformation → Arrival State → Audio
| Transition Method | What to Specify |
|---|---|
| Dive or reverse movement | Camera direction, speed change, and when the next scene begins |
| Character rotation | Pose, rotation direction, and how clothing or background changes continuously |
| Foreground occlusion | When the foreground object fills the frame and the composition that follows |
| Object morph | Corresponding shapes, materials, and the transformation process |
| Push/pull or focus change | Camera movement, focus target, and continuous spatial relationship |
Basic Template:
@Video 1 is the before-transition clip. Use its <ending subject, action, composition, camera direction, and audio>.
@Video 2 is the after-transition clip. Use its <opening subject, composition, camera direction, and audio>.
Keep <character identity, product structure, scene, and primary action> stable in the original portions of @Video 1 and @Video 2.
At the end of @Video 1, <subject or foreground object> triggers the transition through <action>.
The camera <movement direction and speed change>, while <shape, material, light, or space> gradually transforms into <corresponding element> at the start of @Video 2.
The transition ends naturally at @Video 2's opening composition, preserving continuity in <subject position, camera direction, and motion trend>.
Audio transitions smoothly from <before audio> to <after audio>.
Example:
@Video 1 is the before-transition clip. Use its rainy night street, red umbrella, slow push-in, and rain sound.
@Video 2 is the after-transition clip. Use its circular gallery skylight, upward camera movement, and quiet interior reverberation.
Keep the people, street, gallery structure, and primary actions in the two original videos stable.
At the end of @Video 1, the red umbrella approaches the camera and gradually fills the entire frame, triggering the transition.
The camera continues moving forward. The umbrella's circular edge gradually becomes the skylight's metal ring, and the red fabric transitions into white daylight passing through the skylight.
The transition ends naturally at @Video 2's upward-looking opening composition, with the camera movement changing smoothly from forward motion to an upward rise.
The rain gradually fades into footsteps reverberating inside the gallery.
The goal of a seamless transition is visual and audio continuity. A prompt may ask to preserve the primary content of both source videos, but a generated bridge is not a pixel-identical edit splice.
3.4 Emotional Direction and Observable Performance
Emotion words such as "tense," "warm," or "oppressive" communicate an overall direction, but they leave more room for interpretation in the performance. For more stable control of acting, add directly visible or audible cues such as eye movement, brow tension, mouth movement, breathing, gaze direction, and hand movement.
You do not need to list every facial detail. For a single emotional transition, two to four clear cues are usually enough. Use event-triggered stages only when the emotion changes several times.
Single Emotional Transition: Default Structure
The overall emotion shifts from <starting emotion> to <ending emotion>.
After <triggering event>, <subject> first shows <immediate observable reaction>.
Then, <eyes, brows, mouth, breathing, gaze, or hand movement> gradually <changes>.
Finally, <subject> expresses <target emotion> through <restrained or explicit outward behavior>.
Multi-Stage Emotion: Progress Through Triggering Events
When <subject> hears or sees <first triggering event>, <first observable reaction>.
When <second triggering event> occurs, <change in expression, gaze, or breathing>.
After confirming <critical information>, the emotion that <subject> tries to restrain or conceal gradually becomes visible through <observable behavior>.
Finally, <subject's final action, expression, or manner of speaking>.
Example
Applause marking the end of the performance comes from behind the stage. The young actor's fingers suddenly stop on the program, the gaze turns slowly toward the curtain, and the shoulders remain tense.
After confirming that the curtain call is over, the actor exhales softly. The shoulders gradually relax, a restrained smile appears, and the eyes slowly well with tears, but the actor never turns to leave.
3.5 Professional Cinematography Terms
Basic camera language and popular camera techniques can be written directly in the prompt. When a term is uncommon, may have several interpretations, or requires precise control, also state which subject it applies to, how the image changes, and the intended visible result.
Basic Camera Language
| Type | Common Supported Terms |
|---|---|
| Shot size | extreme wide shot, wide shot, medium shot, close-up, extreme close-up |
| Camera movement | push in, pull out, pan, lateral move, follow shot, orbit, dive, dolly out, tilt up, handheld shake |
| Camera position and viewpoint | low angle, overhead view, first-person view |
Popular Camera Techniques
The following popular camera techniques can be used directly. If the frame contains several subjects, still state which subject the camera follows or revolves around, where the movement begins, and where it ends.
| Technique | What to Specify |
|---|---|
| One-take shot | The subjects, spaces, and events the continuous camera passes through in order |
| Dolly zoom | The subject size to preserve and whether the background appears to move closer or farther away |
| Aerial view | Viewing height, movement direction, and the environmental area to reveal |
| FPV | First-person flight or traversal path, speed, and turns |
| Bullet time | The action to freeze or slow down and the camera's orbit direction |
| Handheld camera | The subject being followed and the amount of shake |
| Bounce speed ramp | Where the action accelerates, decelerates, or rebounds, and its final resting state |
Uncommon Cinematography Terms
For a niche term, a term with inconsistent industry usage, or a term the model may not recognize, keep the term itself and translate it into a directly observable visual change.
💡 Cinematography Term + Target Subject + Visual Change + Foreground/Background Relationship + Direction or Speed
Rack focus: shift focus smoothly from the leaves in the foreground to the person in the background. The leaves gradually blur while the person's face changes from soft to sharp.
For a precise transition, also state the trigger time, occluding object, camera direction, transition method, and the composition or motion trend that should continue afterward.
Cinematography Examples
Example 1. Shallow-depth-of-field portrait: keep <Pastry Chef>'s eyes and face sharp while the glass jars and lights in the background become soft, circular bokeh.
Example 2. Tracking shot: move horizontally at the same speed as <Skateboarder>, keeping the subject sharp while the roadside wall forms horizontal motion blur from right to left.
Example 3. Golden hour: warm, low-angle sunlight enters from behind and to the left of <Hiker>, casting long shadows across the mountain ridge.
Example 4. Natural vignette: darken the four corners gradually while keeping the brightness and skin tone of <Pianist> in the center natural, without a black border.
Example 5. Whip-pan transition: at 5 seconds, move the camera rapidly to the left. Cut when the foreground bookshelf fully covers the frame, then continue moving left at a similar speed in the next scene.
Aperture, focal length, and shutter values can be included, but the intended visible result is usually clearer than a numeric value alone.
4. Pre-Submission Checklist
- Does the prompt clearly state the subject and primary action or event?
- Does every reference material state what to use and what not to use?
- Is every distinct character, product, and prop named and bound to a reference?
- Are references selected by scene instead of being required to appear all at once?
- Does each stage of a long video contain only one primary change and a clear end state?
- Do the number of characters, clothing, prop ownership, and spatial relationships remain consistent?
- For video editing, does the prompt define the sole editing master, edit scope, target quantity, and content to preserve?
- Are abstract emotions and cinematography terms paired with directly visible or audible cues?
- Are first/last frames and multiple keyframes assigned one role per image, and do the first and last images use the same aspect ratio?
- Does the storyboard state which structure to inherit? For blockouts, did you first identify whether the reference is coarse or fine and specify the temporal, structural, material, and style information to inherit?
- Do video editing, first/last-frame generation, and video extension follow their automatically locked aspect-ratio and duration rules?
- For video extension, did you check the boundary image, motion trend, and audio continuity?
- For one-click video, does the prompt define material roles, image order, motion amount, editing style, and audio?
- For seamless transitions, does the prompt define the two videos' roles, trigger action, transition process, and arrival state?
5. Usage Limitations
- Timestamps allocate time to events; they are not frame-accurate edit points.
- Video-editing prompts can improve the probability that critical events align with the source video, but they cannot guarantee frame-by-frame overlap.
- The goal of multi-reference creation is to select and combine the correct materials, not to make every material appear at the same time.
- For subtitles, formulas, signs, product specifications, or frame-level timing that must be completely accurate, use prepared reference materials, video generation, and post-production together.
- Video editing automatically locks the input video's aspect ratio and approximate duration; neither can be set separately. The output may differ from the input by up to approximately 0.3 seconds.
- First-frame or first-and-last-frame generation locks the aspect ratio to the first image, while duration can be set. Mismatched first/last image ratios may stretch the last frame.
- Video extension locks the input video's aspect ratio, while extension duration can be set. The extended segment's volume may differ slightly from the source video.
- For one-click video, if image order or character mapping matters, specify it explicitly in the prompt.
- Seamless video transitions aim for visual and audio continuity; they do not guarantee pixel-identical preservation of both source videos.
6. Guide Disclaimer
The examples in this guide illustrate prompt-writing techniques only. Actual generation results may vary depending on the input materials, task complexity, and generation parameters.



