Wan 3 is Alibaba Tongyi Lab's current video model, generating 2 to 30 seconds in one continuous pass at up to 1080p, with audio produced alongside the picture and ratios from 16:9 through 9:16. Thirty seconds arrives as one generation, so a prompt has to carry a whole shot's worth of time and sound. Six endpoints across two tiers cover a text prompt, a still image, and reference material. On fal, standard runs $0.05 to $0.20 a second by resolution, and Prime mirrors all three endpoints faster for $0.068 to $0.28.
In this guide, I'll walk through prompting all six Wan 3 endpoints, from one line of text up to a reference stack carrying a character sheet, a location plate, a motion clip and a music cue into a single generation, with every prompt written so you can drop it into fal's playground as-is or hand it to the API unchanged.
TL;DR
Wan 3 is Alibaba Tongyi Lab's current video model, generating anywhere from 2 to 30 seconds in one continuous pass at up to 1080p, with audio produced alongside the picture and ratios from 16:9 through 9:16.
The behavior that changes how you prompt it is that length: thirty seconds arrives as one generation, so a prompt has to carry a whole shot's worth of time and sound.
Six endpoints across two tiers cover a text prompt, a still image, and reference material you already own.
In terms of pricing on fal, standard runs $0.05 to $0.20 a second by resolution, and Prime mirrors all three endpoints at a faster turnaround for $0.068 to $0.28.
fal is the best place to run every endpoint of Wan 3 behind a single API key, metered by the second, with no subscription and nothing to provision.
Where can you access Wan 3?
fal is the best place to access Wan 3: all six endpoints are metered by the second, with no plan to join and no spend floor to clear.
Three of them make up the standard tier:
alibaba/wan-3.0/text-to-video takes a prompt and nothing else.
alibaba/wan-3.0/image-to-video takes start_image_url as the opening frame, with an optional end_image_url to fix where the shot finishes.
alibaba/wan-3.0/reference-to-video takes reference_image_urls, reference_video_urls and reference_audio_urls, plus file_url for a document and web_url for a public page.
Each has a Prime twin at the same path with the model name swapped, so alibaba/wan-3.0-prime/text-to-video, alibaba/wan-3.0-prime/image-to-video and alibaba/wan-3.0-prime/reference-to-video.
All six share the same controls: prompt, resolution, aspect_ratio, duration, audio, enable_prompt_expansion, enable_thinking, enable_safety_checker and seed.
Here's a text-to-video call through fal's API:
import { fal } from "@fal-ai/client";
const result = await fal.subscribe("alibaba/wan-3.0/text-to-video", {
input: {
prompt: "A red panda walking through a bamboo forest at sunrise",
},
logs: true,
onQueueUpdate: (update) => {
if (update.status === "IN_PROGRESS") {
update.logs.map((log) => log.message).forEach(console.log);
}
},
});
console.log(result.data);
console.log(result.requestId);
To switch endpoints, you edit the model string and supply whatever media field the new one expects.
That same client covers the other 1,000 or more models on fal, so you have to learn the auth and billing once.
And if you would rather not write code, each of the six has a playground page with a form built from its schema, and Sandbox, Agent, MCP, and Workflows are the surfaces next to it for looser testing and for chaining models.
What does a Wan 3 prompt need?
Wan 3 rewrites your prompt before it renders a frame, so a prompt needs enough clarity to come through that rewrite with its non-negotiable details intact.
enable_prompt_expansion defaults to true on all six endpoints, and the response carries a field named actual_prompt alongside the video, worth comparing against what you typed on the first few generations.
Four things belong in the text:
Concrete detail: a named lens, a light source with a direction, a color you could match to a swatch, a movement with a speed attached to it.
A sound line, since audio generates alongside the picture and a prompt without one still gets a soundtrack, just not one you chose.
Spoken lines in double quotes, with the delivery named beside them. fal's own Wan 3 spokesperson example does the second half of that, closing on a warm conversational voice without scripting the words.
One camera setup per beat: our vertical examples hold a single setup across a whole clip, and its Prime image-to-video prompt gives each numbered clip its own.
A held silence is a sound decision as much as a noise is.
Adjectives like "cinematic" and "epic" get absorbed into the rewrite and give the model nothing to hold.
What are the four prompt shapes of Wan 3?
Ordered by how much of the clock you specify:
The one-liner: a subject, a movement and a lens in one sentence with smart duration on, fast to explore with and weak as soon as one detail has to survive.
The block brief: labelled sections down the page. fal's own image-to-video playground prompt runs a headline spec line then FIRST FRAME and SUBJECT; the labels below carry that shape further into STYLE, CHARACTERS, LOCATION, TIMELINE and SOUND.
The timecoded beat sheet: ranges like 0 to 6s with one beat in each, for anything that has to land on a mark, and where enable_thinking starts to pay off.
The numbered clip list: Clip 1, Clip 2, Clip 3, each with a short title and one movement, the shape of fal's Prime image-to-video playground prompt.
How do you prompt Wan 3 for text-to-video?
Text-to-video like Wan 3 needs a prompt and nothing else, so the only setup is choosing resolution, ratio, duration and audio in the form before you decide how much of the shot to specify.
One line at the default length
A one-line prompt hands Wan 3 a subject and a lens and leaves it to fill in the framing, the mood and the supporting motion.
The playground opens at 5 seconds, which is enough for a shot built on one continuous action.
Prompt (duration: 5, aspect_ratio: adaptive): A skier drops into an untracked couloir at first light, spray blowing off the cornice behind her, long lens from the opposite ridge. Wind, steel edges on hard snow, no music.
Generated using Wan 3 on fal, an AI model from Alibaba.
➡️ With fal's API, you can leave duration to 'null', and it'll act as a smart duration, where it hands the length decision to Wan 3, which reads it out of your prompt and any reference media attached.
That behaviour comes from leaving duration unset in an API call, and the playground form asks for a number, so treat it as an API route.
Thirty seconds as one unbroken move
Wan 3 was built around this shape.
Thirty seconds arriving as one generation means the camera can travel somewhere, and the sound can follow it there.
Thinking stays off for this one, since there's a single continuous move and no discrete sequence of events for the model to reason about.
Prompt (duration: 30, resolution: 1080p, aspect_ratio: 16:9, enable_thinking: off): One unbroken thirty-second Steadicam move through a working film studio backlot at two in the morning. No cuts anywhere in the take. 0 to 6s: the camera walks out of a scene dock past stacked flats and a rail of period costume, catching a grip taping down a cable run, then pushes through hanging plastic sheeting into open air. 6 to 14s: it crosses a New York street standing set with the rain machines at full pressure, water sheeting off fire escapes and pooling in the gutters, a stunt performer in a soaked overcoat waiting on her mark while a focus puller measures to her face. 14 to 22s: it turns left down a service alley between two sound stages, past a catering truck with its shutter half down and an electrician asleep in a folding chair, the rain noise falling away behind the wall. 22 to 30s: it comes out onto the lot road, cranes up over a parked camera truck and holds on the whole backlot under sodium light, rain still falling on one street and nowhere else. Anamorphic, sodium and mercury practicals, deep blacks, fine film grain. Sound: rain machines roaring then dropping off, a radio crackle, boots on wet asphalt, a generator running under everything, no music.
Generated using Wan 3 Prime on fal, an AI model from Alibaba.
💡 That last beat is doing double duty.
Cutting the rain off at the wall gives the audio pass a physical reason for the mix to change, and a physical reason is easier to act on than a mood word.
A block brief with dialogue
Four separate beats and two spoken lines make this a sequence of events, so enable_thinking goes on.
Prompt (duration: 25, resolution: 1080p, aspect_ratio: 16:9, enable_thinking: on): STYLE Prestige drama, large-format digital, cold blue dawn against warm stable practicals, long lenses, very little camera movement, natural grain. CHARACTERS A woman in her fifties in a waxed jacket and mud-caked boots, hair pinned back, face unreadable. A man in his thirties in a quilted gilet holding a rolled sheaf of papers he keeps tapping against his leg. LOCATION A racing yard at first light, mist lying across an all-weather gallop, floodlights still burning on the far rail, a horse being walked in circles behind the two of them. TIMELINE 0 to 6s: wide two-shot across the yard, the horse crossing frame between them. He says, "The owners want an answer today." She keeps watching the horse. 6 to 13s: medium on her, floodlights going out one bank at a time behind her head. She says, "Then tell them what I told you in October." 13 to 20s: close on him, the papers gone still against his leg. He starts to answer and doesn't finish it. 20 to 25s: back to the wide. She's already walked out of frame toward the gallop, and he's alone in the yard with the horse. SOUND Hooves on wet sand, floodlight ballasts cutting out, a rook somewhere off camera, both voices close and unraised, no score.
Generated using Wan 3 on fal, an AI model from Alibaba.
Note: I accidentally left the duration to 30 seconds, and I like how the model handled the last 5 seconds of the clip!
A vertical spot at native 9:16
Setting aspect_ratio to 9:16 generates at portrait dimensions, which holds the composition you wrote.
A landscape frame cropped down afterwards loses the sides of the shot you paid for.
Prompt (duration: 10, resolution: 1080p, aspect_ratio: 9:16): A ten-second vertical spot for a titanium dive watch, no dialogue. The watch lies face up on wet black slate in a dark room, a single hard key raking across it from the left so the brushed case and the polished lugs read as two different metals. A hand in a rolled navy sleeve enters from the bottom of frame, lifts the watch, and turns the bezel one full rotation toward camera so every detent is visible against the lume. The camera pushes from a wide macro into a tight macro on the dial as the rotation finishes, then holds. Water beads slide off the crystal and off the edge of the slate. Sound: the bezel clicking through its detents, water running off stone, a low room tone, no music.
Generated using Wan 3 Prime on fal, an AI model from Alibaba.
falMODEL APIs
The fastest, cheapest and most reliable way to run genAI models. 1 API, 100s of models
How do you turn a still image into a Wan 3 clip?
On Wan 3's image-to-video endpoint, start_image_url becomes frame one and the clip builds forward from it.
Adding end_image_url fixes the closing frame as well.
Your prompt has a second responsibility on this endpoint: naming the elements that must not change.
A clip that wanders away from your still is usually one where the face, the wardrobe and the light were never pinned down in words.
One still, thirty seconds, one physical event
Still prompt (Nano Banana 2): A photorealistic film still from a 1990s American diner, shot on 35mm with heavy grain. A woman in her early twenties sits alone in a red vinyl booth by the window, long jet-black hair, winged eyeliner, deep maroon lipstick, a black leather jacket over a white tee. Checkerboard floor, chrome trim, a neon sign reflected in the glass behind her, a half-finished plate of eggs on the table. Soft mid-morning window light mixed with warm ceiling practicals, shallow focus, 40mm lens, no text.
Image generated by Nano Banana 2 on fal, an AI model from Google.
Video prompt (duration: 30, resolution: 1080p): Hold the exact face, hair, makeup, jacket and booth from the still. Keep the diner, the grain and the window light unchanged the whole way through. 0 to 6s: medium wide. She slides out of the booth, stands, turns, and walks straight into a server carrying a loaded breakfast tray. The impact is sudden and physical. 6 to 14s: the camera orbits the collision as it comes apart in slow motion, plates and a coffee pot climbing into the air, coffee pulling into long ribbons and suspended beads. Time locks completely at the top of the spill. Every face in the room freezes mid-reaction and she's the only thing still moving. She takes one beat to register what she's done, then reaches into the floating debris and lifts two rashers of bacon and a fried egg off a hanging plate. 14 to 23s: tracking alongside her as she walks the length of the frozen diner toward the door, eating, everything else held in mid-air behind her. 23 to 27s: she reaches the door and time snaps back. Everything suspended lands at once, plates, tray, coffee pot and food hitting the floor together, the room reacting in real time. 27 to 30s: medium close. She pauses in the doorway, glances back, raises her eyebrows, and gives a small shrug. Sound: diner room tone with a jukebox under it, the crash of the collision, then near-total silence through the frozen section apart from her footsteps and chewing, then the whole crash landing at once, then the room again.
Generated using Wan 3 on fal, an AI model from Alibaba.
➡️ I'd guard that silence in the middle.
As Wan 3 generates picture and sound together, a held quiet is something you ask for up front, and much harder to carve out afterwards.
Both ends pinned with a start and end frame
end_image_url requires start_image_url, so the closing frame is always something you attach on top of an opening one.
Most of the result gets decided by the two stills before Wan 3 sees a word of your prompt.
My approach: build the start frame first, then feed that same file to Nano Banana 2's editing endpoint and describe the change you want at the far end.
Generating both frames from scratch hands you two different cars and two different horizons, and Wan 3 has to reconcile that mismatch somewhere inside your twenty seconds.
Start frame prompt (Nano Banana 2): A cinematic film still of a mint-condition 1967 convertible in deep oxblood red parked side-on at the center of a cracked white salt flat, low three-quarter front angle from ground level, chrome bumper and wire wheels catching low golden light, a flat mountain range far behind under a clear evening sky, 35mm anamorphic, fine grain, no text.
Image generated by Nano Banana 2 on fal, an AI model from Google.
End frame prompt, run on the start frame with Nano Banana 2's edit endpoint: Keep the exact camera position, the same car in the same place at the same angle, the same mountain range and the same horizon line. Age the car by forty years and bury it: paint sun-bleached to chalk pink and blistered, chrome pitted and dull, windshield gone, body sunk to the door handles in drifted sand with a long wind-shaped tail behind it. Replace the golden evening light with flat white overhead noon. Nothing else in the frame changes.
Image edited using Nano Banana 2 on fal, an AI model from Google.
Video prompt (duration: 20, aspect_ratio: 16:9, resolution: 1080p): One continuous decay from the first frame to the last, camera locked in exactly the same position for the whole clip. The light swings from low golden through the middle of the day and settles at flat noon, shadows shortening under the car as it goes. Sand arrives from frame left in long low drifts, banks against the wheels and climbs the body, the paint goes from wet-looking red to chalk, and the windshield goes somewhere in the middle without the camera acknowledging it. No cuts, no camera move, nothing enters frame apart from the sand. Sound: wind rising steadily across the whole clip, sand ticking against metal, a loose sheet of trim knocking somewhere off camera, no music.
Generated using Wan 3 Prime on fal, an AI model from Alibaba.
How do you use reference-to-video on Wan 3?
Reference-to-video conditions a Wan 3 generation on material you already own: up to 10 reference images, up to 5 video clips totaling at most 15 seconds with each clip at 16fps or higher, and up to 5 audio tracks totaling at most 15 seconds, all in one request.
You'll have to be addressing that material positionally, such as "the subject in Image 1 walks past Video 1".
Let's go over a few examples:
Four images, one shot
Image 1 (Nano Banana 2): A photorealistic three-quarter portrait of a woman in her late twenties against a bare mid-gray studio backdrop, close-cropped dark hair, a small scar through one eyebrow, matte skin finish, an unbranded black crew neck, even soft frontal lighting, sharp focus, 85mm lens, no text.
Image generated by Nano Banana 2 on fal, an AI model from Google.
Image 2 (edit of Image 1): The same woman on the same gray backdrop in full left profile, identical hair, skin and black crew neck, identical soft even lighting, 85mm lens, no text.
Image generated by Nano Banana 2 Edit on fal, an AI model from Google.
Image 3 (Nano Banana 2): A product photograph of a floor-length oxblood wool coat on an invisible mannequin against pure white, wide notch lapels, a double-breasted horn button front, a deep back vent, shot straight on with even studio light and no cast shadow, no text.
Image generated by Nano Banana 2 on fal, an AI model from Google.
Image 4 (Nano Banana 2): An empty multi-story car park deck at dusk photographed from one corner, wet concrete floor, low ceiling with exposed sodium strip lights, the open side showing a city skyline going blue, no cars, no people, wide lens, no text.
Image generated by Nano Banana 2 on fal, an AI model from Google.
Video prompt (duration: 15, resolution: 1080p, aspect_ratio: 16:9): Image 1 and Image 2 are the same woman, front and profile. Image 3 is the coat she wears. Image 4 is the location. She walks the length of the deck in Image 4 toward camera wearing the coat from Image 3, unbuttoned, the vent opening behind her with each stride. The camera tracks backward ahead of her at chest height and slows as she reaches it. She stops, turns her head toward the open side, and holds while the sodium strips overhead come on one after another down the deck. Her face matches Image 1 and Image 2 in every frame. The coat keeps its exact color, lapel shape and button placement. Editorial fashion film, cold blue ambient against warm sodium, hard reflections in the wet floor, 50mm, shallow focus. Sound: footsteps echoing off concrete, the ballast hum of the lights coming on, traffic from the street below, no music.
Generated using Wan 3 Prime on fal, an AI model from Alibaba.
💡 Spending a slot on the profile is worth it.
One frontal reference leaves the side of a face to be invented, and this shot asks her to turn her head.
A character, a movement and a score in one request
This one uses all three reference types, each doing a different job.
The motion plate comes off Wan 3's own text-to-video endpoint, so the whole stack can be built without leaving fal.
Reference image prompt (Nano Banana 2): A character design sheet in modern cinematic anime style on a flat off-white background: a lean young woman in a charcoal high-collared coat over a dark red sash, silver-white hair cut to the jaw, a thin scar along the left cheekbone, dark bound-leg trousers and light boots. Full body front view, hand-drawn line quality with painterly cel shading, no text, no labels.
Image generated by Nano Banana 2 on fal, an AI model from Google.
Reference video prompt, generated on alibaba/wan-3.0/text-to-video at duration: 6: A live-action martial artist in unmarked black training clothes runs four steps across a bare concrete floor, plants hard, turns one full spin with the arms drawn in, then lands square and still. Locked-off wide, flat even overhead light, bare gray wall behind, no props. Room tone only, no music.
Generated using Wan 3 on fal, an AI model from Alibaba.
Reference audio prompt (Seed Audio 1.0, 12 seconds): A twelve-second instrumental cue. Low taiko pulse on a slow four from the top, a single shakuhachi line entering at four seconds and bending upward, string tremolo swelling underneath from eight seconds, hard stop on the final beat. No vocals, no percussion fills, dry room, cinematic trailer weight.
Audio generated by Seed Audio 1.0 on fal, an AI model from ByteDance.
Video prompt (duration: 12, resolution: 1080p, aspect_ratio: 16:9): Image 1 is the character. Video 1 is the movement. Audio 1 is the score. Render the character from Image 1 in her own anime style, never photoreal, on a rain-soaked temple courtyard at dusk: wet flagstones, a half-collapsed wooden colonnade, torn prayer flags moving in the wind. She performs the exact run, plant, spin and landing from Video 1, matched beat for beat, with the coat and sash carrying the momentum of the turn and water lifting off the stones under her feet. The camera orbits with her through the spin and settles low as she lands. Her face, hair, scar and clothing match Image 1 in every frame. Cut the picture to Audio 1: the plant falls on the shakuhachi entry and the landing falls on the final beat. Sound: keep Audio 1 as the score and lay rain, wet footfalls and flag snap over the top of it.
Generated using Wan 3 Prime on fal, an AI model from Alibaba.
➡️ Match duration to the length of your audio reference so the cue doesn't run out before the shot does.
Twelve seconds of audio against a twelve-second output means the final beat has somewhere to land.
Wan 3 meters by the second of generated video, and fal's rate note carries no separate charge for attached references, so the extra material costs nothing.
How do you turn a document or a web page into a Wan 3 video?
A document or a public page reaches Wan 3 through two fields on reference-to-video, file_url and web_url, and both need enable_thinking set to true.
Only reference-to-video exposes those two fields, on standard and Prime alike.
➡️ Text-to-video and image-to-video don't carry them.
web_url reads pages that don't require a login, so anything behind a paywall or an internal SSO wall has to be exported and passed through file_url.
The part to get right is the division of labor.
Your source supplies the facts, and your prompt supplies the direction.
You can hand over a deck with no direction attached and the generation has to invent its own brief from your source.
Name the figures, lines or sections that carry the video, then name what gets dropped.
Two practical notes on this path:
As reference-to-video accepts reference_image_urls and file_url together, a brand deck and a set of product stills can travel in the same request, with the prompt naming which does what.
Any text Wan 3 renders on screen out of a document deserves a frame-by-frame check before it ships, numbers pulled from a table especially.
Which Wan 3 parameters are worth changing from their defaults?
Four of Wan 3's parameters move the output enough to deserve a decision before you generate: duration, aspect_ratio, resolution and enable_prompt_expansion.
Duration and smart duration
duration defaults to 5 and accepts whole seconds from 2 to 30.
Leaving it unset through the API turns on smart duration, where Wan 3 reads a length out of your prompt and any reference media attached.
The playground form asks for a number, so smart duration is an API route, and a one-action prompt is what suits it.
A timecoded prompt already states its own length either way, so set the number in the form and let the two agree.
Aspect ratio
aspect_ratio defaults to adaptive, which lets Wan 3 pick from 16:9, 4:3, 1:1, 3:4 and 9:16 based on the shot.
Adaptive is fine while you're exploring.
For anything with a home to go to, set the ratio before the first generation.
Resolution
480p, 720p and 1080p, defaulting to 1080p.
At $0.05 a second, a full thirty-second test at 480p costs $1.50.
That's cheap enough to run the same prompt three or four times before committing to 1080p, and moving up a tier afterwards restates nothing.
Prompt expansion
enable_prompt_expansion defaults to true and rewrites your prompt before generation.
fal's schema notes that turning it off can save roughly 20 to 60 seconds of latency but is likely to degrade generation quality.
Check the response afterwards for the actual_prompt field.
Comparing it against what you typed shows which of your own words survived the rewrite, and usually why a detail went missing.
Two more controls come up less often.
seed lets you repeat a result while you change a single element of the prompt.
enable_safety_checker defaults to true. Disabling it requires account authorization, and unauthorized requests get checked regardless.
What does Wan 3 Prime do differently?
Wan 3 Prime runs the same three endpoints with the same parameter set at a faster turnaround and a higher per-second rate.
Nothing about the interface changes.
You're getting the same resolutions, same 2 to 30 second range, same ratio list, same audio, thinking, expansion, seed and safety controls, same reference caps, same document and web page fields.
Switching tiers is a model string edit and nothing else.
What changes is the rate: $0.068, $0.14 and $0.28 a second at 480p, 720p and 1080p, against $0.05, $0.10 and $0.20 on standard.
This makes the tier a scheduling call.
I'd reach for Prime with a client in the room, or on a review loop where the wait between versions costs more than the render does.
Standard covers the rest, and on a thirty-second 1080p shot the difference is $6.00 against $8.40, which adds up once you're running variations of one brief.
The pairing that gets the most out of both tiers is to lock the prompt on standard at 480p, where a full thirty seconds costs $1.50, then send the finished prompt to whichever tier the deadline points at.
Every parameter carries across unchanged, so the second call is a copy of the first with two strings edited.
Prime's reference-to-video endpoint also takes the same caps as standard, up to 10 images, 5 clips and 5 audio tracks, so a reference stack built on one tier transfers to the other without rebuilding.
Prompt run on alibaba/wan-3.0-prime/text-to-video (duration: 20, resolution: 1080p, aspect_ratio: 16:9, enable_thinking: on): A twenty-second anime action sequence, hand-drawn line quality with painterly cel shading and heavy speed-line work, no photoreal rendering anywhere. A woman in weightless white and silver plate armor with a long translucent cape holds a narrow bridge of shattered marble suspended above a sea of cloud at dusk, a single sword of white light in one hand. 0 to 7s: she sprints along the bridge toward camera as armored figures in charcoal lacquered plate close from both ends, the camera chained to her shoulder in a fast orbit. 7 to 14s: she takes each one in a single strike without breaking stride, armor shearing, banners cut, bodies falling away into the cloud behind her, rain falling upward in the wind. 14 to 20s: the bridge collapses under the last exchange and she drops through the break, cape snapping taut, the camera falling with her until the cloud closes over the frame. Aurora ribbons overhead, volumetric light through the cloud sea, wet marble catching it, dust drifting in slow arcs. Sound: steel on steel with each clash readable on its own, wind pulling at the cape, marble cracking, a low string swell rising through the last four seconds.
Generated using Wan 3 Prime on fal, an AI model from Alibaba.
How much does Wan 3 cost on fal?
Every Wan 3 endpoint meters by the second of generated video, at a rate set by output resolution, with audio included and no subscription attached.
| Tier | 480p | 720p | 1080p |
|---|---|---|---|
| Wan 3 text-to-video, image-to-video, reference-to-video | $0.05/s | $0.10/s | $0.20/s |
| Wan 3 Prime text-to-video, image-to-video, reference-to-video | $0.068/s | $0.14/s | $0.28/s |
Here's the math worked out at the lengths you'd actually use:
| Clip | Wan 3 | Wan 3 Prime |
|---|---|---|
| 5 seconds at 720p | $0.50 | $0.70 |
| 10 seconds at 1080p | $2.00 | $2.80 |
| 30 seconds at 480p | $1.50 | $2.04 |
| 30 seconds at 1080p | $6.00 | $8.40 |
The rate covers audio, and none of the six endpoints carries an audio surcharge.
Reference images, clips and audio tracks aren't priced separately either, so the only line item is the video you generate.
Recently Added
Start creating with Wan 3 on fal
Wan 3 gives you six endpoints on one key, metered by the second, with nothing to subscribe to.
From that single integration you can sketch a shot out of one sentence, put a still into motion, fix both ends of a transformation, hand the model a character plus some of your own footage, or turn a brand deck into a spot.
Creating a fal account costs nothing, and all six run in the browser from the playground if you want a look before touching the API.
Wan 3 FAQs
What does enable_thinking do on Wan 3?
enable_thinking is an optional mode on Wan 3, off by default, that has the model reason about composition, staging and motion before it renders a frame.
You can turn it on for prompts describing a sequence of events over one continuous action, and for the document and web page inputs, which require it because those sources have to be read and interpreted before they can become a shot.
The cost is pretty much generation time, as it's not billed separately.
Can I use Wan 3 clips in commercial work?
Yes.
Anything the fal API generates is cleared for commercial projects, a paid campaign and a client deliverable included.
fal's terms of service carry the full licensing detail.
How long can a Wan 3 clip be?
Anywhere from 2 to 30 seconds in a single generation, on all six endpoints.
duration defaults to 5 seconds and the playground form asks for a number in that range.
Through the API, leaving duration unset turns on smart duration, where Wan 3 picks a length from your prompt and any reference media attached.
What resolutions and aspect ratios does Wan 3 support?
480p, 720p and 1080p, with 1080p as the default.
Aspect ratios cover 16:9, 4:3, 1:1, 3:4 and 9:16, plus an adaptive setting that lets Wan 3 choose the ratio suiting the shot.
Does Wan 3 generate audio?
Yes.
Picture and sound come out of one generation, and fal doesn't meter the audio separately.
The audio field defaults to true on every Wan 3 endpoint.
A named sound bed in the prompt converts that default into a decision, and a held silence is as promptable as a noise.
What can you use as a reference with Wan 3?
Up to 10 reference images, up to 5 video clips totaling at most 15 seconds with each clip at 16fps or higher, and up to 5 audio tracks totaling at most 15 seconds, in a single request.
With enable_thinking on, Wan 3's reference-to-video endpoint also reads a document you upload or a public web page you link.
What is the difference between Wan 3 and Wan 3 Prime?
Wan 3 Prime mirrors all three endpoints with an identical parameter set at a faster turnaround and a higher per-second rate: $0.068, $0.14 and $0.28 against $0.05, $0.10 and $0.20.
Everything else, including the reference caps and the document and web page fields, is the same on both tiers.
Is Wan 3 available as open weights?
No.
Alibaba's published Wan weights currently stop at Wan 2.2 under Apache 2.0, and the releases since then have been commercial API models.
Wan 3 runs on fal as an API model.
Which other Wan models can you run on fal?
Wan 2.7 is available with text-to-video, image-to-video, reference-to-video and instruction-based video editing at 720p and 1080p, native audio, and durations from 2 to 15 seconds.
Wan 2.6 is there as well, alongside open-weight Wan 2.1 and Wan 2.2 endpoints and the Wan LoRA trainers.
Why should you use fal to run Wan 3?
One client covers all six Wan 3 endpoints, with the queue and webhook handling done for you and no GPU fleet to look after.
Nothing has to be subscribed to, output is cleared for commercial use, and the same key opens the rest of fal's catalogue of 1,000 or more models.
How do you use Wan 3 on fal?
This guide leans on two routes, and there are others:
Any endpoint runs in the browser from its playground page without code, while the @fal-ai/client SDK moves those same calls into production.
Sandbox is there for loose experimentation, and Workflows are for chaining several models into one job.
Agents and terminal work go through fal's MCP server and CLI.
![How To Use Wan 3: Prompts & Workflows [2026]](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0aa7ccdb%2FstMjxr_awCqm3GXMRo_pX.jpg/tr:w-1920,q-80/stMjxr_awCqm3GXMRo_pX.webp)





















![How To Use Nano Banana 2 Lite: Prompts & Workflows [2026] | fal](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0aa263d4%2FgskEo8CDZbFmxFaNjXC_6.jpg/tr:w-1080,q-80/gskEo8CDZbFmxFaNjXC_6.webp)
