Wan 3 vs. Wan 2.7: What's the difference in 2026?

A head-to-head of Alibaba's Wan 3.0 Prime and Wan 2.7 on fal, six matched text, image, and reference-to-video tests, plus the spec and pricing differences between them.

John OzuysalSep 30, 202616 min read
Wan 3 vs. Wan 2.7: What's the difference in 2026?

Wan 3.0 Prime and Wan 2.7 both come from Alibaba, and each family covers the same three generation endpoint types on fal. Wan 3.0 Prime runs $0.28 per second at 1080p, with its resolution list starting at 480p at $0.068. Wan 2.7 charges $0.10 at 720p and $0.15 at 1080p on text and image to video. Five seconds of 1080p comes to $1.40 on Wan 3.0 Prime and $0.75 on Wan 2.7 text to video. Wan 3.0 Prime reference to video accepts audio references and, with enable_thinking on, documents and webpages; Wan 2.7 offers a negative prompt, a driving audio track, clip continuation, and an edit-video endpoint.

I ran six matched tests on fal, two on each endpoint type the families share, to show how Wan 3.0 Prime and Wan 2.7 handle the same brief.

TL;DR

Wan 3.0 Prime and Wan 2.7 are AI video generators that both come from Alibaba, and each family covers the same three generation endpoint types on fal.

Wan 3.0 Prime runs $0.28 per second at 1080p, and its resolution list starts at 480p, priced at $0.068 per second.

Wan 2.7 charges $0.10 per second at 720p and $0.15 at 1080p on text and image to video.

Wan 2.7 reference to video bills $0.10 per second, counting any input video seconds along with the output.

Five seconds of 1080p comes to $1.40 on Wan 3.0 Prime and $0.75 on Wan 2.7 text to video.

Standard Wan 3.0 accepts the same inputs as Wan 3.0 Prime for $0.20 per second at 1080p.

Wan 3.0 Prime reference to video accepts audio references, and with enable_thinking on it can also read a document or a public webpage.

Wan 2.7 offers a negative prompt field, a driving audio track on text and image-to-video, clip continuation and a dedicated edit-video endpoint.

How does Wan 3.0 Prime compare to Wan 2.7?

Wan 3.0 Prime adds audio references and document or webpage URLs to reference to video, while Wan 2.7 runs $0.13 less per second at 1080p on text and image to video.

Wan 2.7 also has an edit-video endpoint, and its image-to-video schema supports clip continuation through video_url.

Most field names appear in both schemas. The practical differences come from fields that only one side defines.

Here's how they compare:

Wan 3.0 PrimeWan 2.7
DeveloperAlibabaAlibaba
Best forMixed-media reference work and 480p draftsLower-cost 1080p output and edits to existing footage
Video endpoints covered hereText to video, image to video, reference to videoText to video, image to video, reference to video, edit video
Resolutions480p, 720p, 1080p720p, 1080p
Default resolution1080p1080p
DurationInteger seconds, default 5, or null for smart duration2 to 15 seconds on text and image to video, 2 to 10 on reference to video
Aspect ratios, text to videoadaptive, 16:9, 4:3, 1:1, 3:4, 9:1616:9, 9:16, 1:1, 4:3, 3:4
Default aspect ratioadaptive16:9 on text and reference to video
Aspect ratio control, image to video✅ aspect_ratio, adaptive by default❌ Not exposed
First frame fieldstart_image_url, requiredimage_url
End frame✅ end_image_url, needs a start frame✅ end_image_url
Video continuation❌ Not exposed✅ video_url on image to video, 2 to 10 second clips
Generated audio✅ audio flag, on by default✅ Background music when no audio_url is sent, per the text-to-video schema
Audio file as input✅ On reference to video only✅ audio_url on text and image to video, WAV or MP3 up to 30 seconds
Negative prompt❌ Not exposed✅ Up to 500 characters
Prompt expansion✅ enable_prompt_expansion, on by default✅ enable_prompt_expansion on text and image to video
Reasoning pass✅ enable_thinking❌ Not exposed
Prompt length limitNot stated in the schema5,000 characters
Reference imagesUp to 10No count stated, 20 MB each
Reference videosUp to 5, 15 seconds combined, 16 fps minimumNo count stated, 100 MB each
Reference audioUp to 5, 15 seconds combined❌ No.
Document and webpage references✅ file_url and web_url, both need enable_thinking❌ Not exposed
Reference addressingPositional, such as Image 1 and Video 1Not documented in the schema
Multi-shot control❌ No dedicated field in the schema✅ Natural language on text to video, multi_shots on reference to video
Rewritten prompt returned✅ actual_prompt✅ actual_prompt
Safety checker toggle✅ API only, disabling needs account authorization✅ API only, disabling needs account authorization
Billing unitPer second of outputPer second of output, plus input video seconds on reference to video

Where can you run Wan 3.0 Prime and Wan 2.7?

fal hosts the Wan 3.0 Prime and Wan 2.7 endpoints covered here, and each one runs in the browser playground or through the API on a single key.

Sandbox links appear on the Wan 3.0 Prime pages as well.

The queue and webhook calls in fal's JavaScript client work identically for both families, and the input object is the only part you build per family.

Local start frames and reference files can go through fal.storage.upload, which returns a URL either family accepts.

Here's how that looks like:

javascript
import { fal } from "@fal-ai/client";

const result = await fal.subscribe("fal-ai/wan/v2.7/reference-to-video", {
  input: {
    prompt: "A person walking through a beautiful garden, cinematic style."
  },
  logs: true,
  onQueueUpdate: (update) => {
    if (update.status === "IN_PROGRESS") {
      update.logs.map((log) => log.message).forEach(console.log);
    }
  },
});

console.log(result.data);
console.log(result.requestId);

Wan 3.0 Prime vs. Wan 2.7: text-to-video tests

The first test is a product close-up built around one mechanical event, and the second is a two-shot creator clip with a spoken line:

Test 1: An automatic umbrella opening in the rain (16:9)

I picked an automatic umbrella because the opening happens in one fast motion and ends in a fixed shape.

The prompt specifies eight ribs and eight panels, and I counted both in each open dome.

There's a second count in the rain, which should drip from each of the eight rib tips once the fabric is taut.

Prompt: Five seconds, 16:9, a product film for an unbranded automatic umbrella on a wet city street at night. Style: premium accessories commercial, sodium streetlight from above and behind, cool blue shop-window spill from camera right, shallow depth of field, steady rain. [0 to 2 seconds] A low angle on a gloved hand holding the closed umbrella upright, its black canopy furled tight around the shaft. The matte graphite handle has a single silver button on the grip, and the thumb presses it. [2 to 5 seconds] The canopy springs open in one motion into a dome of eight ribs and eight panels, the fabric snapping taut. Rain beads on the canopy and runs to the rib tips, falling in eight separate streams. The camera rises slowly until the full dome fills the frame. Throughout: eight ribs and eight panels once open, and nothing printed anywhere on the umbrella. Audio: steady rain on asphalt, the button click, the canopy snapping open, rain drumming on the taut fabric. No music or voice.

Generated using Wan 3.0 Prime on fal, an AI model from Alibaba.

Generated using Wan 2.7 on fal, an AI model from Alibaba.

Test 2: A UGC skincare review with a product cutaway (9:16)

This prompt compresses a creator-style review into two shots.

In the first, she says one line to the phone camera, and her mouth shapes need to track the words. The second cuts to the serum bottle on the sink.

A change to its cap or glass between shots would give the cut away.

Prompt: Five seconds, 9:16, a selfie-style UGC skincare review filmed on a phone, two shots. Shot 1, 0 to 3 seconds. A woman in her late twenties with shoulder-length dark curls and a cream knit sweater holds the phone at arm's length in a sunlit bathroom with white tiles behind her. In her other hand she raises an unbranded frosted glass serum bottle with a black dropper cap beside her cheek. She smiles into the lens and says, "I didn't expect a serum to feel this light." Shot 2, 3 to 5 seconds. Hard cut to a close-up of the same bottle standing on the edge of the white sink. Her fingers lift the dropper out and squeeze a single clear drop back into the bottle. Handheld phone framing with a slight natural sway in shot 1, a steadier macro in shot 2. The bottle stays identical across both shots, frosted glass with a black dropper cap and no label or text. Audio: her voice close to the phone microphone, small bathroom room tone, the soft squeeze of the dropper bulb, a faint glass tick as the bottle meets the sink. No music.

Generated using Wan 3.0 Prime on fal, an AI model from Alibaba.

Generated using Wan 2.7 on fal, an AI model from Alibaba.

💡 Switching off enable_prompt_expansion can cut roughly 20 to 60 seconds of latency at a likely cost in output quality. I'd only turn it off for draft runs on prompts that are already detailed.

Wan 3.0 Prime vs. Wan 2.7: image-to-video tests

Both image-to-video endpoints use your image as the first frame and take an optional last frame through end_image_url.

The first-frame field is start_image_url on Wan 3.0 Prime and image_url on Wan 2.7, a rename to handle before one request body can serve both.

Test 3: A candle lit by hand in a dim room (1:1)

Here I wanted to see whether the lighting in the scene responds to the flame.

When the wick catches, the amber glass and the linen under the jar should brighten and turn warmer as the flame grows.

The rest of the room ought to stay as dim as the start frame.

A spent match goes into the prompt too, and its smoke needs to rise and thin out as the hand pulls away.

Start frame, made with GPT Image 2.5 Flare: A photoreal 1:1 product still of an unbranded candle in a thick amber glass tumbler, its flat wooden lid set down beside it, on a raw linen runner over a dark oak table. The candle is unlit, with a flat cream wax surface and one cotton wick standing upright. A single wooden match lies next to the lid. Soft dusk window light from camera left, with a blurred bookshelf far behind in deep shadow. No labels, no text, no logos, no people.

Generated using GPT Image 2.5 Flare on fal, an AI model from OpenAI.

Video prompt: A hand reaches in from the right edge of the frame and picks up the match. It strikes the match on a matchbox held just out of shot and lowers the flame to the wick, which catches within about a second. The hand pulls back and shakes the match out, trailing a thin line of smoke from the match head as it leaves frame. The candle flame settles into a steady teardrop and lights the amber glass from inside, warming the linen and the table around the base. A small melt pool starts to form around the wick while the glass and the lid keep their shape. Keep the camera fixed on a tripod for the whole clip. Audio: a quiet room, the match striking, the soft flare of ignition, the match going out with a flick. No score and no dialogue.

Generated using Wan 3.0 Prime on fal, an AI model from Alibaba.

Generated using Wan 2.7 on fal, an AI model from Alibaba.

Test 4: An orchid bud opening between two fixed frames (3:4)

This one pins down both ends of the clip with images.

The end frame changes only the largest bud, which is closed in the first frame and open in the last.

Everything between them is up to the model.

I'm watching whether the petals unfold in a believable order and whether the four other flowers on the spike stay put.

I built the end frame by editing the start frame, so the pot and the camera position don't move between them.

Start frame, made with GPT Image 2.5 Flare: A photoreal 3:4 vertical product still of a white phalaenopsis orchid in a matte terracotta pot on a pale concrete shelf. One arching flower spike carries two open white blooms with magenta throats on its lower half and three closed green-white buds toward the tip, the bud nearest the open blooms the largest. Two broad glossy leaves spread from the base of the plant. Soft even north light, a warm off-white wall behind. No labels, no text, no logos, no people.

Generated using GPT Image 2.5 Flare on fal, an AI model from OpenAI.

End frame, an edit of the start frame in GPT Image 2.5 Flare Edit: Leave the pot, the shelf, the wall, the leaves, the flower spike, the two open blooms, the two smallest buds and the lens position untouched. Open only the largest bud, the one nearest the open blooms, into a full white bloom with a magenta throat that matches the two blooms below it in size and color. No labels, no text, no logos, no people.

Edited using GPT Image 2.5 Flare Edit on fal, an AI model from OpenAI.

Video prompt: Time-lapse from a fixed tripod position. The largest closed bud, the one nearest the open blooms, swells and then unfolds petal by petal until it matches the two blooms below it. The two smaller buds toward the tip stay closed, and the two open blooms and both leaves hold still. The flower spike keeps its arch, and the light stays constant with no change in shadow direction. Audio: a near-silent room with a faint ambient hum. No music.

Generated using Wan 3.0 Prime on fal, an AI model from Alibaba.

Generated using Wan 2.7 on fal, an AI model from Alibaba.

falMODEL APIs

The fastest, cheapest and most reliable way to run genAI models. 1 API, 100s of models

falSERVERLESS

Scale custom models and apps to thousands of GPUs instantly

falCOMPUTE

A fully controlled GPU cloud for enterprise AI training + research

Wan 3.0 Prime vs. Wan 2.7: reference-to-video tests

Wan 3.0 Prime documents positional addressing for references, so a prompt can point to Image 1 or Video 1 by number.

The Wan 2.7 schema documents no addressing convention.

To cover both, the prompts below also describe each reference in words, and the text went to both models unchanged.

Test 5: A pet harness fitted onto a reference dog (16:9)

The reference pack for this test holds two stills.

One shows a corgi, the other a harness laid flat, and the prompt asks for the dog wearing that harness while it trots along a beach.

I designed the harness to be asymmetric, giving it a reflective strip on the left side strap only and a single D-ring on the back.

A mirrored strip or a second D-ring would be easy to spot in the result.

Image 1, made with GPT Image 2.5 Sunburst: A photoreal 16:9 studio portrait of a red and white Pembroke Welsh corgi standing in three-quarter view against a pale gray paper backdrop. A white blaze runs up the center of the face, and the dog's left ear has a small notch at the tip. The coat is short and glossy. Soft even light, full body in frame, no collar, no harness, no props, no text.

Generated using GPT Image 2.5 Sunburst on fal, an AI model from OpenAI.

Image 2, made with GPT Image 2.5 Sunburst: A photoreal 16:9 product still of an unbranded dog harness laid flat on a pale gray paper backdrop, shot from directly above. Deep teal webbing with contrast orange stitching along every edge and a padded chest plate. One brushed steel D-ring is centered on the back strap. A reflective silver strip runs along the side strap that goes on the dog's left, and the right side strap has none. Two black side-release buckles, one on each side strap. Even soft light, no dog, no leash, no labels, no logos, no text.

Generated using GPT Image 2.5 Sunburst on fal, an AI model from OpenAI.

Video prompt: The dog comes from Image 1 and the harness from Image 2. The corgi from Image 1 wears the harness from Image 2 and trots along the wet sand of a wide beach at golden hour, moving toward the camera and slightly to the right. Match the dog to Image 1, including the white blaze and the notched tip of the left ear. The harness should match Image 2 at every step, down to the teal webbing, the orange stitching, the single steel D-ring on the back and the reflective strip on the dog's left side only. The camera tracks backward at the dog's shoulder height, with a low sun behind the dog putting a rim light on the coat. Wet sand kicks up at each step, and small waves wash in at the edge of the frame. No leash, and no people or other animals anywhere on the beach. Audio: surf, paws on wet sand, the D-ring jingling lightly with each stride, gulls far off. No music or narration.

Generated using Wan 3.0 Prime on fal, an AI model from Alibaba.

Generated using Wan 2.7 on fal, an AI model from Alibaba.

Test 6: A desk fan carried from a reference clip into a new room (16:9)

In this pack, the product arrives as a video with no still of it.

The fan and its left-to-right sweep come from Video 1, and the room comes from a single still.

We will want to look at the blade count and the wire cage first, then the airflow.

Papers on the desk are meant to lift only when the fan head points left, with the curtain moving only on the right-hand pass.

Image 1, made with GPT Image 2.5 Flare: A photoreal 16:9 interior still of a small home office in late afternoon, with a pale ash desk against a white wall and a tall window on the right hung with a sheer linen curtain. A loose stack of printed papers rests on the left side of the desk next to a ceramic mug, and the middle of the desk surface is clear. Warm sun through the window, soft shadows, no people, no text, no logos.

Generated using GPT Image 2.5 Flare on fal, an AI model from OpenAI.

Video 1, a five-second clip made with H3 Max: Five seconds, 16:9. An unbranded retro desk fan with five translucent sage-green blades inside a chrome wire cage, running against a white studio backdrop. It stands on a round chrome base with a small rotary switch on the front. The fan head oscillates slowly from left to right and back once across the clip while the blades spin at full speed. Even studio light, locked camera, no people, no text.

Generated using H3 Max on fal, an AI model from fal's collaboration with MiniMax.

Video prompt: Video 1 shows the fan, and Image 1 shows the room. The fan from Video 1 stands in the clear area in the middle of the desk from Image 1, running at full speed. The fan head sweeps left to right and back once across the five seconds, at the same speed and through the same arc as in Video 1. Carry over the five translucent sage-green blades, the chrome wire cage, the round chrome base and the rotary switch from Video 1 unchanged. When the fan head points left, the top sheets of the paper stack lift and flutter, and when it points right, the sheer curtain billows toward the window. The camera pushes in slowly from a wide view of the desk to a medium shot of the fan. The room and its window light stay as Image 1 shows them. Audio: the fan motor hum, soft air movement, paper rustling on the left, the curtain brushing the window frame. No music or voiceover.

Generated using Wan 3.0 Prime on fal, an AI model from Alibaba.

Generated using Wan 2.7 on fal, an AI model from Alibaba.

How do Wan 3.0 Prime and Wan 2.7 compare on pricing?

At 720p and 1080p, Wan 2.7 bills $0.10 and $0.15 per second against $0.14 and $0.28 on Wan 3.0 Prime, per fal's standard playground rates as of September 2026.

Wan 3.0 Prime also lists a 480p rate of $0.068 per second.

Both families charge per second of output, so the rates compare directly once the resolution is fixed.

Wan 2.7 reference to video is the exception, because it's $0.10 per second and also counts the duration of any input video.

The table below prices the six tests next to the base rates.

Wan 3.0 PrimeWan 2.7
Billing unitPer second of outputPer second of output
480p rate$0.068 per secondNot offered
720p rate$0.14 per second$0.10 per second
1080p rate$0.28 per second$0.15 per second
Reference to video rateSame as text and image to video$0.10 per second, input video seconds included
Each of Tests 1 to 4$1.40$0.75
Test 5, two reference stills$1.40$0.50
Test 6, one five-second reference clip$1.40$1.00
All six tests$8.40$4.50
15 seconds at 1080p$4.20$2.25

Input assets are billed separately on their own models and are left out of the totals:

Wan 2.7 edit-video bills $0.10 per second of output at 720p and $0.15 at 1080p.

On every Wan 3.0 Prime endpoint, reference-to-video included, the playground pricing note charges for seconds of generated video.

On the standard Wan 3.0 tier, five seconds at 1080p costs $1.00 at $0.20 per second, and 480p drops to $0.05 per second.

💡 I'd settle the prompt at 480p before rendering the final take at 1080p. That's $0.34 for five seconds on Wan 3.0 Prime and $0.25 on the standard Wan 3.0 tier.

Which one should you use: Wan 3.0 Prime or Wan 2.7?

If I were you, I'd choose based on what you start with and what you deliver.

Wan 3.0 Prime suits briefs that arrive as mixed media or source documents.

For 1080p volume work, and for jobs that begin with footage or a soundtrack you already have, Wan 2.7 is the closer match.

Reach for Wan 3.0 Prime when

  • The brief starts as a document or a public product page, and you want reference to video to read it directly with file_url or web_url and enable_thinking on.
  • You have a voice or sound cue on file, and reference to video can take up to five audio clips with 15 seconds combined.
  • You draft at 480p.
  • Frame shape and clip length are open, and an adaptive aspect ratio or a null duration can leave both to the model.
  • A complicated prompt could use the optional reasoning pass behind enable_thinking.
  • You need to choose the output aspect ratio for image-to-video, and Wan 3.0 Prime exposes an aspect_ratio field there.

Reach for Wan 2.7 when

  • 1080p is the delivery spec, and volume sets the budget, at $0.15 per second for text-to-video and image-to-video.
  • A recorded voiceover or music track should drive the video, and audio_url accepts it as a WAV or MP3 of up to 30 seconds.
  • You're restyling or changing footage you already have, since edit-video accepts clips of 2 to 10 seconds with a text instruction.
  • Original sound needs to stay in after the edit, and audio_setting accepts origin for that.
  • A clip needs a continuation through video_url on image-to-video.
  • Negative prompts are already part of your pipeline.
  • Stills should come from the same family, with Wan 2.7 text to image at $0.03 per image and the Pro tier at $0.075.

💡 Wan 2.7 edit-video accepts MP4 or MOV clips of 2 to 10 seconds and up to 100 MB. A five-second Wan 3.0 Prime render fits within those limits for a restyle pass.

Recently Added

Get started with Wan 3.0 Prime and Wan 2.7 on fal

Wan 3.0 Prime and Wan 2.7 are available on fal today, and one API key covers every endpoint below.

Wan 3.0 Prime runs at alibaba/wan-3.0-prime/text-to-video, alibaba/wan-3.0-prime/image-to-video and alibaba/wan-3.0-prime/reference-to-video.

Standard Wan 3.0 runs at alibaba/wan-3.0/text-to-video, alibaba/wan-3.0/image-to-video and alibaba/wan-3.0/reference-to-video.

Wan 2.7 video runs at fal-ai/wan/v2.7/text-to-video, fal-ai/wan/v2.7/image-to-video, fal-ai/wan/v2.7/reference-to-video and fal-ai/wan/v2.7/edit-video.

Check out fal to get started and create your free account.

Wan 3.0 Prime vs. Wan 2.7 FAQs

How does Wan 3.0 Prime differ from the standard Wan 3.0 tier?

On fal, Wan 3.0 and Wan 3.0 Prime take the same input fields on every endpoint, and they bill at different rates.

Wan 3.0 costs $0.20 per second at 1080p against $0.28 on Wan 3.0 Prime.

The 480p rates are $0.05 and $0.068 per second.

A request body written for one tier runs on the other after an endpoint change.

All six tests in this article ran on Wan 3.0 Prime.

Can Wan 2.7 edit a clip generated with Wan 3.0 Prime?

Yes, as long as the clip runs between 2 and 10 seconds and stays under 100 MB as an MP4 or MOV file.

Wan 2.7 edit-video takes a text instruction and an optional reference image, and it outputs at 720p or 1080p.

Setting audio_setting to origin keeps the Wan 3.0 Prime soundtrack in the edited result.

What are the audio input limits on Wan 2.7 and Wan 3.0 Prime?

Wan 2.7 takes one WAV or MP3 of up to 30 seconds and 15 MB through audio_url on text and image-to-video.

The minimum length is 3 seconds for text-to-video and 2 seconds for image-to-video.

Wan 3.0 Prime reads up to five reference audio files through reference_audio_urls, with a 15-second combined limit.

Neither Wan 2.7 reference to video nor Wan 3.0 Prime text or image to video defines an audio input field.

Which fields change when a request moves from Wan 2.7 to Wan 3.0 Prime?

The endpoint string changes first, and on image to video image_url becomes start_image_url.

negative_prompt and audio_url come out of the request, since Wan 3.0 Prime text and text-to-video define neither.

video_url for continuation and multi_shots on reference to video have no Wan 3.0 Prime equivalent either.

Wan 3.0 Prime types duration as an integer that also accepts null for smart duration.

On Wan 2.7, duration is an enum from 2 to 15 seconds, or 2 to 10 on reference to video.

Wan 3.0 Prime adds audio and enable_thinking as optional fields, and the queue and file upload calls stay the same across both families.

About the author
John Ozuysal

Founder of House of Growth. 2x entrepreneur, 1x exit, mentor at 500, Plug and Play, and Techstars.

Build with generative media on fal

Hundreds of production-ready image, video, and audio models behind one API.