Alibaba logo
alibaba/wan-3.0-prime/reference-to-video

Wan 3.0 Prime Reference-to-Video combines reference images, videos, and audio into a unified video with fast generation and strong multimodal coherence. It follows character identity, visual style, movement, and sound cues across references to create controlled, consistent, and production-ready results.
Inference
Commercial use
Partner

Prompt examples

Examples are generated using the Wan 3.0 Prime. You can customize them by clicking on the "Playground" button.

Create one seamless 30-second cinematic action sequence using the exact sky courier from Image 1. Preserve her identity and appearance throughout: short silver braids, translucent amber visor, cobalt flight jacket with the white comet patch, burnt orange scarf, black utility trousers, facial features, body proportions, and clothing details. Use the athletic full-body choreography, grounded movement, and left-to-right tracking-camera language from Video 1 as motion reference, then naturally extend the performance into a complete story. 0-5 seconds: a low tracking shot races beside Nova as she sprints across the roof of a maglev train flying above a vast bioluminescent canyon at night; match the rhythm and physicality of Video 1, with strong wind pulling her scarf and jacket. 5-10 seconds: the camera rises into a fast three-quarter orbit as she clears the gap between two train cars, lands with believable weight, and keeps running while blue light streaks across her visor. 10-15 seconds: three small pursuit drones sweep in behind her; she ducks beneath one, pivots around another, and slides under a glowing overhead signal as sparks scatter across the roof. 15-20 seconds: without cutting, the camera swings in front of her and travels backward as the train enters a narrow crystalline canyon; luminous rock walls rush past, reflections moving naturally over her face and clothing. 20-25 seconds: she reaches the lead car, turns, and drops to one knee as the spoken line from Audio 1 plays clearly in the original voice; preserve the supplied voice, timing, and emotion while the train, wind, and distant mechanical ambience remain underneath it. 25-30 seconds: on the final words, Nova looks upward and hundreds of warm golden lantern drones ignite in waves across the sky; the camera cranes rapidly away to reveal the entire train crossing the glowing canyon, ending on a majestic wide shot with Nova still clearly recognizable at the center. Continuous spatial continuity, fluid cinematic camera movement, physically plausible motion, consistent character identity, rich cobalt-and-amber color contrast, detailed wind interaction, synchronized production sound, train hum, rushing air, drone passes, subtle rising orchestral score, no duplicated character, no morphing, no wardrobe changes, no captions, no logos, no text.
seed48291
aspect_ratio16:9
Playground