.Create a pixel-art side-scrolling platformer gameplay animation, presented as authentic 16-bit / 32-bit retro game footage. The game is a fast, continuous San Francisco-themed endless runner. A heroic pixel-art adventurer automatically runs from left to right while the camera smoothly tracks the character, keeping the hero roughly centered on screen. The player cannot stop or move around obstacles. Every obstacle that enters the hero’s running path must require either jumping over it or ducking underneath it. The hero must never simply run past, behind, in front of, or through an obstacle. GAMEPLAY The hero continuously runs forward through a beautiful pixel-art version of San Francisco. Obstacles approach clearly from the right side of the screen with enough visual anticipation to understand what is happening. There are only two ways to avoid hazards: JUMP: Used for anything occupying the ground-level running path. DUCK: Used for anything crossing directly above the running path at head or upper-body height. Every gameplay obstacle must physically intersect the hero’s path if no action is taken. Examples of San Francisco-themed obstacles: * A cheerful corgi runs directly across the hero’s path. The hero jumps completely over it. * A startup founder or venture capitalist charges directly toward the hero holding out a giant term sheet. The hero jumps completely over the person. * An autonomous delivery robot rolls directly into the running path. Jump over it. * An abandoned electric scooter lies horizontally across the sidewalk. Jump over it. * A shopping cart rolls directly into the running path. Jump over it. * A low-flying delivery drone crosses directly at the hero’s head level. Duck underneath it. * A flock of pigeons flies directly across the path at head height. Duck underneath them. * A strange startup prototype drone flies directly across the running path. Duck underneath it. * Low construction scaffolding extends across the sidewalk. Duck underneath it. Keep obstacles visually distinct and spaced far enough apart that every jump or duck reads clearly as a deliberate gameplay action. SAN FRANCISCO ENVIRONMENT The world should be unmistakably San Francisco while retaining the beautiful fantasy-adventure atmosphere of the original concept. The environment scrolls continuously behind the hero using multi-layer parallax scrolling: Foreground: lush grass, sidewalk edges, rocks, flowers, tiny plants, fire hydrants, cable-car rails, fallen leaves, utility covers and small particles moving quickly. Midground: colorful Victorian houses, steep streets, cafés, trees, cable cars, fences, stone walls, staircases, construction scaffolding and occasional fantasy waterfalls moving at medium speed. Background: enormous green hills mixed with recognizable San Francisco geography, dense neighborhoods, downtown skyscrapers, the Transamerica Pyramid, distant castle-like silhouettes, the Bay, the Golden Gate Bridge and enormous mountains beyond the city, moving slowly. Very distant atmospheric layer: colorful California sky, enormous clouds, drifting coastal fog, distant birds and subtle atmospheric pixel animation. Blend recognizable San Francisco architecture with a lush fantasy-adventure world. Victorian houses can sit beside ancient stone ruins. Cable-car tracks can wind through grassy hills. Waterfalls can descend beside city staircases. The Golden Gate Bridge can appear dramatically between mountains and fog. Background pedestrians, cars, buildings and scenery are decorative only. Anything presented as an actual gameplay hazard must enter the hero’s exact physical running path and trigger a jump or duck. HERO ANIMATION Use a highly polished 8–12 frame running cycle with expressive movement, bouncing hair and clothing, an animated cape, scarf or jacket, natural arm and leg motion, and tiny dust particles appearing beneath the boots. During gameplay, transition cleanly between: RUN → JUMP → LAND → RUN → DUCK → RUN Jump animations must visibly carry the hero completely over each ground obstacle. Duck animations must visibly lower the hero enough for airborne obstacles to pass directly over the character. Keep the hero’s design, face, clothing, colors, proportions and sprite dimensions 100% consistent throughout every frame. SCORE AND GAME UI Include a small, authentic retro-game score counter in the upper-right corner. The score continuously increases as the hero survives and travels farther. Use a simple pixel font and numbers, such as: SCORE 004820 The score should visibly increase during gameplay. Keep the interface minimal so it feels like an actual endless-runner game rather than a cinematic animation. Do not add menus, tutorial text, dialogue boxes or unnecessary interface elements. MUSIC AND SOUND Include energetic retro video-game music throughout the gameplay. The soundtrack should sound like a polished SNES / early PlayStation-era adventure game: upbeat, catchy and energetic, with chiptune-inspired synths, punchy retro drums, melodic bass and a bright adventurous lead melody. Add synchronized retro sound effects: * Soft rhythmic footsteps while running. * A short arcade-style jump sound whenever the hero jumps. * A quick swoosh during ducking. * Landing sound when the hero touches the ground. * Drone buzzing as drones approach. * Pigeon wing flaps. * Small environmental sounds for passing cable cars or city activity. * A subtle score or point sound after successfully clearing an obstacle. Music should remain the dominant audio layer, with sound effects mixed cleanly underneath. PIXEL-ART REQUIREMENTS Crisp hand-crafted pixel art with clearly visible individual pixels, a limited retro color palette and extremely detailed sprite work. The visual quality should resemble a premium modern indie pixel-art platformer, combining nostalgic SNES / PlayStation-era game art with highly polished modern animation. Use detailed environmental pixel art, expressive sprite animation, colorful lighting and rich parallax depth. Absolutely no 3D rendering, no vector graphics, no smooth painted edges, no realistic textures, no motion blur and no anti-aliased artwork. Everything on screen must look intentionally drawn as pixel art. CAMERA Classic side-view 2D platformer perspective. Fixed camera height. Smooth horizontal tracking. No zooming. No cuts. No camera rotation. The hero remains approximately centered while the environment moves continuously from right to left. Obstacles always approach from the right and cross the hero’s exact running line. Never allow the hero to avoid an obstacle simply because it passes in the foreground or background. If it is a gameplay obstacle, the hero must jump over it or duck underneath it. OVERALL FEEL Make the game feel like a beautiful fantasy-adventure platformer crossed with a simple offline endless runner, except the entire world is a playful pixel-art interpretation of San Francisco. Prioritize readable gameplay, detailed scenery, satisfying running animation, clear jumps and ducks, recognizable San Francisco details, a continuously increasing score, and catchy retro game music. Format: 16:9 widescreen, side-scrolling 2D gameplay, 60 FPS if supported, continuous gameplay, retro soundtrack and sound effects, minimal score HUD, no logos.
moreAI Video ModelsProduction-ready video generation APIs
Build with the latest AI video models on fal. Generate production-ready video from text, images, and references through fast, scalable APIs.
Choose the right video model for your workflow
Compare current video generation models from leading providers, selected from the active fal catalog.

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

Wan 3.0 Prime Text-to-Video transforms written prompts into polished videos with accelerated generation, fluid motion, strong scene fidelity, and coherent visual storytelling. Built for fast creative iteration, it brings complex ideas to life while preserving visual detail and cinematic consistency throughout each shot.

Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.

Generate high quality 1080p videos using Kling's Turbo 3.0 model, with improved lipsync and multishot generation capabilities.

Kling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling

Generate film-grade videos from text prompts with native audio, up to 1080p and 15 seconds, using PixVerse C1.

Veo 3.1 by Google, the most advanced AI video generation model in the world. With sound on!

Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.

The Avatar X API offers access to Mirage's most advanced generation model yet, delivering industry-leading identity preservation and expressivity in AI video

Generate high quality 1080p videos from images using Kling's Turbo 3.0 model, with improved lipsync and multishot generation capabilities.

FLUX 3 is Black Forest Labs' frontier video model. This endpoint builds video from a sequence of keyframes, generating the motion between each anchor point for precise control over how a shot progresses.

Animate images into cinematic videos with PixVerse C1, supporting 1080p resolution and native audio generation.

Veo 3.1 is the latest state-of-the art video generation model from Google DeepMind
From an idea to a production video
Use the playground for a quick test or call the same model from your application through the API.
See what the latest AI video models can create
Explore generations from the featured models, then open an example to inspect its inputs and run it yourself.
minecraft creative mode but in real world. in new york city.
moreThe near man says flatly: "Did you hear about H3 Max?" A beat of silence, wind moving across the lot. The far man's mouth pulls into a slow half-smile and he answers, warm and certain: "Dude, I love it." The near man exhales and shakes his head once, almost smiling. Native audio, lip-synced English dialogue, distant traffic hum, wind. No music, no subtitles, no camera movement.
moreTight close-up portrait of a 20 years old Gen Z fashion model, early twenties, dewy glass-skin complexion with a light dusting of faux freckles, glossy bitten-lip stain in a muted berry tone, fluffy laminated brows, and a single pearl-studded graphic liner flick in chrome silver across one eyelid. Her hair is slicked back into a sleek low bun with two soft face-framing tendrils, small silver butterfly clips catching the light. She wears oversized vintage-style chrome chandelier earrings and a sheer mesh high-neck top layered under a cropped leather moto jacket. The camera holds an extreme close-up on her face, then slowly arcs around her in a subtle 45-degree orbit as she tilts her chin down, cuts her eyes directly into the lens with a confident deadpan stare, and exhales softly. A gentle breeze lifts the loose strands of hair across her cheekbone. Lighting is soft neon-tinged — cool lavender key light from the left blending into a warm peach rim light from the right, creating a duotone gradient across her skin. Background is an out-of-focus wash of deep magenta and teal bokeh. Cinematic editorial fashion aesthetic, shallow depth of field, 85mm lens compression, subtle film grain, photorealistic, 24fps, reminiscent of a modern i-D Magazine or Vogue Beauty film.
moreThree different products, text slides and shows the actual product in the middle for each scene
moreA dynamic, ultra-realistic medium close-up of a young man juggling a soccer ball inside a high-concept editorial studio. Frame him tightly from the shins to just above his head, keeping his face, torso, raised leg, and the soccer ball prominent in the composition. Avoid a full-body or wide establishing shot. The subject should fill most of the frame, with minimal empty space around him. He wears plain, everyday athletic clothing with no visible logos or branding: a fitted solid-color T-shirt, simple athletic shorts, plain white crew socks, and neutral training shoes. His expression is focused and calm, eyes locked on the ball, with realistic tension through his shoulders, core, and raised leg. His arms extend naturally for balance. Capture him performing recognizable classic soccer juggling moves in a smooth sequence: controlled alternating foot touches, a thigh juggle, an inside-foot pop, an outside-foot touch, a knee stall, and a clean around-the-world motion before bringing the ball back under control. The movement should feel technically accurate, fluid, and athletic rather than exaggerated. Keep the soccer ball close to his body and within the central area of the frame. The camera is positioned low and close at approximately waist height, angled slightly upward for a dramatic sports-editorial perspective. It slowly pushes in and arcs around him in a subtle 45-degree orbit while maintaining the tight crop from shins to head. Use very subtle motion blur on the moving foot and ball, while keeping his face, clothing, and upper body crisp and detailed. Lighting is soft and neon-tinged, with a cool lavender key light from the left blending into a warm peach rim light from the right, creating a smooth duotone gradient across his skin, clothing, and the soccer ball. The background is an out-of-focus wash of deep magenta and teal bokeh with faint studio lights, glossy reflections, and a light atmospheric haze. The floor is dark and polished, only partially visible beneath his lower legs. Include realistic skin texture, light perspiration, fine fabric creases, subtle movement in the shirt and shorts, visible muscle tension, and small scuffs on the soccer ball. Overall mood: modern, skillful, confident, athletic, minimal, and cinematic. Luxury sports editorial aesthetic, tight framing, shallow depth of field, 85mm lens compression, subtle film grain, photorealistic detail, 24fps, reminiscent of a modern i-D Magazine or Vogue fashion film.
moreAn aging warrior-monk in scorched ceremonial armor, a split ceramic mask and a humming prayer-band on his forearm + 0–4s: climbs silently along a cliffside terrace as insect-like surveyor drones comb the sun-bleached ruins, standing alone under as the camera rises over the ruined monastery. Sun-blasted post-apocalyptic sci-fi action, dust, heat haze, practical debris, sweeping crane tracking, silence broken only by wind.
moreAlfred (calm, measured tone): "Batman, there’s a roadblock ahead. You need to reroute." Batman (gruff, determined voice): "Copy that, Alfred. Taking the alternate path."
moreBLACK MIRROR TIDEFLATS / THE COLLAPSING SEA-CATHEDRAL Style: 8K cinematic. Photorealistic — no 3D render, no game engine, no game-cutscene aesthetic. Cinematography: naturalistic master cinematography. Lighting: high-contrast dramatic light. Cold storm-blue dusk across the flats vs. blinding sodium-orange furnace glow pouring from the toppling structure. The wet mirror surface doubles every light source into a second inverted world. Color: 60:30:10 — dominant gunmetal blue and wet slate (the tideflats and sky) / secondary sodium amber and rust-red (burning floodlights, corroded steel) / accent magnesium white and cobalt blue (the failing reactor core and its arc-flashes). Camera: physical cine lens. 180° shutter motion blur. Continuous handheld operator shake, motivated by the ground concussions rolling through the shallow water. Physics: gravity and inertia respected. Steel trusses buckle and fold under their own mass; spray, ash and salt mist obey atmospheric drag. SUBJECTS: @pilot: a lone figure in a salt-crusted pressure suit crewing a skeletal wind-skiff — a low wheeled land-yacht with a torn carbon sail — hunched at the tiller, cracked gold-tinted visor reflecting the falling structure, throwing twin walls of white spray from its outriggers. @cathedral: a colossal derelict sea-rig, a kilometre of rusted iron lattice and floodlight towers, listing and buckling as it collapses into the shallows, shedding burning platform decks and cobalt arc-flash from its ruptured core. LOCATION: an endless tidal flat of black volcanic glass under two inches of standing water — a perfect mirror to the horizon — broken by rotting mooring pylons and drifting curtains of salt mist, the skyline dominated by the burning collapse. ACTION — ONE CONTINUOUS TAKE, single unbroken handheld tracking shot, NO cuts, 15 seconds. 0:00–0:05 Medium-Wide on @pilot: camera tracks alongside the skiff as it carves between rotting pylons that burst into splinters and spray behind the wheels; the horizon tears open as @cathedral enters frame, leaning, trailing burning deck plates. 0:05–0:10 continuous zoom-in / tilt up: camera glides past the skiff and tilts up through the pylons onto the collapse; a core vessel detonates, the shockwave ripping the standing water into a flat sheet and bending the flare into a horizontal streak across frame. 0:10–0:15 Extreme Close-Up on @cathedral: telephoto-compressed view of the disintegrating rig against the sky, steel plate peeling and glowing white-hot, raining burning debris down toward the camera. CONSTRAINTS: 16:9 anamorphic widescreen. ONE CONTINUOUS SHOT — absolutely NO hard cuts. Telephoto compression during the zoom. No slow-motion. Photoreal throughout. AUDIO (NO MUSIC): deep structural groan of failing steel, the hiss and slap of water under the hull, tearing-metal screech from above, the rigging-hum and canvas crack of the skiff at speed, wind roar across open flats.
moreTight close-up portrait of a Gen Z fashion model, early twenties, dewy glass-skin complexion with a light dusting of faux freckles, glossy bitten-lip stain in a muted berry tone, fluffy laminated brows, and a single pearl-studded graphic liner flick in chrome silver across one eyelid. Her hair is slicked back into a sleek low bun with two soft face-framing tendrils, small silver butterfly clips catching the light. She wears oversized vintage-style chrome chandelier earrings and a sheer mesh high-neck top layered under a cropped leather moto jacket. The camera holds an extreme close-up on her face, then slowly arcs around her in a subtle 45-degree orbit as she tilts her chin down, cuts her eyes directly into the lens with a confident deadpan stare, and exhales softly. A gentle breeze lifts the loose strands of hair across her cheekbone. Lighting is soft neon-tinged — cool lavender key light from the left blending into a warm peach rim light from the right, creating a duotone gradient across her skin. Background is an out-of-focus wash of deep magenta and teal bokeh. Cinematic editorial fashion aesthetic, shallow depth of field, 85mm lens compression, subtle film grain, photorealistic, 24fps, reminiscent of a modern i-D Magazine or Vogue Beauty film.
moreHe explodes upward out of the black. Both arms drive toward the falling apple, body twisting, white linen sleeves snapping taut then bunching as his shoulders roll. The apple tumbles fast, end over end, its stem whipping. He mistimes it — the fruit glances off his fingertips and spins away; he lunges after it, throwing his weight sideways, hair swinging across his face, eyes flashing wide. He catches it hard against his palm, the impact shoving his arm back and down, linen collapsing in loose folds. He pulls it in against his chest, breathing hard, and breaks into a grin as he looks straight into the lens. The candlelight jumps with every movement, flaring across his forearms and guttering into darkness behind him.
moreWhat can you build with AI video models?
Generate and transform video without managing model infrastructure.
Turn concepts into moving scenes
Generate cinematic shots, storyboards, and visual sequences from scripts, prompts, and reference frames.
Produce campaign video faster
Create social clips, product videos, and localized campaign variations without managing model infrastructure.
Build video generation into your product
Add text-to-video, image animation, reference-guided generation, and video utilities to automated workflows.
Prepare generated video for production
Trim, resize, blend, and reverse video with hosted utility endpoints that compose with your generation workflow.

FFMPEG Utility for Trim Video

FFMPEG Utilities to Scale Videos

FFMPEG Utility for Blending Videos

FFMPEG Utility to Reverse Videos
Common questions about AI video models
Which AI video model should I choose?
Choose based on your input, output quality, speed, duration, audio, and control requirements. Open a featured model to compare its schema, pricing, and examples before integrating it.
Can these models generate video from text and images?
Yes. fal hosts text-to-video, image-to-video, reference-to-video, and keyframe-to-video endpoints. The supported inputs and controls vary by model.
Can I use AI video models through an API?
Yes. Every featured model has a hosted fal API and an interactive playground. Open a model to review its request schema, pricing, and code examples, or create an API key from your dashboard.
Do AI-generated videos include audio?
Some models generate synchronized audio natively, while others return silent video. Check the selected endpoint's schema and model description for its audio capabilities.
What video utilities are available?
fal provides hosted utilities for common post-processing tasks including trimming, scaling, blending, and reversing video. They can be composed with generation endpoints in a production workflow.