Skip to main content
fal Agent is a conversational creative partner. You describe what you want to make. The agent chooses the models, writes the prompts, and runs the generations. It keeps your references consistent from one step to the next. It works across the full fal model catalog. One chat can move from a still image to a video, to a sound bed, and to a 3D asset without a change of tool. fal Agent runs in the browser at fal.ai/agent. It is a product for creators, studios, and teams who want production-grade results without hand-picking a model for every step.
fal Agent is in Early Access. Access comes with a fal Agent credit tier or an enterprise agreement. See Access and pricing.

What fal Agent does

  • Chooses and runs models. The agent picks the model for the task, writes the inputs, and repairs failed runs. You can pin a model when you want a specific one.
  • Plans multi-step work. Larger requests become an editable plan card. You can rename steps, reorder them, pin models, and add approval checkpoints before anything runs.
  • Keeps context. Projects hold chats, media, documents, characters, and a shared memory. The agent recalls decisions and preferences across sessions.
  • Edits with code. A Python sandbox with ffmpeg, Pillow, and a video compositing renderer handles deterministic edits, captions, and cuts.
  • Reads the web and video. The agent can search the web, read pages, download public media, and answer questions about a video with timestamps.
  • Arranges video sequences. Clips become an editable timeline with audio layers. You export the final cut as one file.
  • Follows skills. Built-in skills from fal encode creative guidelines for characters, cinematography, UGC ads, photo editing, and more. You can write your own.

How a turn works

Every message you send starts a turn. The agent reasons about the request, calls tools, and streams its work into the chat.
1

Understand

The agent reads your message, your attachments, the active skills, and the project memory. If the request is ambiguous, it can ask a short multiple-choice question.
2

Plan

For multi-step work, the agent renders a plan card. You can edit the plan or run it as written. Simple requests skip this step.
3

Run

Each generation is a normal fal model run. The agent prepares the input and submits the run to the fal queue. Results stream into the chat and the media rail as they land.
4

Continue

When a run lands, the agent can continue with the next step automatically. Every remaining step is visible in the queue, and you can stop, reorder, or edit it at any time.

What you pay for

Model runs are billed as normal fal usage against your credits. Nothing else is billed. The agent’s reasoning, the sandbox, web search, and video understanding are free to use for now. A spend chip next to the composer shows the running total for the chat, and spending caps ask for confirmation before expensive runs.

Start here

Quickstart

Run your first chat in five minutes

Access and pricing

Credit tiers, enterprise access, and what is billed

Plan cards

Edit, approve, and version the agent’s plan

Projects

Shared memory, documents, and media across chats

Skills

Built-in creative guidelines and your own skills

Tools

Sandbox, web, video understanding, sequences, and more