Every model on fal runs on Serverless.

Serverless is fal's inference runtime, built for generative media workloads. Bring a container, a Python app, or an existing inference server, and run it on managed GPUs with autoscaling, logs, and analytics built in.

Access is reviewed by our team. Serverless is for teams running production workloads, and a team member follows up on every request.

Requests by app

How Serverless works

3 steps from application code to a production endpoint.

However your application arrives, it lands on the same managed infrastructure.

1

Bring what you have

Your Dockerfile, your weights, your dependencies. Migration guides for Modal, Baseten, and RunPod, or wrap a model in a fal.App class.

2

Deploy on managed GPUs

fal deploy builds, pushes, warms, and serves your app behind a stable endpoint on H100s through B300s.

3

Run and scale

Runners start on demand, scale with traffic, and return to zero. Logs, analytics, and request traces are built in.

What's included

An app.py defining a fal App, beside a terminal running fal deploy

Demo applications

Open an application to see its metrics, logs, and runner activity.

View all

Tutorials & guides

Step-by-step guides for deploying and migrating your own workloads.

View all

Ready to deploy your own?

Bring an existing container, migrate an inference server, or build with fal.App — on managed GPUs, with logs and analytics built in.

Talk to an engineer