The inference platform for generative media and world models.

fal is an AI inference platform built for image, video, audio, 3D and world models. Call 1,400+ model endpoints through one inference API, or run your own models and pipelines on serverless GPUs that autoscale from zero, with reserved capacity when you need a guaranteed floor.

Scale
Billions
Requests served a year
Developers
2M+
Developers and teams building on fal
Model endpoints
1,400+
Image, video, audio and 3D endpoints behind one API and one key
GPUs
Up to B300
From RTX PRO 6000 to B300 and GB200, chosen per app
Trusted by teams at

Call our inference API, or serve your own model

Most teams start by calling a hosted model through the inference API. When the model is yours, they deploy it on fal Serverless. Both run on the same inference runtime and the same queue, so the client and the request stay the same.

fal Model APIs01

Inference API

Call hosted image, video, audio and 3D models with one key. Built for consumer and creative apps shipping to millions.

1,400+model endpoints behind one inference API

  • Day-zero access to new models
  • Switch models in one line
  • Pay per use, nothing while idle
See the inference API →
fal Serverless02

Serverless inference

Deploy your own model, pipeline or container on the runtime behind every fal Model API. Built for model labs and media teams.

0 → 1,000sof GPUs, scaling with your traffic

  • Your own queue and GPU runners
  • Custom scaling and hardware
  • Reserved capacity with burst
Explore the Serverless demo →

1,400+ model endpoints behind one inference API

The models teams build on, from the labs that make them, ready to call with one key. No GPUs to provision and no deploy step.

Same client for every model
import { fal } from "@fal-ai/client"; const result = await fal.subscribe("meshy/v7.1/text-to-3d", {  input: {    prompt: "A rustic, antique wooden treasure chest with a curved, domed lid, constructed from weathered, dark brown planks exhibiting prominent wood grain and subtle distress. It's heavily reinforced with broad, dark grey, oxidized metal bands secured by numerous circular rivets. Ornate, dark iron decorative elements featuring swirling foliate patterns and dragon motifs adorn the corners and lid. A prominent, circular, intricately carved metal lock plate with a central keyhole dominates the front, flanked by two large, dark metallic pull rings.",    target_polycount: 30000,  },}); console.log(result.data.model_glb.url);
javascriptmeshy/v7.1/text-to-3d
GLB · Drag to rotate
Meshy 7.1
Meshy · Example output
Try in Playground

Every new model, one API

New models from leading labs go live on fal as they are released, behind the same client and request shape.

  • Day-zero access to new releases
  • Switch models in one line

Pay for what you use

Most models are priced per output, and the rest per second of GPU time. When nothing runs, nothing is billed.

  • No GPUs to reserve or manage
  • Prices listed on every model page

Reliable by default

Queued requests go through a managed queue. It absorbs traffic spikes, and requests that hit a transient failure go back into the queue and run again.

  • Automatic retries on transient failures
  • Requests queue at your concurrency limit, not fail

Ready for enterprise

Organizations, access controls and data controls for teams shipping to production.

  • Organizations with per-model access for API and UI
  • Output retention per request, deletion on demand
  • No training on enterprise data

Serverless inference for your own models

The same runtime every fal Model API runs on. Deploy a model, pipeline or container, and fal handles the queue, the GPUs and the scaling.

Keep your Dockerfile and routes. Add your port to pyproject.toml, then run fal deploy.

1DeployShip your code with one command
Dockerfile
FROM pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime WORKDIR /appCOPY . .RUN pip install -r requirements.txt # Your existing FLUX server, unchangedCMD ["python", "server.py", "--host", "0.0.0.0", "--port", "8000"]
pyproject.toml
[tool.fal.apps.flux-server]machine_type = "GPU-H100"exposed_port = 8000keep_alive = 300max_concurrency = 1000 [tool.fal.apps.flux-server.image]dockerfile = "Dockerfile"
Terminalzsh
$
2CallYour users' requests hit the endpoint
const result = await fal.subscribe("acme/flux-server/generate", {  input: { prompt: "a knight resting by a campfire" },});
Waiting for deployPOST fal.run/acme/flux-server/generate
    3ScaleScale to thousands of GPUs, with infrastructure handled by fal
    Outputs0

    Elastic capacity

    Keep a reserved baseline of warm runners for steady traffic, and burst onto the shared fleet when demand spikes. Billed per second while runners are up.

    • Warm baseline, so steady traffic rarely waits on a cold start
    • Burst capacity for spikes and launches
    • Choose B300, GB200, B200, H200, H100 or RTX PRO 6000

    Full observability

    Requests, latency, errors, runners and cold starts in one dashboard, with logs for every request. Add your own OpenTelemetry traces.

    • Log drains and custom OpenTelemetry traces
    • Prometheus metrics for Grafana or Datadog

    Migrate in minutes

    Bring your own container image or an existing Docker server. Guides cover moving from Modal, RunPod, Baseten and Replicate.

    • Custom containers and existing servers
    • Step-by-step guides per platform

    Faster cold starts

    fal's optimization kit cuts startup time. Files on /data are cached automatically across three layers, FlashPack streams weights from disk to GPU, and compiled kernels can be shared across runners.

    • FlashPack, fal's open-source weight loader
    • Local NVMe, datacenter and global cache layers

    Distribute your model to enterprises and millions of developers

    Serve a model on fal, then sell it on fal. Publish to the marketplace and reach the developers and enterprises already calling the inference API, with fal's team co-selling enterprise deals.

    Talk to us about publishing
    Model labs on fal Serverless
    HeyGen logo
    HeyGen
    Bria logo
    Bria
    VEED logo
    VEED
    Creatify logo
    Creatify
    Krea logo
    Krea
    fal model gallery
    Google logo
    OpenAI logo
    Black Forest Labs logo
    Bytedance logo
    Minimax logo
    xAI logo
    Alibaba logo
    Kling logo
    ElevenLabs logo
    1,400+ endpoints
    2M+Developers
    EnterprisesCo-sold by fal
    Model labs on fal Serverless
    HeyGen logo
    HeyGen
    Bria logo
    Bria
    VEED logo
    VEED
    Creatify logo
    Creatify
    Krea logo
    Krea
    fal model gallery
    Google logo
    OpenAI logo
    Black Forest Labs logo
    Bytedance logo
    Minimax logo
    xAI logo
    Alibaba logo
    Kling logo
    ElevenLabs logo
    1,400+ endpoints
    2M+Developers
    EnterprisesCo-sold by fal
    1. 01

      Listed next to 1,400+ model endpoints

      Your model sits in the fal gallery and API alongside the models teams already use, in front of the developers and enterprise teams building on fal.

    2. 02

      Co-marketed at launch

      fal launches your model with you, across its channels and community, with GTM support on joint calls and custom demos.

    3. 03

      A new revenue channel

      Callers pay with their own fal key. You report billable units per request, fal sets the unit price with you, and there is no billing or infrastructure to run.

    fal doesn't act like an infrastructure vendor; it acts like a go-to-market partner. Our engineers deploy, iterate and release on their own schedule, and the time from a model being ready in research to being live for developers is measured in days, not weeks. Every launch fal amplifies brings a wave of new developers and teams. It's the reason we chose to build both on fal and with fal.
    Misha Feinstein, CTO, BriaBria

    Inference for image, video, audio, 3D and world models

    Each modality puts a different load on inference. Image wants throughput, video wants GPU memory, audio wants streaming, and world models want a live connection.

    The inference platform for world models

    fal WMA, the World Model Accelerator, streams world models over a live WebRTC connection for interactive worlds, with the same inference, serverless compute and distribution as every other model on fal.

    • Optimized for causal and bidirectional diffusion
    • Scale from 1 to 1,000 GPUs across clouds and regions
    • WebRTC with signaling and TURN relay handled

    Inference at fal's own scale

    Every fal Model API runs on the runtime you deploy to. Billions of requests served a year across 1,300+ production endpoints harden it before it reaches yours.

    H3 Max: from training to inference on fal

    fal Research post-trained H3 Max on fal Compute and serves it on fal Serverless across multiple nodes behind one endpoint, published as a Model API on the inference API. It ranks first for image-to-video on Artificial Analysis.

    Read the case study →
    Elasticity is what's changed how we work. Standing up a new model fleet on fal Serverless, sizing it against real traffic, and scaling without hardware procurement helped make our Suno v6 rollout possible.
    Georg Kucsko, CTO and Cofounder, SunoSuno

    Inference platform questions

    01What is an inference platform?

    An inference platform runs trained AI models in production and serves their outputs over an API. It manages the GPUs, autoscaling, request queues and monitoring, so the team building the product does not have to.

    02What is AI inference?

    AI inference is running a trained model on new inputs to produce outputs, such as generating an image from a prompt. Training builds the model once; inference runs every time a user makes a request, so its speed and cost decide the product's experience and margins.

    03What is the best inference platform for world models?

    fal runs world models through fal WMA, the World Model Accelerator. It streams causal and bidirectional diffusion models over WebRTC, scales from 1 to 1,000 GPUs across clouds and regions, and handles signaling and TURN relay. H3 Max Director and Abot World run on it today, and fal Worlds lets you browse them and start a live session.

    04What is the best inference platform for generative media?

    fal is built specifically for generative media inference: an inference API for 1,400+ image, video, audio and 3D model endpoints, serverless inference for your own models with reserved capacity and burst, real-time WebSocket and WebRTC transport, and a marketplace to distribute the models you serve.

    05What is serverless inference?

    Serverless inference runs a model on GPUs that start when requests arrive and stop when they do not. On fal Serverless you set floors, caps and warm buffers per app, and you are billed per second while a runner is alive.

    06What is the difference between the inference API and serverless inference?

    The inference API calls models fal already hosts, priced per output or per GPU second. Serverless inference runs your own model, pipeline or container on fal's GPUs, billed per second of runner time. Both use the same client, queue and calling patterns.

    07Does fal support new models on launch day?

    Yes. New models from leading labs go live on fal as they are released, with an API endpoint and a playground on day one.

    08What are concurrency limits on the inference API?

    Every account has a concurrency limit on how many requests run at once. New accounts start at 2 and the limit rises automatically with purchased credits, up to 40. Requests above the limit wait in the queue. Contact sales for higher limits.

    09How long does fal keep generated outputs?

    Generated files are stored on the fal CDN for at least 7 days by default. You can set the retention period per request with the X-Fal-Object-Lifecycle-Preference header, and delete a request's stored payloads and output files through the Platform API.

    10Can I reserve GPU capacity for inference?

    Yes. A reservation guarantees a baseline of GPUs for your Serverless endpoints, and traffic above it bursts onto fal's shared fleet automatically. Contact sales to size one.

    11How do I deploy a custom model for inference on fal?

    Write a fal.App class in Python with @fal.endpoint routes, bring a container image, or expose an existing HTTP server with exposed_port. Test with fal run and deploy with fal deploy.

    12How does fal reduce inference cold starts?

    Weights on /data are cached between startups automatically. FlashPack, fal's open-source loader, streams tensors from disk to GPU, and compiled kernels can be shared across runners. For latency-critical apps, min_concurrency and concurrency_buffer keep runners warm.

    13How is inference priced on fal?

    Inference API models are priced per output unit or, for some models, per GPU second, as listed on each model page, and time waiting in the queue is not billed. Serverless inference is billed per second while a runner is alive, by GPU type; time spent pending or pulling the image is not billed. Serverless GPU list prices per hour: B300 $12.99, GB200 $9.99, B200 $7.99, H200 $6.00, H100 $4.50, RTX PRO 6000 $4.00, with lower rates for volume and commitments. Reserved capacity is priced through sales. See all pricing

    14Does fal support real-time inference?

    Yes, for supported models. Real-time endpoints keep a persistent WebSocket open to a runner, models with a streaming endpoint send outputs over server-sent events, and fal WMA streams over WebRTC for interactive applications.

    15Can I sell a model I serve on fal?

    Yes. Publish a Serverless app to the fal marketplace. Callers pay with their own fal key for their own usage, and the fal team configures unit pricing and the model card with you.

    16Should I use fal for training too?

    For large training runs, fal Compute provides dedicated GPU clusters on reserved terms. Contact sales to size one. You can deploy the trained weights to fal Serverless for inference.

    Start your partnership with fal.

    Tell us what you are serving, the traffic you expect and the latency you need. The team comes back with a plan for the inference API, serverless inference or a reserved floor, and stays with you from the first deploy to launch day.

    Contact form