Skip to main content
Every model on the fal Marketplace is a fal.App running on Serverless. You can publish your own app to the marketplace, making it callable by any fal user who authenticates with their own API key and pays for their own usage. Publishing involves deploying your app in shared auth mode, defining billable units so callers are charged correctly, and working with the fal team to get listed. Before publishing, your app should be deployed and stable, scaled for your expected traffic, and have health checks and analytics enabled.

Set your app to shared mode

Shared mode means callers must authenticate with their own fal API key. Each caller is billed for their own usage — the app owner is not charged for caller compute.
Or via the CLI:
Shared mode requires admin enablement on your account. Contact the fal team to get it enabled before deploying with app_auth = "shared".

Define billable units

When your app runs in shared mode, you control how callers are charged by setting the x-fal-billable-units response header. This tells the platform how many billing units to charge for each request.
The cost per unit is configured separately by the fal team in the billing system — your code only reports how many units each request consumed. If you don’t set the header, the platform falls back to per-second billing based on the GPU machine type used. Reported units change what a request costs only once the fal team has configured a unit price for that endpoint. Until then the platform charges every request by duration at the machine rate, even when it carries the header or you report units for it. Ask the fal team to configure pricing when you are ready to test unit billing.

Common patterns

Per image (flat rate) Charge one unit per generated image, regardless of resolution:
Per megapixel Scale the charge with output resolution. Higher resolutions cost proportionally more:
Per video second Charge based on the duration of generated video:
Per text chunk For text-based models (TTS, LLMs), scale with input length:
Flat per request Charge a fixed amount regardless of input or output:

Report units after the request

Some apps only know the unit count after the response has started. A streaming transcription counts audio seconds as they arrive, for example. For these apps, set the x-fal-billable-units-webhook header to 1 instead of a unit count.
The platform then records no charge for the request and waits for your report. When the work is done, report the final unit count from your own backend:
  • $REQUEST_ID is the x-fal-request-id header that your app receives with the request. See Request Headers.
  • $FAL_KEY must be an API key of the account that owns the app. A key from another account gets a 403 response.
  • The first report returns 200. A second report for the same request returns 409.
  • The platform does not charge the caller for a request that ended with a 5xx status, even if you report units for it.
Report every request that you flag with x-fal-billable-units-webhook. The platform never charges a flagged request until your report arrives. This also applies to a WebSocket session that ends with a gateway error.

WebSocket endpoints

A WebSocket endpoint sends one HTTP response, the 101 upgrade. Set the billing headers on that response. In a fal app, pass them to accept():
Apps that run their own server set the same headers on the upgrade response. If the unit count is known when the connection opens, set x-fal-billable-units on the upgrade response instead. Most WebSocket apps count units during the session, so they use the webhook header and report the total when the session ends. Without either header, the platform charges a WebSocket session by its duration at the machine rate. If a session ends with a gateway error, the platform ignores the unit count from the upgrade response and charges that session by duration. A session flagged with the webhook header still waits for your report. Report each session once. A second report for the same request returns 409, which means the platform has already recorded the session, not that the report failed. The request’s cost reflects your reported units only after a unit price exists for the endpoint.

Get listed on the Marketplace

Once your app is deployed in shared mode with billable units configured, contact the fal team to get listed on fal.ai/models. The team will configure:
  • Model card — name, description, category, and example inputs/outputs
  • Pricing — the cost per billable unit, visible on your model’s page and at fal.ai/pricing
  • Visibility — when and how your model appears in the marketplace
After listing, callers see your model’s pricing on its page and the X-Fal-Billable-Units header in every response, so they can track their usage programmatically.

Deploy to Production

Deployment strategies and authentication modes.

Model APIs Pricing

How marketplace model pricing works from the caller’s perspective.