fal-ai/vggt-1b

Turn images or video into a detailed 3D scene with depth, camera poses, and a colored point cloud.
Inference
Commercial use

About

Reconstruct 3D scene geometry from images and/or a video with VGGT-1B.

A single feed-forward pass jointly estimates camera extrinsics and intrinsics, per-frame depth and depth confidence for every input view.

1. Calling the API#

Install the client#

The client provides a convenient way to interact with the model API.

npm install --save @fal-ai/client

Setup your API Key#

Set FAL_KEY as an environment variable in your runtime.

export FAL_KEY="YOUR_API_KEY"

Submit a request#

The client API handles the API submit protocol. It will handle the request status updates and return the result when the request is completed.

import { fal } from "@fal-ai/client";

const result = await fal.subscribe("fal-ai/vggt-1b", {
  input: {},
  logs: true,
  onQueueUpdate: (update) => {
    if (update.status === "IN_PROGRESS") {
      update.logs.map((log) => log.message).forEach(console.log);
    }
  },
});
console.log(result.data);
console.log(result.requestId);

2. Authentication#

The API uses an API Key for authentication. It is recommended you set the FAL_KEY environment variable in your runtime when possible.

API Key#

In an environment where you cannot set environment variables, you can configure the API key manually on the client.
import { fal } from "@fal-ai/client";

fal.config({
  credentials: "YOUR_FAL_KEY"
});

3. Queue#

Submit a request#

The client API provides a convenient way to submit requests to the model.

import { fal } from "@fal-ai/client";

const { request_id } = await fal.queue.submit("fal-ai/vggt-1b", {
  input: {},
  webhookUrl: "https://optional.webhook.url/for/results",
});

Fetch request status#

You can fetch the status of a request to check if it is completed or still in progress.

import { fal } from "@fal-ai/client";

const status = await fal.queue.status("fal-ai/vggt-1b", {
  requestId: "764cabcf-b745-4b3e-ae38-1200304cf45b",
  logs: true,
});

Get the result#

Once the request is completed, you can fetch the result. See the Output Schema for the expected result format.

import { fal } from "@fal-ai/client";

const result = await fal.queue.result("fal-ai/vggt-1b", {
  requestId: "764cabcf-b745-4b3e-ae38-1200304cf45b"
});
console.log(result.data);
console.log(result.requestId);

4. Files#

Some attributes in the API accept file URLs as input. Whenever that's the case you can pass your own URL or a Base64 data URI.

Data URI (base64)#

You can pass a Base64 data URI as a file input. The API will handle the file decoding for you. Keep in mind that for large files, this alternative although convenient can impact the request performance.

Hosted files (URL)#

You can also pass your own URLs as long as they are publicly accessible. Be aware that some hosts might block cross-site requests, rate-limit, or consider the request as a bot.

Uploading files#

We provide a convenient file storage that allows you to upload files and use them in your requests. You can upload files using the client API and use the returned URL in your requests.

import { fal } from "@fal-ai/client";

const file = new File(["Hello, World!"], "hello.txt", { type: "text/plain" });
const url = await fal.storage.upload(file);

Read more about file handling in our file upload guide.

5. Schema#

Input#

image_urls list<string>

Input images to reconstruct. Accepts JPEG, PNG, WEBP and the other formats Pillow decodes. Images are padded to a square and resized so the longest side is 518 pixels.

video_url string

Optional input video. Frames are sampled every frame_sampling_rate-th frame (the first and last frames are always included) and appended after image_urls.

frame_sampling_rate integer

Sample every n-th frame of video_url. The first and last frames of the video are always included. Ignored when no video is given. Default value: 24

alpha_blend_onto AlphaBlendOntoEnum

How to handle images with an alpha channel. 'white'/'black' composite onto that background, 'mean' composites onto the ImageNet mean RGB, and 'keep' discards alpha and keeps the original pixel values. Default value: "white"

Possible enum values: keep, white, black, mean

export_prediction_data boolean

Return one JSON file per frame with the raw model outputs (camera pose encoding, depth, depth confidence, mask). Default value: true

export_depth_maps boolean

Return one 16-bit grayscale PNG depth raster per frame. Pixel values map linearly from depth_range onto 0-65535. Default value: true

export_point_cloud boolean

Return a GLB scene holding the fused coloured point cloud plus a cone mesh per estimated camera. Default value: true

array_encoding ArrayEncodingEnum

How arrays are encoded inside the prediction JSON files. 'base64' emits {data, shape, dtype} objects; 'list' emits nested JSON lists, which are far larger. Default value: "base64"

Possible enum values: base64, list

exclude_keys list<Enum>

Keys to omit from each prediction JSON file.

confidence_percentile float

Drop this percentage of the lowest-confidence points before building the point cloud. 0 keeps every point. Default value: 50

max_points integer

Upper bound on the number of points written to the GLB. Points above the default 250,000-point budget are evenly subsampled to limit GLB serialization and upload latency. Default value: 250000

enable_safety_checker boolean

Enable safety checking of the input images and video. Default value: true

{
  "image_urls": [
    "https://storage.googleapis.com/falserverless/example_inputs/dog.png"
  ],
  "frame_sampling_rate": 24,
  "alpha_blend_onto": "white",
  "export_prediction_data": true,
  "export_depth_maps": true,
  "export_point_cloud": true,
  "array_encoding": "base64",
  "confidence_percentile": 50,
  "max_points": 250000,
  "enable_safety_checker": true
}

Output#

point_cloud File

GLB scene with the fused point cloud and camera cones (when export_point_cloud is true).

depth_maps list<Image>

One 16-bit grayscale PNG depth raster per frame, in frame order (when export_depth_maps is true).

prediction_data list<File>

One JSON file of raw predictions per frame, in frame order (when export_prediction_data is true).

num_frames integer* required

Number of frames reconstructed (images plus sampled video frames).

depth_range list<float>* required

The [min, max] depth used to quantize depth_maps. depth = min + (pixel / 65535) * (max - min).

extrinsics list<list<list<float>>>* required

Per-frame 3x4 camera-from-world extrinsics in OpenCV convention (x-right, y-down, z-forward).

intrinsics list<list<list<float>>>* required

Per-frame 3x3 pinhole intrinsics, in pixels of the 518x518 model frame.

timings Timings

Wall-clock seconds per stage.

{
  "depth_maps": [
    {
      "url": "",
      "content_type": "image/png",
      "file_name": "z9RV14K95DvU.png",
      "file_size": 4404019,
      "width": 1024,
      "height": 1024
    }
  ],
  "prediction_data": [
    {
      "url": "",
      "content_type": "image/png",
      "file_name": "z9RV14K95DvU.png",
      "file_size": 4404019
    }
  ]
}

Other types#

Image#

url string* required

The URL where the file can be downloaded from.

content_type string

The mime type of the file.

file_name string

The name of the file. It will be auto-generated if not provided.

file_size integer

The size of the file in bytes.

width integer

The width of the image in pixels.

height integer

The height of the image in pixels.

File#

url string* required

The URL where the file can be downloaded from.

content_type string

The mime type of the file.

file_name string

The name of the file. It will be auto-generated if not provided.

file_size integer

The size of the file in bytes.