fal-ai/vggt-1b
About
Reconstruct 3D scene geometry from images and/or a video with VGGT-1B.
A single feed-forward pass jointly estimates camera extrinsics and intrinsics, per-frame depth and depth confidence for every input view.
1. Calling the API#
Install the client#
The client provides a convenient way to interact with the model API.
npm install --save @fal-ai/clientMigrate to @fal-ai/client
The @fal-ai/serverless-client package has been deprecated in favor of @fal-ai/client. Install the new package and update your imports — see client setup.
Setup your API Key#
Set FAL_KEY as an environment variable in your runtime.
export FAL_KEY="YOUR_API_KEY"Submit a request#
The client API handles the API submit protocol. It will handle the request status updates and return the result when the request is completed.
import { fal } from "@fal-ai/client";
const result = await fal.subscribe("fal-ai/vggt-1b", {
input: {},
logs: true,
onQueueUpdate: (update) => {
if (update.status === "IN_PROGRESS") {
update.logs.map((log) => log.message).forEach(console.log);
}
},
});
console.log(result.data);
console.log(result.requestId);2. Authentication#
The API uses an API Key for authentication. It is recommended you set the FAL_KEY environment variable in your runtime when possible.
API Key#
import { fal } from "@fal-ai/client";
fal.config({
credentials: "YOUR_FAL_KEY"
});Protect your API Key
When running code on the client-side (e.g. in a browser, mobile app or GUI applications), make sure to not expose your FAL_KEY. Instead, use a server-side proxy to make requests to the API. For more information, check out our server-side integration guide.
3. Queue#
Submit a request#
The client API provides a convenient way to submit requests to the model.
import { fal } from "@fal-ai/client";
const { request_id } = await fal.queue.submit("fal-ai/vggt-1b", {
input: {},
webhookUrl: "https://optional.webhook.url/for/results",
});Fetch request status#
You can fetch the status of a request to check if it is completed or still in progress.
import { fal } from "@fal-ai/client";
const status = await fal.queue.status("fal-ai/vggt-1b", {
requestId: "764cabcf-b745-4b3e-ae38-1200304cf45b",
logs: true,
});Get the result#
Once the request is completed, you can fetch the result. See the Output Schema for the expected result format.
import { fal } from "@fal-ai/client";
const result = await fal.queue.result("fal-ai/vggt-1b", {
requestId: "764cabcf-b745-4b3e-ae38-1200304cf45b"
});
console.log(result.data);
console.log(result.requestId);4. Files#
Some attributes in the API accept file URLs as input. Whenever that's the case you can pass your own URL or a Base64 data URI.
Data URI (base64)#
You can pass a Base64 data URI as a file input. The API will handle the file decoding for you. Keep in mind that for large files, this alternative although convenient can impact the request performance.
Hosted files (URL)#
You can also pass your own URLs as long as they are publicly accessible. Be aware that some hosts might block cross-site requests, rate-limit, or consider the request as a bot.
Uploading files#
We provide a convenient file storage that allows you to upload files and use them in your requests. You can upload files using the client API and use the returned URL in your requests.
import { fal } from "@fal-ai/client";
const file = new File(["Hello, World!"], "hello.txt", { type: "text/plain" });
const url = await fal.storage.upload(file);Auto uploads
The client will auto-upload the file for you if you pass a binary object (e.g. File, Data).
Read more about file handling in our file upload guide.
5. Schema#
Input#
Input images to reconstruct. Accepts JPEG, PNG, WEBP and the other formats Pillow decodes. Images are padded to a square and resized so the longest side is 518 pixels.
video_url stringOptional input video. Frames are sampled every frame_sampling_rate-th frame (the first and last frames are always included) and appended after image_urls.
frame_sampling_rate integerSample every n-th frame of video_url. The first and last frames of the video are always included. Ignored when no video is given. Default value: 24
alpha_blend_onto AlphaBlendOntoEnumHow to handle images with an alpha channel. 'white'/'black' composite onto that background, 'mean' composites onto the ImageNet mean RGB, and 'keep' discards alpha and keeps the original pixel values. Default value: "white"
Possible enum values: keep, white, black, mean
export_prediction_data booleanReturn one JSON file per frame with the raw model outputs (camera pose encoding, depth, depth confidence, mask). Default value: true
export_depth_maps booleanReturn one 16-bit grayscale PNG depth raster per frame. Pixel values map linearly from depth_range onto 0-65535. Default value: true
export_point_cloud booleanReturn a GLB scene holding the fused coloured point cloud plus a cone mesh per estimated camera. Default value: true
array_encoding ArrayEncodingEnumHow arrays are encoded inside the prediction JSON files. 'base64' emits {data, shape, dtype} objects; 'list' emits nested JSON lists, which are far larger. Default value: "base64"
Possible enum values: base64, list
Keys to omit from each prediction JSON file.
confidence_percentile floatDrop this percentage of the lowest-confidence points before building the point cloud. 0 keeps every point. Default value: 50
max_points integerUpper bound on the number of points written to the GLB. Points above the default 250,000-point budget are evenly subsampled to limit GLB serialization and upload latency. Default value: 250000
enable_safety_checker booleanEnable safety checking of the input images and video. Default value: true
{
"image_urls": [
"https://storage.googleapis.com/falserverless/example_inputs/dog.png"
],
"frame_sampling_rate": 24,
"alpha_blend_onto": "white",
"export_prediction_data": true,
"export_depth_maps": true,
"export_point_cloud": true,
"array_encoding": "base64",
"confidence_percentile": 50,
"max_points": 250000,
"enable_safety_checker": true
}Output#
GLB scene with the fused point cloud and camera cones (when export_point_cloud is true).
One 16-bit grayscale PNG depth raster per frame, in frame order (when export_depth_maps is true).
One JSON file of raw predictions per frame, in frame order (when export_prediction_data is true).
num_frames integer* requiredNumber of frames reconstructed (images plus sampled video frames).
The [min, max] depth used to quantize depth_maps. depth = min + (pixel / 65535) * (max - min).
Per-frame 3x4 camera-from-world extrinsics in OpenCV convention (x-right, y-down, z-forward).
Per-frame 3x3 pinhole intrinsics, in pixels of the 518x518 model frame.
Wall-clock seconds per stage.
{
"depth_maps": [
{
"url": "",
"content_type": "image/png",
"file_name": "z9RV14K95DvU.png",
"file_size": 4404019,
"width": 1024,
"height": 1024
}
],
"prediction_data": [
{
"url": "",
"content_type": "image/png",
"file_name": "z9RV14K95DvU.png",
"file_size": 4404019
}
]
}Other types#
Image#
url string* requiredThe URL where the file can be downloaded from.
content_type stringThe mime type of the file.
file_name stringThe name of the file. It will be auto-generated if not provided.
file_size integerThe size of the file in bytes.
width integerThe width of the image in pixels.
height integerThe height of the image in pixels.
File#
url string* requiredThe URL where the file can be downloaded from.
content_type stringThe mime type of the file.
file_name stringThe name of the file. It will be auto-generated if not provided.
file_size integerThe size of the file in bytes.