# Vidu Q4 Reference to Video

> Vidu Q4 generates videos from up to 12 reference images and 3 voice clips, keeping characters, objects, and voices consistent across up to 16 seconds at up to 4K.


## Overview

- **Endpoint**: `https://fal.run/fal-ai/vidu/q4/reference-to-video`
- **Model ID**: `fal-ai/vidu/q4/reference-to-video`
- **Category**: image-to-video
- **Kind**: inference
**Description**: Vidu Q4 Reference-to-Video builds a scene from your references. You tag up to 12 images (characters, products, objects, locations) and up to 3 voice clips directly in the prompt, and the model keeps each subject and voice consistent throughout the clip. It supports optional native audio with dialogue and sound effects, durations from 3 to 16 seconds, five aspect ratios, and output from 540p up to 4K, which makes it well suited for branded content, character-driven storytelling, and ads.

**Tags**: reference-to-video, multi-shot, 4k



## Pricing

For every second of video you generate, you will be charged **$0.0315** at **540p**, **$0.0665** at **720p**, **$0.084** at **1080p**, **$0.133** at **2K** or **$0.273** at **4K**. Turning on **audio** does not change the price. For example, a **5s** video at **720p** (the default) will cost **$0.3325**.

Note: these are promotional rates, **30%** off until **November 30**, after which **540p** is **$0.045**/second, **720p** is **$0.095**/second, **1080p** is **$0.12**/second, **2K** is **$0.19**/second and **4K** is **$0.39**/second.

For more details, see [fal.ai pricing](https://fal.ai/pricing).

## API Information

This model can be used via our HTTP API or more conveniently via our client libraries.
See the input and output schema below, as well as the usage examples.


### Input Schema

The API accepts the following input parameters:


- **`prompt`** (`string`, _required_):
  Text prompt for video generation, max 5000 characters. Refer to references by their position in the lists: [@reference_image_1], [@reference_image_2], ... for images and [reference_audio_1], [reference_audio_2], ... for audio clips.
  - Examples: "[@reference_image_1] walks along the beach at sunset holding [@reference_image_2], then turns to the camera and speaks in the voice of [reference_audio_1]."

- **`reference_image_urls`** (`list<string>`, _optional_):
  URLs or base64 data URIs of up to 12 reference images (png, jpeg, jpg, webp) used to keep subjects consistent.
  - Array of string
  - Examples: ["https://storage.googleapis.com/falserverless/web-examples/vidu/new-examples/reference1.png","https://storage.googleapis.com/falserverless/web-examples/vidu/new-examples/reference2.png"]

- **`reference_audio_urls`** (`list<string>`, _optional_):
  URLs or base64 data URIs of up to 3 mp3 reference audio clips (3-12 seconds each, max 50 MB) used to keep voices consistent.
  - Array of string

- **`duration`** (`integer`, _optional_):
  Duration of the video in seconds (3-16) Default value: `5`
  - Default: `5`
  - Range: `3` to `16`

- **`seed`** (`integer`, _optional_):
  Random seed for reproducibility. If None, a random seed is chosen.

- **`aspect_ratio`** (`AspectRatioEnum`, _optional_):
  The aspect ratio of the output video Default value: `"16:9"`
  - Default: `"16:9"`
  - Options: `"16:9"`, `"9:16"`, `"4:3"`, `"3:4"`, `"1:1"`

- **`resolution`** (`ResolutionEnum`, _optional_):
  Output video resolution Default value: `"720p"`
  - Default: `"720p"`
  - Options: `"540p"`, `"720p"`, `"1080p"`, `"2K"`, `"4K"`

- **`audio`** (`boolean`, _optional_):
  Whether to generate audio (dialogue and sound effects) with the video. When false, the output video is silent.
  - Default: `false`



**Required Parameters Example**:

```json
{
  "prompt": "[@reference_image_1] walks along the beach at sunset holding [@reference_image_2], then turns to the camera and speaks in the voice of [reference_audio_1]."
}
```

**Full Example**:

```json
{
  "prompt": "[@reference_image_1] walks along the beach at sunset holding [@reference_image_2], then turns to the camera and speaks in the voice of [reference_audio_1].",
  "reference_image_urls": [
    "https://storage.googleapis.com/falserverless/web-examples/vidu/new-examples/reference1.png",
    "https://storage.googleapis.com/falserverless/web-examples/vidu/new-examples/reference2.png"
  ],
  "duration": 5,
  "aspect_ratio": "16:9",
  "resolution": "720p"
}
```


### Output Schema

The API returns the following output format:

- **`video`** (`File`, _required_):
  The generated video from references using the Q4 Preview model
  - Examples: {"url":"https://v3b.fal.media/files/b/0a8c9189/n9z3uUDPqmU2msAtqr25-_output.mp4"}

- **`cover_image`** (`Image`, _optional_):
  Cover image of the generated video

- **`seed`** (`integer`, _required_):
  The random seed used for generation



**Example Response**:

```json
{
  "video": {
    "url": "https://v3b.fal.media/files/b/0a8c9189/n9z3uUDPqmU2msAtqr25-_output.mp4"
  }
}
```


## Usage Examples

### cURL

```bash
curl --request POST \
  --url https://fal.run/fal-ai/vidu/q4/reference-to-video \
  --header "Authorization: Key $FAL_KEY" \
  --header "Content-Type: application/json" \
  --data '{
     "prompt": "[@reference_image_1] walks along the beach at sunset holding [@reference_image_2], then turns to the camera and speaks in the voice of [reference_audio_1]."
   }'
```

### Python

Ensure you have the Python client installed:

```bash
pip install fal-client
```

Then use the API client to make requests:

```python
import fal_client

def on_queue_update(update):
    if isinstance(update, fal_client.InProgress):
        for log in update.logs:
           print(log["message"])

result = fal_client.subscribe(
    "fal-ai/vidu/q4/reference-to-video",
    arguments={
        "prompt": "[@reference_image_1] walks along the beach at sunset holding [@reference_image_2], then turns to the camera and speaks in the voice of [reference_audio_1]."
    },
    with_logs=True,
    on_queue_update=on_queue_update,
)
print(result)
```

### JavaScript

Ensure you have the JavaScript client installed:

```bash
npm install --save @fal-ai/client
```

Then use the API client to make requests:

```javascript
import { fal } from "@fal-ai/client";

const result = await fal.subscribe("fal-ai/vidu/q4/reference-to-video", {
  input: {
    prompt: "[@reference_image_1] walks along the beach at sunset holding [@reference_image_2], then turns to the camera and speaks in the voice of [reference_audio_1]."
  },
  logs: true,
  onQueueUpdate: (update) => {
    if (update.status === "IN_PROGRESS") {
      update.logs.map((log) => log.message).forEach(console.log);
    }
  },
});
console.log(result.data);
console.log(result.requestId);
```


## Additional Resources

### Documentation

- [Model Playground](https://fal.ai/models/fal-ai/vidu/q4/reference-to-video)
- [API Documentation](https://fal.ai/models/fal-ai/vidu/q4/reference-to-video/api)
- [OpenAPI Schema](https://fal.ai/api/openapi/queue/openapi.json?endpoint_id=fal-ai/vidu/q4/reference-to-video)

### fal.ai Platform

- [Platform Documentation](https://fal.ai/docs/documentation)
- [Python Client](https://fal.ai/docs/api-reference/client-libraries/python)
- [JavaScript Client](https://fal.ai/docs/api-reference/client-libraries/javascript)

### Other agent-readable surfaces

This file covers one model. To find anything else:

- [Platform overview](https://fal.ai/llms.txt): Entry points and representative endpoint IDs
- [Documentation index](https://fal.ai/docs/llms.txt): Every documentation page
- [Full documentation text](https://fal.ai/docs/llms-full.txt): The whole documentation inlined
- Any other model: `https://fal.ai/models/<endpoint-id>/llms.txt`
