Minimax logo
minimax/h3/ref2va/trainer

Train a MiniMax H3 LoRA with reference conditioning, so different modalities animate into video with audio; captions optional.
Training
Commercial use

About

Reference-conditioned (ref2va) LoRA training; captions optional.

1. Calling the API#

Install the client#

The client provides a convenient way to interact with the model API.

npm install --save @fal-ai/client

Setup your API Key#

Set FAL_KEY as an environment variable in your runtime.

export FAL_KEY="YOUR_API_KEY"

Submit a request#

The client API handles the API submit protocol. It will handle the request status updates and return the result when the request is completed.

import { fal } from "@fal-ai/client";

const result = await fal.subscribe("minimax/h3/ref2va/trainer", {
  input: {
    training_data_url: ""
  },
  logs: true,
  onQueueUpdate: (update) => {
    if (update.status === "IN_PROGRESS") {
      update.logs.map((log) => log.message).forEach(console.log);
    }
  },
});
console.log(result.data);
console.log(result.requestId);

2. Authentication#

The API uses an API Key for authentication. It is recommended you set the FAL_KEY environment variable in your runtime when possible.

API Key#

In case your app is running in an environment where you cannot set environment variables, you can set the API Key manually as a client configuration.
import { fal } from "@fal-ai/client";

fal.config({
  credentials: "YOUR_FAL_KEY"
});

3. Queue#

Submit a request#

The client API provides a convenient way to submit requests to the model.

import { fal } from "@fal-ai/client";

const { request_id } = await fal.queue.submit("minimax/h3/ref2va/trainer", {
  input: {
    training_data_url: ""
  },
  webhookUrl: "https://optional.webhook.url/for/results",
});

Fetch request status#

You can fetch the status of a request to check if it is completed or still in progress.

import { fal } from "@fal-ai/client";

const status = await fal.queue.status("minimax/h3/ref2va/trainer", {
  requestId: "764cabcf-b745-4b3e-ae38-1200304cf45b",
  logs: true,
});

Get the result#

Once the request is completed, you can fetch the result. See the Output Schema for the expected result format.

import { fal } from "@fal-ai/client";

const result = await fal.queue.result("minimax/h3/ref2va/trainer", {
  requestId: "764cabcf-b745-4b3e-ae38-1200304cf45b"
});
console.log(result.data);
console.log(result.requestId);

4. Files#

Some attributes in the API accept file URLs as input. Whenever that's the case you can pass your own URL or a Base64 data URI.

Data URI (base64)#

You can pass a Base64 data URI as a file input. The API will handle the file decoding for you. Keep in mind that for large files, this alternative although convenient can impact the request performance.

Hosted files (URL)#

You can also pass your own URLs as long as they are publicly accessible. Be aware that some hosts might block cross-site requests, rate-limit, or consider the request as a bot.

Uploading files#

We provide a convenient file storage that allows you to upload files and use them in your requests. You can upload files using the client API and use the returned URL in your requests.

import { fal } from "@fal-ai/client";

const file = new File(["Hello, World!"], "hello.txt", { type: "text/plain" });
const url = await fal.storage.upload(file);

Read more about file handling in our file upload guide.

5. Schema#

Input#

training_data_url string* required

URL to zip archive with videos or images of your subject. Try to use at least 10 files, although more is better.

Supported video formats: .mp4, .mov, .avi, .mkv Supported image formats: .png, .jpg, .jpeg

Note: The dataset must contain ONLY videos OR ONLY images - mixed datasets are not supported.

Optional per-clip extras, matched by base name: a caption text file (clip01.txt), and ORDERED reference sidecars clip01.ref_1.<ext> .. clip01.ref_4.<ext> where the extension picks the modality — images (.png/.jpg/.jpeg), videos (.mp4/.mov/.avi/.mkv/.webm; their soundtrack becomes an audio reference too), or audio (.wav/.mp3/.flac/.m4a/.ogg/.aac). Clips without reference sidecars train against a frame of the clip itself as the reference.

rank RankEnum

The rank of the LoRA adaptation. Higher values increase capacity but use more memory. Default value: "32"

Possible enum values: 8, 16, 32, 64, 128

number_of_steps integer

The number of training steps. Default value: 2000

learning_rate float

Learning rate for optimization. Higher values can lead to faster training but may cause overfitting. Default value: 0.0002

number_of_frames integer

Number of frames per training sample. Must satisfy frames % 17 == 5 (e.g., 22, 39, 56, 73, 90, 107, 124). Default value: 73

frame_rate integer

Target frames per second for the video. Default value: 24

resolution ResolutionEnum

Resolution to use for training. Higher resolutions require more memory. Default value: "medium"

Possible enum values: low, medium, high

aspect_ratio AspectRatioEnum

Aspect ratio to use for training. Matches the aspect ratios supported by the MiniMax H3 API. Default value: "16:9"

Possible enum values: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16

trigger_phrase string

A phrase that will trigger the LoRA style. Will be prepended to captions during training. Default value: ""

reference_conditioning_p float

Probability of conditioning on the reference images during training. The remaining probability mass trains unconditioned text-to-video-audio, which keeps prompt-only generation stable. Default value: 0.9

resume_from_lora_url string

URL of a previously trained LoRA (.safetensors) from this trainer to WARM-START from: its weights initialize the adapter and training continues for number_of_steps more steps. Weights only — the optimizer and learning-rate schedule start fresh.

auto_scale_input boolean

If true, videos will be automatically scaled to the target frame count and fps. This option has no effect on image datasets.

split_input_into_scenes boolean

If true, videos above a certain duration threshold will be split into scenes. Note: reference sidecars are dropped for clips that get split. Default value: true

split_input_duration_threshold float

The duration threshold in seconds. If a video is longer than this, it will be split into scenes. Default value: 30

debug_dataset boolean

When enabled, the trainer returns a downloadable archive of your preprocessed training data for manual inspection. Use this to verify that your videos, images, captions and reference images were processed correctly before committing to a full training run.

strict_dataset boolean

Reject dataset-quality fallbacks during preprocessing, including captions that would use a blank prompt, missing reference sidecars, and videos without a readable audio track. By default these cases are reported in logs and use the documented fallback.

{
  "training_data_url": "",
  "rank": 32,
  "number_of_steps": 2000,
  "learning_rate": 0.0002,
  "number_of_frames": 73,
  "frame_rate": 24,
  "resolution": "medium",
  "aspect_ratio": "16:9",
  "trigger_phrase": "",
  "reference_conditioning_p": 0.9,
  "auto_scale_input": false,
  "split_input_into_scenes": true,
  "split_input_duration_threshold": 30,
  "strict_dataset": false
}

Output#

lora_file File* required

URL to the trained LoRA weights (.safetensors).

config_file File* required

Configuration used for setting up inference endpoints.

debug_dataset File

A downloadable archive containing the preprocessed training data. Only present when debug_dataset is enabled in the input.

{
  "lora_file": {
    "url": "",
    "content_type": "image/png",
    "file_name": "z9RV14K95DvU.png",
    "file_size": 4404019
  },
  "config_file": {
    "url": "",
    "content_type": "image/png",
    "file_name": "z9RV14K95DvU.png",
    "file_size": 4404019
  }
}

Other types#

File#

url string* required

The URL where the file can be downloaded from.

content_type string

The mime type of the file.

file_name string

The name of the file. It will be auto-generated if not provided.

file_size integer

The size of the file in bytes.

Related Models