Customize your input with more control.
Customize your input with more control.
Hint: Drag and drop files from your computer, images from web pages, paste from clipboard (Ctrl/Cmd+V), or provide a URL.
Customize your input with more control.
The cost of training depends on the number of steps. The formula is: 0.005 * steps. With 1000 steps, your request will cost $5.00.
Fine-tune your training parameters and start right now.
/text-to-video-audio)The /text-to-video-audio endpoint trains a LoRA adapter for the MiniMax H3 (Hailuo-03) video+audio model on your own clips, so the model learns a new subject, character, object, or visual style that you can then summon at inference time with a text prompt. This endpoint trains the pure text-to-video objective: no keyframe conditioning is used, and the resulting LoRA is intended for H3's text-to-video mode.
Key features:
Provide a single .zip archive (linked via training_data_url) containing your clips and captions:
.mp4, .mov, .avi, .mkv.txt file with the same base name as each media file (e.g. clip01.mp4 + clip01.txt).Captions are required on this endpoint — every media file needs a matching .txt, or you must set a trigger_phrase (which then stands in as the caption for uncaptioned files). If neither is present the request fails with a 422. If you want caption-less training, use the /image-to-video-audio or /first-last-frame-to-video-audio endpoints instead.
The archive must contain video clips only — image datasets are rejected (422), as are mixed image/video archives. Aim for at least 10 clips; more is generally better. Files in subfolders are fine — clips with the same name in different subfolders are kept distinct automatically.
Minimum clip length: with auto_scale_input off (the default), each video must already have at least number_of_frames frames (default 73, ≈ 3.0 s at 24 fps). Shorter clips are silently skipped, and if every clip is too short the request fails (422). Turn on auto_scale_input to resample shorter clips to the target frame count instead.
training_data_url (required)Type: string
URL to the .zip archive of training clips and captions. See "Dataset Format" above.
trigger_phraseType: string
Default: ""
A phrase prepended to every caption during training. At inference, including this phrase activates the learned concept. Leave empty when teaching a general style you always want applied.
rankType: integer (one of 8, 16, 32, 64, 128)
Default: 32
LoRA capacity. Higher values can capture more detail but use more memory and are more prone to overfitting on small datasets.
| Value | Use Case |
|---|---|
| 8–16 | Small datasets, subtle styles, lower risk of overfitting |
| 32 | Balanced default |
| 64–128 | Larger datasets or complex subjects with lots of variation |
number_of_stepsType: integer
Default: 2000 (range 1–6000)
How many optimization steps to run. More steps means more learning but also more time and a higher chance of overfitting. Note that billing has a 100-step floor (see Billing).
learning_rateType: number
Default: 0.0002
How aggressively the model updates each step. The default is a sensible starting point; raise it cautiously and lower it if results look unstable or degraded.
number_of_framesType: integer
Default: 73 (range 22–124)
Frames per training clip. H3's video VAE requires frames % 17 == 5 — valid counts are 22, 39, 56, 73, 90, 107, 124. Other values are adjusted with 17 * (value // 17) + 5 and clamped to the supported range rather than using nearest-value rounding; for example, 100 becomes 90 while 123 becomes 124.
frame_rateType: integer
Default: 24 (range 8–60)
Target frames per second for training clips. H3's native output rate is a fixed 24 fps, so keep the default unless you have a specific reason to deviate.
resolutionType: string (one of low, medium, high)
Default: medium
Training resolution bucket. Combined with aspect_ratio this picks the exact pixel size:
| Resolution | 21:9 | 16:9 | 4:3 | 1:1 | 3:4 | 9:16 |
|---|---|---|---|---|---|---|
| low | 672×288 | 512×288 | 512×384 | 512×512 | 384×512 | 288×512 |
| medium | 896×384 | 768×448 | 672×512 | 768×768 | 512×672 | 448×768 |
| high | 1120×480 | 960×544 | 896×672 | 960×960 | 672×896 | 544×960 |
H3's native output is 2K; LoRA training at these sub-native resolutions is standard practice across the video-trainer family and carries over to full-resolution inference.
aspect_ratioType: string (one of 21:9, 16:9, 4:3, 1:1, 3:4, 9:16)
Default: 16:9
Aspect ratio for training clips — the same set the H3 API supports. See the table above.
auto_scale_inputType: boolean
Default: false
When true, videos are automatically fit to the target frame count and frame rate.
split_input_into_scenesType: boolean
Default: true
When true, videos longer than the duration threshold are automatically split into separate scenes (shots) before training.
split_input_duration_thresholdType: number
Default: 30.0 (range 1.0–60.0)
Videos longer than this many seconds are eligible for scene splitting.
lora_file — the trained LoRA weights (.safetensors). This is the main artifact.config_file — a small JSON containing the trigger phrase as instance_prompt and training_type set to t2va, for setting up inference.debug_dataset — a downloadable archive of your preprocessed data, only present when debug_dataset is enabled.The endpoint does not accept validation samples and does not return preview inference; evaluate the LoRA by loading it into the inference endpoint.
debug_datasetType: boolean
Default: false
When enabled, returns an archive of the preprocessed training data so you can verify your videos, images, and captions were processed correctly before committing to a longer run. The archive includes preprocessing_report.json, a bounded summary of any dataset fallbacks.
strict_datasetType: boolean
Default: false
Preprocessing always logs a summary when it substitutes silence for an unreadable or absent video soundtrack. Enable this option to reject those fallbacks with a 422 instead. Captions longer than 4,096 tokenizer tokens are rejected with a 422 in either mode, before the GPU-heavy encoders load.
A successful run is billed max(100, number_of_steps) step units at $0.00500 per step (category: training). The default 2,000-step run bills 2,000 units = $10.00; the 6,000-step maximum bills $30.00; requests under 100 steps are floored to 100 units = $0.50. Requests that fail before training completes (input-validation errors / HTTP 422, or dataset-download failures) are billed 0 units. If the training node is interrupted, the run resumes from its latest checkpoint under the same request — you are billed once, on completion.
number_of_steps on H3's packed video+audio+text sequence. Resume checkpoints are written continuously..zip is unpacked; macOS __MACOSX metadata folders are ignored..txt of the same base name.auto_scale_input, clips are resampled to the target frame rate and frame count.split_input_into_scenes, clips longer than the threshold are cut into separate shots, each becoming its own training sample.A LoRA is a small set of adapter weights layered on top of the frozen base model — here, H3's attention projections. Training only updates these adapters, so the result is a compact file you load alongside the base H3 model at inference. The base model's general capabilities are preserved; the LoRA nudges it toward your subject or style.
tronl0g0 a glowing blue logo spinning on a desk.Good caption: a red sports car drives along a coastal highway at sunset
Weak caption: car (too sparse to anchor the concept)
tronl0g0) so it does not collide with words the model already knows, and include it in every caption..txt caption per file.If split_input_into_scenes is on, one long video becomes several shorter clips that all share the original caption. If different parts of the video show different things, the shared caption may not describe each split accurately. For precise captions, pre-split your clips and disable scene splitting.
Prompt the trained LoRA the same way you captioned it. If you trained with a trigger phrase, include that phrase at inference. If your captions were short and descriptive, short descriptive prompts will behave most predictably.
json{ "training_data_url": "https://example.com/my_dataset.zip", "trigger_phrase": "tronl0g0", "rank": 32, "number_of_steps": 2000, "learning_rate": 0.0002, "number_of_frames": 73, "frame_rate": 24, "resolution": "medium", "aspect_ratio": "16:9" }
number_of_steps, raise rank, or improve caption quality..txt on this endpoint (or set a trigger_phrase).% 17 == 5 rule (they are snapped automatically — check the logs if durations look off).