Customize your input with more control.
Customize your input with more control.
Hint: Drag and drop files from your computer, images from web pages, paste from clipboard (Ctrl/Cmd+V), or provide a URL.
Customize your input with more control.
The cost of training depends on the number of steps. The formula is: 0.01 * steps. With 1000 steps, your request will cost $10.00.
Fine-tune your training parameters and start right now.
/image-to-video-audio)The /image-to-video-audio endpoint trains a LoRA adapter for the MiniMax H3 (Hailuo-03) video+audio model with first-frame conditioning: during training, each sample is conditioned on its opening frame with probability first_frame_conditioning_p (default 0.5), teaching the LoRA to animate from a starting image. The result works in both H3's image-to-video and text-to-video modes.
Key features:
.txt train against the trigger phrase (or a blank prompt).<stem>.first_frame.png next to a clip to condition on that image instead of the clip's literal first frame.Provide a single .zip archive (linked via training_data_url) containing your clips and, optionally, captions and sidecar keyframes:
.mp4, .mov, .avi, .mkv.txt file with the same base name as each media file (e.g. clip01.mp4 + clip01.txt). Clips without a caption train on the trigger phrase, or a blank prompt if none is set.<stem>.first_frame.png (also .jpg/.jpeg, any case) next to clip01.mp4 replaces the clip's own first frame as the conditioning image. Useful when the ideal conditioning image differs from the literal opening frame (e.g. a clean product shot vs. a motion-blurred frame).The archive must contain video clips only — image datasets are rejected (422), as are mixed image/video archives (sidecar keyframes are conditioning images, not dataset media, and stay allowed). Aim for at least 10 clips; more is generally better. Files in subfolders are fine — clips with the same name in different subfolders are kept distinct, and each sidecar binds to the clip in its own folder.
Minimum clip length: with auto_scale_input off (the default), each video must already have at least number_of_frames frames (default 73, ≈ 3.0 s at 24 fps). Shorter clips are silently skipped, and if every clip is too short the request fails (422). Turn on auto_scale_input to resample shorter clips to the target frame count instead.
training_data_url (required)Type: string
URL to the .zip archive of training clips (and optional captions/sidecars). See "Dataset Format" above.
trigger_phraseType: string
Default: ""
A phrase prepended to every caption during training (and used as the whole caption for clips without one). At inference, including this phrase activates the learned concept.
first_frame_conditioning_pType: number
Default: 0.5 (range 0.0–1.0)
Probability that a training sample is conditioned on its first frame (or its sidecar keyframe). Higher values weight the LoRA toward image-to-video behavior; the remaining probability mass trains unconditioned text-to-video. At 0.0 no image conditioning is trained.
rankType: integer (one of 8, 16, 32, 64, 128)
Default: 32
LoRA capacity. Higher values can capture more detail but use more memory and are more prone to overfitting on small datasets.
| Value | Use Case |
|---|---|
| 8–16 | Small datasets, subtle styles, lower risk of overfitting |
| 32 | Balanced default |
| 64–128 | Larger datasets or complex subjects with lots of variation |
number_of_stepsType: integer
Default: 2000 (range 1–6000)
How many optimization steps to run. More steps means more learning but also more time and a higher chance of overfitting. Note that billing has a 100-step floor (see Billing).
learning_rateType: number
Default: 0.0002
How aggressively the model updates each step. The default is a sensible starting point; raise it cautiously and lower it if results look unstable or degraded.
number_of_framesType: integer
Default: 73 (range 22–124)
Frames per training clip. H3's video VAE requires frames % 17 == 5 — valid counts are 22, 39, 56, 73, 90, 107, 124. Other values are adjusted with 17 * (value // 17) + 5 and clamped to the supported range rather than using nearest-value rounding; for example, 100 becomes 90 while 123 becomes 124.
frame_rateType: integer
Default: 24 (range 8–60)
Target frames per second for training clips. H3's native output rate is a fixed 24 fps, so keep the default unless you have a specific reason to deviate.
resolutionType: string (one of low, medium, high)
Default: medium
Training resolution bucket. Combined with aspect_ratio this picks the exact pixel size:
| Resolution | 21:9 | 16:9 | 4:3 | 1:1 | 3:4 | 9:16 |
|---|---|---|---|---|---|---|
| low | 672×288 | 512×288 | 512×384 | 512×512 | 384×512 | 288×512 |
| medium | 896×384 | 768×448 | 672×512 | 768×768 | 512×672 | 448×768 |
| high | 1120×480 | 960×544 | 896×672 | 960×960 | 672×896 | 544×960 |
H3's native output is 2K; LoRA training at these sub-native resolutions is standard practice across the video-trainer family and carries over to full-resolution inference.
aspect_ratioType: string (one of 21:9, 16:9, 4:3, 1:1, 3:4, 9:16)
Default: 16:9
Aspect ratio for training clips — the same set the H3 API supports. See the table above.
auto_scale_inputType: boolean
Default: false
When true, videos are automatically fit to the target frame count and frame rate.
split_input_into_scenesType: boolean
Default: true
When true, videos longer than the duration threshold are automatically split into separate scenes (shots) before training. Note: sidecar keyframes are dropped for clips that get split (the keyframe can't be matched to a specific scene) — disable this if your dataset uses sidecars on long clips.
split_input_duration_thresholdType: number
Default: 30.0 (range 1.0–60.0)
Videos longer than this many seconds are eligible for scene splitting.
lora_file — the trained LoRA weights (.safetensors). This is the main artifact.config_file — a small JSON containing the trigger phrase as instance_prompt and training_type set to i2va, for setting up inference.debug_dataset — a downloadable archive of your preprocessed data, only present when debug_dataset is enabled.The endpoint does not accept validation samples and does not return preview inference; evaluate the LoRA by loading it into the inference endpoint.
debug_datasetType: boolean
Default: false
When enabled, returns an archive of the preprocessed training data so you can verify your videos, captions, and keyframes were processed correctly before committing to a longer run. The archive includes preprocessing_report.json, a bounded summary of any dataset fallbacks.
strict_datasetType: boolean
Default: false
Preprocessing always logs a summary when it uses a blank prompt, substitutes the clip's first frame for a missing sidecar, or substitutes silence for an unreadable or absent video soundtrack. Enable this option to reject those fallbacks with a 422 instead. Captions longer than 4,096 tokenizer tokens are rejected with a 422 in either mode, before the GPU-heavy encoders load.
A successful run is billed max(100, number_of_steps) step units at $0.01000 per step (category: training). The default 2,000-step run bills 2,000 units = $20.00; the 6,000-step maximum bills $60.00; requests under 100 steps are floored to 100 units = $1.00. Requests that fail before training completes (input-validation errors / HTTP 422, or dataset-download failures) are billed 0 units. If the training node is interrupted, the run resumes from its latest checkpoint under the same request — you are billed once, on completion.
number_of_steps; per sample, first-frame conditioning is applied with probability first_frame_conditioning_p. Resume checkpoints are written continuously.When a sample is conditioned, the conditioning image — the sidecar <stem>.first_frame.png if present, else the clip's own first frame — is encoded and packed into the training sequence at a near-clean signal level, so the model learns to generate the clip given that image. This is the exact mechanism behind H3's image-to-video inference mode; training it with your data teaches the LoRA how your subject moves when animated from a still.
A LoRA is a small set of adapter weights layered on top of the frozen base model — here, H3's attention projections. Training only updates these adapters, so the result is a compact file you load alongside the base H3 model at inference. The base model's general capabilities are preserved; the LoRA nudges it toward your subject or style.
Captions are optional here, but they still steer the LoRA when present:
tronl0g0 the logo assembles itself from glowing particles.trigger_phrase so all clips share a consistent anchor; fully caption-less training (blank prompts) works but gives you less prompt control at inference.0.5 (default) balances image-to-video and text-to-video ability.0.8–1.0 when the LoRA will only ever be used for image-to-video.0.2–0.3 when text-prompted generation matters more but you still want first-frame support.If split_input_into_scenes is on, one long video becomes several shorter clips that share the original caption — and any sidecar keyframe for that video is dropped. For precise captions and reliable sidecar binding, pre-split your clips and disable scene splitting.
json{ "training_data_url": "https://example.com/my_dataset.zip", "trigger_phrase": "tronl0g0", "rank": 32, "number_of_steps": 2000, "learning_rate": 0.0002, "first_frame_conditioning_p": 0.5, "number_of_frames": 73, "frame_rate": 24, "resolution": "medium", "aspect_ratio": "16:9", "split_input_into_scenes": false }
number_of_steps, raise rank, or improve captions.first_frame_conditioning_p is well above 0 and that your clips' first frames actually show the subject.split_input_into_scenes..first_frame.png sidecars.% 17 == 5 rule (they are snapped automatically — check the logs if durations look off).