Customize your input with more control.
Customize your input with more control.
Hint: Drag and drop files from your computer, images from web pages, paste from clipboard (Ctrl/Cmd+V), or provide a URL.
Customize your input with more control.
The cost of training depends on the number of steps. The formula is: 0.015 * steps. With 1000 steps, your request will cost $15.00.
Fine-tune your training parameters and start right now.
/reference-to-video-audio)The /reference-to-video-audio endpoint trains a LoRA adapter for the MiniMax H3 (Hailuo-03) Ref2VA model — the reference-conditioned variant behind minimax/h3/reference-to-video. During training, each sample is conditioned on the subject's reference images with probability reference_conditioning_p (default 0.9): the references enter the model both as frozen visual condition blocks and as pictures inside the Qwen3-VL prompt, exactly the way the released ref2va inference presents them. The result is a LoRA that strengthens H3's ability to keep your subject consistent when generating from reference images.
Key features:
<Picture i> vision blocks in the prompt)..txt train against the trigger phrase (or a blank prompt).<stem>.ref_1.<ext> … <stem>.ref_4.<ext> next to each clip — images, videos, or audio files. Clips without reference sidecars train against a middle frame of the clip itself as a self-reference.resume_from_lora_url (weights only; the optimizer and learning-rate schedule start fresh).References can be images, videos, or audio files — a reference video contributes both its frames (visual reference) and, when present, its original soundtrack (audio reference); standalone audio references condition the soundtrack side only.
Provide a single .zip archive (linked via training_data_url) containing your clips and, optionally, captions and reference sidecars:
.mp4, .mov, .avi, .mkv.txt file with the same base name as each media file (e.g. clip01.mp4 + clip01.txt).clip01.ref_1.<ext> … clip01.ref_4.<ext> next to clip01.mp4 — ORDERED references of the subject appearing in that clip. The extension picks the modality: images (.png/.jpg/.jpeg), videos (.mp4/.mov/.avi/.mkv/.webm; the soundtrack, when present, becomes an audio reference too), or audio (.wav/.mp3/.flac/.m4a/.ogg/.aac). Visual references must have an aspect ratio between 1:4 and 4:1; reference videos are normalized to 24 fps and truncated to the training bucket's frame count.The archive must contain video clips only — image datasets are rejected (422), as are mixed image/video archives (reference sidecars are conditioning inputs, not dataset media, and stay allowed). Aim for at least 10 clips. Files in subfolders are fine; sidecars bind to the clip in their own folder.
Minimum clip length: with auto_scale_input off (the default), each video must already have at least number_of_frames frames (default 73, ≈ 3.0 s at 24 fps). Shorter clips are silently skipped; if every clip is too short the request fails (422). Turn on auto_scale_input to resample shorter clips instead.
training_data_url (required)Type: string
URL to the .zip archive. See "Dataset Format" above.
trigger_phraseType: string
Default: ""
A phrase prepended to every caption during training (and used as the whole caption for clips without one). Include it in inference prompts to activate the learned subject.
reference_conditioning_pType: number
Default: 0.9 (range 0.0–1.0)
Probability that a training sample is conditioned on its reference images. The remaining probability mass trains unconditioned text-to-video-audio, which keeps prompt-only generation stable. At 0.0 no reference conditioning is trained (a plain t2va run).
rankType: integer (one of 8, 16, 32, 64, 128)
Default: 32
LoRA capacity. Higher values can capture more detail but use more memory and are more prone to overfitting on small datasets.
number_of_stepsType: integer
Default: 2000 (range 1–6000)
How many optimization steps to run. Billing has a 100-step floor (see Billing).
learning_rateType: number
Default: 0.0002
How aggressively the model updates each step.
number_of_framesType: integer
Default: 73 (range 22–124)
Frames per training clip. H3's video VAE requires frames % 17 == 5 — valid counts are 22, 39, 56, 73, 90, 107, 124. Other values are adjusted with 17 * (value // 17) + 5 and clamped to the supported range rather than using nearest-value rounding; for example, 100 becomes 90 while 123 becomes 124.
frame_rateType: integer
Default: 24 (range 8–60)
Target frames per second (H3's native rate is a fixed 24 fps).
resolutionType: string (one of low, medium, high)
Default: medium
Training resolution bucket (same tables as the sibling minimax-h3-trainer; combined with aspect_ratio). Reference images are resized independently with the release's aspect-preserving recipe at the bucket's short edge.
aspect_ratioType: string (one of 21:9, 16:9, 4:3, 1:1, 3:4, 9:16)
Default: 16:9
auto_scale_inputType: boolean
Default: false
When true, videos are automatically fit to the target frame count and frame rate.
split_input_into_scenesType: boolean
Default: true
When true, videos longer than the duration threshold are split into scenes. Note: reference sidecars are dropped for clips that get split — disable this if your dataset uses sidecars on long clips.
split_input_duration_thresholdType: number
Default: 30.0 (range 1.0–60.0)
resume_from_lora_urlType: string
Default: null
URL of a previously trained ref2va LoRA (.safetensors) to initialize from. The checkpoint's adapter layout must match this trainer's (rank and target modules are validated; a mismatch fails loudly). This is a weights-only warm start: the optimizer state and learning-rate schedule start fresh, and steps count from 1 — use it to train a finished LoRA further (e.g. another few hundred steps on new data).
lora_file — the trained LoRA weights (.safetensors).config_file — a small JSON containing the trigger phrase as instance_prompt and training_type set to ref2va, for setting up inference.debug_dataset — a downloadable archive of your preprocessed data, only present when debug_dataset is enabled.The endpoint does not accept validation samples and does not return preview inference; evaluate the LoRA by loading it into the inference endpoint.
debug_datasetType: boolean
Default: false
Returns an archive of the preprocessed training data so you can verify clips, captions, and reference images before committing to a longer run. The archive includes preprocessing_report.json, a bounded summary of any dataset fallbacks.
strict_datasetType: boolean
Default: false
Preprocessing always logs a summary when it uses a blank prompt, a middle-frame self-reference in place of missing reference sidecars, or silence for an unreadable or absent video soundtrack. Enable this option to reject those fallbacks with a 422 instead. Captions longer than 4,096 tokenizer tokens are rejected with a 422 in either mode, before the GPU-heavy encoders load.
A successful run is billed max(100, number_of_steps) step units at $0.01500 per step (category: training). The default 2,000-step run bills 2,000 units = $30.00; the 6,000-step maximum bills $90.00; requests under 100 steps are floored to 100 units = $1.50. Requests that fail before training completes (input-validation errors / HTTP 422, dataset-download failures) are billed 0 units. If the training node is interrupted, the run resumes from its latest checkpoint under the same request — you are billed once, on completion.
<Picture i> vision blocks alongside the caption — the release's dual encoding.number_of_steps; per sample, reference conditioning applies with probability reference_conditioning_p. The frozen reference blocks sit before the target in the packed sequence at the 0.999 anchor timestep, exactly as ref2va inference packs them. Resume checkpoints are written continuously.ref_1..ref_4 when possible.reference_conditioning_p high (0.8–1.0) — the whole point of this mode is reference-conditioned generation; the small unconditioned share just stabilizes prompt following.split_input_into_scenes so sidecars bind reliably.json{ "training_data_url": "https://example.com/my_subject_dataset.zip", "trigger_phrase": "subj3ct", "rank": 32, "number_of_steps": 2000, "learning_rate": 0.0002, "reference_conditioning_p": 0.9, "number_of_frames": 73, "resolution": "medium", "aspect_ratio": "16:9", "split_input_into_scenes": false }
reference_conditioning_p, improve reference image quality, or add more varied clips of the subject.reference_conditioning_p slightly (e.g. 0.8) so more unconditioned steps train.split_input_into_scenes.