Customize your input with more control.
Customize your input with more control.
Hint: Drag and drop files from your computer, images from web pages, paste from clipboard (Ctrl/Cmd+V), or provide a URL.
Customize your input with more control.
The cost of training depends on the number of steps. The formula is: 0.00675 * steps. With 1000 steps, your request will cost $6.75.
Fine-tune your training parameters and start right now.
This guide explains how to use the Ideogram V4 trainer to fine-tune a LoRA adapter for text-to-image generation. The result is a small LoRA weights file you can apply at inference time to teach Ideogram V4 a new subject, character, object, or visual style.
The Ideogram V4 trainer fine-tunes a LoRA (Low-Rank Adaptation) adapter on top of the Ideogram V4 image model. Instead of retraining the whole model, it learns a compact set of additional weights that nudge the base model toward the look of your dataset.
Key features:
images_data_url (required)Type: string
URL to a zip archive containing your training images and (optionally) caption files.
Supported image formats: .png, .jpg, .jpeg, .webp. Files in any other format (including .heic, .heif, and .avif) are not processed and are silently skipped, so convert them to one of the supported formats before zipping.
Captions: Add a .txt file with the same base name as each image. The caption file holds the text that describes that image.
sunset.png sunset.txt # Contains: "a watercolor painting of a sunset over the ocean"
If an image has no caption file, the trainer falls back to default_caption (see below). If neither is present, training stops with an error telling you which image is missing a caption.
Notes:
_mask (for example sunset_mask.png) are ignored, so you can leave masks in the archive without affecting training.cat/01.png and dog/01.png) collide and one is lost. Give each image (and its caption) a unique name.default_captionType: string
Default: none
A fallback caption used for any image that does not have its own .txt file. This is convenient when every image in your dataset shares the same description (for example, a single style or a single subject).
If you provide caption files for some images and a default_caption as well, each image uses its own caption when present and falls back to default_caption only when its caption file is missing.
stepsType: integer (100 - 40,000)
Default: 1000
Total number of training steps. More steps means the adapter is exposed to your dataset more times, learning it more strongly, up to the point of overfitting.
| Steps | Use Case |
|---|---|
| 500-1500 | Styles, or small/simple datasets |
| 1500-3000 | A specific subject, character, or object |
| 3000+ | Larger or more varied datasets (watch for overfitting) |
learning_rateType: float (1e-6 to 1e-2)
Default: 0.0001
How large each training update is. The default works well for most cases. Raise it cautiously for faster learning at the risk of instability; lower it for gentler, slower learning.
resolutionType: string
Default: "auto"
The pixel dimensions images are trained at. Images are center-cropped to this size without upscaling, so every image in your dataset must be at least as large as the selected resolution.
You can pass:
auto (the default) to let the trainer pick the largest size your dataset supports (see "What Happens to Your Data")WIDTHxHEIGHT string (for example 1280x768), where both numbers must be divisible by 16| Preset | Dimensions (W×H) | Shape |
|---|---|---|
square | 1024×1024 | Square |
landscape | 1536×1024 | Landscape |
portrait | 1024×1536 | Portrait |
widescreen | 1920×1088 | Wide |
ultrawide | 2048×768 | Very wide |
phone_wallpaper | 1024×1792 | Tall |
social_banner | 1584×400 | Banner |
Limits for custom and auto resolutions: width up to 2048, height up to 1792, and total area up to about 2,088,960 pixels. A custom size beyond these limits is rejected. Every named preset is within these limits by construction, so any preset is always accepted.
Guidance: Pick the resolution that matches what you intend to generate later. If your dataset images are smaller than every preset, use auto so the trainer can size the crop to your images instead of rejecting them.
output_lora_formatType: string ("fal" or "comfy")
Default: "fal"
Naming scheme for the keys inside the produced weights file.
| Value | Use Case |
|---|---|
fal | Use the adapter with fal's Ideogram V4 inference endpoint |
comfy | Use the adapter in ComfyUI's Ideogram V4 workflow |
The two files contain the same trained weights; only the internal key names differ. Choose the one that matches where you will load the adapter.
steps you requested.Archive extraction: The zip is unpacked, including any nested folders. Mac-specific junk entries (__MACOSX, files beginning with ._) are ignored. Only .png, .jpg, .jpeg, and .webp images are used; files in other formats are skipped.
File matching: Each image is paired with a caption file that has the exact same base name (name.png ↔ name.txt). Images whose name ends in _mask are skipped entirely. If a caption file is missing, the image uses default_caption; if there is no default_caption either, training stops and tells you which caption is missing.
Orientation: Images are auto-rotated to their correct upright orientation based on their embedded orientation metadata before anything else happens.
Image fitting: Each image is scaled to cover the target resolution while keeping its aspect ratio, then center-cropped to the exact width and height. Images are never upscaled, so any image smaller than the chosen resolution is rejected with a message telling you to pick a smaller resolution or upload larger images. Very large images (beyond the platform's maximum pixel limit) are also rejected.
Auto resolution: When resolution is auto, the trainer inspects every image, finds the smallest width and smallest height across the dataset, and rounds each down to a multiple of 16. That becomes the training size, so the largest possible crop is used without upscaling any image. Note that auto does not shrink the result to fit the limits: if your smallest images are large enough that the inferred size exceeds the caps (width 2048, height 1792, or area ~2,088,960 px) — common for high-resolution datasets — training fails with an error asking you to choose a preset or a smaller custom resolution. Pick an explicit resolution in that case.
Captions: Caption text is read as-is from your .txt files (or taken from default_caption). Captions are not rewritten, summarized, or auto-generated. Whatever you write is exactly what the adapter learns to associate with each image.
A LoRA adapter is a small set of extra weights layered onto the base Ideogram V4 model. Training only updates these extra weights, leaving the base model untouched. This keeps the output file small and makes it easy to switch the learned look on or off at inference time. The base model itself is never modified or redistributed.
Your dataset is the single biggest factor in the result. A quality checklist:
resolution: auto so they are not rejected.Suggested starting sizes:
| Goal | Image Count |
|---|---|
| A consistent style | 10-30 images |
| A specific subject / character / object | 10-30 images |
| A complex or highly varied concept | 30+ images |
Good caption:
a photo of sks dog sitting on a wooden porch, soft afternoon light
Weak caption:
dog
A trigger phrase is a distinctive word or short phrase you include in every caption so you can summon the learned concept at generation time.
sks dog, tok_woman, mybrand logo) in every caption. At inference, include that same trigger phrase in your prompt to invoke the subject.in the style of mystyle) on every caption, or rely on default_caption to apply one shared description to the whole set. At inference, add that phrase to steer the style.The adapter learns the relationship between your caption wording and your images. At generation time, prompt in the same style and format you used in training, including your trigger phrase. If you trained with short tag-like captions, short prompts will work best; if you trained with full descriptive sentences, prompt with full sentences.
json{ "images_data_url": "https://your-host/dataset.zip", "steps": 1000, "resolution": "auto", "learning_rate": 0.0001, "default_caption": "a photo of sks subject", "output_lora_format": "fal" }
Run this first, then generate a few images with the resulting LoRA and adjust steps (and optionally learning_rate) based on what you see.
Judge results by applying the trained LoRA at inference time and looking at what it generates. The job returns the LoRA weights and a config file only — no intermediate metrics or previews — so evaluate the finished adapter directly.
Signs of overfitting (trained too hard):
steps, add more variety to the dataset, or apply the LoRA at a lower scale at inference time.Signs of underfitting (trained too little):
steps, make sure your trigger phrase is in every caption, or add more representative images.Training failures:
.txt and no default_caption was set. Add captions or set default_caption.resolution: auto or pick a smaller size.WIDTHxHEIGHT — or the size auto inferred from large images — exceeds the supported limits, or a custom size is not divisible by 16. Pick a preset, or a smaller custom size that is a multiple of 16..png/.jpg/.jpeg/.webp, re-create the zip, and confirm it opens on your computer.auto when in doubt.fal for the fal endpoint and comfy for ComfyUI.