Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
pixart-sigma

Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

diffusion
text-to-image
A natural and expressive Brazilian Portuguese text-to-speech model optimized for clarity and fluency.
kokoro/brazilian-portuguese

A natural and expressive Brazilian Portuguese text-to-speech model optimized for clarity and fluency.

speech
text-to-audio
Create seamless cinematic transitions between two images with PixVerse C1, with native audio and up to 1080p.
Pixverse logo
pixverse/c1/transition

Create seamless cinematic transitions between two images with PixVerse C1, with native audio and up to 1080p.

video-generation
transition
pixverse
image-to-video
Wan-2.2 text-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts. This endpoint supports LoRAs made for Wan 2.2.
Alibaba logo
wan/v2.2-a14b/text-to-video/lora

Wan-2.2 text-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts. This endpoint supports LoRAs made for Wan 2.2.

text-to-video
Turn up to five reference images into one continuous, consistent video with Bernini-R, with smooth, stable camera motion and no scene cuts.
Bytedance logo
bernini-r/reference-to-video

Turn up to five reference images into one continuous, consistent video with Bernini-R, with smooth, stable camera motion and no scene cuts.

reference
video
stylized
image-to-video
FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.
Black Forest Labs logo
flux/schnell/redux

FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.

style transfer
image-to-image
Vector font generation with VecGlypher. Create custom glyphs from text descriptions or reference images—outputs clean SVG paths directly without raster-to-vector conversion.
vecglypher/image-to-svg

Vector font generation with VecGlypher. Create custom glyphs from text descriptions or reference images—outputs clean SVG paths directly without raster-to-vector conversion.

image-to-image
Split 3D models into parts with Hunyuan 3D
hunyuan-3d/v3.1/part

Split 3D models into parts with Hunyuan 3D

3d
hunyuan
mesh
3d-to-3d
Generate 3D models from text descriptions using Tripo P1.
tripo3d/p1/text-to-3d

Generate 3D models from text descriptions using Tripo P1.

3d
3d-generation
tripo
text-to-3d
Interpolate images with RIFE - Real-Time Intermediate Flow Estimation
rife

Interpolate images with RIFE - Real-Time Intermediate Flow Estimation

interpolation
image-to-image
Super fast text-to-image endpoint for the FLUX.1 Kontext [dev] model with LoRA support, enabling rapid and high-quality image generation using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.
Black Forest Labs logo
flux-kontext-lora/text-to-image

Super fast text-to-image endpoint for the FLUX.1 Kontext [dev] model with LoRA support, enabling rapid and high-quality image generation using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.

text-to-image
Generate high quality video clips with different effects using PixVerse v5
Pixverse logo
pixverse/v5/effects

Generate high quality video clips with different effects using PixVerse v5

image-to-video
FLUX1.1 [pro] ultra fine-tuned is the newest version of FLUX1.1 [pro] with a fine-tuned LoRA, maintaining professional-grade image quality while delivering up to 2K resolution with improved photo realism.
Black Forest Labs logo
flux-pro/v1.1-ultra-finetuned

FLUX1.1 [pro] ultra fine-tuned is the newest version of FLUX1.1 [pro] with a fine-tuned LoRA, maintaining professional-grade image quality while delivering up to 2K resolution with improved photo realism.

high-res
realism
text-to-image
Blend products into backgrounds with automatic perspective and lighting correction
Alibaba logo
qwen-image-edit-plus-lora-gallery/integrate-product

Blend products into backgrounds with automatic perspective and lighting correction

stylized
transform
image-to-image
SCAIL-2 is an end-to-end character animation model that drives a reference character from a source video without relying on intermediate pose representations like skeleton maps.
scail-2

SCAIL-2 is an end-to-end character animation model that drives a reference character from a source video without relying on intermediate pose representations like skeleton maps.

stylized
transform
video-to-video
Generate professional headshot photos with customizable backgrounds.
image-apps-v2/headshot-photo

Generate professional headshot photos with customizable backgrounds.

headshot
profile-photo
image-to-image
An efficent SDXL multi-controlnet image-to-image model.
sdxl-controlnet-union/image-to-image

An efficent SDXL multi-controlnet image-to-image model.

diffusion
controlnet
composition
image-to-image
Add a background to images with white/clean background
Black Forest Labs logo
flux-2-lora-gallery/add-background

Add a background to images with white/clean background

stylized
transform
image-to-image
Accelerated image generation with Ideogram V2A Turbo. Create high-quality visuals, posters, and logos with enhanced speed while maintaining Ideogram's signature quality.
Ideogram logo
ideogram/v2a/turbo

Accelerated image generation with Ideogram V2A Turbo. Create high-quality visuals, posters, and logos with enhanced speed while maintaining Ideogram's signature quality.

realism
typography
text-to-image
Generate long videos in 720p/30fps from text using LongCat Video Distilled
longcat-video/distilled/text-to-video/720p

Generate long videos in 720p/30fps from text using LongCat Video Distilled

text-to-video
Reframe entire videos scene-by-scene using Wan VACE 2.1
Alibaba logo
wan-vace-apps/long-reframe

Reframe entire videos scene-by-scene using Wan VACE 2.1

video-to-video
Infinitalk model generates a talking avatar video from an image and audio file. The avatar lip-syncs to the provided audio with natural facial expressions.
infinitalk/video-to-video

Infinitalk model generates a talking avatar video from an image and audio file. The avatar lip-syncs to the provided audio with natural facial expressions.

video-to-video
Kandinsky 5.0 Pro is a diffusion model for fast, high-quality text-to-video generation.
kandinsky5-pro/text-to-video

Kandinsky 5.0 Pro is a diffusion model for fast, high-quality text-to-video generation.

text-to-video
Professional motion deblur powered by Topaz Labs. Themis 2 restores clarity to fast-moving, motion-blurred footage at source resolution. Best for sports and action footage.
Topaz Labs logo
topaz/deblur/video

Professional motion deblur powered by Topaz Labs. Themis 2 restores clarity to fast-moving, motion-blurred footage at source resolution. Best for sports and action footage.

deblur
video
video-to-video
Showing 961 to 984 of 1493 results