Add details to faces, enhance face features, remove blur.
image-editing/realism

Add details to faces, enhance face features, remove blur.

stylized
transform
realism
image-to-image
Utilize Flux.1 [dev] Controlnet to generate high-quality images with precise control over composition, style, and structure through advanced edge detection and guidance mechanisms.
Black Forest Labs logo
flux-lora-canny

Utilize Flux.1 [dev] Controlnet to generate high-quality images with precise control over composition, style, and structure through advanced edge detection and guidance mechanisms.

controlnet
detection
lora
image-to-image
Replace or dub audio on an existing video with fast audio-only lip-sync.
Heygen logo
heygen/v3/lipsync/speed

Replace or dub audio on an existing video with fast audio-only lip-sync.

stylized
transform
lipsync
video-to-video
VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.
Alibaba logo
wan-vace-14b/depth

VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

image-to-video
text-to-video
video-to-video
Adjust and enhance videos with Ray-2 Reframe. This advanced tool seamlessly reframes videos to your desired aspect ratio, intelligently inpainting missing regions to ensure realistic visuals and coherent motion, delivering exceptional quality and creative flexibility.
Luma AI logo
luma-dream-machine/ray-2/reframe

Adjust and enhance videos with Ray-2 Reframe. This advanced tool seamlessly reframes videos to your desired aspect ratio, intelligently inpainting missing regions to ensure realistic visuals and coherent motion, delivering exceptional quality and creative flexibility.

reframe
outpaint
video-to-video
Super fast endpoint for the FLUX.1 [dev] inpainting model with LoRA support, enabling rapid and high-quality image inpaingting using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.
Black Forest Labs logo
flux-krea-lora/inpainting

Super fast endpoint for the FLUX.1 [dev] inpainting model with LoRA support, enabling rapid and high-quality image inpaingting using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.

lora
personalization
image-to-image
Generate high-quality video with audio from audio, text and images using LTX-2.3
LTX logo
ltx-2.3-quality/audio-to-video

Generate high-quality video with audio from audio, text and images using LTX-2.3

audio-to-video
All-in-one image AI with JoyAI-Image. Understand, create, and edit images through natural language—the model's deep visual understanding powers more accurate generation and precise editing in a unified system.
joyai-image-edit

All-in-one image AI with JoyAI-Image. Understand, create, and edit images through natural language—the model's deep visual understanding powers more accurate generation and precise editing in a unified system.

image-editing
image-to-image
Vidu Reference-to-Image creates images by using a reference images and combining them with a prompt.
vidu/q2/reference-to-image

Vidu Reference-to-Image creates images by using a reference images and combining them with a prompt.

images-to-imag
reference-to-image
image-to-image
Kandinsky 5.0 Pro is a diffusion model for fast, high-quality image-to-video generation.
kandinsky5-pro/image-to-video

Kandinsky 5.0 Pro is a diffusion model for fast, high-quality image-to-video generation.

image-to-video
EchoMimic V3 generates a talking avatar model from a picture, audio and text prompt.
echomimic-v3

EchoMimic V3 generates a talking avatar model from a picture, audio and text prompt.

echomimic
talking-head
audio-to-video
FLUX1.1 [pro] Redux is a high-performance endpoint for the FLUX1.1 [pro] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.
Black Forest Labs logo
flux-pro/v1.1/redux

FLUX1.1 [pro] Redux is a high-performance endpoint for the FLUX1.1 [pro] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.

style transfer
image-to-image
Generate 3D models from a single image with Hi3D.
hitem3d/hi3d/image-to-3d

Generate 3D models from a single image with Hi3D.

3d
mesh
image-to-3d
Generate long, expressive multi-voice speech using Microsoft's powerful TTS
vibevoice

Generate long, expressive multi-voice speech using Microsoft's powerful TTS

multi-speaker
podcast
text-to-speech
FLUX LoRA Image-to-Image is a high-performance endpoint that transforms existing images using FLUX models, leveraging LoRA adaptations to enable rapid and precise image style transfer, modifications, and artistic variations.
Black Forest Labs logo
flux-krea-lora/image-to-image

FLUX LoRA Image-to-Image is a high-performance endpoint that transforms existing images using FLUX models, leveraging LoRA adaptations to enable rapid and precise image style transfer, modifications, and artistic variations.

lora
style transfer
image-to-image
Super fast endpoint for the FLUX.1 [schnell] model with subject input capabilities, enabling rapid and high-quality image generation for personalization, specific styles, brand identities, and product-specific outputs.
Black Forest Labs logo
flux-subject

Super fast endpoint for the FLUX.1 [schnell] model with subject input capabilities, enabling rapid and high-quality image generation for personalization, specific styles, brand identities, and product-specific outputs.

personalization
customization
text-to-image
FireRed Image Edit is FireRed's state of the art open source editing model, re-trained from Qwen Image Edit 2509.
firered-image-edit

FireRed Image Edit is FireRed's state of the art open source editing model, re-trained from Qwen Image Edit 2509.

image-editing
firered
image-to-image
Hunyuan Video 1.5 is Tencent's latest and best video model
hunyuan-video-v1.5/text-to-video

Hunyuan Video 1.5 is Tencent's latest and best video model

hunyuan-video
text-to-video
Detect speech presence and timestamps with accuracy and speed using the ultra-lightweight Silero VAD model
silero-vad

Detect speech presence and timestamps with accuracy and speed using the ultra-lightweight Silero VAD model

vad
silero
voice-activity-detection
audio-to-text
Bagel is a 7B parameter multimodal model from Bytedance-Seed that can generate both images and text.
bagel/edit

Bagel is a 7B parameter multimodal model from Bytedance-Seed that can generate both images and text.

image-editing
image-to-image
Customizing Realistic Human Photos via Stacked ID Embedding
photomaker

Customizing Realistic Human Photos via Stacked ID Embedding

editing
customization
realism
image-to-image
Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs.
nvidia/cosmos-3-super/image-to-video

Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs.

stylized
transform
lipsync
image-to-video
Vidu Start-End to Video generates smooth transition videos between specified start and end images.
vidu/start-end-to-video

Vidu Start-End to Video generates smooth transition videos between specified start and end images.

motion
transition
image-to-video
Generate images from text, an image, a mask and custom LoRA using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.
Alibaba logo
z-image/turbo/inpaint/lora

Generate images from text, an image, a mask and custom LoRA using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

inpainting
image-to-image
Showing 913 to 936 of 1493 results