Easily adjust the perspective of any image to different angles.
image-apps-v2/perspective

Easily adjust the perspective of any image to different angles.

change-angle
perspective
image-to-image
Generate video clips from your prompts using Kling 1.6 (std)
Kling logo
kling-video/v1.6/standard/effects

Generate video clips from your prompts using Kling 1.6 (std)

text-to-video
Bagel is a 7B parameter multimodal model from Bytedance-Seed that can generate both text and images.
bagel/understand

Bagel is a 7B parameter multimodal model from Bytedance-Seed that can generate both text and images.

image-to-text
vlm
image-to-json
Transform your photos into cool plushies while keeping the original characters likeness
image-editing/plushie-style

Transform your photos into cool plushies while keeping the original characters likeness

stylized
transform
image-to-image
Meshy-5 multi image generates realistic and production ready 3D models from multiple images.
meshy/v5/multi-image-to-3d

Meshy-5 multi image generates realistic and production ready 3D models from multiple images.

multi-image-to-3d
image-to-3d
Generate 3D models from your images using Trellis 2. A native 3D generative model enabling versatile and high-quality 3D asset creation.
trellis-2/retexture

Generate 3D models from your images using Trellis 2. A native 3D generative model enabling versatile and high-quality 3D asset creation.

image-to-3d
image-to-3d
Stable Cascade: Image generation on a smaller & cheaper latent space.
stable-cascade

Stable Cascade: Image generation on a smaller & cheaper latent space.

diffusion
lcm
text-to-image
Remove unwanted elements (objects, people, text) while maintaining image consistency
Alibaba logo
qwen-image-edit-plus-lora-gallery/remove-element

Remove unwanted elements (objects, people, text) while maintaining image consistency

stylized
transform
image-to-image
Run SDXL at the speed of light
fast-sdxl/inpainting

Run SDXL at the speed of light

diffusion
high-res
lora
image-to-image
Run SDXL at the speed of light
fast-lcm-diffusion/image-to-image

Run SDXL at the speed of light

lcm
diffusion
turbo
image-to-image
Stable Audio 3 Medium audio inpainting is a 1.4 billion parameter latent diffusion model that fills in or reworks selected segments of a stereo track guided by text prompts, supporting single- and multi-segment editing.
stable-audio-3/medium/audio-inpainting

Stable Audio 3 Medium audio inpainting is a 1.4 billion parameter latent diffusion model that fills in or reworks selected segments of a stereo track guided by text prompts, supporting single- and multi-segment editing.

music
editing
restoration
audio-to-audio
Generate images from your prompts using Luma Photon Flash. Photon Flash is the most creative, personalizable, and intelligent visual models for creatives, bringing a step-function change in the cost of high-quality image generation.
Luma AI logo
luma-photon/flash

Generate images from your prompts using Luma Photon Flash. Photon Flash is the most creative, personalizable, and intelligent visual models for creatives, bringing a step-function change in the cost of high-quality image generation.

text-to-image
DreamOmni2 is a unified multimodal model for text and image guided image editing.
dreamomni2/edit

DreamOmni2 is a unified multimodal model for text and image guided image editing.

image-to-image
VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.
Alibaba logo
wan-vace-14b/outpainting

VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

image-to-video
text-to-video
video-to-video
Leffa Pose Transfer is an endpoint for changing pose of an image with a reference image.
leffa/pose-transfer

Leffa Pose Transfer is an endpoint for changing pose of an image with a reference image.

pose
utility
image-to-image
CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs.
csm-1b

CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs.

conversational
text to speech
text-to-audio
Audio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.
sam-audio/visual-separate

Audio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.

sam-audio
video-to-audio
PixVerse Extend model is a video extending tool for your videos using with high-quality video extending techniques
Pixverse logo
pixverse/extend

PixVerse Extend model is a video extending tool for your videos using with high-quality video extending techniques

utility
editing
video-to-video
Generate long videos from prompts using LTX Video-0.9.8 13B Distilled and custom LoRA
LTX logo
ltxv-13b-098-distilled

Generate long videos from prompts using LTX Video-0.9.8 13B Distilled and custom LoRA

video
ltx-video
text-to-video
FFMPEG Utility to Reverse Videos
workflow-utilities/reverse-video

FFMPEG Utility to Reverse Videos

video-to-video
Hunyuan World 1.0 turns a single image into a panorama or a 3D world. It creates realistic scenes from the image, allowing you to explore and view it from different angles.
hunyuan_world/image-to-world

Hunyuan World 1.0 turns a single image into a panorama or a 3D world. It creates realistic scenes from the image, allowing you to explore and view it from different angles.

image-to-3d
Generate video with audio from text using LTX-2 Distilled
LTX logo
ltx-2-19b/distilled/text-to-video

Generate video with audio from text using LTX-2 Distilled

text-to-video
Choose the Nth image from an image URL list for workflows.
workflow-utilities/pick-image-by-index

Choose the Nth image from an image URL list for workflows.

workflow
AuraFlow v0.3 is an open-source flow-based text-to-image generation model that achieves state-of-the-art results on GenEval. The model is currently in beta.
aura-flow

AuraFlow v0.3 is an open-source flow-based text-to-image generation model that achieves state-of-the-art results on GenEval. The model is currently in beta.

typography
style
text-to-image
Showing 937 to 960 of 1493 results