Adjust and enhance videos with Ray-2 Reframe. This advanced tool seamlessly reframes videos to your desired aspect ratio, intelligently inpainting missing regions to ensure realistic visuals and coherent motion, delivering exceptional quality and creative flexibility.
Luma AI logo
luma-dream-machine/ray-2-flash/reframe

Adjust and enhance videos with Ray-2 Reframe. This advanced tool seamlessly reframes videos to your desired aspect ratio, intelligently inpainting missing regions to ensure realistic visuals and coherent motion, delivering exceptional quality and creative flexibility.

reframe
outpaint
flash
video-to-video
FLUX Control LoRA Canny is a high-performance endpoint that uses a control image using a Canny edge map to transfer structure to the generated image and another initial image to guide color.
Black Forest Labs logo
flux-control-lora-canny/image-to-image

FLUX Control LoRA Canny is a high-performance endpoint that uses a control image using a Canny edge map to transfer structure to the generated image and another initial image to guide color.

lora
style transfer
image-to-image
Nucleus-Image is a text-to-image generation model built on a sparse mixture-of-experts (MoE) diffusion transformer architecture.
nucleus-image

Nucleus-Image is a text-to-image generation model built on a sparse mixture-of-experts (MoE) diffusion transformer architecture.

stylized
transform
typography
text-to-image
Remove background from videos filmed using chromakey, with automatic green spill suppression for clean, professional edges.
Veed logo
veed/video-background-removal/green-screen

Remove background from videos filmed using chromakey, with automatic green spill suppression for clean, professional edges.

video-to-video
Optimize 3D mesh topology with Hunyuan 3D Smart Topology.
hunyuan-3d/v3.1/smart-topology

Optimize 3D mesh topology with Hunyuan 3D Smart Topology.

3d
hunyuan
topology
3d-to-3d
Wan-2.2 video-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts and source videos.
Alibaba logo
wan/v2.2-a14b/video-to-video

Wan-2.2 video-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts and source videos.

video-to-video
Automatically splits a 3D model into semantic parts for editing, texturing, and rigging.
tripo3d/tripo/segment

Automatically splits a 3D model into semantic parts for editing, texturing, and rigging.

stylized
transform
3d-to-3d
Unified image generation with HiDream-O1-Image. Create, edit, and personalize high-resolution images up to 2K—single native model handles text-to-image, editing, and custom subjects without external components.
hidream-o1-image/dev

Unified image generation with HiDream-O1-Image. Create, edit, and personalize high-resolution images up to 2K—single native model handles text-to-image, editing, and custom subjects without external components.

text-to-image
Inpaint high-quality video using LTX-2.3
LTX logo
ltx-2.3-quality/inpaint

Inpaint high-quality video using LTX-2.3

video-to-video
Turn images or video into a detailed 3D scene with depth, camera poses, and a colored point cloud.
new
vggt-1b

Turn images or video into a detailed 3D scene with depth, camera poses, and a colored point cloud.

image-to-json
Image To Image Model using Boogu-Image
boogu-image/edit

Image To Image Model using Boogu-Image

image-to-image
Generate high-quality video with audio from reference video, text and images using LTX-2.3
LTX logo
ltx-2.3-quality/reference-video-to-video

Generate high-quality video with audio from reference video, text and images using LTX-2.3

video-to-video
Cohere Transcribe turns your business audio into accurate text, ready for search, analytics, and automation
cohere-transcribe

Cohere Transcribe turns your business audio into accurate text, ready for search, analytics, and automation

speech
transcribe
stt
speech-to-text
A blazing fast FLUX dev LoRA trainer for subjects and styles.
Black Forest Labs logo
turbo-flux-trainer

A blazing fast FLUX dev LoRA trainer for subjects and styles.

training
Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks
florence-2-large/caption

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks

captioning
multimodal
vision
Turn simple sketches into detailed, fully-textured 3D models. Instantly convert your concept designs into formats ready for Unity, Unreal, and Blender.
hunyuan3d-v3/text-to-3d

Turn simple sketches into detailed, fully-textured 3D models. Instantly convert your concept designs into formats ready for Unity, Unreal, and Blender.

text-to-3d
Outpaint high-quality video using LTX-2.3
LTX logo
ltx-2.3-quality/outpaint

Outpaint high-quality video using LTX-2.3

outpaint
outpainting
video-to-video
Generate video clips from your prompts using MiniMax model
Minimax logo
minimax/video-01-live

Generate video clips from your prompts using MiniMax model

motion
transformation
text-to-video
Pixverse Transition
Pixverse logo
pixverse/v5.5/transition

Pixverse Transition

image-to-video
Audio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.
sam-audio/span-separate

Audio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.

sam-audio
audio-to-audio
Qwen-Image (Image-to-Image) transforms and edits input images with high fidelity, enabling precise style transfer, enhancement, and creative modification.
Alibaba logo
qwen-image/image-to-image

Qwen-Image (Image-to-Image) transforms and edits input images with high fidelity, enabling precise style transfer, enhancement, and creative modification.

image-to-image
Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks
florence-2-large/ocr

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks

ocr
multimodal
vision
Edit any video with a natural-language instruction using Bernini-R, changing objects, weather, background, or camera angle while keeping the rest of the scene intact.
Bytedance logo
bernini-r/edit-video

Edit any video with a natural-language instruction using Bernini-R, changing objects, weather, background, or camera angle while keeping the rest of the scene intact.

edit
transform
stylized
video-to-video
Stable Audio 3 Medium Base is the foundational 1.4 billion parameter text-to-audio checkpoint generating stereo music up to 6 minutes, intended as the unmodified base for custom fine-tuning workflows.
stable-audio-3/medium/base/text-to-audio

Stable Audio 3 Medium Base is the foundational 1.4 billion parameter text-to-audio checkpoint generating stereo music up to 6 minutes, intended as the unmodified base for custom fine-tuning workflows.

music
audio
stereo
text-to-audio
Showing 793 to 816 of 1493 results