
Easily adjust the perspective of any image to different angles.

Generate video clips from your prompts using Kling 1.6 (std)

Bagel is a 7B parameter multimodal model from Bytedance-Seed that can generate both text and images.

Transform your photos into cool plushies while keeping the original characters likeness

Meshy-5 multi image generates realistic and production ready 3D models from multiple images.

Generate 3D models from your images using Trellis 2. A native 3D generative model enabling versatile and high-quality 3D asset creation.

Stable Cascade: Image generation on a smaller & cheaper latent space.

Remove unwanted elements (objects, people, text) while maintaining image consistency

Run SDXL at the speed of light

Run SDXL at the speed of light

Stable Audio 3 Medium audio inpainting is a 1.4 billion parameter latent diffusion model that fills in or reworks selected segments of a stereo track guided by text prompts, supporting single- and multi-segment editing.

Generate images from your prompts using Luma Photon Flash. Photon Flash is the most creative, personalizable, and intelligent visual models for creatives, bringing a step-function change in the cost of high-quality image generation.

DreamOmni2 is a unified multimodal model for text and image guided image editing.

VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

Leffa Pose Transfer is an endpoint for changing pose of an image with a reference image.

CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs.

Audio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.

PixVerse Extend model is a video extending tool for your videos using with high-quality video extending techniques

Generate long videos from prompts using LTX Video-0.9.8 13B Distilled and custom LoRA

FFMPEG Utility to Reverse Videos

Hunyuan World 1.0 turns a single image into a panorama or a 3D world. It creates realistic scenes from the image, allowing you to explore and view it from different angles.

Generate video with audio from text using LTX-2 Distilled

Choose the Nth image from an image URL list for workflows.

AuraFlow v0.3 is an open-source flow-based text-to-image generation model that achieves state-of-the-art results on GenEval. The model is currently in beta.