Model Catalog
Every model currently on the Segmind AI Gateway — 670 models with their API slugs, grouped by task. Regenerated automatically from the live catalog.
Every model on the AI Gateway, with the slug you pass
to https://api.segmind.com/v1/{'{slug}'} (or /v2 for async) and to
segmind.run() in the Python SDK. 670 models,
regenerated automatically from the live catalog — if a model is listed here, it
is callable today.
Model-specific parameters live on each model's page at
segmind.com/models, or in the
parameters_schema field of the
catalog API.
Recently added
| Model | Slug | Description | Added |
|---|---|---|---|
| Gemini Omni 1.1 | gemini-omni-1.1 | Text-to-video with synchronized native audio, up to 4K. | 2026-08 |
| Gemini Omni 1.1 Video Extend | gemini-omni-1.1-video-extend | Extend short video clips into longer seamless scenes. | 2026-08 |
| Gemini Omni 1.1 Video Edit | gemini-omni-1.1-video-edit | Edit videos with a text prompt, subject preserved. | 2026-08 |
| Lyria 3 Pro | lyria-3-pro | Full-length text-to-music songs with vocals and lyrics. | 2026-08 |
| Lyria 3 | lyria-3 | Generate 30-second songs with vocals from text or images. | 2026-08 |
| Gemini 3.7 Flash | gemini-3.7-flash | Fast multimodal LLM for coding, agents, and long-document analysis. | 2026-08 |
| Kokoro 82M | kokoro-82m | Text-to-speech with 54 multilingual voices. | 2026-08 |
| Whisper Large V3 | whisper-large-v3 | Transcribe speech-to-text in 99 languages with timestamps. | 2026-08 |
| Wan 3.0 Video | wan3.0-video | Generate 30-second 1080p video with native audio. | 2026-08 |
| Wan 2.6 Image to Video Flash | wan2.6-i2v-flash | Animate photos into 15-second 1080p video with native audio. | 2026-08 |
| Grok Imagine Image 2 | grok-imagine-image-2 | Text-to-image and image editing with crisp, legible text. | 2026-08 |
| LTX 2.5 Pro | ltx-2.5-pro | Generate 1080p video with native audio and multi-shot scenes. | 2026-08 |
| LTX 2.5 Fast | ltx-2.5-fast | Text-to-video and image-to-video with native audio, up to 4K. | 2026-08 |
| Qwen Image 3.0 | qwen-image-3 | Generate and edit legible in-image text, up to 2K. | 2026-08 |
| Seedream 5.0 Pro Layer Decomposition | seedream-5-pro-layer-decomposition | Split any image into editable transparent PNG layers. | 2026-08 |
| Bria Extract Object | bria-extract-object | Extract any named object into a transparent PNG cutout. | 2026-08 |
| Seedance 2.5 | seedance-2.5 | Generate cinematic multi-shot AI videos up to 30 seconds with synchronized native audio from text, images, or references. | 2026-08 |
| FLUX 3 Draft Enhance | flux-3-draft-enhance | Upscale AI video drafts to Full-HD with native audio. | 2026-08 |
| FLUX 3 Extend Video | flux-3-extend-video | Extend clips into seamless video continuations with synchronized audio. | 2026-08 |
| FLUX 3 Image to Video | flux-3-image-to-video | Animate images into 20-second clips with synchronized native audio. | 2026-08 |
| FLUX 3 Text to Video | flux-3-text-to-video | Cinematic text-to-video with native lip-synced audio, up to 20s. | 2026-08 |
| Qwen3.8 Max | qwen-3.8-max | Multimodal reasoning and agentic coding with 1M-token context. | 2026-08 |
| Grok Imagine Video 1.5 Reference to Video | grok-imagine-video-1.5-reference-to-video | Character-consistent video from up to 7 reference images. | 2026-08 |
| Grok Imagine Video 1.5 Image to Video | grok-imagine-video-1.5-image-to-video | Animate a still image into 1080p video with synced audio. | 2026-08 |
| Grok Imagine Video 1.5 Text to Video | grok-imagine-video-1.5-text-to-video | Text-to-video clips up to 1080p with native synchronized audio. | 2026-08 |
this is first creation
| Model | Slug | Description | Added |
|---|---|---|---|
| Demo | Image-shakil | This is Demo creation | 2024-10 |
Audio & speech synthesis
| Model | Slug | Description | Added |
|---|---|---|---|
| Ace Step Music | ace-step-music | ACE-Step generates high-quality music rapidly, enhancing the creative process for developers and artists worldwide. | 2025-05 |
| Chatterbox TTS | chatterbox-tts | Chatterbox transforms text into rich, natural speech with adjustable emotional expressiveness for diverse applications. | 2025-07 |
| Chatterbox Turbo TTS | chatterbox-turbo-tts | Ultra-fast, human-quality TTS with emotional expression. | 2025-12 |
| Dia (Text to Speech) | dia | Dia by Nari Labs is an advanced open-weights TTS model that brings scripts to life with natural speech, emotions, and nonverbal cues. Easil… | 2025-04 |
| ElevenLabs Dubbing | dubbing | Instantly dubs audio and video into 29 languages while preserving each speaker's original voice. | 2024-07 |
| Elevenlabs Dialogue | elevenlabs-dialogue | Immersive, emotionally expressive multi-speaker audio dialogue. | 2025-11 |
| Gemini TTS 2.5 Flash | gemini-2.5-flash-tts | Fast, lifelike text-to-speech with expressive emotional tones. | 2025-12 |
| Gemini TTS 2.5 Pro | gemini-2.5-pro-tts | Human-like speech synthesis with rich expressive emotional depth. | 2025-12 |
| Gemini 3.1 Flash TTS | gemini-3.1-flash-tts | Expressive, controllable TTS with 70+ language support. | 2026-05 |
| Grok Text-to-Speech | grok-tts | Convert text to speech in 20 languages with five voices. | 2026-06 |
| Kokoro 82M | kokoro-82m | Text-to-speech with 54 multilingual voices. | 2026-08 |
| Lyria 2 | lyria-2 | Lyria 2 by Google DeepMind is an advanced model that generates high-fidelity 48kHz stereo instrumental music from text prompts or lyrics, o… | 2025-05 |
| Lyria 3 | lyria-3 | Generate 30-second songs with vocals from text or images. | 2026-08 |
| Lyria 3 Pro | lyria-3-pro | Full-length text-to-music songs with vocals and lyrics. | 2026-08 |
| Meta MusicGen Medium | meta-musicgen-medium | MusicGen: Transform text into music with AI. Create unique, high-quality audio from simple descriptions. Experience the future of music gen… | 2024-10 |
| MyShell Text To Speech | myshell-tts | MyShell's Voice Cloning and Text to Speech - Transform your audio content with realistic, personalized voices. Experience high-quality, eff… | 2024-10 |
| Openvoice | openvoice | OpenVoice is a versatile voice cloning model that supports multiple languages and offers precise tone replication, flexible style control,… | 2024-09 |
| 3B Orpheus TTS (0.1) | orpheus-3b-0.1 | Orpheus TTS is an open-source text-to-speech (TTS) system powered by the Llama 3B language model, designed for high-quality and customizabl… | 2025-03 |
| Sam Audio Large | sam-audio-large | Isolate any described sound from mixed audio tracks. | 2026-02 |
| Seed Audio 1.0 | seed-audio-1.0 | Generate full audio scenes: dialogue, music, effects, voice cloning. | 2026-06 |
| BytePlus Seed Speech TTS | seed-speech-tts | Natural multilingual text-to-speech and voiceovers from text. | 2026-06 |
| Sonilo Text to Audio | sonilo-text-to-audio | Commercial-safe music and sound effects from text prompts. | 2026-08 |
| Elevenlabs Sound Generation | sound-generation | Eleven Labs' Sound Generation API provides a robust development tool for programmatically generating audio content using artificial intelli… | 2024-06 |
| Elevenlabs Text To Speech | tts-eleven-labs | ElevenLabs TTS transforms text into captivating, human-like speech for diverse applications. | 2024-06 |
| VeenaMax TTS | veena-max-tts | VeenaMAX transforms text into expressive, real-time speech across multiple Indian languages for seamless communication. | 2025-09 |
| Veena TTS | veena-tts | Veena transforms text into high-fidelity, expressive speech in Hindi and English for real-time applications. | 2025-07 |
Audio to audio
| Model | Slug | Description | Added |
|---|---|---|---|
| Elevenlabs Audio Isolation | elevenlabs-audio-isolation | Extract clear speech from noisy audio and video. | 2025-11 |
| Elevenlabs Speech To Speech | sts-eleven-labs | Eleven Labs Speech-to-Speech offers AI-powered voice conversion for content creators, media professionals, and anyone seeking to modify or… | 2024-06 |
Image editing & transformation
| Model | Slug | Description | Added |
|---|---|---|---|
| AI Product Photo Editor | ai-product-photo-editor | AI Product Photo Editor leverages advanced image-based ML techniques to generate high-quality product visuals using text prompts, product i… | 2024-07 |
| AI Product Photography | ai-product-photography | Elevate your product imagery with our AI-powered photography model. Create stunning, professional-quality photos that boost engagement and… | 2024-11 |
| IDM + Faceswap (updated) | alle-v2 | IDM+Faceswap | 2024-07 |
| Aura Flow | aura-flow | Largest completely open sourced flow-based generation model that is capable of text-to-image generation | 2024-07 |
| Automatic Mask Generator | automatic-mask-generator | Automatic Mask Generator is a powerful tool that automates the creation of precise masks for inpainting | 2024-06 |
| Profile Photo Style Transfer | become-image | Turn any image of a face into artwork using Stable Diffusion Controlnet and IPAdapter | 2024-06 |
| Background Removal | bg-removal | This model removes the background image from any image | 2023-09 |
| Background Removal V2 | bg-removal-v2 | This model removes the background image from any image | 2024-03 |
| Bria Blur Background | bria-blur-background | Bria AI Image Editing API v2 enables precise and context-aware image manipulation for stunning visual outcomes. | 2025-08 |
| Bria Enhance Image | bria-enhance-image | Bria AI creates precise, high-quality image enhancements and manipulations for diverse creative applications. | 2025-08 |
| Bria Erase Foreground | bria-erase-foreground | Seamlessly removes foreground subjects and regenerates backgrounds for flawless image editing. | 2025-08 |
| Bria Eraser | bria-eraser | AI object removal with seamless context-aware inpainting. | 2025-08 |
| Bria Expand Image | bria-expand-image | Bria Expand enables precise image manipulation and enhancement with generative AI, trained exclusively on licensed data for safe, risk-free… | 2025-08 |
| Bria Extract Object | bria-extract-object | Extract any named object into a transparent PNG cutout. | 2026-08 |
| Bria FIBO 1.5 Image Edit | bria-fibo-image-edit | Edit images with structured JSON, masks, and multi-image references. | 2026-01 |
| Bria Generative Fill | bria-gen-fill | Bria AI enables precise generative image editing for seamless creative enhancements and transformations. | 2025-08 |
| Bria Increase Resolution | bria-increase-resolution | Seamlessly upscale and manipulate images while preserving the highest fidelity and safety standards. | 2025-08 |
| Lifestyle Product Shot by Image | bria-lifestyle-shot-by-image | Transforms ordinary product images into stunning, marketing-ready visuals for eCommerce success. | 2025-08 |
| Bria Lifestyle Product Shot by Text | bria-lifestyle-shot-by-text | Transform isolated product images into dynamic lifestyle scenes with AI-driven contextual realism. | 2025-08 |
| Bria Product Cutout | bria-product-cutout | Automates precise product cutouts and background removal for professional eCommerce imagery at scale. | 2025-08 |
| Bria Product Packshot | bria-product-packshot | Transform product photos into professional, market-ready images with intelligent enhancements and background removal. | 2025-08 |
| Bria Product Shadow | bria-product-shadow | Bria Product Shadow enhances product images with realistic shadows for professional eCommerce presentations. | 2025-08 |
| Bria RMBG 2.0 | bria-remove-background | Effortlessly extract backgrounds with unmatched precision, powered by models trained exclusively on licensed data for safe and risk-free co… | 2025-08 |
| Bria Generate Background | bria-replace-background | Transform images through advanced background editing and generative content creation for diverse applications. | 2025-08 |
| Caricature Style | caricature-style | Transform everyday photos into lively, whimsical caricature illustrations that highlight individual features with playful exaggeration. | 2025-05 |
| Clarity Upscaler | clarity-upscaler | High resolution creative image Upscaler and Enhancer. A free Magnific alternative. | 2024-06 |
| ClarityAI Creative Upscaler | clarityai-creative-upscaler | Creative image upscaling with fine detail enhancement. | 2025-10 |
| ClarityAI Crystal Upscaler | clarityai-crystal-upscaler | Upscale images up to 200x with enhanced detail and vibrancy. | 2025-10 |
| ClarityAI Flux Upscaler | clarityai-flux-upscaler | Transform low-resolution images into stunning high-quality visuals. | 2025-10 |
| Codeformer | codeformer | CodeFormer is a robust face restoration algorithm for old photos or AI-generated faces. | 2023-09 |
| Consistent Character | consistent-character | Create images of a given character in different poses | 2024-06 |
| Consistent Character With Pose | consistent-character-with-pose | Create images of a given character in different poses | 2024-09 |
| ESRGAN | esrgan | ERGAN is an Image Super-Resolution (upscaler) model that enhances images with stunning, high-quality upscaling while preserving the exact c… | 2023-09 |
| Expression Editor | expression-editor | Expression Editor uses reference images to accurately generate new images with desired expressions. Perfect for digital art, memes, and mar… | 2024-09 |
| Face Detailer | face-detailer | Restore characters' faces to their original glory with Face Detailer. Enhance facial details, eliminate distortion, and upscale images for… | 2024-10 |
| face-to-many | face-to-many | Turn a face into 3D, emoji, pixel art, video game, claymation or toy | 2024-05 |
| face-to-sticker | face-to-sticker | Turn a face into a sticker | 2024-05 |
| Segmind FaceSwap Comic v1 | faceswap-comic | FaceSwap Comic v1 is an AI-powered face swapping model designed to blend real faces into illustrated or cartoon-style images while preservi… | 2025-05 |
| Faceswap V3 Moonfrog | faceswap-moonfrog-v3 | Take a picture/gif and replace the face in it with a face of your choice. You only need one image of the desired face. No dataset, no train… | 2024-07 |
| Faceswap V2 | faceswap-v2 | Take a picture/gif and replace the face in it with a face of your choice. You only need one image of the desired face. No dataset, no train… | 2024-04 |
| Faceswap V3 | faceswap-v3 | Face Swap V3 is a cutting-edge tool that empowers you to seamlessly swap faces in images. With customizable features and advanced technolog… | 2024-10 |
| Faceswap V3 Multifaceswap | faceswap-v3-multifaceswap | Faceswap V3 Multifaceswap enables realistic face swapping in images, preserving lighting and expressions for professional results. | 2025-07 |
| Segmind Faceswap v4 | faceswap-v4 | Segmind FaceSwap v4 enables fast and precise face or head swapping between images with customizable options for style, output format, and i… | 2025-03 |
| Segmind Faceswap v5 | faceswap-v5 | Ultra-fast face and head swapping in images. | 2026-01 |
| Flux 2 Flex | flux-2-flex | Consistent-style photorealistic images using reference inputs. | 2025-11 |
| Flux-2 Klein-4b | flux-2-klein-4b | Sub-second photorealistic image generation and editing. | 2026-01 |
| Flux-2 Klein-9b | flux-2-klein-9b | Ultra-fast photorealistic image generation on consumer GPUs. | 2026-01 |
| Flux 2 Max | flux-2-max | Photorealistic images with maximum consistency and fine detail. | 2025-12 |
| Flux 2 Pro | flux-2-pro | High-quality photorealistic images with cross-output consistency. | 2025-11 |
| Flux Canny Dev | flux-canny-dev | Open-weight edge-guided image generation. Control structure and composition using Canny edge detection. | 2024-11 |
| Flux Canny Pro | flux-canny-pro | Professional edge-guided image generation. Control structure and composition using Canny edge detection | 2024-11 |
| Flux Controlnets | flux-controlnet | Flux ControlNets is a collection of models that gives you precise control over image generation. By integrating ControlNet with Flux.1, the… | 2024-09 |
| Flux Depth Dev | flux-depth-dev | Open-weight depth-aware image generation. Edit images while preserving spatial relationships. | 2024-11 |
| Flux Depth Pro | flux-depth-pro | Professional depth-aware image generation. Edit images while preserving spatial relationships. | 2024-11 |
| Flux Fill Dev | flux-fill-dev | Open-weight inpainting model for editing and extending images. Guidance-distilled from FLUX.1 Fill Dev | 2024-11 |
| Flux Fill Pro | flux-fill-pro | Professional inpainting and outpainting model with state-of-the-art performance. Edit or extend images with natural, seamless results. | 2024-11 |
| Flux.1 Image To Image | flux-img2img | Flux Image-To-Image model by Black Forest Labs is an advanced deep learning tool designed for transforming images based on specific textual… | 2024-08 |
| Flux Inpaint | flux-inpaint | Flux Inpainting is a powerful image editing tool designed to effortlessly edit and enhance your images. It's perfect for tasks like removin… | 2024-09 |
| Flux Ipadapter | flux-ipadapter | Flux IP Adapter is a cutting-edge AI model that lets you to create stunning, customized images. With its advanced style adaptation capabili… | 2024-09 |
| FLUX.1 Kontext [dev] | flux-kontext-dev | FLUX.1 Kontext [dev] creates coherent and editable images by integrating text and visual cues for iterative design. | 2025-06 |
| Flux Kontext Max | flux-kontext-max | FLUX.1 Kontext [max] transforms textual descriptions into stunning, high-fidelity images with seamless typography integration. | 2025-05 |
| Flux Kontext Pro | flux-kontext-pro | FLUX.1 Kontext Pro transforms text prompts into high-quality, customized images with remarkable efficiency and precision. | 2025-05 |
| Flux Krea Dev | flux-krea-dev | FLUX.1 Krea generates stunning, photorealistic images with fine-tuned aesthetic control for diverse creative applications. | 2025-08 |
| Flux Pulid | flux-pulid | Flux PuLID: Customize AI-generated images with your unique identity. Seamlessly integrate faces into text-to-image models for realistic and… | 2024-09 |
| Flux Redux Dev | flux-redux-dev | Open-weight image variation model. Create new versions while preserving key elements of your original. | 2024-11 |
| Flux Redux Schnell | flux-redux-schnell | Fast, efficient image variation model for rapid iteration and experimentation. | 2024-11 |
| Fooocus Outpainting | focus-outpaint | Fooocus Outpainting transforms ordinary images into extraordinary works of art by seamlessly expanding their boundaries. | 2024-02 |
| Fooocus | fooocus | Fooocus enables high-quality image generation effortlessly, combining the best of Stable Diffusion and Midjourney. | 2024-06 |
| GPT Image 1 Edit | gpt-image-1-edit | Edit and compose images using natural language with GPT Image 1 Edit, OpenAI’s powerful inpainting and multi-reference editing model. Perfe… | 2025-04 |
| GPT Image 1 Edit Mini | gpt-image-1-edit-mini | Affordable text-driven image generation and editing. | 2025-10 |
| GPT Image 1.5 Edit | gpt-image-1.5-edit | Precise image editing via natural language instructions. | 2025-12 |
| Grok Imagine Image | grok-imagine-image | Text-to-image generation and editing, up to 2K resolution. | 2026-06 |
| Grok Imagine Image 2 | grok-imagine-image-2 | Text-to-image and image editing with crisp, legible text. | 2026-08 |
| HeyGen Generate Look | heygen-generate-look | Change avatar outfits and backgrounds while keeping the same face. | 2026-07 |
| HiDream-I1 (Fast) | hidream-l1-fast | HiDream-I1 is a next-generation, open-source image generative foundation model designed for text-to-image synthesis, especially for renderi… | 2025-04 |
| Higgsfield Soul 2.0 | higgsfield-soul-2 | Generate fashion-editorial photorealistic photos from text or reference. | 2026-07 |
| Higgsfield Text 2 Image Soul | higgsfield-text2image-soul | SOUL AI transforms text into stunning, customizable visuals with unparalleled style control and precision. | 2025-09 |
| HyperSwap Image Faceswap by FaceFusion Labs | hyperswap-image-faceswap-by-facefusion-labs | High-quality face swapping built for real production workflows. | 2026-03 |
| Relighting | ic-light | Prompts to auto-magically relight your images. | 2024-06 |
| Icon Overlay | icon-overlay | icon overlay | 2024-11 |
| Ideogram 2a Image to Image | ideogram-2a-img-2-img | Ideogram Image to Image: Transform your images with ease! Enhance, modify, or create entirely new visuals using advanced AI. Perfect for ar… | 2025-03 |
| Ideogram 3 Reframe | ideogram-3-reframe | Ideogram 3.0's Reframe effortlessly adapts images to diverse formats, enhancing visual content creation for any platform. | 2025-05 |
| Ideogram 3 Remix | ideogram-3-remix | Ideogram 3 Remix enables versatile image transformation, enhancing creativity through customizable design iterations. | 2025-05 |
| Ideogram 3 Replace Background | ideogram-3-replace-background | Effortlessly replace backgrounds in images, enhancing visual storytelling and creativity with precision and speed. | 2025-05 |
| Ideogram Character | ideogram-character | Achieve perfect character consistency across multiple generations from a single reference image. | 2025-08 |
| Ideogram Image To Image | ideogram-img-2-img | Ideogram Image to Image: Transform your images with ease! Enhance, modify, or create entirely new visuals using advanced AI. Perfect for ar… | 2024-12 |
| Ideogram Reframe | ideogram-reframe | Transform your images with Ideogram Reframe! Easily reframe square images to your chosen resolution. | 2025-03 |
| Ideogram Turbo Image To Image | ideogram-turbo-img-2-img | Transform images instantly with Ideogram Turbo Image to Image! Fast AI for quick edits & creative remixes. | 2025-03 |
| Ideogram V4 Remix | ideogram-v4-remix | Restyle any image into posters with legible in-image text. | 2026-07 |
| IDM VTON | idm-vton | Best-in-class clothing virtual try on in the wild | 2024-06 |
| illusion-diffusion-hq | illusion-diffusion-hq | Monster Labs QrCode ControlNet on top of SD Realistic Vision v5.1 | 2024-06 |
| Minimax-image-01 | image-01 | Generate high-fidelity images from text with precise control & stunning quality with Minimax Image-01. | 2025-03 |
| Infinite You | infinite-you | InfiniteYou generates high-fidelity portraits preserving identity while aligning with creative text prompts. | 2025-07 |
| Inpaint Mask Maker | inpaint-mask-maker | Real-Time Open-Vocabulary Object Detection | 2024-06 |
| Insta Depth | insta-depth | InstantID aims to generate customized images with various poses or styles from only a single reference ID image while ensuring high fidelity | 2024-04 |
| InstantID | instantid | InstantID aims to generate customized images with various poses or styles from only a single reference ID image while ensuring high fidelity | 2024-02 |
| IP-adapter Depth XL | ip-sdxl-depth | IP Adapter Depth XL is built on the SDXL framework. This model integrates the IP Adapter and Depth preprocessor to offer unparalleled contr… | 2023-11 |
| Kling V3 Image 2 Image | kling-3-image2image | Transform images into photorealistic, production-ready visuals. | 2026-03 |
| Kling O1 | kling-o1 | Text-to-video creation with precise AI-driven motion control. | 2026-01 |
| Kolors | kolors | Kolors is a cutting-edge text-to-image model that bridges language and visual art. Transform your textual ideas into photorealistic images… | 2024-07 |
| Luma Uni-1 | luma-uni-1 | Reasoning-first text-to-image and natural-language image editing. | 2026-06 |
| Luma Uni-1 Max | luma-uni-1-max | Generate and edit images from plain-text instructions. | 2026-06 |
| Magic Eraser | magic-eraser | LaMA Object Removal- AI Magic Eraser | 2024-06 |
| material-transfer | material-transfer | Transfer a material from an image to a subject | 2024-05 |
| Multi Image Kontext Max | multi-image-kontext-max | FLUX.1 Kontext [max] creates stunning, photorealistic images from text prompts and input images seamlessly. | 2025-06 |
| Multi Image Kontext Pro | multi-image-kontext-pro | Transform text into stunning, professional-grade images with precise editing capabilities. | 2025-07 |
| Nano Banana 2 | nano-banana-2 | Fast photorealistic images — ideal for marketing and ads. | 2026-02 |
| Nano Banana 2 Lite | nano-banana-2-lite | Generate and edit 1K images in about four seconds. | 2026-07 |
| Nano Banana Pro | nano-banana-pro | High-fidelity images with accurate multilingual text rendering. | 2025-11 |
| Nomos Image Upscaler 4k | nomos-upscaler | This upscaling model is ideal for enhancing amateur to professional photos, excelling with subjects like cats, hair, and party scenes. It h… | 2025-05 |
| Omini Control | ominicontrol | OminiControl is an innovative framework that optimizes Diffusion Transformer models for versatile image generation tasks. | 2024-12 |
| Omni Zero | omni-zero | Omni-Zero: A diffusion pipeline for zero-shot stylized portrait creation. | 2024-06 |
| Pruna P Image Edit | p-image-edit | Multi-image editing with AI-guided precision and control. | 2025-11 |
| Pruna P Image Try-On | p-image-try-on | Dress photos in multiple garments with photorealistic virtual try-on. | 2026-06 |
| PuLID | pulid-base | Novel tuning-free ID customization method for text-to-image generation. | 2024-06 |
| Qwen Image Edit | qwen-image-edit | Transform images effortlessly through semantic context and pixel-perfect appearance changes. | 2025-08 |
| Qwen Image Edit Fast | qwen-image-edit-fast | Qwen-Image-Edit enables precise bilingual image editing for seamless localization and professional content creation. | 2025-08 |
| Qwen Image Edit Plus | qwen-image-edit-plus | Multi-image editing with precise text-guided transformations. | 2025-10 |
| Qwen Image Edit Plus Add People Lora | qwen-image-edit-plus-add-people | Generate realistic multi-character scenes with natural interactions. | 2025-11 |
| Qwen Image Edit Plus Blend It | qwen-image-edit-plus-blend-it | Product placement into backgrounds with precise lighting match. | 2025-11 |
| Qwen Image Edit Plus Eigen Banana | qwen-image-edit-plus-eigen-banana | Precise text-guided image transformation and creative editing. | 2025-11 |
| Qwen Image Edit Plus Eraser | qwen-image-edit-plus-eraser | Remove unwanted objects while preserving realistic backgrounds. | 2025-11 |
| Qwen Image Edit Plus Face To Portrait | qwen-image-edit-plus-face-to-portrait | Cropped face into full identity-preserving portrait photo. | 2025-11 |
| Qwen Image Edit Plus Group Photo | qwen-image-edit-plus-group-photo | Merge individual portraits into realistic group photos. | 2025-11 |
| Qwen Image Edit Plus Multi Lora | qwen-image-edit-plus-multi-lora | Multi-image editing with superior identity and style control. | 2025-11 |
| Qwen Image Edit Plus Multiple Angles | qwen-image-edit-plus-multiple-angle | Transform image perspective with natural language prompts. | 2025-11 |
| Qwen Image Edit Plus Next Scene | qwen-image-edit-plus-next-scene | Create cinematic sequences with seamless visual continuity. | 2025-11 |
| Qwen Image Edit Plus Product Photography | qwen-image-edit-plus-product-photography | Transform white-background products into immersive lifestyle scenes. | 2025-11 |
| Qwen Image Edit Plus Relight | qwen-image-edit-plus-relight | Advanced image relighting using natural language prompts. | 2025-11 |
| Qwen Image Edit Plus Remove Lighting | qwen-image-edit-plus-remove-lighting | Remove artificial lighting effects and restore natural tones. | 2025-11 |
| Qwen Image Edit Plus Texture Apply | qwen-image-edit-plus-texture-apply | Apply precise textures to images using natural language. | 2025-11 |
| Qwen Image Edit Plus Texture Extract | qwen-image-edit-plus-texture-extract | Extract seamless, tileable textures from photographs. | 2025-11 |
| Runway Gen 4 Image | runway-gen4-image | Runway's Gen-4 Image API enables precise, multimodal image generation for innovative creative and technical applications. | 2025-05 |
| Segment Anything Model | sam-img2img | The Segment Anything Model (SAM) produces high quality object masks from input prompts such as points or boxes, and it can be used to gener… | 2023-09 |
| Sam V2 Image | sam-v2-image | SAM v2, the next-gen segmentation model from Meta AI, revolutionizes computer vision. Building on SAM's success, it excels at accurately se… | 2024-08 |
| Sam3 Image | sam3-image | Precise object segmentation and tracking in images. | 2025-11 |
| ControlNet Canny | sd1.5-controlnet-canny | This model corresponds to the ControlNet conditioned on Canny edges. | 2023-09 |
| ControlNet Depth | sd1.5-controlnet-depth | This model corresponds to the ControlNet conditioned on Depth estimation. | 2023-09 |
| ControlNet Openpose | sd1.5-controlnet-openpose | This model corresponds to the ControlNet conditioned on Human Pose Estimation. | 2023-09 |
| ControlNet Scribble | sd1.5-controlnet-scribble | This model corresponds to the ControlNet conditioned on Scribble images. | 2023-09 |
| ControlNet Soft Edge | sd1.5-controlnet-softedge | This model corresponds to the ControlNet conditioned on Soft Edge. | 2023-09 |
| Stable Diffusion img2img | sd1.5-img2img | This model uses diffusion-denoising mechanism as first proposed by SDEdit, Stable Diffusion is used for text-guided image-to-image translat… | 2023-09 |
| SD Outpainting | sd1.5-outpaint | Stable Diffusion Outpainting can extend any image in any direction | 2023-09 |
| Faceswap | sd2.1-faceswapper | Take a picture/gif and replace the face in it with a face of your choice. You only need one image of the desired face. No dataset, no train… | 2023-09 |
| SD3 Medium Canny Controlnet | sd3-med-canny | Stable Diffusion 3 (SD3) Medium Canny ControlNet uses Canny edge detection to provide fine-grained control over the generated outputs. | 2024-07 |
| SD3 Medium Pose Controlnet | sd3-med-pose | Stable Diffusion 3 (SD3) Pose ControlNet is a large generative image model tailored for generating images based on text prompts while using… | 2024-07 |
| SD3 Medium Tile Controlnet | sd3-med-tile | SD3 Medium Tile ControlNet is a large generative image model designed for generating detailed images based on textual prompts and tile-base… | 2024-07 |
| SDXL Controlnet | sdxl-controlnet | SDXL ControlNet gives unprecedented control over text-to-image generation. SDXL ControlNet models Introduces the concept of conditioning in… | 2024-07 |
| SDXL Img2Img | sdxl-img2img | SDXL Img2Img is used for text-guided image-to-image translation. This model uses the weights from Stable Diffusion to generate new images f… | 2024-07 |
| SDXL-Openpose | sdxl-openpose | This model leverages SDXL to generate the images with ControlNet conditioned on Human Pose Estimation. | 2023-11 |
| Seedream 4.0 (4k) | seedream-4 | Seedream 4.0 generates high-resolution, professional-grade visuals with superior text rendering for impactful design. | 2025-09 |
| Seedream 4.5 | seedream-4.5 | Photorealistic image generation with precise text understanding. | 2025-12 |
| Seedream 5.0 Pro Layer Decomposition | seedream-5-pro-layer-decomposition | Split any image into editable transparent PNG layers. | 2026-08 |
| Seedream 5.0 Lite: Image-to-Image | seedream-v5-lite-image-to-image | Transform images intelligently with detailed text prompts. | 2026-02 |
| Segmind SegSwap v0.1 | seg-swap | Swap Objects Instantly. The Segmind SegSwap v0.1 model enables dynamic and precise image editing by allowing users to remove, replace, or a… | 2025-03 |
| Segmind SegFit v1.1 | segfit-v1.1 | Segmind's Fashion and Immersive Try-on model. SegFIT offers effortless AI virtual try-on from just a product image. No models needed! Boost… | 2025-04 |
| Segmind SegFit v1.2 | segfit-v1.2 | SegFit v1.2 creates hyper-realistic virtual try-on images, transforming fashion retail engagement and conversion rates. | 2025-06 |
| Segmind SegFit v1.3 | segfit-v1.3 | SegFit v1.3 enables hyper-realistic virtual try-ons, enhancing online fashion retail experiences without physical photoshoots. | 2025-07 |
| Segmind Relighting | segmind-relighting | Prompts to auto-magically relight your images. | 2025-03 |
| Segmind Relighting V2 | segmind-relighting-v2 | Transform images with customizable, photorealistic lighting for unparalleled visual creativity and authenticity. | 2025-05 |
| Segmind SceneCraft v0.1 | segmind-scenecraft-v01 | SceneCraft transforms plain or existing product images into visually rich, photorealistic scenes. Whether starting from a white background… | 2025-04 |
| Skin Contrast Upscaler | skin-contrast-upscaler | Enhances skin detail in images while preserving background quality for professional photography and art. | 2025-05 |
| Smart Banner Resizer | smart-banner-resizer | Recompose one image into multiple ad and banner sizes. | 2026-05 |
| SSD-Depth | ssd-depth | This model leverages SSD-1B to generate the images with ControlNet conditioned on Depth Estimation | 2023-11 |
| SSD Img2Img | ssd-img2img | This model uses SSD-1B to generate images by passing a text prompt and an initial image to condition the generation | 2023-11 |
| Story Diffusion | storydiffusion | Story Diffusion turns your written narratives into stunning image sequences. | 2024-07 |
| IPAdapter Style Transfer | style-transfer | Style & Composition Transfer with Stable Diffusion IP Adapter | 2024-06 |
| Image Superimpose | superimpose | Superimpose model lets you to create captivating visuals by seamlessly overlaying one image on top of another. It streamlines your image la… | 2024-07 |
| Image Superimpose V2 | superimpose-v2 | Superimpose V2 elevates image editing! Seamlessly layer images with background removal, precise positioning, and flexible resizing options.… | 2024-07 |
| Supir Photo-Realistic Image Restoration | supir | SUPIR restores and enhances images to stunning, photo-realistic quality with advanced AI techniques. | 2025-04 |
| Text Overlay | text-overlay | Elevate your visuals withText Overlay Model. Easily add customized text to any image, perfect for social media, marketing, and blogs. Enjoy… | 2024-09 |
| Topaz Labs Image Upscale | topaz-image-upscale | Topaz Labs image upscale is an industry-leading AI photo upscaler designed to increase the resolution of photos while preserving and enhanc… | 2025-04 |
| Transparent Background Maker | transparent-background-maker | Transform your images with Transparent Background Maker. Quickly remove backgrounds using AI technology, supporting PNG and JPG formats. Id… | 2024-11 |
| Image Mask | utility-image-mask | Build / refine binary masks from JSON-described shapes (rect/polygon/ellipse). Optional dilate/erode/feather/invert; multi-shape merge. | 2026-04 |
| Image Transform Pipeline | utility-image-transform | Apply an ordered pipeline of resize / crop / rotate / flip in a single call. Replaces the four separate tools. | 2026-04 |
| Word2img | w2imgsd1.5-img2img | Create beautifully designed words using Segmind’s word to image for your marketing purposes | 2023-09 |
Image generation
| Model | Slug | Description | Added |
|---|---|---|---|
| LordSwaminaryan | 675ed494b9-thejagstudio-LordSwaminaryan | 2024-08 | |
| sdxl_lora_architecture_siheyuan | 9c54110684-frank-chieng-sdxl_lora_architecture_siheyuan | 2024-04 | |
| FLUX.1 | acn-0effee052a | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FG-v3 | aio-fg-full-fc1d87eafb | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| Aladdin | Aladdin5k-701ba0d580 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| FLUX.1 | alexV2-2a5661e8ef | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-08 |
| FLUX.1 | AlexV2-655afc6166 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-08 |
| FLUX.1 | alexV2FastFlux-e99e4b5aa1 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-08 |
| FLUX.1 | alice-in-wonderland-ab7255451b | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-12 |
| FLUX.1 | ANIME-GEN-V4-9dae59678b | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| FLUX.1 | anlora-e8d79af2af | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| FLUX.1 | ap-fc0862b464 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-12 |
| FLUX.1 | arlora-9ceba116a8 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| FLUX.1 | armp-13752d05df | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| ClayAnimationRedmond | artificialguybr-ClayAnimationRedmond | Clay Animation Redmond based on SDXL 1.0, excels at creating mesmerizing clay animation images with unparalleled ease and precision. | 2023-10 |
| FLUX.1 | Asherflex-c25117ed67 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| FLUX.1 | awf-66afd16a65 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-12 |
| FLUX.1 | awmcn-617fbb4446 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | awp-29909cf89b | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-12 |
| FLUX.1 | Ayalora-5ff7f8f2fd | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| watercolor_style_lora_sdxl | b0445a2335-ostris-watercolor_style_lora_sdxl | 2024-03 | |
| Background Eraser | background-eraser | Background Eraser helps in flawless background removal with exceptional accuracy. | 2024-06 |
| FLUX.1 | baiqiang-8812bdefb9 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | Balloonblowing-d8d4c47115 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| FLUX.1 | bed-and-chair-combined-8dd2e74338 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| FLUX.1 | bizhen-3683a8e947 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| Bria FIBO 1.5 | bria-fibo-generate | Generate photorealistic images with structured JSON prompt control. | 2025-11 |
| Bria 3.2 Text to Image | bria-text-to-image | Bria 3.2 AI transforms natural language into stunning visuals for diverse creative applications — with Base, Fast, and HD modes to match yo… | 2025-08 |
| Bria Vector Graphics | bria-text-to-vector-graphics | Bria Vision enables high-quality text-to-image and text-to-vector graphic generation for versatile commercial use. | 2025-08 |
| aether-bubbles-foam-lora-for-sdxl | c42bf9a66d-joachimsallstrom-aether-bubbles-foam-lora-for-sdxl | 2024-02 | |
| FLUX.1 | cf-6ccbc88846 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | cfoly-c65493430d | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| blacklight-makeup-sdxl-lora | chillpixel-blacklight-makeup-sdxl-lora | Blacklight Makeup SDXL LoRA is fine-tuned to generate makeup designs that are not only visually striking but also perfectly suited for blac… | 2023-10 |
| Chroma | chroma | Chroma is an open-source, 8.9B parameter text-to-image model (based on FLUX.1-schnell) designed for diverse and uncensored content generati… | 2025-05 |
| FLUX.1 | cloth-finetune-flux-91b5cc1278 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| FLUX.1 | crazy-8eae969efe | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| FLUX.1 | ctmn-806488193b | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | Delibrate_V2-8dbc3ba5a9 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| FLUX.1 | dianying-549c4441a6 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | Dinaone-ExteriorFloor-c404fe7cc8 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | djwx-635a3531b4 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| dog-example-sdxl-lora | dminhk-dog-example-sdxl-lora | Dog Example SDXL LoRA, a specialized AI model within the Stable Diffusion XL framework, uniquely trained to enhance canine imagery. | 2023-10 |
| FLUX.1 | doctor-dolittle-0018c78851 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-12 |
| FLUX.1 | dslora-b6e9a6e529 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| FLUX.1 | dw_bedroom_1_5k-71b62c19fb | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| FLUX.1 | dw_bedroom_2k-559dc9f020 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| FLUX.1 | egyptian-d9f4de37d4 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| Fast Flux.1 Schnell | fast-flux-schnell | Fast Flux.1 Schnell by Segmind is an optimized text-to-image model designed for developers needing faster image generation. It offers high… | 2024-08 |
| sdxl-lora-index-modern-luxury-1 | fb4ab6b705-naphatmanu-sdxl-lora-index-modern-luxury-1 | 2024-02 | |
| FLUX.1 | felora-94021b7888 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| FLUX.1 | fg_v2-5a5c3da2b3 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| FLUX.1 | fg-individual-15-5b442665e9 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| Forge Vision | fg15k-4879e3ca32 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-12 |
| FLUX.1 | fg5kFull-2fffd1e2b0 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| FLUX.1 | fgResize5k-b2981fb7b8 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| flux-pro-1.1 | flux-1.1-pro | Flux Pro 1.1 is a cutting-edge image generation tool offering exceptional speed, quality, and customization. Ideal for digital artists, des… | 2024-10 |
| Flux-1.1 Pro Ultra | flux-1.1-pro-ultra | Create stunning visuals effortlessly with Flux 1.1 Pro Ultra. Experience unparalleled image quality and speed. | 2024-11 |
| Flux.1 Dev | flux-dev | Flux Dev is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-08 |
| Flux Dev Finetuned | flux-dev-finetuned | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-07 |
| FLUX.1 | flux-hanurama-22f3d043c0 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| FLUX.1 | flux-pixar-27334bad28 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-04 |
| Flux .1 Pro | flux-pro | Flux Pro is a state-of-the-art image generation with top of the line prompt following, visual quality, image detail and output diversity. | 2024-08 |
| Flux Realism Lora with Upscale | flux-realism-lora | Flux Realism Lora with upscale, developed by XLabs AI is a cutting-edge model designed to generate realistic images from textual descriptio… | 2024-08 |
| Flux.1 Schnell | flux-schnell | Flux Schnell is a state-of-the-art text-to-image generation model engineered for speed and efficiency. | 2024-08 |
| FLUX.1 | Flux-Turbo-1b4d78be3f | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| FLUX.1 | Flux-Turbo-37e8b8f594 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| FLUX.1 | fluxanimals-05320abd47 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-03 |
| FLUX.1 | fluxdbb-a972e916ef | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| FLUX.1 | FluxPixar-30c83df57d | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-04 |
| FLUX.1 | FluxTurbo-f28792fc23 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| FLUX.1 | fulora-f3df2e804d | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| FLUX.1 | fylora-fb68c65109 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| FLUX.1 | Gen-v2-5c9cdcda8e | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| FLUX.1 | Ghibil-27c901962a | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-04 |
| cyborg_style_xl | goofyai-cyborg_style_xl | Cyborg Style SDXL specializes in generating cyborg-themed artwork based on science fiction and futuristic aesthetics. | 2023-10 |
| GPT Image 1 | gpt-image-1 | Create high-quality AI-generated images from text prompts using OpenAI's GPT Image 1 model. Ideal for product design, content creation, and… | 2025-04 |
| GPT Image 1 Mini | gpt-image-1-mini | High-quality image generation from text, fast and affordable. | 2025-10 |
| GPT Image 1.5 | gpt-image-1.5 | Stunning photorealistic images with exceptional instruction-following. | 2025-12 |
| GPT Image 2 | gpt-image-2 | Generate photorealistic images with legible multilingual text and 2K output. | 2026-04 |
| lora-sdxl-notion-illustration | gvrizzo-lora-sdxl-notion-illustration | 2024-01 | |
| FLUX.1 | haosc-610e769fa8 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| Ideogram 2a Text To Image | ideogram-2a-txt-2-img | Create captivating designs, realistic images & innovative logos with Ideogram 2a text-to-image. | 2025-03 |
| Ideogram 3.0 | ideogram-3 | Ideogram 3.0 revolutionizes content creation with photorealistic text-to-image generation and diverse aesthetic styles. | 2025-05 |
| Ideogram 4.0 | ideogram-4 | Generate 2K posters and logos with accurate text rendering. | 2026-06 |
| Ideogram Turbo Text To Image | ideogram-turbo-txt-2-img | Create stunning images in seconds with Ideogram Turbo Text to Image. Fast AI model for quick ideation & text rendering. | 2025-03 |
| Ideogram Text To Image | ideogram-txt-2-img | Ideogram Text to Image: Turn your ideas into stunning visuals instantly with this powerful AI tool. Create captivating designs, realistic i… | 2024-09 |
| Ideogram V4 Fast | ideogram-v4-fast | Generate posters and logos with accurate in-image text. | 2026-07 |
| Imagen 3 | imagen | Imagen 3 is Google DeepMind's highest quality text-to-image model. Generates detailed images with enhanced lighting, diverse styles, and im… | 2025-02 |
| Imagen 4 | imagen-4 | Imagen 4 is Google’s most advanced AI image generation model, creating detailed, photorealistic or abstract images from text prompts. It ex… | 2025-05 |
| Imagen 4 Fast | imagen-4-fast | Fast photorealistic image generation for bulk and iteration. | 2026-05 |
| Imagen 4 Ultra | imagen-4-ultra | Photorealistic images with native 2K resolution and precise text. | 2026-05 |
| sdxl-khuze-nocrop-1e-4-1200 | jayashri710-sdxl-khuze-nocrop-1e-4-1200 | 2024-01 | |
| sdxl-lora-khuze-1e-4-1200-512x512images | jayashri710-sdxl-lora-khuze-1e-4-1200-512x512images | 2024-01 | |
| FLUX.1 | jdq-396bb5a145 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | Jenny1-6546baecae | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| Juggernaut Lightning Flux | juggernaut-lightning-flux | Juggernaut Lightning Flux: Blazing fast (<300ms!) & powerful inference with enhanced visuals. | 2025-03 |
| Juggernaut Pro Flux | juggernaut-pro-flux | Juggernaut Pro FLUX: Create stunningly realistic AI images with unprecedented detail and sharpness. | 2025-03 |
| FLUX.1 | jzzs-a0fc97d5df | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | jzzsd-ef9fd76525 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| punk-collage | KappaNeuro-punk-collage | Punk Collage Model offers a unique way to create digital collages that resonate with the punk culture's raw energy and subversive charm. | 2023-10 |
| lora-sdxl-watercolor | kchoi-lora-sdxl-watercolor | 2024-01 | |
| FLUX.1 | KellyTest-66fbd687a2 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-06 |
| Kling V3 Text to Image | kling-3-text2image | Photorealistic, print-ready images from text prompts. | 2026-03 |
| FLUX.1 | liangnv-b2e39c7ec4 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | M-EdenRock-bed-2d45decb6a | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| FLUX.1 | M-HubbaArmChair-and-M-RicochetFabric-bed-18fed2c922 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| FLUX.1 | M-HubbaArmChair-fff004d426 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| FLUX.1 | M-RicochetFabric-bed-9180ccb8f7 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| FLUX.1 | Marimekko-mrk3903-1e8cab847c | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-03 |
| FLUX.1 | meizhuang-2ae6bc0009 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| sdxl-ugly-sonic-lora | minimaxir-sdxl-ugly-sonic-lora | SDXL Ugly Sonic LoRA excels at generating quirky and iconic version of one of the most beloved movie characters - Sonic the hedgehog. | 2023-10 |
| sdxl-wrong-lora | minimaxir-sdxl-wrong-lora | SDXL Wrong LoRA is engineered with a focus on delivering images of higher detail, color saturation and vibrance, bringing images to life wi… | 2023-10 |
| FLUX.1 | mnsh-86fb5b8ef4 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | mnzs-b5e1ea1473 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| SDXL - Multi Lora | multi-lora | SDXL Model with multiple LoRa loading support. | 2024-04 |
| Nano Banana | nano-banana | Gemini Image Editor preserves authentic subject identity while enabling seamless image editing and manipulation. | 2025-08 |
| FLUX.1 | Narinder-dfac11ecfb | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-05 |
| FLUX.1 | nav_007_Krishna-6fce6aad5b | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| lego-minifig-xl | nerijs-lego-minifig-xl | LEGO Minifig XL is designed to generate LEGO images and excels in creating detailed and accurate representations of LEGO minifigures and it… | 2023-10 |
| SDXL-StickerSheet-Lora | Norod78-SDXL-StickerSheet-Lora | SDXL StickerSheet LoRA is expertly fine-tuned on a comprehensive collection of sticker images, enabling it to produce a wide variety of sti… | 2023-10 |
| crayon_style_lora_sdxl | ostris-crayon_style_lora_sdxl | Crayon Style - SDXL LoRA is a unique model designed to convert any text prompt into a vibrant, crayon-style drawing. | 2023-10 |
| ikea-instructions-lora-sdxl | ostris-ikea-instructions-lora-sdxl | Ikea Instructions LoRA SDXL model is fine-tuned on IKEA diagrams and specializes in generating clear, concise, and easy-to-follow visual in… | 2023-10 |
| stained-glass-style-sdxl | ostris-stained-glass-style-sdxl | Stained Glass Style SDXL is trained extensively on diverse stained glass images and can replicate the essence of stained glass in digital a… | 2023-10 |
| Pruna P Image | p-image | p-image generates high-quality images from text prompts in seconds, optimizing for speed and fidelity. | 2025-11 |
| Pruna P Image Ideogram | p-image-ideogram | Sub-second text-to-image with legible in-image text. | 2026-07 |
| FLUX.1 | pixar-3d-5ab56ddd76 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-02 |
| FLUX.1 | playzippyIramayanakids-7f162e1446 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| FLUX.1 | prabhas-5f173e43a4 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| FLUX.1 | quanshen-7d0d5c4bdb | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| Qwen Image | qwen-image | Qwen-Image revolutionizes image generation and editing with seamless multilingual text integration and photorealistic detail. | 2025-08 |
| Qwen Image 2512 | qwen-image-2512 | Photorealistic image generation with precise text description following. | 2025-12 |
| Qwen Image 3.0 | qwen-image-3 | Generate and edit legible in-image text, up to 2K. | 2026-08 |
| Qwen Image Fast | qwen-image-fast | Qwen-Image expertly generates stunning images with complex text integration, especially for Chinese typography. | 2025-08 |
| sdxl-lora-lower-decks-aesthetic | ra100-sdxl-lora-lower-decks-aesthetic | SDXL LoRA Lower Decks Aesthetic model, inspired by the unique style of “Star Trek: Lower Decks.” generates artwork in the distinctive anima… | 2023-10 |
| Rama | Rama-558043a138 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| Ramayan Model | ramayanaMulti5k-69c1cc4ae9 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| lora-dog-SSD-1B | ramsrigouthamg-lora-dog-SSD-1B | LoRA Dog SSD-1B specializes in generating photorealistic images of dogs | 2023-10 |
| Recraft V3 | recraft-v3 | Recraft V3, the latest iteration of Recraft AI, offers a significant advancement in AI-driven image generation. This state-of-the-art model… | 2024-11 |
| Recraft V3 Svg | recraft-v3-svg | Recraft V3 SVG generates high-quality, customizable vector graphics with precision and ease. Perfect for logos, infographics, illustrations… | 2024-11 |
| Reve 2 | reve-2 | Generate and edit 4K images with sharp in-image text. | 2026-07 |
| FLUX.1 | rxmx-324bd4c1e2 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | SanaAI-542b8dbe01 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-10 |
| FLUX.1 | SBIC-Stone-fbcc5912c5 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | sclora-274da0a25f | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| Cyber Realistic | sd1.5-cyberrealistic | The most versatile photorealistic model that blends various models to achieve the amazing realistic images. | 2023-09 |
| Edge of Realism | sd1.5-edgeofrealism | This model corresponds to the Stable Diffusion Edge of Realism checkpoint for detailed images at the cost of a super detailed prompt | 2023-09 |
| Epic Realism | sd1.5-epicrealism | This model corresponds to the Stable Diffusion Epic Realism checkpoint for detailed images at the cost of a super detailed prompt | 2023-09 |
| Juggernaut Final | sd1.5-juggernaut | The most versatile photorealistic model that blends various models to achieve the amazing realistic images. | 2023-09 |
| Realistic Vision | sd1.5-realisticvision | This model corresponds to the Stable Diffusion Realistic Vision checkpoint for detailed images at the cost of a super detailed prompt | 2023-09 |
| Reliberate | sd1.5-reliberate | This model corresponds to the Stable Diffusion Reliberate checkpoint for detailed images at the cost of a super detailed prompt | 2023-09 |
| Colossus Lightning SDXL | sdxl1.0-colossus-lightning | Colossus Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps. | 2024-03 |
| Dreamshaper SDXL | sdxl1.0-dreamshaper | The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software. | 2023-10 |
| DreamShaper Lightning SDXL | sdxl1.0-dreamshaper-lightning | DreamShaper Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps. | 2024-03 |
| Dynavis Lightning SDXL | sdxl1.0-dyanvis-lightning | Dynavis Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps. | 2024-03 |
| Juggernaut Lightning SDXL | sdxl1.0-juggernaut-lightning | Juggernaut Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps. | 2024-03 |
| NewReality Lightning SDXL | sdxl1.0-newreality-lightning | NewReality Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps. | 2024-03 |
| NightVis Lightning SDXL | sdxl1.0-nightvis-lightning | NightVis Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps. | 2024-03 |
| ProtoVision Lightning SDXL | sdxl1.0-protovis-lightning | ProtoVision Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps. | 2024-03 |
| RealDream Lightning | sdxl1.0-realdream-lightning | RealDream is a sophisticated image generation model utilizing SDXL Lightning architecture. It creates incredibly realistic images from text… | 2024-07 |
| Realdream Pony V9 | sdxl1.0-realdream-pony-v9 | Real Dream Pony V9 is an advanced image generation model based on the Stable Diffusion XL (SDXL) architecture, excelling in photorealism. | 2024-07 |
| Realism Lightning SDXL | sdxl1.0-realism-lightning | Realism Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps. | 2024-03 |
| Realvis SDXL | sdxl1.0-realvis | The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software. | 2023-10 |
| Realvis Lightning SDXL | sdxl1.0-realvis-lightning | Realvis Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps. | 2024-03 |
| Samaritan 3D XL | sdxl1.0-samaritan-3d | Samaritan 3D XL leverages the robust capabilities of the SDXL framework, ensuring high-quality, detailed 3D character renderings. | 2023-12 |
| Samaritan Lightning SDXL | sdxl1.0-samaritan-lightning | Samaritan Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps. | 2024-03 |
| Copax Timeless SDXL | sdxl1.0-timeless | The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software. | 2023-10 |
| Stable Diffusion XL 1.0 | sdxl1.0-txt2img | The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software | 2023-09 |
| WildCard Lightning SDXL | sdxl1.0-wildcard-lightning | WildCard Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps. | 2024-03 |
| Zavychroma SDXL | sdxl1.0-zavychroma | The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software. | 2023-10 |
| Seedream 5.0 Pro | seedream-5-pro | Region-precise image editing with native multilingual text. | 2026-07 |
| Seedream 5.0 Lite: Text-to-Image | seedream-v5-lite-text-to-image | Fast, affordable instruction-following image generation. | 2026-02 |
| Segmind-Vega | segmind-vega | The Segmind-Vega Model is a distilled version of the Stable Diffusion XL (SDXL), offering a remarkable 70% reduction in size and an impress… | 2023-12 |
| Segmind-VegaRT | segmind-vega-rt-v1 | Segmind-VegaRT a distilled consistency adapter for Segmind-Vega that allows to reduce the number of inference steps to only between 2 - 8 s… | 2023-12 |
| FLUX.1 | Selfie-Aesthetic-f9c0f409c5 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | Selfie-Aesthetic1-9e21b5a57f | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | sf-31e1b1ee3c | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | sheb-b899ceaa9f | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-04 |
| FLUX.1 | sheb-b9150b5b83 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-04 |
| FLUX.1 | Shravya_ai-f44dd1e647 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| Simple Vector Flux Lora | Simple_Vector_Flux | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| FLUX.1 | siria-e3fb0ee32f | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-03 |
| Sita | Sita-b382eb5c53 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| Flux.1 Schnell | softpasty-flux-dev-a4edbd379c | Flux Schnell is a state-of-the-art text-to-image generation model engineered for speed and efficiency. | 2024-09 |
| FLUX.1 | SpeedyMary-b361f2fd1d | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-02 |
| SSD-1B | ssd-1b | SSD-1B efficiently generates high-quality, diverse images from text prompts in real-time. | 2023-10 |
| Stable Diffusion 3 Medium Text to Image | stable-diffusion-3-medium-txt2img | Stable Diffusion is a type of latent diffusion model that can generate images from text. It was created by a team of researchers and engine… | 2024-06 |
| Stable Diffusion 3.5 Large Text to Image | stable-diffusion-3.5-large-txt2img | Stable Diffusion 3.5 Large offers exceptional customizability, efficient performance on consumer hardware, and diverse image outputs that a… | 2024-10 |
| Stable Diffusion 3.5 Turbo Text to Image | stable-diffusion-3.5-turbo-txt2img | Stable Diffusion 3.5 Turbo offers exceptional customizability, efficient performance on consumer hardware, and diverse image outputs that a… | 2024-10 |
| FLUX | Test 001 -4b5d788e6c | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-08 |
| FLUX.1 | thumbnail-bc5f00ef3c | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | tianmei-e09a3dfa92 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | v6lora-1745486a5e | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-09 |
| FLUX.1 | veo-dd69bf026e | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-11 |
| Wan 2.7 Image Generation | wan2.7-image | 2K image generation with precise multilingual text rendering. | 2026-04 |
| Wan 2.7 Image Generation Pro | wan2.7-image-pro | 4K images with chain-of-thought reasoning and multilingual text. | 2026-04 |
| FLUX.1 | whnm-339c18f951 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | winnie-the-pooh-c3ccfd99b1 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2024-12 |
| FLUX.1 | wxns-9cf0075aee | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | xse-1f865b3596 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | xz-c445d830b2 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | yujia-0d97d92ca5 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| Flux.1 Schnell | yunxi-aac5d8354b | Flux Schnell is a state-of-the-art text-to-image generation model engineered for speed and efficiency. | 2025-04 |
| Z Image Turbo | z-image-turbo | Photorealistic images in under one second, bilingual text. | 2025-11 |
| FLUX.1 | zh-66dcb61456 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | zscj-7c64578707 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | zsjz-7ca8eff909 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
| FLUX.1 | zsnh-7634404088 | Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions | 2025-01 |
Image t o image
| Model | Slug | Description | Added |
|---|---|---|---|
| Stable Diffusion 3 Medium Image to Image | sd3-med-img2img | Stable Diffusion 3 Medium image-to-image is a cutting-edge AI tool that uses advanced image-to-image technology to transform one image into… | 2024-07 |
Image to data
| Model | Slug | Description | Added |
|---|---|---|---|
| Image Metadata | utility-image-metadata | Read image metadata: dimensions, format, EXIF (with GPS decoded to decimal), ICC profile name, raw XMP. Returns JSON, not an image. | 2026-04 |
Image to video
| Model | Slug | Description | Added |
|---|---|---|---|
| AI Face Swap (image and video) | ai-face-swap | AI Face Swap: Effortlessly replace faces online. Fine-tune swaps with advanced controls for age, gender, and resolution. | 2025-01 |
| Cog videoX Image To Video | cog-video-5b-i2v | CogVideoX image-to-video is a cutting-edge AI model that converts static images into dynamic, high-quality videos. Perfect for content crea… | 2024-09 |
| FLUX 3 Image to Video | flux-3-image-to-video | Animate images into 20-second clips with synchronized native audio. | 2026-08 |
| Grok Imagine Video 1.5 Image to Video | grok-imagine-video-1.5-image-to-video | Animate a still image into 1080p video with synced audio. | 2026-08 |
| Grok Imagine Video 1.5 (Preview) | grok-imagine-video-1.5-preview | Image-to-video with native synchronized audio, up to 720p. | 2026-06 |
| Grok Imagine Video 1.5 Reference to Video | grok-imagine-video-1.5-reference-to-video | Character-consistent video from up to 7 reference images. | 2026-08 |
| Hailuo 02 Fast | hailuo-02-fast | Transform any static image into a captivating, high-quality video clip effortlessly. | 2025-08 |
| Hailuo 2.3 | hailuo-2.3 | Hyper-realistic videos from text with fluid character motion. | 2025-10 |
| Hailuo 2.3 Fast | hailuo-2.3-fast | Professional-quality videos from text and images at speed. | 2025-10 |
| Hallo | hallo | Hallo lets you create portrait videos from single images. | 2024-06 |
| HappyHorse 1.0 | happyhorse | Cinematic 1080p text-to-video with native audio and lip-sync. | 2026-04 |
| HappyHorse 1.1 | happyhorse-1.1 | Generate cinematic video with synchronized native audio and multilingual lip-sync from text, an image, or reference images. | 2026-06 |
| Heygen Avatar IV | heygen-avatar-iv | Single photo into a lifelike talking avatar video. | 2025-12 |
| Higgsfield Image 2 Video | higgsfield-image2video | Transform static images into dynamic, motion-rich videos with unparalleled control and creative depth. | 2025-09 |
| Higgsfield Speech 2 Video | higgsfield-speech2video | Transform images and audio into dynamic, lip-synced videos for engaging digital content. | 2025-09 |
| InfiniteTalk | infinite-talk | Full-body animation from images synchronized perfectly to audio. | 2025-10 |
| Kling AI 1.6 Image to Video | kling-1.6-image2video | Kling AI 1.6 Image-to-Video is a powerful AI tool that transforms static images into captivating, animated videos. Create high-quality cont… | 2025-01 |
| Kling 2 | kling-2 | Kling 2.0 is an advanced AI video generator (5 and 10 seconds) that creates cinematic, dynamic videos from text or images with lifelike mot… | 2025-04 |
| Kling 2.1 AI Video Generator | kling-2.1 | Kling 2.1 offers hyper-realistic video generation with improved motion, sharper 1080p visuals, and instant restyling capabilities. Its cost… | 2025-05 |
| Kling 2.5 Turbo | kling-2.5-turbo | Kling AI 2.5 Turbo generates fluid, cinematic videos from text and images, enhancing content creation and storytelling. | 2025-09 |
| Kling 2.6 | kling-2.6 | Still images into immersive cinematic videos with synchronized audio. | 2025-12 |
| Kling 3.0 Pro Image-to-Video | kling-3-pro-image2video | Animated 1080p videos from images with dynamic motion. | 2026-02 |
| Kling 3.0 Standard Image-to-Video | kling-3-standard-image2video | Controlled cinematic 1080p videos from starting images. | 2026-02 |
| Kling bloombloom | kling-bloombloom | Kling AI transforms text and images into dynamic, high-quality video content with realistic motion and sound. | 2025-05 |
| Kling dizzydizzy | kling-dizzydizzy | Kling DizzyDizzy transforms static content into dynamic, high-resolution videos, enhancing engagement and storytelling for creators. | 2025-05 |
| Kling Expansion | kling-expansion | Unleash dynamic visuals with Kling Expansion! Effortlessly inflate and stretch elements for surreal and captivating effects. | 2025-04 |
| Kling fuzzyfuzzy | kling-fuzzyfuzzy | Transform your photos instantly into adorable, plush-toy-like visuals with Kling fuzzyfuzzy effect. | 2025-04 |
| Kling Heart Gesture | kling-heart-gesture | Express affection visually with Kling AI's heart gesture effect! Input two portraits and instantly create heartwarming videos featuring a d… | 2025-04 |
| Kling Hug | kling-hug | Create heartwarming videos instantly with Kling hug effect! Generate tender embracing animations. | 2025-04 |
| Kling AI Image to Video | kling-image2video | Kling AI Image-to-Video is a powerful AI tool that transforms static images into captivating, animated videos. Create high-quality content… | 2024-10 |
| Kling Kiss | kling-kiss | Create a heartfelt video in seconds with Kling kiss effect! Input two portraits and instantly generate a kissing animation. | 2025-04 |
| Kling O1 Image 2 Video | kling-o1-image-to-video | Physics-driven animations from images for creative storytelling. | 2026-01 |
| Kling O1 Reference Image 2 Video | kling-o1-reference-image-to-video | Identity-preserving videos from static images with character reference. | 2026-01 |
| Kling O3 Image To Video | kling-o3-image2video | Images to cinematic videos with precise motion control. | 2026-03 |
| Kling Squish | kling-squish | Transform your visuals with Kling AI squish effect! Easily compress and distort images/videos for playful, exaggerated effects. | 2025-04 |
| Kling V1 Pro AI Avatar | kling-v1-pro-ai-avatar | Dynamic AI avatars with synchronized speech from image. | 2025-10 |
| Kling V1 Standard AI Avatar | kling-v1-standard-ai-avatar | Lifelike AI avatars with precise lip-sync for presentations. | 2025-10 |
| Kling V2 Pro Avatar | kling-v2-pro-avatar | Talking avatar videos from image and audio, high quality. | 2025-12 |
| Kling Avatar V2 Standard | kling-v2-standard-avatar | Lifelike video avatars with precise lip synchronization. | 2025-12 |
| Live Portrait | live-portrait | Live Portrait animates static images using a reference driving video through implicit key point based framework, bringing a portrait to lif… | 2024-07 |
| Live Portrait video to video | live-portrait-video-to-video | Experience the magic of Live Portrait’s Video-to-Video Model! Transform your static images into dynamic videos seamlessly. | 2024-07 |
| LTX 2 Fast | ltx-2-fast | Fast, high-quality text-to-video generation by Lightricks. | 2025-10 |
| LTX 2 Pro | ltx-2-pro | High-quality video generation with advanced motion control. | 2025-10 |
| LTX 2.5 Fast | ltx-2.5-fast | Text-to-video and image-to-video with native audio, up to 4K. | 2026-08 |
| LTX 2.5 Pro | ltx-2.5-pro | Generate 1080p video with native audio and multi-shot scenes. | 2026-08 |
| LTX Video | ltx-video | LTX-Video is the first DiT-based video generation model capable of generating high-quality videos in real-time. It produces 24 FPS videos a… | 2024-12 |
| Luma Ray 3.2 | luma-ray-3-2 | Cinematic text-to-video and image-to-video clips up to 1080p. | 2026-06 |
| MiniMax AI (Hailuo) | minimax-ai | With Video-01 by MiniMax, create high-definition videos at 720p resolution and 25fps, featuring cinematic camera movement effects based on… | 2024-12 |
| MiniMax Hailuo H3 Image to Video | minimax-h3-image-to-video | Animate a still image into 2K video up to 15s. | 2026-07 |
| MiniMax Hailuo H3 Reference to Video | minimax-h3-reference-to-video | Keep characters and products consistent in 2K reference-to-video. | 2026-07 |
| Minimax Hailou 2 | minimax-hailuo-2 | Generate breathtaking 1080P cinematic videos from text or images with ultra-realistic motion and physics. | 2025-07 |
| Luma Modify Video | modify-video | Transform videos seamlessly with high-fidelity generative edits while preserving original actor performances. | 2025-07 |
| Motion Control SVD | motionctrl-svd | Motion Control SVD is an innovative deep learning framework that breathes life into static images. By intelligently managing both camera an… | 2024-07 |
| Muscle Surge | muscle-surge | Instantly add muscle and strength to your videos with Pixverse Muscle Surge effect! | 2025-04 |
| Pruna P Video Animate | p-video-animate | Transfer video motion and audio onto any still image. | 2026-07 |
| Pruna P Video Avatar | p-video-avatar | Animate any portrait into a lip-synced talking avatar. | 2026-06 |
| Pruna P Video Replace | p-video-replace | Swap on-screen video characters while preserving motion and audio. | 2026-07 |
| Pixverse 4.5 Effects | pixverse-4.5-effects | PixVerse 4.5 transforms photos and text into stunning animated videos for impactful storytelling and marketing. | 2025-05 |
| Pixverse 4.5 Video | pixverse-4.5-video | Pixverse 4.5 transforms static images and text into dynamic, engaging videos for captivating social media content. | 2025-05 |
| Pixverse 5 Extend | pixverse-5-extend | Seamlessly extend and continue AI-generated videos. | 2025-10 |
| Pixverse 5 Transition | pixverse-5-transition | Seamless AI-generated video transitions between scenes. | 2025-10 |
| Pixverse 5 Video | pixverse-5-video | Cinematic videos from text and images with photorealism. | 2025-10 |
| Pixverse Image to Video | pixverse-image2video | Animate your photos effortlessly with Pixverse Image to Video AI! Upload, add motion prompts and styles. | 2025-04 |
| Pixverse Mimic | pixverse-mimic | Transfer motion from reference videos onto still images. | 2026-05 |
| Pixverse V6 | pixverse-v6 | 15-second AI videos with native audio and cinematic controls. | 2026-04 |
| Luma Ray flash 2 (720p) | ray-flash-2-720p | Generate stunning 720p videos from text with the Luma ray-flash-2-720p model. Faster & cheaper than Ray 2, offering realistic motion & deta… | 2025-03 |
| Runway Gen Alpha Turbo Image to Video | runway-gen3-alphaturbo | Runway Gen-3 AlphaTurbo is a cutting-edge AI tool that transforms static images into dynamic videos with exceptional fidelity and motion | 2024-10 |
| Runway Gen 4 Turbo | runway-gen4-turbo | Generate videos faster and cheaper with Runway Gen-4 Turbo! Create high-quality text, image, and combined video generation for rapid conten… | 2025-04 |
| SadTalker | sadtalker | Audio-based Lip Synchronization for Talking Head Video | 2024-06 |
| Wan Scail | scail | Professional character animations from reference images. | 2025-12 |
| Seedance 1.0 Pro Fast | seedance-1.0-pro-fast | Cinematic videos from text and images at ultra speed. | 2025-10 |
| Seedance 1.5 Pro | seedance-1.5-pro | Synchronized video and audio generation for dynamic storytelling. | 2025-12 |
| Seedance 2.0 | seedance-2.0 | Cinematic AI videos with native audio and multi-shot narratives. | 2026-04 |
| Seedance 2.0 Fast | seedance-2.0-fast | Professional-grade video creation model with native audio, similar to SeeDance 2.0 but faster and cheaper. | 2026-04 |
| Seedance 2.0 Mini | seedance-2.0-mini | Fast text-to-video and image-to-video with synchronized audio. | 2026-06 |
| Seedance 2.5 | seedance-2.5 | Generate cinematic multi-shot AI videos up to 30 seconds with synchronized native audio from text, images, or references. | 2026-08 |
| Seedance 1.0 Pro | seedance-pro | Seedance Pro transforms text and images into engaging 720p dynamic videos with cinematic storytelling. | 2025-07 |
| Sora 2 | sora-2 | Stunning dynamic videos from detailed text descriptions. | 2025-10 |
| Sora 2 Pro | sora-2-pro | Cinematic-quality videos from text with temporal consistency. | 2025-10 |
| Stable Video Diffusion | svd | Takes image as input and returns a video. | 2023-12 |
| Tooncrafter | tooncrafter | Create videos from illustrated input images | 2024-06 |
| V Express | v-express | V-Express lets you create portrait videos from single images. | 2024-06 |
| VEED Fabric 1.0 | veed-fabric-1.0 | Animate any image into a realistic talking video, lip-synced to your audio or generated from a text script. | 2026-07 |
| Google Veo 2 Image To Video | veo-2-image2video | Discover Google Veo 2, an AI-powered image-to-video model with 4K resolution, realistic motion, and cinematic effects for creators and deve… | 2025-03 |
| Veo 3.1 | veo-3.1 | Static images into high-quality videos with synchronized audio. | 2025-10 |
| Veo 3.1 Fast | veo-3.1-fast | Transforms static images into dynamic 1080p videos with synchronized audio and natural motion. | 2025-10 |
| Wan Video Effects | video-effects | Transform your videos with diverse video effects. Start creating captivating videos today. | 2025-03 |
| HyperSwap: Video Faceswap by FaceFusion Labs | video-faceswap-by-facefusion-labs | Realistic face swapping in videos from a single image. | 2026-03 |
| Video Frame Interpolation | video-frame-interpolation | FILM synthesizes smooth, high-quality intermediate frames for fluid motion in videos with significant movement. | 2025-09 |
| Video Stitch | video-stitch | Revolutionize your video editing with the Video Stitch Model. Seamlessly stitch clips, add captivating audio, and create professional-looki… | 2024-10 |
| Video Tryon | video-tryon | Video Tryon is Segmind’s next-generation AI video model for instant virtual try-on, allowing users to visualize any outfit on any person in… | 2025-09 |
| Video Tryon V2 | video-tryon-v2 | Video Tryon is Segmind’s next-generation AI video model for instant virtual try-on, allowing users to visualize any outfit on any person in… | 2025-11 |
| Video Watermark Remover | video-watermark-remover | Remove watermarks from any video instantly with AI. | 2025-10 |
| Video Faceswap | videofaceswap | Video Faceswap is a powerful tool for creators, filmmakers, and meme enthusiasts. With this innovative technology, you can effortlessly rep… | 2024-07 |
| Wan 2.2 Image to Video Fast | wan-2.2-i2v-fast | Transforms simple text prompts into breathtaking cinematic-quality videos in minutes. | 2025-08 |
| Wan 2.2 Image to Video Flash | wan-2.2-i2v-flash | Convert a single image into a coherent dynamic video. | 2026-03 |
| Wan 2.5 Image to Video | wan-2.5-i2v | Wan2.5-Preview creates stunning, high-resolution videos with flawless audio synchronization from multiple inputs. | 2025-09 |
| Wan 2.6 Image To Video | wan-2.6-i2v | Transform images into high-quality videos with audio sync. | 2025-12 |
| Wan Animate | wan-animate | Animate characters and replace video subjects seamlessly. | 2025-10 |
| Wan 2.1 720p image to video | wan2.1-i2v-720p | Create high-quality 720p videos with excellent visual quality and a broad spectrum of motion from static images. | 2025-02 |
| Wan 2.6 Image to Video Flash | wan2.6-i2v-flash | Animate photos into 15-second 1080p video with native audio. | 2026-08 |
| Wan 2.7 Image to Video | wan2.7-i2v | Animate any image into cinematic 1080P video with audio. | 2026-04 |
| Wan 2.7 Reference to Video | wan2.7-r2v | Character-consistent multi-subject videos from reference images. | 2026-04 |
| Wan 3.0 Video | wan3.0-video | Generate 30-second 1080p video with native audio. | 2026-08 |
| Warmth of Jesus | warmth-of-jesus | Experience the viral "Warmth of Jesus" effect on PixVerse! Transform your images into heartwarming videos of Jesus embracing people. | 2025-04 |
Image to3d
| Model | Slug | Description | Added |
|---|---|---|---|
| Hunyuan3D-2 | hunyuan-3d-2 | Hunyuan3D 2.0 enables the creation of high-quality 3D models with intricate details. Produce assets that are visually appealing and suitabl… | 2025-01 |
| Hunyuan3d-2.1 | hunyuan3d-2.1 | Transform 2D images into photorealistic, high-fidelity 3D assets effortlessly. | 2025-08 |
| Hunyuan-3d 2mv | hunyuan3d-2mv | Hunyuan3D-2mv is finetuned from Hunyuan3D-2 to support multiview controlled shape generation. | 2025-03 |
| Sam 3D Body | sam-3d-body | Reconstruct 3D human body meshes from a single photo. | 2025-12 |
| Sam 3D Object | sam-3d-objects | Single 2D image into detailed 3D object models. | 2025-12 |
Image understanding
| Model | Slug | Description | Added |
|---|---|---|---|
| Bria Fibo Structured Prompt | bria-fibo-generate-structured-prompt | Convert complex inputs into structured JSON prompts for generation. | 2025-11 |
| Bria Mask Generator | bria-mask-generator | Bria AI Get Masks automatically generates accurate object masks for advanced image editing and enhancement. | 2025-09 |
| Bria Prompt Enhancer | bria-prompt-enhancer | Bria AI generates high-quality, commercially safe images tailored to diverse creative needs. | 2025-09 |
| Google Translate | google-translate | Translate effortlessly with the powerful Google Translation AI model. | 2025-04 |
| Ideogram Describe | ideogram-describe | Ideogram describe can effortlessly generate detailed prompts from images. Perfect for refining creations or replicating styles. | 2025-03 |
| Image Converter | image-converter | Convert images between formats instantly. | 2025-10 |
| Image resizer | image-resizer | Resize images to any dimension quickly and precisely. | 2025-10 |
| LLAVA 1.6 7B | llava-v1.6 | LLaVa translates images into text descriptions & captions. | 2024-06 |
| NSFW Checker | nsfw-checker | Detect NSFW and other inappropriate content in images. Returns a boolean has_nsfw_concepts flag, an overall NSFW score (0-1), and the full… | 2026-05 |
| Sam V2.1 Hiera Large | sam-v21-hiera-large | Meta's next-gen segmentation model for images and video. | 2025-10 |
| Video Speed Change | video-speed-change | Speed up or slow down any video precisely. | 2025-10 |
Inpainting
| Model | Slug | Description | Added |
|---|---|---|---|
| Fooocus Inpainting | focus-inpaint | Fooocus Inpainting is a powerful image generation model that allows you to selectively edit and enhance images. | 2024-02 |
| Controlnet Inpainting | inpaint-auto | This model is capable of generating photo-realistic images given any text input, with the extra capability of inpainting and controlling th… | 2024-01 |
| Stable Diffusion Inpainting | sd1.5-inpainting | Stable Diffusion Inpainting is a latent text-to-image diffusion model capable of generating photo-realistic images given any text input, wi… | 2023-09 |
| SDXL Inpaint | sdxl-inpaint | This model is capable of generating photo-realistic images given any text input, with the extra capability of inpainting the pictures by us… | 2023-11 |
| Try-On Diffusion | try-on-diffusion | Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on | 2024-03 |
Language models (LLMs)
| Model | Slug | Description | Added |
|---|---|---|---|
| Claude 4 Sonnet | claude-4-sonnet | Advanced coding and multi-step agentic reasoning model. | 2025-10 |
| Claude 4.5 Sonnet | claude-4.5-sonnet | Claude Sonnet 4.5 empowers developers with advanced coding and reasoning for complex software solutions. | 2025-09 |
| Claude Opus 4.7 | claude-opus-4.7 | Anthropic's most capable AI model excelling at agentic coding, complex reasoning, and high-resolution vision with a 1M-token context window. | 2026-04 |
| DeepSeek Chat | deepseek-chat | DeepSeek V3 combines cutting-edge AI technology with practical usability. Featuring a 671B parameter architecture, enhanced reasoning capab… | 2025-01 |
| DeepSeek R1 | deepseek-reasoner | DeepSeek-R1 is a cutting-edge AI reasoning model that combines reinforcement learning with supervised fine-tuning. Excels in complex proble… | 2025-01 |
| Gemini 2.5 Flash | gemini-2.5-flash | Multimodal AI with transparent reasoning, fast and affordable. | 2025-10 |
| Gemini 2.5 Flash Lite | gemini-2.5-flash-lite | Fastest Gemini 2.5 model for high-volume text and vision tasks. | 2026-05 |
| Gemini 2.5 PRO | gemini-2.5-pro | Complex multimodal reasoning across diverse inputs and formats. | 2025-10 |
| Gemini 3 Flash | gemini-3-flash | Frontier-class reasoning and multimodal AI at scale. | 2026-05 |
| Gemini 3 Pro | gemini-3-pro | Autonomous multimodal AI for complex reasoning and coding. | 2025-11 |
| Gemini 3.1 Flash Lite | gemini-3.1-flash-lite | Ultra-fast, affordable LLM for high-volume AI pipelines. | 2026-05 |
| Gemini 3.1 Pro | gemini-3.1-pro | Frontier reasoning across text, images, video, and code. | 2026-05 |
| Gemini 3.7 Flash | gemini-3.7-flash | Fast multimodal LLM for coding, agents, and long-document analysis. | 2026-08 |
| GLM 5.2 | glm-5.2 | 1M-token open-weight LLM for long-horizon coding. | 2026-07 |
| GPT 4 | gpt-4 | GPT-4 outperforms both previous large language models and as of 2023, most state-of-the-art systems (which often have benchmark-specific tr… | 2024-05 |
| GPT 4 turbo | gpt-4-turbo | GPT-4 outperforms both previous large language models and as of 2023, most state-of-the-art systems (which often have benchmark-specific tr… | 2024-05 |
| GPT 4o | gpt-4o | GPT-4o (“o” for “omni”) is our most advanced model. It is multimodal (accepting text or image inputs and outputting text), and it has the s… | 2024-05 |
| GPT 5 | gpt-5 | GPT-5 automates complex coding tasks with integrated tools for seamless software development and deployment. | 2025-08 |
| GPT 5 Mini | gpt-5-mini | Rapid high-quality AI across text, images, and files. | 2025-10 |
| GPT 5 Nano | gpt-5-nano | Ultra-fast LLM responses for real-time AI applications. | 2025-10 |
| GPT 5.1 | gpt-5.1 | Precise code review and developer workflow assistant. | 2025-12 |
| GPT 5.2 | gpt-5.2 | Advanced reasoning with multimodal input for precise tasks. | 2025-12 |
| GPT 5.4 | gpt-5.4 | Most powerful GPT for frontier reasoning and multimodal tasks. | 2026-03 |
| GPT 5.4 Mini | gpt-5.4-mini | Fastest efficient model for coding and computer-use tasks. | 2026-03 |
| GPT 5.4 Nano | gpt-5.4-nano | Flagship-class AI for classification and extraction tasks. | 2026-03 |
| GPT 5.5 | gpt-5.5 | Frontier reasoning and coding with 1M-token context window. | 2026-05 |
| Kimi K2 Instruct 0905 | kimi-k2-instruct-0905 | Deep contextual understanding and complex code generation. | 2025-11 |
| Llama 3 8b | llama-v3-8b-instruct | Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned gener… | 2024-05 |
| Llama 3.1 70b | llama-v3p1-70b-instruct | Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned gener… | 2024-07 |
| Llama 3.1 8b | llama-v3p1-8b-instruct | Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned gener… | 2024-07 |
| Llama 4 Maverick Instruct Basic | llama4-maverick-instruct-basic | Llama 4 Maverick Instruct Basic is a 400B parameter powerhouse with 128 experts for unparalleled text and image understanding. | 2025-04 |
| Llama 4 Scout Instruct Basic | llama4-scout-instruct-basic | Unlock powerful multimodal AI with Llama 4 Scout basic, a 17 billion active parameters model offering leading text & image understanding. | 2025-04 |
| MiniMax M3 | minimax-m3 | Reason over 1M-token context for coding and agents. | 2026-07 |
| Mixtral 8x22b | mixtral-8x22b-instruct | Mistral MoE 8x22B Instruct v0.1 model with Sparse Mixture of Experts. Fine tuned for instruction following. | 2024-05 |
| Nemotron 3 Ultra | nemotron-3-ultra | 1M-token reasoning for coding agents and deep research. | 2026-07 |
| OpenAI o3 | o3 | Frontier reasoning model for complex coding, math, and science. | 2026-03 |
| OpenAI o3 Mini | o3-mini | Cost-efficient reasoning model for coding, math, and science. | 2026-03 |
| O4 Mini | o4-mini | OpenAI o4-mini enhances decision-making by processing text and images with advanced reasoning capabilities. | 2025-07 |
| QVQ Max | qvq-max | Chain-of-thought visual reasoning for math, charts, and diagrams. | 2026-03 |
| Qwen3.8 Max | qwen-3.8-max | Multimodal reasoning and agentic coding with 1M-token context. | 2026-08 |
| Qwen Flash | qwen-flash | Fastest low-cost LLM with 1M context for high-volume tasks. | 2026-03 |
| Qwen Plus | qwen-plus | Mid-tier 1M context LLM for summarization and content tasks. | 2026-03 |
| Qwen2 VL 72B Instruct | qwen2-vl-72b-instruct | Qwen2-VL-72B-Instruct is a state-of-the-art multimodal model excelling in image and video understanding, with advanced capabilities for tex… | 2025-02 |
| Qwen 3 Coder Flash | qwen3-coder-flash | High-volume code generation with 1M token context window. | 2026-03 |
| Qwen 3 Coder Plus | qwen3-coder-plus | Generates, debugs, and refactors entire codebases efficiently. | 2026-03 |
| Qwen 3 Max | qwen3-max | 1T-parameter LLM with hybrid reasoning and 262K context. | 2026-03 |
| Qwen 3 VL Flash | qwen3-vl-flash | Fast, affordable vision-language model with 262K context OCR. | 2026-03 |
| Qwen 3 VL Plus | qwen3-vl-plus | Powerful visual QA and document analysis from images. | 2026-03 |
| Qwen 3.5 Flash | qwen3.5-flash | Fast multimodal AI processing text, images, and video affordably. | 2026-03 |
| Qwen 3.5 Plus | qwen3.5-plus | Multimodal 1M context AI for image, video, and text. | 2026-03 |
| QwQ Plus | qwq-plus | Deep chain-of-thought reasoning for math, code, and logic. | 2026-03 |
Text to embed
| Model | Slug | Description | Added |
|---|---|---|---|
| Gemini Embedding 001 | gemini-embedding-001 | MTEB #1 text embeddings for RAG, search, and clustering. | 2026-05 |
| Gemini Embedding 2 | gemini-embedding-2 | Natively multimodal embeddings — text, image, audio, video and PDF mapped into one vector space, with 8 task-specific modes. | 2026-05 |
| Text Embedding 3 Large | text-embedding-3-large | Text-embedding-3-large is a robust language model by OpenAI designed for generating high-dimensional text embeddings for a wide range of na… | 2024-08 |
| Text Embedding 3 Small | text-embedding-3-small | Text-embedding-3-small is a compact and efficient model developed for generating high-quality text embeddings. These embeddings are numeric… | 2024-08 |
Transcription
| Model | Slug | Description | Added |
|---|---|---|---|
| Elevenlabs Transcript | eleven-labs-transcript | Transcribe audio to accurate text in 99 languages with speaker diarization and word-level timestamps. | 2025-03 |
| Elevenlabs Dialogue With Timing | elevenlabs-dialogue-with-timestamps | Multi-speaker dialogue with expressive timestamps included. | 2025-11 |
| Elevenlabs Forced Alignment | elevenlabs-forced-alignment | Precise audio-text synchronization with word-level timestamps. | 2025-11 |
| Elevenlabs Voice Cloning | elevenlabs-voice-clone | Hyper-realistic voice cloning from short audio samples. | 2025-11 |
| Elevenlabs Voice Design | elevenlabs-voice-design | Generate unique synthetic voices without audio samples. | 2025-11 |
| TTS Elevenlabs With Timing | tts-elevenlabs-with-timestamps | Emotionally expressive TTS with word-level timestamp output. | 2025-11 |
| Whisper Large V3 | whisper-large-v3 | Transcribe speech-to-text in 99 languages with timestamps. | 2026-08 |
Video editing
| Model | Slug | Description | Added |
|---|---|---|---|
| Bria Video Eraser | bria-erase-video | Remove unwanted objects from videos while preserving audio. | 2025-12 |
| Bria Increase Video Resolution | bria-increase-video-resolution | Transform your videos with AI-powered upscaling and seamless background removal for professional quality. | 2025-09 |
| Bria Video Background Removal 3.0 | bria-video-background-removal-3.0 | Remove video backgrounds with flicker-free, transparent alpha output. | 2026-06 |
| Esrgan Video Upscaler | esrgan-video-upscaler | ESRGAN Video Upscaler: Experience sharper, clearer 4k videos with ESRGAN. This AI-powered video upscaler boosts resolution and reduces arti… | 2024-09 |
| FLUX 3 Draft Enhance | flux-3-draft-enhance | Upscale AI video drafts to Full-HD with native audio. | 2026-08 |
| FLUX 3 Extend Video | flux-3-extend-video | Extend clips into seamless video continuations with synchronized audio. | 2026-08 |
| Gemini Omni 1.1 Video Edit | gemini-omni-1.1-video-edit | Edit videos with a text prompt, subject preserved. | 2026-08 |
| Gemini Omni 1.1 Video Extend | gemini-omni-1.1-video-extend | Extend short video clips into longer seamless scenes. | 2026-08 |
| Heygen Video Translate | heygen-video-translate | Translate videos to multiple languages with natural lip-sync. | 2025-10 |
| Kling 2.6 Pro Motion Control | kling-2.6-pro-motion-control | Transfer motion from videos to animate custom characters. | 2025-12 |
| Kling 2.6 Standard Motion Control | kling-2.6-standard-motion-control | Precise motion transfer from reference videos to characters. | 2025-12 |
| Kling O1 Video 2 Video Edit | kling-o1-video-to-video-edit | Edit any video with precise natural language commands. | 2026-01 |
| Kling O1 Video 2 Video Reference | kling-o1-video-to-video-reference | Video style transfer using reference character images. | 2026-01 |
| Kling O3 Video To Video Edit | kling-o3-video2video-edit | Text-based video editor — swap backgrounds, characters, restyle scenes. | 2026-03 |
| Kling O3 Video To Video Reference | kling-o3-video2video-reference | Swap characters and restyle videos using reference images. | 2026-03 |
| LTX Retake Video | ltx-retake-video | Precise segment-level video edits maintaining full scene continuity. | 2025-12 |
| Multi Video Merge | multi-video-merge | Merge multiple videos into a single combined output. | 2025-10 |
| OpusClip - Clips From Video | opus-clips-from-video | Turn long videos into captioned vertical shorts. | 2026-07 |
| Pixverse Lipsync | pixverse-lipsync | PixVerse Lipsync expertly synchronizes lip movements to audio for flawless video content creation. | 2025-07 |
| Runway Gen4 Aleph | runway-gen4-aleph | Runway Aleph revolutionizes video editing with intelligent automation for seamless object and environment manipulation. | 2025-08 |
| Sam V2 Video | sam-v2-video | SAM v2 Video by Meta AI, allows promptable segmentation of objects in videos. | 2024-08 |
| Sam3 Video | sam3-video | Real-time video segmentation and multi-object tracking. | 2025-11 |
| Sonilo Video to Video | sonilo-video-to-video | Add frame-synced AI music and sound effects to video. | 2026-08 |
| Sync.so Lipsync 2 Pro | sync.so-lipsync-2-pro | Lipsync-2-Pro seamlessly synchronizes lips in videos for instant, high-quality multilingual content creation. | 2025-09 |
| Sync.so React 1 | sync.so-react-1 | Edit video actors' emotions with realistic re-expression. | 2025-12 |
| Topaz Labs Video Upscale | topaz-video-upscale | Topaz Video AI upscales, enhances, denoises, stabilizes, and increases frame rates in video footage, transforming low-quality or standard-d… | 2025-04 |
| VEED Lipsync v2 | veed-2-lipsync | Dub talking-head videos with emotion-matched lip-sync. | 2026-07 |
| VEED Lipsync | veed-lipsync | Re-syncs the lips of any talking-head video to a new speech audio track for realistic dubbing and localization. | 2026-07 |
| VEED Subtitles | veed-subtitles | Automatically transcribes and burns styled, translated subtitles into any video with 30 presets and a single API call. | 2026-07 |
| VEED Video Background Removal | veed-video-background-removal | Remove any video's background with no green screen, or cleanly key chroma footage, using AI matting. | 2026-07 |
| Video Audio Merge | video-audio-merge | Effortlessly merge audio and video with our intuitive Video Audio Merge model. Create stunning multimedia content with precise timing, fade… | 2024-10 |
| Video Captioner | video-captioner | With Video Captioner create accurate, customizable subtitles for your videos effortlessly. | 2024-10 |
| Video Concatenate | video-concatenate | Merge videos with custom layouts, spacing, and audio. | 2025-12 |
| Video Editor Agent | video-editor-agent | General-purpose AI media agent: describe the transformation in natural language and it runs ffmpeg in a sandbox (transcode, resize, extract… | 2026-07 |
| Video Loop | video-loop | Effortlessly loop videos for engaging social media & storytelling with our Video Loop. | 2025-03 |
| Video Slicer | video-slicer | Video Slicer | 2025-04 |
| Video Split | video-split | Utility node: Video Split. 1->N split; returns videos[] array | 2026-07 |
| Wan 2.7 Video Editing | wan2.7-videoedit | Edit existing videos precisely using natural language text instructions. | 2026-04 |
Video generation
| Model | Slug | Description | Added |
|---|---|---|---|
| Cog Video X 5B | cog-video-5b-t2v | CogVideo is a groundbreaking AI model that turns text into high-quality videos. Create realistic scenes, animations, and more with ease. Id… | 2024-09 |
| FLUX 3 Text to Video | flux-3-text-to-video | Cinematic text-to-video with native lip-synced audio, up to 20s. | 2026-08 |
| Gemini Omni 1.1 | gemini-omni-1.1 | Text-to-video with synchronized native audio, up to 4K. | 2026-08 |
| Gemini Omni Flash | gemini-omni-flash | Text-to-video and image-to-video with synchronized native audio. | 2026-06 |
| Grok Imagine Video | grok-imagine-video | Text-to-video and image-to-video with native synchronized audio. | 2026-06 |
| Grok Imagine Video 1.5 Text to Video | grok-imagine-video-1.5-text-to-video | Text-to-video clips up to 1080p with native synchronized audio. | 2026-08 |
| HeyGen Avatar V | heygen-avatar-v | Studio-quality talking-avatar videos from text or audio. | 2026-05 |
| Kling AI 1.6 Text to Video | kling-1.6-text2video | Kling AI 1.6 Text-to-Video is a cutting-edge AI tool that transforms text into stunning, lifelike videos. Create professional-quality conte… | 2025-01 |
| Kling 3.0 Pro Text-to-Video | kling-3-pro-text2video | Cinematic 1080p videos with realistic audio from text. | 2026-02 |
| Kling 3.0 Standard Text-to-Video | kling-3-standard-text2video | Stunning 1080p cinematic videos from simple text prompts. | 2026-02 |
| Kling O3 Text-to-Video | kling-o3-text2video | 15-second cinematic AI videos with native audio. | 2026-03 |
| Kling AI Text to Video | kling-text2video | Kling AI Text-to-Video is a cutting-edge AI tool that transforms text into stunning, lifelike videos. Create professional-quality content e… | 2024-10 |
| LTX-2-19B I2V | ltx-2-19b-i2v | Synchronized 4K audio-video generation from images, fast. | 2026-01 |
| LTX-2-19B T2V | ltx-2-19b-t2v | Synchronized video and audio from text, multiple input types. | 2026-01 |
| Minimax AI Director | minimax-ai-director | Minimax video-01-director: Create high-quality videos with control camera movements precisely using text prompts. | 2025-02 |
| MiniMax Hailuo H3 Text to Video | minimax-h3-text-to-video | Text-to-video: cinematic 2K clips with native audio. | 2026-07 |
| Pixverse Text to Video | pixverse-text2video | Effortlessly create captivating videos from text with Pixverse text to video AI! Customize style, duration, and more. | 2025-04 |
| Timeline | timeline | Utility node: Timeline. declarative multi-track video compositor | 2026-07 |
| VEED Avatars | veed-avatars | Generate UGC-style talking avatar videos from text or audio using 28 stock presenters with realistic lip-sync. | 2026-07 |
| Google Veo 2 | veo-2 | Create stunning, realistic videos with Veo 2, Google's state-of-the-art AI video generation model. Experience enhanced quality & cinematic… | 2025-03 |
| Google Veo 3 | veo-3 | Veo 3 revolutionizes video creation with advanced text-to-video generation and realistic audio synthesis for cinematic content. | 2025-06 |
| Veo 3 Fast | veo-3-fast | Veo 3 Fast rapidly creates high-quality, 8-second videos with synchronized audio for diverse content needs. | 2025-07 |
| Wan 2.2 Text to Video Fast | wan-2.2-t2v-fast | Wan2.2 transforms text and images into high-quality video clips with cinematic flair. | 2025-08 |
| Wan 2.5 Text to Video | wan-2.5-t2v | Wan2.5-Preview generates synchronized multimedia content, merging text, image, video, and audio seamlessly. | 2025-09 |
| Wan_2.1 Text to Video | wan2.1-t2v | Create visually impressive and feature varied, lifelike motion videos with Wan2.1 using text prompts. | 2025-03 |
| Wan 2.7 Text to Video | wan2.7-t2v | 1080P cinematic videos with audio sync and multi-shot control. | 2026-04 |
Video to audio
| Model | Slug | Description | Added |
|---|---|---|---|
| Sonilo Video to Audio | sonilo-video-to-audio | Generate video-synced music and sound effects from footage. | 2026-08 |
Video to image
| Model | Slug | Description | Added |
|---|---|---|---|
| Start & End Frame Extractor | start-end-frame-extractor | Extract first and last frames from any video. | 2025-10 |
Video to text
| Model | Slug | Description | Added |
|---|---|---|---|
| HeyGen Avatar V — Create Avatar | heygen-avatar-v-create | Train a Digital Twin avatar from reference video. | 2026-05 |
Voice
| Model | Slug | Description | Added |
|---|---|---|---|
| Kling Create Voice | kling-create-voice | Clone any voice from a single audio sample. | 2026-02 |