SegmindSegmind / Docs

Model Catalog

Every model currently on the Segmind AI Gateway — 670 models with their API slugs, grouped by task. Regenerated automatically from the live catalog.

Every model on the AI Gateway, with the slug you pass to https://api.segmind.com/v1/{'{slug}'} (or /v2 for async) and to segmind.run() in the Python SDK. 670 models, regenerated automatically from the live catalog — if a model is listed here, it is callable today.

Model-specific parameters live on each model's page at segmind.com/models, or in the parameters_schema field of the catalog API.

Recently added

ModelSlugDescriptionAdded
Gemini Omni 1.1gemini-omni-1.1Text-to-video with synchronized native audio, up to 4K.2026-08
Gemini Omni 1.1 Video Extendgemini-omni-1.1-video-extendExtend short video clips into longer seamless scenes.2026-08
Gemini Omni 1.1 Video Editgemini-omni-1.1-video-editEdit videos with a text prompt, subject preserved.2026-08
Lyria 3 Prolyria-3-proFull-length text-to-music songs with vocals and lyrics.2026-08
Lyria 3lyria-3Generate 30-second songs with vocals from text or images.2026-08
Gemini 3.7 Flashgemini-3.7-flashFast multimodal LLM for coding, agents, and long-document analysis.2026-08
Kokoro 82Mkokoro-82mText-to-speech with 54 multilingual voices.2026-08
Whisper Large V3whisper-large-v3Transcribe speech-to-text in 99 languages with timestamps.2026-08
Wan 3.0 Videowan3.0-videoGenerate 30-second 1080p video with native audio.2026-08
Wan 2.6 Image to Video Flashwan2.6-i2v-flashAnimate photos into 15-second 1080p video with native audio.2026-08
Grok Imagine Image 2grok-imagine-image-2Text-to-image and image editing with crisp, legible text.2026-08
LTX 2.5 Proltx-2.5-proGenerate 1080p video with native audio and multi-shot scenes.2026-08
LTX 2.5 Fastltx-2.5-fastText-to-video and image-to-video with native audio, up to 4K.2026-08
Qwen Image 3.0qwen-image-3Generate and edit legible in-image text, up to 2K.2026-08
Seedream 5.0 Pro Layer Decompositionseedream-5-pro-layer-decompositionSplit any image into editable transparent PNG layers.2026-08
Bria Extract Objectbria-extract-objectExtract any named object into a transparent PNG cutout.2026-08
Seedance 2.5seedance-2.5Generate cinematic multi-shot AI videos up to 30 seconds with synchronized native audio from text, images, or references.2026-08
FLUX 3 Draft Enhanceflux-3-draft-enhanceUpscale AI video drafts to Full-HD with native audio.2026-08
FLUX 3 Extend Videoflux-3-extend-videoExtend clips into seamless video continuations with synchronized audio.2026-08
FLUX 3 Image to Videoflux-3-image-to-videoAnimate images into 20-second clips with synchronized native audio.2026-08
FLUX 3 Text to Videoflux-3-text-to-videoCinematic text-to-video with native lip-synced audio, up to 20s.2026-08
Qwen3.8 Maxqwen-3.8-maxMultimodal reasoning and agentic coding with 1M-token context.2026-08
Grok Imagine Video 1.5 Reference to Videogrok-imagine-video-1.5-reference-to-videoCharacter-consistent video from up to 7 reference images.2026-08
Grok Imagine Video 1.5 Image to Videogrok-imagine-video-1.5-image-to-videoAnimate a still image into 1080p video with synced audio.2026-08
Grok Imagine Video 1.5 Text to Videogrok-imagine-video-1.5-text-to-videoText-to-video clips up to 1080p with native synchronized audio.2026-08

this is first creation

ModelSlugDescriptionAdded
DemoImage-shakilThis is Demo creation2024-10

Audio & speech synthesis

ModelSlugDescriptionAdded
Ace Step Musicace-step-musicACE-Step generates high-quality music rapidly, enhancing the creative process for developers and artists worldwide.2025-05
Chatterbox TTSchatterbox-ttsChatterbox transforms text into rich, natural speech with adjustable emotional expressiveness for diverse applications.2025-07
Chatterbox Turbo TTSchatterbox-turbo-ttsUltra-fast, human-quality TTS with emotional expression.2025-12
Dia (Text to Speech)diaDia by Nari Labs is an advanced open-weights TTS model that brings scripts to life with natural speech, emotions, and nonverbal cues. Easil…2025-04
ElevenLabs DubbingdubbingInstantly dubs audio and video into 29 languages while preserving each speaker's original voice.2024-07
Elevenlabs Dialogueelevenlabs-dialogueImmersive, emotionally expressive multi-speaker audio dialogue.2025-11
Gemini TTS 2.5 Flashgemini-2.5-flash-ttsFast, lifelike text-to-speech with expressive emotional tones.2025-12
Gemini TTS 2.5 Progemini-2.5-pro-ttsHuman-like speech synthesis with rich expressive emotional depth.2025-12
Gemini 3.1 Flash TTSgemini-3.1-flash-ttsExpressive, controllable TTS with 70+ language support.2026-05
Grok Text-to-Speechgrok-ttsConvert text to speech in 20 languages with five voices.2026-06
Kokoro 82Mkokoro-82mText-to-speech with 54 multilingual voices.2026-08
Lyria 2lyria-2Lyria 2 by Google DeepMind is an advanced model that generates high-fidelity 48kHz stereo instrumental music from text prompts or lyrics, o…2025-05
Lyria 3lyria-3Generate 30-second songs with vocals from text or images.2026-08
Lyria 3 Prolyria-3-proFull-length text-to-music songs with vocals and lyrics.2026-08
Meta MusicGen Mediummeta-musicgen-mediumMusicGen: Transform text into music with AI. Create unique, high-quality audio from simple descriptions. Experience the future of music gen…2024-10
MyShell Text To Speechmyshell-ttsMyShell's Voice Cloning and Text to Speech - Transform your audio content with realistic, personalized voices. Experience high-quality, eff…2024-10
OpenvoiceopenvoiceOpenVoice is a versatile voice cloning model that supports multiple languages and offers precise tone replication, flexible style control,…2024-09
3B Orpheus TTS (0.1)orpheus-3b-0.1Orpheus TTS is an open-source text-to-speech (TTS) system powered by the Llama 3B language model, designed for high-quality and customizabl…2025-03
Sam Audio Largesam-audio-largeIsolate any described sound from mixed audio tracks.2026-02
Seed Audio 1.0seed-audio-1.0Generate full audio scenes: dialogue, music, effects, voice cloning.2026-06
BytePlus Seed Speech TTSseed-speech-ttsNatural multilingual text-to-speech and voiceovers from text.2026-06
Sonilo Text to Audiosonilo-text-to-audioCommercial-safe music and sound effects from text prompts.2026-08
Elevenlabs Sound Generationsound-generationEleven Labs' Sound Generation API provides a robust development tool for programmatically generating audio content using artificial intelli…2024-06
Elevenlabs Text To Speechtts-eleven-labsElevenLabs TTS transforms text into captivating, human-like speech for diverse applications.2024-06
VeenaMax TTSveena-max-ttsVeenaMAX transforms text into expressive, real-time speech across multiple Indian languages for seamless communication.2025-09
Veena TTSveena-ttsVeena transforms text into high-fidelity, expressive speech in Hindi and English for real-time applications.2025-07

Audio to audio

ModelSlugDescriptionAdded
Elevenlabs Audio Isolationelevenlabs-audio-isolationExtract clear speech from noisy audio and video.2025-11
Elevenlabs Speech To Speechsts-eleven-labsEleven Labs Speech-to-Speech offers AI-powered voice conversion for content creators, media professionals, and anyone seeking to modify or…2024-06

Image editing & transformation

ModelSlugDescriptionAdded
AI Product Photo Editorai-product-photo-editorAI Product Photo Editor leverages advanced image-based ML techniques to generate high-quality product visuals using text prompts, product i…2024-07
AI Product Photographyai-product-photographyElevate your product imagery with our AI-powered photography model. Create stunning, professional-quality photos that boost engagement and…2024-11
IDM + Faceswap (updated)alle-v2IDM+Faceswap2024-07
Aura Flowaura-flowLargest completely open sourced flow-based generation model that is capable of text-to-image generation2024-07
Automatic Mask Generatorautomatic-mask-generatorAutomatic Mask Generator is a powerful tool that automates the creation of precise masks for inpainting2024-06
Profile Photo Style Transferbecome-imageTurn any image of a face into artwork using Stable Diffusion Controlnet and IPAdapter2024-06
Background Removalbg-removalThis model removes the background image from any image2023-09
Background Removal V2bg-removal-v2This model removes the background image from any image2024-03
Bria Blur Backgroundbria-blur-backgroundBria AI Image Editing API v2 enables precise and context-aware image manipulation for stunning visual outcomes.2025-08
Bria Enhance Imagebria-enhance-imageBria AI creates precise, high-quality image enhancements and manipulations for diverse creative applications.2025-08
Bria Erase Foregroundbria-erase-foregroundSeamlessly removes foreground subjects and regenerates backgrounds for flawless image editing.2025-08
Bria Eraserbria-eraserAI object removal with seamless context-aware inpainting.2025-08
Bria Expand Imagebria-expand-imageBria Expand enables precise image manipulation and enhancement with generative AI, trained exclusively on licensed data for safe, risk-free…2025-08
Bria Extract Objectbria-extract-objectExtract any named object into a transparent PNG cutout.2026-08
Bria FIBO 1.5 Image Editbria-fibo-image-editEdit images with structured JSON, masks, and multi-image references.2026-01
Bria Generative Fillbria-gen-fillBria AI enables precise generative image editing for seamless creative enhancements and transformations.2025-08
Bria Increase Resolutionbria-increase-resolutionSeamlessly upscale and manipulate images while preserving the highest fidelity and safety standards.2025-08
Lifestyle Product Shot by Imagebria-lifestyle-shot-by-imageTransforms ordinary product images into stunning, marketing-ready visuals for eCommerce success.2025-08
Bria Lifestyle Product Shot by Textbria-lifestyle-shot-by-textTransform isolated product images into dynamic lifestyle scenes with AI-driven contextual realism.2025-08
Bria Product Cutoutbria-product-cutoutAutomates precise product cutouts and background removal for professional eCommerce imagery at scale.2025-08
Bria Product Packshotbria-product-packshotTransform product photos into professional, market-ready images with intelligent enhancements and background removal.2025-08
Bria Product Shadowbria-product-shadowBria Product Shadow enhances product images with realistic shadows for professional eCommerce presentations.2025-08
Bria RMBG 2.0bria-remove-backgroundEffortlessly extract backgrounds with unmatched precision, powered by models trained exclusively on licensed data for safe and risk-free co…2025-08
Bria Generate Backgroundbria-replace-backgroundTransform images through advanced background editing and generative content creation for diverse applications.2025-08
Caricature Stylecaricature-styleTransform everyday photos into lively, whimsical caricature illustrations that highlight individual features with playful exaggeration.2025-05
Clarity Upscalerclarity-upscalerHigh resolution creative image Upscaler and Enhancer. A free Magnific alternative.2024-06
ClarityAI Creative Upscalerclarityai-creative-upscalerCreative image upscaling with fine detail enhancement.2025-10
ClarityAI Crystal Upscalerclarityai-crystal-upscalerUpscale images up to 200x with enhanced detail and vibrancy.2025-10
ClarityAI Flux Upscalerclarityai-flux-upscalerTransform low-resolution images into stunning high-quality visuals.2025-10
CodeformercodeformerCodeFormer is a robust face restoration algorithm for old photos or AI-generated faces.2023-09
Consistent Characterconsistent-characterCreate images of a given character in different poses2024-06
Consistent Character With Poseconsistent-character-with-poseCreate images of a given character in different poses2024-09
ESRGANesrganERGAN is an Image Super-Resolution (upscaler) model that enhances images with stunning, high-quality upscaling while preserving the exact c…2023-09
Expression Editorexpression-editorExpression Editor uses reference images to accurately generate new images with desired expressions. Perfect for digital art, memes, and mar…2024-09
Face Detailerface-detailerRestore characters' faces to their original glory with Face Detailer. Enhance facial details, eliminate distortion, and upscale images for…2024-10
face-to-manyface-to-manyTurn a face into 3D, emoji, pixel art, video game, claymation or toy2024-05
face-to-stickerface-to-stickerTurn a face into a sticker2024-05
Segmind FaceSwap Comic v1faceswap-comicFaceSwap Comic v1 is an AI-powered face swapping model designed to blend real faces into illustrated or cartoon-style images while preservi…2025-05
Faceswap V3 Moonfrogfaceswap-moonfrog-v3Take a picture/gif and replace the face in it with a face of your choice. You only need one image of the desired face. No dataset, no train…2024-07
Faceswap V2faceswap-v2Take a picture/gif and replace the face in it with a face of your choice. You only need one image of the desired face. No dataset, no train…2024-04
Faceswap V3faceswap-v3Face Swap V3 is a cutting-edge tool that empowers you to seamlessly swap faces in images. With customizable features and advanced technolog…2024-10
Faceswap V3 Multifaceswapfaceswap-v3-multifaceswapFaceswap V3 Multifaceswap enables realistic face swapping in images, preserving lighting and expressions for professional results.2025-07
Segmind Faceswap v4faceswap-v4Segmind FaceSwap v4 enables fast and precise face or head swapping between images with customizable options for style, output format, and i…2025-03
Segmind Faceswap v5faceswap-v5Ultra-fast face and head swapping in images.2026-01
Flux 2 Flexflux-2-flexConsistent-style photorealistic images using reference inputs.2025-11
Flux-2 Klein-4bflux-2-klein-4bSub-second photorealistic image generation and editing.2026-01
Flux-2 Klein-9bflux-2-klein-9bUltra-fast photorealistic image generation on consumer GPUs.2026-01
Flux 2 Maxflux-2-maxPhotorealistic images with maximum consistency and fine detail.2025-12
Flux 2 Proflux-2-proHigh-quality photorealistic images with cross-output consistency.2025-11
Flux Canny Devflux-canny-devOpen-weight edge-guided image generation. Control structure and composition using Canny edge detection.2024-11
Flux Canny Proflux-canny-proProfessional edge-guided image generation. Control structure and composition using Canny edge detection2024-11
Flux Controlnetsflux-controlnetFlux ControlNets is a collection of models that gives you precise control over image generation. By integrating ControlNet with Flux.1, the…2024-09
Flux Depth Devflux-depth-devOpen-weight depth-aware image generation. Edit images while preserving spatial relationships.2024-11
Flux Depth Proflux-depth-proProfessional depth-aware image generation. Edit images while preserving spatial relationships.2024-11
Flux Fill Devflux-fill-devOpen-weight inpainting model for editing and extending images. Guidance-distilled from FLUX.1 Fill Dev2024-11
Flux Fill Proflux-fill-proProfessional inpainting and outpainting model with state-of-the-art performance. Edit or extend images with natural, seamless results.2024-11
Flux.1 Image To Imageflux-img2imgFlux Image-To-Image model by Black Forest Labs is an advanced deep learning tool designed for transforming images based on specific textual…2024-08
Flux Inpaintflux-inpaintFlux Inpainting is a powerful image editing tool designed to effortlessly edit and enhance your images. It's perfect for tasks like removin…2024-09
Flux Ipadapterflux-ipadapterFlux IP Adapter is a cutting-edge AI model that lets you to create stunning, customized images. With its advanced style adaptation capabili…2024-09
FLUX.1 Kontext [dev]flux-kontext-devFLUX.1 Kontext [dev] creates coherent and editable images by integrating text and visual cues for iterative design.2025-06
Flux Kontext Maxflux-kontext-maxFLUX.1 Kontext [max] transforms textual descriptions into stunning, high-fidelity images with seamless typography integration.2025-05
Flux Kontext Proflux-kontext-proFLUX.1 Kontext Pro transforms text prompts into high-quality, customized images with remarkable efficiency and precision.2025-05
Flux Krea Devflux-krea-devFLUX.1 Krea generates stunning, photorealistic images with fine-tuned aesthetic control for diverse creative applications.2025-08
Flux Pulidflux-pulidFlux PuLID: Customize AI-generated images with your unique identity. Seamlessly integrate faces into text-to-image models for realistic and…2024-09
Flux Redux Devflux-redux-devOpen-weight image variation model. Create new versions while preserving key elements of your original.2024-11
Flux Redux Schnellflux-redux-schnellFast, efficient image variation model for rapid iteration and experimentation.2024-11
Fooocus Outpaintingfocus-outpaintFooocus Outpainting transforms ordinary images into extraordinary works of art by seamlessly expanding their boundaries.2024-02
FooocusfooocusFooocus enables high-quality image generation effortlessly, combining the best of Stable Diffusion and Midjourney.2024-06
GPT Image 1 Editgpt-image-1-editEdit and compose images using natural language with GPT Image 1 Edit, OpenAI’s powerful inpainting and multi-reference editing model. Perfe…2025-04
GPT Image 1 Edit Minigpt-image-1-edit-miniAffordable text-driven image generation and editing.2025-10
GPT Image 1.5 Editgpt-image-1.5-editPrecise image editing via natural language instructions.2025-12
Grok Imagine Imagegrok-imagine-imageText-to-image generation and editing, up to 2K resolution.2026-06
Grok Imagine Image 2grok-imagine-image-2Text-to-image and image editing with crisp, legible text.2026-08
HeyGen Generate Lookheygen-generate-lookChange avatar outfits and backgrounds while keeping the same face.2026-07
HiDream-I1 (Fast)hidream-l1-fastHiDream-I1 is a next-generation, open-source image generative foundation model designed for text-to-image synthesis, especially for renderi…2025-04
Higgsfield Soul 2.0higgsfield-soul-2Generate fashion-editorial photorealistic photos from text or reference.2026-07
Higgsfield Text 2 Image Soulhiggsfield-text2image-soulSOUL AI transforms text into stunning, customizable visuals with unparalleled style control and precision.2025-09
HyperSwap Image Faceswap by FaceFusion Labshyperswap-image-faceswap-by-facefusion-labsHigh-quality face swapping built for real production workflows.2026-03
Relightingic-lightPrompts to auto-magically relight your images.2024-06
Icon Overlayicon-overlayicon overlay2024-11
Ideogram 2a Image to Imageideogram-2a-img-2-imgIdeogram Image to Image: Transform your images with ease! Enhance, modify, or create entirely new visuals using advanced AI. Perfect for ar…2025-03
Ideogram 3 Reframeideogram-3-reframeIdeogram 3.0's Reframe effortlessly adapts images to diverse formats, enhancing visual content creation for any platform.2025-05
Ideogram 3 Remixideogram-3-remixIdeogram 3 Remix enables versatile image transformation, enhancing creativity through customizable design iterations.2025-05
Ideogram 3 Replace Backgroundideogram-3-replace-backgroundEffortlessly replace backgrounds in images, enhancing visual storytelling and creativity with precision and speed.2025-05
Ideogram Characterideogram-characterAchieve perfect character consistency across multiple generations from a single reference image.2025-08
Ideogram Image To Imageideogram-img-2-imgIdeogram Image to Image: Transform your images with ease! Enhance, modify, or create entirely new visuals using advanced AI. Perfect for ar…2024-12
Ideogram Reframeideogram-reframeTransform your images with Ideogram Reframe! Easily reframe square images to your chosen resolution.2025-03
Ideogram Turbo Image To Imageideogram-turbo-img-2-imgTransform images instantly with Ideogram Turbo Image to Image! Fast AI for quick edits & creative remixes.2025-03
Ideogram V4 Remixideogram-v4-remixRestyle any image into posters with legible in-image text.2026-07
IDM VTONidm-vtonBest-in-class clothing virtual try on in the wild2024-06
illusion-diffusion-hqillusion-diffusion-hqMonster Labs QrCode ControlNet on top of SD Realistic Vision v5.12024-06
Minimax-image-01image-01Generate high-fidelity images from text with precise control & stunning quality with Minimax Image-01.2025-03
Infinite Youinfinite-youInfiniteYou generates high-fidelity portraits preserving identity while aligning with creative text prompts.2025-07
Inpaint Mask Makerinpaint-mask-makerReal-Time Open-Vocabulary Object Detection2024-06
Insta Depthinsta-depthInstantID aims to generate customized images with various poses or styles from only a single reference ID image while ensuring high fidelity2024-04
InstantIDinstantidInstantID aims to generate customized images with various poses or styles from only a single reference ID image while ensuring high fidelity2024-02
IP-adapter Depth XLip-sdxl-depthIP Adapter Depth XL is built on the SDXL framework. This model integrates the IP Adapter and Depth preprocessor to offer unparalleled contr…2023-11
Kling V3 Image 2 Imagekling-3-image2imageTransform images into photorealistic, production-ready visuals.2026-03
Kling O1kling-o1Text-to-video creation with precise AI-driven motion control.2026-01
KolorskolorsKolors is a cutting-edge text-to-image model that bridges language and visual art. Transform your textual ideas into photorealistic images…2024-07
Luma Uni-1luma-uni-1Reasoning-first text-to-image and natural-language image editing.2026-06
Luma Uni-1 Maxluma-uni-1-maxGenerate and edit images from plain-text instructions.2026-06
Magic Erasermagic-eraserLaMA Object Removal- AI Magic Eraser2024-06
material-transfermaterial-transferTransfer a material from an image to a subject2024-05
Multi Image Kontext Maxmulti-image-kontext-maxFLUX.1 Kontext [max] creates stunning, photorealistic images from text prompts and input images seamlessly.2025-06
Multi Image Kontext Promulti-image-kontext-proTransform text into stunning, professional-grade images with precise editing capabilities.2025-07
Nano Banana 2nano-banana-2Fast photorealistic images — ideal for marketing and ads.2026-02
Nano Banana 2 Litenano-banana-2-liteGenerate and edit 1K images in about four seconds.2026-07
Nano Banana Pronano-banana-proHigh-fidelity images with accurate multilingual text rendering.2025-11
Nomos Image Upscaler 4knomos-upscalerThis upscaling model is ideal for enhancing amateur to professional photos, excelling with subjects like cats, hair, and party scenes. It h…2025-05
Omini ControlominicontrolOminiControl is an innovative framework that optimizes Diffusion Transformer models for versatile image generation tasks.2024-12
Omni Zeroomni-zeroOmni-Zero: A diffusion pipeline for zero-shot stylized portrait creation.2024-06
Pruna P Image Editp-image-editMulti-image editing with AI-guided precision and control.2025-11
Pruna P Image Try-Onp-image-try-onDress photos in multiple garments with photorealistic virtual try-on.2026-06
PuLIDpulid-baseNovel tuning-free ID customization method for text-to-image generation.2024-06
Qwen Image Editqwen-image-editTransform images effortlessly through semantic context and pixel-perfect appearance changes.2025-08
Qwen Image Edit Fastqwen-image-edit-fastQwen-Image-Edit enables precise bilingual image editing for seamless localization and professional content creation.2025-08
Qwen Image Edit Plusqwen-image-edit-plusMulti-image editing with precise text-guided transformations.2025-10
Qwen Image Edit Plus Add People Loraqwen-image-edit-plus-add-peopleGenerate realistic multi-character scenes with natural interactions.2025-11
Qwen Image Edit Plus Blend Itqwen-image-edit-plus-blend-itProduct placement into backgrounds with precise lighting match.2025-11
Qwen Image Edit Plus Eigen Bananaqwen-image-edit-plus-eigen-bananaPrecise text-guided image transformation and creative editing.2025-11
Qwen Image Edit Plus Eraserqwen-image-edit-plus-eraserRemove unwanted objects while preserving realistic backgrounds.2025-11
Qwen Image Edit Plus Face To Portraitqwen-image-edit-plus-face-to-portraitCropped face into full identity-preserving portrait photo.2025-11
Qwen Image Edit Plus Group Photoqwen-image-edit-plus-group-photoMerge individual portraits into realistic group photos.2025-11
Qwen Image Edit Plus Multi Loraqwen-image-edit-plus-multi-loraMulti-image editing with superior identity and style control.2025-11
Qwen Image Edit Plus Multiple Anglesqwen-image-edit-plus-multiple-angleTransform image perspective with natural language prompts.2025-11
Qwen Image Edit Plus Next Sceneqwen-image-edit-plus-next-sceneCreate cinematic sequences with seamless visual continuity.2025-11
Qwen Image Edit Plus Product Photographyqwen-image-edit-plus-product-photographyTransform white-background products into immersive lifestyle scenes.2025-11
Qwen Image Edit Plus Relightqwen-image-edit-plus-relightAdvanced image relighting using natural language prompts.2025-11
Qwen Image Edit Plus Remove Lightingqwen-image-edit-plus-remove-lightingRemove artificial lighting effects and restore natural tones.2025-11
Qwen Image Edit Plus Texture Applyqwen-image-edit-plus-texture-applyApply precise textures to images using natural language.2025-11
Qwen Image Edit Plus Texture Extractqwen-image-edit-plus-texture-extractExtract seamless, tileable textures from photographs.2025-11
Runway Gen 4 Imagerunway-gen4-imageRunway's Gen-4 Image API enables precise, multimodal image generation for innovative creative and technical applications.2025-05
Segment Anything Modelsam-img2imgThe Segment Anything Model (SAM) produces high quality object masks from input prompts such as points or boxes, and it can be used to gener…2023-09
Sam V2 Imagesam-v2-imageSAM v2, the next-gen segmentation model from Meta AI, revolutionizes computer vision. Building on SAM's success, it excels at accurately se…2024-08
Sam3 Imagesam3-imagePrecise object segmentation and tracking in images.2025-11
ControlNet Cannysd1.5-controlnet-cannyThis model corresponds to the ControlNet conditioned on Canny edges.2023-09
ControlNet Depthsd1.5-controlnet-depthThis model corresponds to the ControlNet conditioned on Depth estimation.2023-09
ControlNet Openposesd1.5-controlnet-openposeThis model corresponds to the ControlNet conditioned on Human Pose Estimation.2023-09
ControlNet Scribblesd1.5-controlnet-scribbleThis model corresponds to the ControlNet conditioned on Scribble images.2023-09
ControlNet Soft Edgesd1.5-controlnet-softedgeThis model corresponds to the ControlNet conditioned on Soft Edge.2023-09
Stable Diffusion img2imgsd1.5-img2imgThis model uses diffusion-denoising mechanism as first proposed by SDEdit, Stable Diffusion is used for text-guided image-to-image translat…2023-09
SD Outpaintingsd1.5-outpaintStable Diffusion Outpainting can extend any image in any direction2023-09
Faceswapsd2.1-faceswapperTake a picture/gif and replace the face in it with a face of your choice. You only need one image of the desired face. No dataset, no train…2023-09
SD3 Medium Canny Controlnetsd3-med-cannyStable Diffusion 3 (SD3) Medium Canny ControlNet uses Canny edge detection to provide fine-grained control over the generated outputs.2024-07
SD3 Medium Pose Controlnetsd3-med-poseStable Diffusion 3 (SD3) Pose ControlNet is a large generative image model tailored for generating images based on text prompts while using…2024-07
SD3 Medium Tile Controlnetsd3-med-tileSD3 Medium Tile ControlNet is a large generative image model designed for generating detailed images based on textual prompts and tile-base…2024-07
SDXL Controlnetsdxl-controlnetSDXL ControlNet gives unprecedented control over text-to-image generation. SDXL ControlNet models Introduces the concept of conditioning in…2024-07
SDXL Img2Imgsdxl-img2imgSDXL Img2Img is used for text-guided image-to-image translation. This model uses the weights from Stable Diffusion to generate new images f…2024-07
SDXL-Openposesdxl-openposeThis model leverages SDXL to generate the images with ControlNet conditioned on Human Pose Estimation.2023-11
Seedream 4.0 (4k)seedream-4Seedream 4.0 generates high-resolution, professional-grade visuals with superior text rendering for impactful design.2025-09
Seedream 4.5seedream-4.5Photorealistic image generation with precise text understanding.2025-12
Seedream 5.0 Pro Layer Decompositionseedream-5-pro-layer-decompositionSplit any image into editable transparent PNG layers.2026-08
Seedream 5.0 Lite: Image-to-Imageseedream-v5-lite-image-to-imageTransform images intelligently with detailed text prompts.2026-02
Segmind SegSwap v0.1seg-swapSwap Objects Instantly. The Segmind SegSwap v0.1 model enables dynamic and precise image editing by allowing users to remove, replace, or a…2025-03
Segmind SegFit v1.1segfit-v1.1Segmind's Fashion and Immersive Try-on model. SegFIT offers effortless AI virtual try-on from just a product image. No models needed! Boost…2025-04
Segmind SegFit v1.2segfit-v1.2SegFit v1.2 creates hyper-realistic virtual try-on images, transforming fashion retail engagement and conversion rates.2025-06
Segmind SegFit v1.3segfit-v1.3SegFit v1.3 enables hyper-realistic virtual try-ons, enhancing online fashion retail experiences without physical photoshoots.2025-07
Segmind Relightingsegmind-relightingPrompts to auto-magically relight your images.2025-03
Segmind Relighting V2segmind-relighting-v2Transform images with customizable, photorealistic lighting for unparalleled visual creativity and authenticity.2025-05
Segmind SceneCraft v0.1segmind-scenecraft-v01SceneCraft transforms plain or existing product images into visually rich, photorealistic scenes. Whether starting from a white background…2025-04
Skin Contrast Upscalerskin-contrast-upscalerEnhances skin detail in images while preserving background quality for professional photography and art.2025-05
Smart Banner Resizersmart-banner-resizerRecompose one image into multiple ad and banner sizes.2026-05
SSD-Depthssd-depthThis model leverages SSD-1B to generate the images with ControlNet conditioned on Depth Estimation2023-11
SSD Img2Imgssd-img2imgThis model uses SSD-1B to generate images by passing a text prompt and an initial image to condition the generation2023-11
Story DiffusionstorydiffusionStory Diffusion turns your written narratives into stunning image sequences.2024-07
IPAdapter Style Transferstyle-transferStyle & Composition Transfer with Stable Diffusion IP Adapter2024-06
Image SuperimposesuperimposeSuperimpose model lets you to create captivating visuals by seamlessly overlaying one image on top of another. It streamlines your image la…2024-07
Image Superimpose V2superimpose-v2Superimpose V2 elevates image editing! Seamlessly layer images with background removal, precise positioning, and flexible resizing options.…2024-07
Supir Photo-Realistic Image RestorationsupirSUPIR restores and enhances images to stunning, photo-realistic quality with advanced AI techniques.2025-04
Text Overlaytext-overlayElevate your visuals withText Overlay Model. Easily add customized text to any image, perfect for social media, marketing, and blogs. Enjoy…2024-09
Topaz Labs Image Upscaletopaz-image-upscaleTopaz Labs image upscale is an industry-leading AI photo upscaler designed to increase the resolution of photos while preserving and enhanc…2025-04
Transparent Background Makertransparent-background-makerTransform your images with Transparent Background Maker. Quickly remove backgrounds using AI technology, supporting PNG and JPG formats. Id…2024-11
Image Maskutility-image-maskBuild / refine binary masks from JSON-described shapes (rect/polygon/ellipse). Optional dilate/erode/feather/invert; multi-shape merge.2026-04
Image Transform Pipelineutility-image-transformApply an ordered pipeline of resize / crop / rotate / flip in a single call. Replaces the four separate tools.2026-04
Word2imgw2imgsd1.5-img2imgCreate beautifully designed words using Segmind’s word to image for your marketing purposes2023-09

Image generation

ModelSlugDescriptionAdded
LordSwaminaryan675ed494b9-thejagstudio-LordSwaminaryan2024-08
sdxl_lora_architecture_siheyuan9c54110684-frank-chieng-sdxl_lora_architecture_siheyuan2024-04
FLUX.1acn-0effee052aFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FG-v3aio-fg-full-fc1d87eafbFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
AladdinAladdin5k-701ba0d580Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
FLUX.1alexV2-2a5661e8efFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-08
FLUX.1AlexV2-655afc6166Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-08
FLUX.1alexV2FastFlux-e99e4b5aa1Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-08
FLUX.1alice-in-wonderland-ab7255451bFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-12
FLUX.1ANIME-GEN-V4-9dae59678bFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
FLUX.1anlora-e8d79af2afFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
FLUX.1ap-fc0862b464Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-12
FLUX.1arlora-9ceba116a8Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
FLUX.1armp-13752d05dfFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
ClayAnimationRedmondartificialguybr-ClayAnimationRedmondClay Animation Redmond based on SDXL 1.0, excels at creating mesmerizing clay animation images with unparalleled ease and precision.2023-10
FLUX.1Asherflex-c25117ed67Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
FLUX.1awf-66afd16a65Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-12
FLUX.1awmcn-617fbb4446Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1awp-29909cf89bFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-12
FLUX.1Ayalora-5ff7f8f2fdFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
watercolor_style_lora_sdxlb0445a2335-ostris-watercolor_style_lora_sdxl2024-03
Background Eraserbackground-eraserBackground Eraser helps in flawless background removal with exceptional accuracy.2024-06
FLUX.1baiqiang-8812bdefb9Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1Balloonblowing-d8d4c47115Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
FLUX.1bed-and-chair-combined-8dd2e74338Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
FLUX.1bizhen-3683a8e947Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
Bria FIBO 1.5bria-fibo-generateGenerate photorealistic images with structured JSON prompt control.2025-11
Bria 3.2 Text to Imagebria-text-to-imageBria 3.2 AI transforms natural language into stunning visuals for diverse creative applications — with Base, Fast, and HD modes to match yo…2025-08
Bria Vector Graphicsbria-text-to-vector-graphicsBria Vision enables high-quality text-to-image and text-to-vector graphic generation for versatile commercial use.2025-08
aether-bubbles-foam-lora-for-sdxlc42bf9a66d-joachimsallstrom-aether-bubbles-foam-lora-for-sdxl2024-02
FLUX.1cf-6ccbc88846Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1cfoly-c65493430dFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
blacklight-makeup-sdxl-lorachillpixel-blacklight-makeup-sdxl-loraBlacklight Makeup SDXL LoRA is fine-tuned to generate makeup designs that are not only visually striking but also perfectly suited for blac…2023-10
ChromachromaChroma is an open-source, 8.9B parameter text-to-image model (based on FLUX.1-schnell) designed for diverse and uncensored content generati…2025-05
FLUX.1cloth-finetune-flux-91b5cc1278Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
FLUX.1crazy-8eae969efeFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
FLUX.1ctmn-806488193bFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1Delibrate_V2-8dbc3ba5a9Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
FLUX.1dianying-549c4441a6Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1Dinaone-ExteriorFloor-c404fe7cc8Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1djwx-635a3531b4Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
dog-example-sdxl-loradminhk-dog-example-sdxl-loraDog Example SDXL LoRA, a specialized AI model within the Stable Diffusion XL framework, uniquely trained to enhance canine imagery.2023-10
FLUX.1doctor-dolittle-0018c78851Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-12
FLUX.1dslora-b6e9a6e529Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
FLUX.1dw_bedroom_1_5k-71b62c19fbFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
FLUX.1dw_bedroom_2k-559dc9f020Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
FLUX.1egyptian-d9f4de37d4Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
Fast Flux.1 Schnellfast-flux-schnellFast Flux.1 Schnell by Segmind is an optimized text-to-image model designed for developers needing faster image generation. It offers high…2024-08
sdxl-lora-index-modern-luxury-1fb4ab6b705-naphatmanu-sdxl-lora-index-modern-luxury-12024-02
FLUX.1felora-94021b7888Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
FLUX.1fg_v2-5a5c3da2b3Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
FLUX.1fg-individual-15-5b442665e9Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
Forge Visionfg15k-4879e3ca32Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-12
FLUX.1fg5kFull-2fffd1e2b0Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
FLUX.1fgResize5k-b2981fb7b8Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
flux-pro-1.1flux-1.1-proFlux Pro 1.1 is a cutting-edge image generation tool offering exceptional speed, quality, and customization. Ideal for digital artists, des…2024-10
Flux-1.1 Pro Ultraflux-1.1-pro-ultraCreate stunning visuals effortlessly with Flux 1.1 Pro Ultra. Experience unparalleled image quality and speed.2024-11
Flux.1 Devflux-devFlux Dev is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-08
Flux Dev Finetunedflux-dev-finetunedFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-07
FLUX.1flux-hanurama-22f3d043c0Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
FLUX.1flux-pixar-27334bad28Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-04
Flux .1 Proflux-proFlux Pro is a state-of-the-art image generation with top of the line prompt following, visual quality, image detail and output diversity.2024-08
Flux Realism Lora with Upscaleflux-realism-loraFlux Realism Lora with upscale, developed by XLabs AI is a cutting-edge model designed to generate realistic images from textual descriptio…2024-08
Flux.1 Schnellflux-schnellFlux Schnell is a state-of-the-art text-to-image generation model engineered for speed and efficiency.2024-08
FLUX.1Flux-Turbo-1b4d78be3fFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
FLUX.1Flux-Turbo-37e8b8f594Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
FLUX.1fluxanimals-05320abd47Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-03
FLUX.1fluxdbb-a972e916efFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
FLUX.1FluxPixar-30c83df57dFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-04
FLUX.1FluxTurbo-f28792fc23Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
FLUX.1fulora-f3df2e804dFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
FLUX.1fylora-fb68c65109Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
FLUX.1Gen-v2-5c9cdcda8eFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
FLUX.1Ghibil-27c901962aFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-04
cyborg_style_xlgoofyai-cyborg_style_xlCyborg Style SDXL specializes in generating cyborg-themed artwork based on science fiction and futuristic aesthetics.2023-10
GPT Image 1gpt-image-1Create high-quality AI-generated images from text prompts using OpenAI's GPT Image 1 model. Ideal for product design, content creation, and…2025-04
GPT Image 1 Minigpt-image-1-miniHigh-quality image generation from text, fast and affordable.2025-10
GPT Image 1.5gpt-image-1.5Stunning photorealistic images with exceptional instruction-following.2025-12
GPT Image 2gpt-image-2Generate photorealistic images with legible multilingual text and 2K output.2026-04
lora-sdxl-notion-illustrationgvrizzo-lora-sdxl-notion-illustration2024-01
FLUX.1haosc-610e769fa8Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
Ideogram 2a Text To Imageideogram-2a-txt-2-imgCreate captivating designs, realistic images & innovative logos with Ideogram 2a text-to-image.2025-03
Ideogram 3.0ideogram-3Ideogram 3.0 revolutionizes content creation with photorealistic text-to-image generation and diverse aesthetic styles.2025-05
Ideogram 4.0ideogram-4Generate 2K posters and logos with accurate text rendering.2026-06
Ideogram Turbo Text To Imageideogram-turbo-txt-2-imgCreate stunning images in seconds with Ideogram Turbo Text to Image. Fast AI model for quick ideation & text rendering.2025-03
Ideogram Text To Imageideogram-txt-2-imgIdeogram Text to Image: Turn your ideas into stunning visuals instantly with this powerful AI tool. Create captivating designs, realistic i…2024-09
Ideogram V4 Fastideogram-v4-fastGenerate posters and logos with accurate in-image text.2026-07
Imagen 3imagenImagen 3 is Google DeepMind's highest quality text-to-image model. Generates detailed images with enhanced lighting, diverse styles, and im…2025-02
Imagen 4imagen-4Imagen 4 is Google’s most advanced AI image generation model, creating detailed, photorealistic or abstract images from text prompts. It ex…2025-05
Imagen 4 Fastimagen-4-fastFast photorealistic image generation for bulk and iteration.2026-05
Imagen 4 Ultraimagen-4-ultraPhotorealistic images with native 2K resolution and precise text.2026-05
sdxl-khuze-nocrop-1e-4-1200jayashri710-sdxl-khuze-nocrop-1e-4-12002024-01
sdxl-lora-khuze-1e-4-1200-512x512imagesjayashri710-sdxl-lora-khuze-1e-4-1200-512x512images2024-01
FLUX.1jdq-396bb5a145Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1Jenny1-6546baecaeFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
Juggernaut Lightning Fluxjuggernaut-lightning-fluxJuggernaut Lightning Flux: Blazing fast (<300ms!) & powerful inference with enhanced visuals.2025-03
Juggernaut Pro Fluxjuggernaut-pro-fluxJuggernaut Pro FLUX: Create stunningly realistic AI images with unprecedented detail and sharpness.2025-03
FLUX.1jzzs-a0fc97d5dfFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1jzzsd-ef9fd76525Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
punk-collageKappaNeuro-punk-collagePunk Collage Model offers a unique way to create digital collages that resonate with the punk culture's raw energy and subversive charm.2023-10
lora-sdxl-watercolorkchoi-lora-sdxl-watercolor2024-01
FLUX.1KellyTest-66fbd687a2Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-06
Kling V3 Text to Imagekling-3-text2imagePhotorealistic, print-ready images from text prompts.2026-03
FLUX.1liangnv-b2e39c7ec4Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1M-EdenRock-bed-2d45decb6aFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
FLUX.1M-HubbaArmChair-and-M-RicochetFabric-bed-18fed2c922Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
FLUX.1M-HubbaArmChair-fff004d426Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
FLUX.1M-RicochetFabric-bed-9180ccb8f7Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
FLUX.1Marimekko-mrk3903-1e8cab847cFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-03
FLUX.1meizhuang-2ae6bc0009Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
sdxl-ugly-sonic-loraminimaxir-sdxl-ugly-sonic-loraSDXL Ugly Sonic LoRA excels at generating quirky and iconic version of one of the most beloved movie characters - Sonic the hedgehog.2023-10
sdxl-wrong-loraminimaxir-sdxl-wrong-loraSDXL Wrong LoRA is engineered with a focus on delivering images of higher detail, color saturation and vibrance, bringing images to life wi…2023-10
FLUX.1mnsh-86fb5b8ef4Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1mnzs-b5e1ea1473Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
SDXL - Multi Loramulti-loraSDXL Model with multiple LoRa loading support.2024-04
Nano Banananano-bananaGemini Image Editor preserves authentic subject identity while enabling seamless image editing and manipulation.2025-08
FLUX.1Narinder-dfac11ecfbFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-05
FLUX.1nav_007_Krishna-6fce6aad5bFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
lego-minifig-xlnerijs-lego-minifig-xlLEGO Minifig XL is designed to generate LEGO images and excels in creating detailed and accurate representations of LEGO minifigures and it…2023-10
SDXL-StickerSheet-LoraNorod78-SDXL-StickerSheet-LoraSDXL StickerSheet LoRA is expertly fine-tuned on a comprehensive collection of sticker images, enabling it to produce a wide variety of sti…2023-10
crayon_style_lora_sdxlostris-crayon_style_lora_sdxlCrayon Style - SDXL LoRA is a unique model designed to convert any text prompt into a vibrant, crayon-style drawing.2023-10
ikea-instructions-lora-sdxlostris-ikea-instructions-lora-sdxlIkea Instructions LoRA SDXL model is fine-tuned on IKEA diagrams and specializes in generating clear, concise, and easy-to-follow visual in…2023-10
stained-glass-style-sdxlostris-stained-glass-style-sdxlStained Glass Style SDXL is trained extensively on diverse stained glass images and can replicate the essence of stained glass in digital a…2023-10
Pruna P Imagep-imagep-image generates high-quality images from text prompts in seconds, optimizing for speed and fidelity.2025-11
Pruna P Image Ideogramp-image-ideogramSub-second text-to-image with legible in-image text.2026-07
FLUX.1pixar-3d-5ab56ddd76Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-02
FLUX.1playzippyIramayanakids-7f162e1446Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
FLUX.1prabhas-5f173e43a4Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
FLUX.1quanshen-7d0d5c4bdbFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
Qwen Imageqwen-imageQwen-Image revolutionizes image generation and editing with seamless multilingual text integration and photorealistic detail.2025-08
Qwen Image 2512qwen-image-2512Photorealistic image generation with precise text description following.2025-12
Qwen Image 3.0qwen-image-3Generate and edit legible in-image text, up to 2K.2026-08
Qwen Image Fastqwen-image-fastQwen-Image expertly generates stunning images with complex text integration, especially for Chinese typography.2025-08
sdxl-lora-lower-decks-aestheticra100-sdxl-lora-lower-decks-aestheticSDXL LoRA Lower Decks Aesthetic model, inspired by the unique style of “Star Trek: Lower Decks.” generates artwork in the distinctive anima…2023-10
RamaRama-558043a138Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
Ramayan ModelramayanaMulti5k-69c1cc4ae9Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
lora-dog-SSD-1Bramsrigouthamg-lora-dog-SSD-1BLoRA Dog SSD-1B specializes in generating photorealistic images of dogs2023-10
Recraft V3recraft-v3Recraft V3, the latest iteration of Recraft AI, offers a significant advancement in AI-driven image generation. This state-of-the-art model…2024-11
Recraft V3 Svgrecraft-v3-svgRecraft V3 SVG generates high-quality, customizable vector graphics with precision and ease. Perfect for logos, infographics, illustrations…2024-11
Reve 2reve-2Generate and edit 4K images with sharp in-image text.2026-07
FLUX.1rxmx-324bd4c1e2Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1SanaAI-542b8dbe01Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-10
FLUX.1SBIC-Stone-fbcc5912c5Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1sclora-274da0a25fFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
Cyber Realisticsd1.5-cyberrealisticThe most versatile photorealistic model that blends various models to achieve the amazing realistic images.2023-09
Edge of Realismsd1.5-edgeofrealismThis model corresponds to the Stable Diffusion Edge of Realism checkpoint for detailed images at the cost of a super detailed prompt2023-09
Epic Realismsd1.5-epicrealismThis model corresponds to the Stable Diffusion Epic Realism checkpoint for detailed images at the cost of a super detailed prompt2023-09
Juggernaut Finalsd1.5-juggernautThe most versatile photorealistic model that blends various models to achieve the amazing realistic images.2023-09
Realistic Visionsd1.5-realisticvisionThis model corresponds to the Stable Diffusion Realistic Vision checkpoint for detailed images at the cost of a super detailed prompt2023-09
Reliberatesd1.5-reliberateThis model corresponds to the Stable Diffusion Reliberate checkpoint for detailed images at the cost of a super detailed prompt2023-09
Colossus Lightning SDXLsdxl1.0-colossus-lightningColossus Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.2024-03
Dreamshaper SDXLsdxl1.0-dreamshaperThe SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software.2023-10
DreamShaper Lightning SDXLsdxl1.0-dreamshaper-lightningDreamShaper Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.2024-03
Dynavis Lightning SDXLsdxl1.0-dyanvis-lightningDynavis Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.2024-03
Juggernaut Lightning SDXLsdxl1.0-juggernaut-lightningJuggernaut Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.2024-03
NewReality Lightning SDXLsdxl1.0-newreality-lightningNewReality Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.2024-03
NightVis Lightning SDXLsdxl1.0-nightvis-lightningNightVis Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.2024-03
ProtoVision Lightning SDXLsdxl1.0-protovis-lightningProtoVision Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.2024-03
RealDream Lightningsdxl1.0-realdream-lightningRealDream is a sophisticated image generation model utilizing SDXL Lightning architecture. It creates incredibly realistic images from text…2024-07
Realdream Pony V9sdxl1.0-realdream-pony-v9Real Dream Pony V9 is an advanced image generation model based on the Stable Diffusion XL (SDXL) architecture, excelling in photorealism.2024-07
Realism Lightning SDXLsdxl1.0-realism-lightningRealism Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.2024-03
Realvis SDXLsdxl1.0-realvisThe SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software.2023-10
Realvis Lightning SDXLsdxl1.0-realvis-lightningRealvis Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.2024-03
Samaritan 3D XLsdxl1.0-samaritan-3dSamaritan 3D XL leverages the robust capabilities of the SDXL framework, ensuring high-quality, detailed 3D character renderings.2023-12
Samaritan Lightning SDXLsdxl1.0-samaritan-lightningSamaritan Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.2024-03
Copax Timeless SDXLsdxl1.0-timelessThe SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software.2023-10
Stable Diffusion XL 1.0sdxl1.0-txt2imgThe SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software2023-09
WildCard Lightning SDXLsdxl1.0-wildcard-lightningWildCard Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.2024-03
Zavychroma SDXLsdxl1.0-zavychromaThe SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software.2023-10
Seedream 5.0 Proseedream-5-proRegion-precise image editing with native multilingual text.2026-07
Seedream 5.0 Lite: Text-to-Imageseedream-v5-lite-text-to-imageFast, affordable instruction-following image generation.2026-02
Segmind-Vegasegmind-vegaThe Segmind-Vega Model is a distilled version of the Stable Diffusion XL (SDXL), offering a remarkable 70% reduction in size and an impress…2023-12
Segmind-VegaRTsegmind-vega-rt-v1Segmind-VegaRT a distilled consistency adapter for Segmind-Vega that allows to reduce the number of inference steps to only between 2 - 8 s…2023-12
FLUX.1Selfie-Aesthetic-f9c0f409c5Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1Selfie-Aesthetic1-9e21b5a57fFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1sf-31e1b1ee3cFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1sheb-b899ceaa9fFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-04
FLUX.1sheb-b9150b5b83Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-04
FLUX.1Shravya_ai-f44dd1e647Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
Simple Vector Flux LoraSimple_Vector_FluxFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
FLUX.1siria-e3fb0ee32fFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-03
SitaSita-b382eb5c53Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
Flux.1 Schnellsoftpasty-flux-dev-a4edbd379cFlux Schnell is a state-of-the-art text-to-image generation model engineered for speed and efficiency.2024-09
FLUX.1SpeedyMary-b361f2fd1dFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-02
SSD-1Bssd-1bSSD-1B efficiently generates high-quality, diverse images from text prompts in real-time.2023-10
Stable Diffusion 3 Medium Text to Imagestable-diffusion-3-medium-txt2imgStable Diffusion is a type of latent diffusion model that can generate images from text. It was created by a team of researchers and engine…2024-06
Stable Diffusion 3.5 Large Text to Imagestable-diffusion-3.5-large-txt2imgStable Diffusion 3.5 Large offers exceptional customizability, efficient performance on consumer hardware, and diverse image outputs that a…2024-10
Stable Diffusion 3.5 Turbo Text to Imagestable-diffusion-3.5-turbo-txt2imgStable Diffusion 3.5 Turbo offers exceptional customizability, efficient performance on consumer hardware, and diverse image outputs that a…2024-10
FLUXTest 001 -4b5d788e6cFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-08
FLUX.1thumbnail-bc5f00ef3cFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1tianmei-e09a3dfa92Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1v6lora-1745486a5eFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-09
FLUX.1veo-dd69bf026eFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-11
Wan 2.7 Image Generationwan2.7-image2K image generation with precise multilingual text rendering.2026-04
Wan 2.7 Image Generation Prowan2.7-image-pro4K images with chain-of-thought reasoning and multilingual text.2026-04
FLUX.1whnm-339c18f951Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1winnie-the-pooh-c3ccfd99b1Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2024-12
FLUX.1wxns-9cf0075aeeFlux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1xse-1f865b3596Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1xz-c445d830b2Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1yujia-0d97d92ca5Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
Flux.1 Schnellyunxi-aac5d8354bFlux Schnell is a state-of-the-art text-to-image generation model engineered for speed and efficiency.2025-04
Z Image Turboz-image-turboPhotorealistic images in under one second, bilingual text.2025-11
FLUX.1zh-66dcb61456Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1zscj-7c64578707Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1zsjz-7ca8eff909Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01
FLUX.1zsnh-7634404088Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions2025-01

Image t o image

ModelSlugDescriptionAdded
Stable Diffusion 3 Medium Image to Imagesd3-med-img2imgStable Diffusion 3 Medium image-to-image is a cutting-edge AI tool that uses advanced image-to-image technology to transform one image into…2024-07

Image to data

ModelSlugDescriptionAdded
Image Metadatautility-image-metadataRead image metadata: dimensions, format, EXIF (with GPS decoded to decimal), ICC profile name, raw XMP. Returns JSON, not an image.2026-04

Image to video

ModelSlugDescriptionAdded
AI Face Swap (image and video)ai-face-swapAI Face Swap: Effortlessly replace faces online. Fine-tune swaps with advanced controls for age, gender, and resolution.2025-01
Cog videoX Image To Videocog-video-5b-i2vCogVideoX image-to-video is a cutting-edge AI model that converts static images into dynamic, high-quality videos. Perfect for content crea…2024-09
FLUX 3 Image to Videoflux-3-image-to-videoAnimate images into 20-second clips with synchronized native audio.2026-08
Grok Imagine Video 1.5 Image to Videogrok-imagine-video-1.5-image-to-videoAnimate a still image into 1080p video with synced audio.2026-08
Grok Imagine Video 1.5 (Preview)grok-imagine-video-1.5-previewImage-to-video with native synchronized audio, up to 720p.2026-06
Grok Imagine Video 1.5 Reference to Videogrok-imagine-video-1.5-reference-to-videoCharacter-consistent video from up to 7 reference images.2026-08
Hailuo 02 Fasthailuo-02-fastTransform any static image into a captivating, high-quality video clip effortlessly.2025-08
Hailuo 2.3hailuo-2.3Hyper-realistic videos from text with fluid character motion.2025-10
Hailuo 2.3 Fasthailuo-2.3-fastProfessional-quality videos from text and images at speed.2025-10
HallohalloHallo lets you create portrait videos from single images.2024-06
HappyHorse 1.0happyhorseCinematic 1080p text-to-video with native audio and lip-sync.2026-04
HappyHorse 1.1happyhorse-1.1Generate cinematic video with synchronized native audio and multilingual lip-sync from text, an image, or reference images.2026-06
Heygen Avatar IVheygen-avatar-ivSingle photo into a lifelike talking avatar video.2025-12
Higgsfield Image 2 Videohiggsfield-image2videoTransform static images into dynamic, motion-rich videos with unparalleled control and creative depth.2025-09
Higgsfield Speech 2 Videohiggsfield-speech2videoTransform images and audio into dynamic, lip-synced videos for engaging digital content.2025-09
InfiniteTalkinfinite-talkFull-body animation from images synchronized perfectly to audio.2025-10
Kling AI 1.6 Image to Videokling-1.6-image2videoKling AI 1.6 Image-to-Video is a powerful AI tool that transforms static images into captivating, animated videos. Create high-quality cont…2025-01
Kling 2kling-2Kling 2.0 is an advanced AI video generator (5 and 10 seconds) that creates cinematic, dynamic videos from text or images with lifelike mot…2025-04
Kling 2.1 AI Video Generatorkling-2.1Kling 2.1 offers hyper-realistic video generation with improved motion, sharper 1080p visuals, and instant restyling capabilities. Its cost…2025-05
Kling 2.5 Turbokling-2.5-turboKling AI 2.5 Turbo generates fluid, cinematic videos from text and images, enhancing content creation and storytelling.2025-09
Kling 2.6kling-2.6Still images into immersive cinematic videos with synchronized audio.2025-12
Kling 3.0 Pro Image-to-Videokling-3-pro-image2videoAnimated 1080p videos from images with dynamic motion.2026-02
Kling 3.0 Standard Image-to-Videokling-3-standard-image2videoControlled cinematic 1080p videos from starting images.2026-02
Kling bloombloomkling-bloombloomKling AI transforms text and images into dynamic, high-quality video content with realistic motion and sound.2025-05
Kling dizzydizzykling-dizzydizzyKling DizzyDizzy transforms static content into dynamic, high-resolution videos, enhancing engagement and storytelling for creators.2025-05
Kling Expansionkling-expansionUnleash dynamic visuals with Kling Expansion! Effortlessly inflate and stretch elements for surreal and captivating effects.2025-04
Kling fuzzyfuzzykling-fuzzyfuzzyTransform your photos instantly into adorable, plush-toy-like visuals with Kling fuzzyfuzzy effect.2025-04
Kling Heart Gesturekling-heart-gestureExpress affection visually with Kling AI's heart gesture effect! Input two portraits and instantly create heartwarming videos featuring a d…2025-04
Kling Hugkling-hugCreate heartwarming videos instantly with Kling hug effect! Generate tender embracing animations.2025-04
Kling AI Image to Videokling-image2videoKling AI Image-to-Video is a powerful AI tool that transforms static images into captivating, animated videos. Create high-quality content…2024-10
Kling Kisskling-kissCreate a heartfelt video in seconds with Kling kiss effect! Input two portraits and instantly generate a kissing animation.2025-04
Kling O1 Image 2 Videokling-o1-image-to-videoPhysics-driven animations from images for creative storytelling.2026-01
Kling O1 Reference Image 2 Videokling-o1-reference-image-to-videoIdentity-preserving videos from static images with character reference.2026-01
Kling O3 Image To Videokling-o3-image2videoImages to cinematic videos with precise motion control.2026-03
Kling Squishkling-squishTransform your visuals with Kling AI squish effect! Easily compress and distort images/videos for playful, exaggerated effects.2025-04
Kling V1 Pro AI Avatarkling-v1-pro-ai-avatarDynamic AI avatars with synchronized speech from image.2025-10
Kling V1 Standard AI Avatarkling-v1-standard-ai-avatarLifelike AI avatars with precise lip-sync for presentations.2025-10
Kling V2 Pro Avatarkling-v2-pro-avatarTalking avatar videos from image and audio, high quality.2025-12
Kling Avatar V2 Standardkling-v2-standard-avatarLifelike video avatars with precise lip synchronization.2025-12
Live Portraitlive-portraitLive Portrait animates static images using a reference driving video through implicit key point based framework, bringing a portrait to lif…2024-07
Live Portrait video to videolive-portrait-video-to-videoExperience the magic of Live Portrait’s Video-to-Video Model! Transform your static images into dynamic videos seamlessly.2024-07
LTX 2 Fastltx-2-fastFast, high-quality text-to-video generation by Lightricks.2025-10
LTX 2 Proltx-2-proHigh-quality video generation with advanced motion control.2025-10
LTX 2.5 Fastltx-2.5-fastText-to-video and image-to-video with native audio, up to 4K.2026-08
LTX 2.5 Proltx-2.5-proGenerate 1080p video with native audio and multi-shot scenes.2026-08
LTX Videoltx-videoLTX-Video is the first DiT-based video generation model capable of generating high-quality videos in real-time. It produces 24 FPS videos a…2024-12
Luma Ray 3.2luma-ray-3-2Cinematic text-to-video and image-to-video clips up to 1080p.2026-06
MiniMax AI (Hailuo)minimax-aiWith Video-01 by MiniMax, create high-definition videos at 720p resolution and 25fps, featuring cinematic camera movement effects based on…2024-12
MiniMax Hailuo H3 Image to Videominimax-h3-image-to-videoAnimate a still image into 2K video up to 15s.2026-07
MiniMax Hailuo H3 Reference to Videominimax-h3-reference-to-videoKeep characters and products consistent in 2K reference-to-video.2026-07
Minimax Hailou 2minimax-hailuo-2Generate breathtaking 1080P cinematic videos from text or images with ultra-realistic motion and physics.2025-07
Luma Modify Videomodify-videoTransform videos seamlessly with high-fidelity generative edits while preserving original actor performances.2025-07
Motion Control SVDmotionctrl-svdMotion Control SVD is an innovative deep learning framework that breathes life into static images. By intelligently managing both camera an…2024-07
Muscle Surgemuscle-surgeInstantly add muscle and strength to your videos with Pixverse Muscle Surge effect!2025-04
Pruna P Video Animatep-video-animateTransfer video motion and audio onto any still image.2026-07
Pruna P Video Avatarp-video-avatarAnimate any portrait into a lip-synced talking avatar.2026-06
Pruna P Video Replacep-video-replaceSwap on-screen video characters while preserving motion and audio.2026-07
Pixverse 4.5 Effectspixverse-4.5-effectsPixVerse 4.5 transforms photos and text into stunning animated videos for impactful storytelling and marketing.2025-05
Pixverse 4.5 Videopixverse-4.5-videoPixverse 4.5 transforms static images and text into dynamic, engaging videos for captivating social media content.2025-05
Pixverse 5 Extendpixverse-5-extendSeamlessly extend and continue AI-generated videos.2025-10
Pixverse 5 Transitionpixverse-5-transitionSeamless AI-generated video transitions between scenes.2025-10
Pixverse 5 Videopixverse-5-videoCinematic videos from text and images with photorealism.2025-10
Pixverse Image to Videopixverse-image2videoAnimate your photos effortlessly with Pixverse Image to Video AI! Upload, add motion prompts and styles.2025-04
Pixverse Mimicpixverse-mimicTransfer motion from reference videos onto still images.2026-05
Pixverse V6pixverse-v615-second AI videos with native audio and cinematic controls.2026-04
Luma Ray flash 2 (720p)ray-flash-2-720pGenerate stunning 720p videos from text with the Luma ray-flash-2-720p model. Faster & cheaper than Ray 2, offering realistic motion & deta…2025-03
Runway Gen Alpha Turbo Image to Videorunway-gen3-alphaturboRunway Gen-3 AlphaTurbo is a cutting-edge AI tool that transforms static images into dynamic videos with exceptional fidelity and motion2024-10
Runway Gen 4 Turborunway-gen4-turboGenerate videos faster and cheaper with Runway Gen-4 Turbo! Create high-quality text, image, and combined video generation for rapid conten…2025-04
SadTalkersadtalkerAudio-based Lip Synchronization for Talking Head Video2024-06
Wan ScailscailProfessional character animations from reference images.2025-12
Seedance 1.0 Pro Fastseedance-1.0-pro-fastCinematic videos from text and images at ultra speed.2025-10
Seedance 1.5 Proseedance-1.5-proSynchronized video and audio generation for dynamic storytelling.2025-12
Seedance 2.0seedance-2.0Cinematic AI videos with native audio and multi-shot narratives.2026-04
Seedance 2.0 Fastseedance-2.0-fastProfessional-grade video creation model with native audio, similar to SeeDance 2.0 but faster and cheaper.2026-04
Seedance 2.0 Miniseedance-2.0-miniFast text-to-video and image-to-video with synchronized audio.2026-06
Seedance 2.5seedance-2.5Generate cinematic multi-shot AI videos up to 30 seconds with synchronized native audio from text, images, or references.2026-08
Seedance 1.0 Proseedance-proSeedance Pro transforms text and images into engaging 720p dynamic videos with cinematic storytelling.2025-07
Sora 2sora-2Stunning dynamic videos from detailed text descriptions.2025-10
Sora 2 Prosora-2-proCinematic-quality videos from text with temporal consistency.2025-10
Stable Video DiffusionsvdTakes image as input and returns a video.2023-12
TooncraftertooncrafterCreate videos from illustrated input images2024-06
V Expressv-expressV-Express lets you create portrait videos from single images.2024-06
VEED Fabric 1.0veed-fabric-1.0Animate any image into a realistic talking video, lip-synced to your audio or generated from a text script.2026-07
Google Veo 2 Image To Videoveo-2-image2videoDiscover Google Veo 2, an AI-powered image-to-video model with 4K resolution, realistic motion, and cinematic effects for creators and deve…2025-03
Veo 3.1veo-3.1Static images into high-quality videos with synchronized audio.2025-10
Veo 3.1 Fastveo-3.1-fastTransforms static images into dynamic 1080p videos with synchronized audio and natural motion.2025-10
Wan Video Effectsvideo-effectsTransform your videos with diverse video effects. Start creating captivating videos today.2025-03
HyperSwap: Video Faceswap by FaceFusion Labsvideo-faceswap-by-facefusion-labsRealistic face swapping in videos from a single image.2026-03
Video Frame Interpolationvideo-frame-interpolationFILM synthesizes smooth, high-quality intermediate frames for fluid motion in videos with significant movement.2025-09
Video Stitchvideo-stitchRevolutionize your video editing with the Video Stitch Model. Seamlessly stitch clips, add captivating audio, and create professional-looki…2024-10
Video Tryonvideo-tryonVideo Tryon is Segmind’s next-generation AI video model for instant virtual try-on, allowing users to visualize any outfit on any person in…2025-09
Video Tryon V2video-tryon-v2Video Tryon is Segmind’s next-generation AI video model for instant virtual try-on, allowing users to visualize any outfit on any person in…2025-11
Video Watermark Removervideo-watermark-removerRemove watermarks from any video instantly with AI.2025-10
Video FaceswapvideofaceswapVideo Faceswap is a powerful tool for creators, filmmakers, and meme enthusiasts. With this innovative technology, you can effortlessly rep…2024-07
Wan 2.2 Image to Video Fastwan-2.2-i2v-fastTransforms simple text prompts into breathtaking cinematic-quality videos in minutes.2025-08
Wan 2.2 Image to Video Flashwan-2.2-i2v-flashConvert a single image into a coherent dynamic video.2026-03
Wan 2.5 Image to Videowan-2.5-i2vWan2.5-Preview creates stunning, high-resolution videos with flawless audio synchronization from multiple inputs.2025-09
Wan 2.6 Image To Videowan-2.6-i2vTransform images into high-quality videos with audio sync.2025-12
Wan Animatewan-animateAnimate characters and replace video subjects seamlessly.2025-10
Wan 2.1 720p image to videowan2.1-i2v-720pCreate high-quality 720p videos with excellent visual quality and a broad spectrum of motion from static images.2025-02
Wan 2.6 Image to Video Flashwan2.6-i2v-flashAnimate photos into 15-second 1080p video with native audio.2026-08
Wan 2.7 Image to Videowan2.7-i2vAnimate any image into cinematic 1080P video with audio.2026-04
Wan 2.7 Reference to Videowan2.7-r2vCharacter-consistent multi-subject videos from reference images.2026-04
Wan 3.0 Videowan3.0-videoGenerate 30-second 1080p video with native audio.2026-08
Warmth of Jesuswarmth-of-jesusExperience the viral "Warmth of Jesus" effect on PixVerse! Transform your images into heartwarming videos of Jesus embracing people.2025-04

Image to3d

ModelSlugDescriptionAdded
Hunyuan3D-2hunyuan-3d-2Hunyuan3D 2.0 enables the creation of high-quality 3D models with intricate details. Produce assets that are visually appealing and suitabl…2025-01
Hunyuan3d-2.1hunyuan3d-2.1Transform 2D images into photorealistic, high-fidelity 3D assets effortlessly.2025-08
Hunyuan-3d 2mvhunyuan3d-2mvHunyuan3D-2mv is finetuned from Hunyuan3D-2 to support multiview controlled shape generation.2025-03
Sam 3D Bodysam-3d-bodyReconstruct 3D human body meshes from a single photo.2025-12
Sam 3D Objectsam-3d-objectsSingle 2D image into detailed 3D object models.2025-12

Image understanding

ModelSlugDescriptionAdded
Bria Fibo Structured Promptbria-fibo-generate-structured-promptConvert complex inputs into structured JSON prompts for generation.2025-11
Bria Mask Generatorbria-mask-generatorBria AI Get Masks automatically generates accurate object masks for advanced image editing and enhancement.2025-09
Bria Prompt Enhancerbria-prompt-enhancerBria AI generates high-quality, commercially safe images tailored to diverse creative needs.2025-09
Google Translategoogle-translateTranslate effortlessly with the powerful Google Translation AI model.2025-04
Ideogram Describeideogram-describeIdeogram describe can effortlessly generate detailed prompts from images. Perfect for refining creations or replicating styles.2025-03
Image Converterimage-converterConvert images between formats instantly.2025-10
Image resizerimage-resizerResize images to any dimension quickly and precisely.2025-10
LLAVA 1.6 7Bllava-v1.6LLaVa translates images into text descriptions & captions.2024-06
NSFW Checkernsfw-checkerDetect NSFW and other inappropriate content in images. Returns a boolean has_nsfw_concepts flag, an overall NSFW score (0-1), and the full…2026-05
Sam V2.1 Hiera Largesam-v21-hiera-largeMeta's next-gen segmentation model for images and video.2025-10
Video Speed Changevideo-speed-changeSpeed up or slow down any video precisely.2025-10

Inpainting

ModelSlugDescriptionAdded
Fooocus Inpaintingfocus-inpaintFooocus Inpainting is a powerful image generation model that allows you to selectively edit and enhance images.2024-02
Controlnet Inpaintinginpaint-autoThis model is capable of generating photo-realistic images given any text input, with the extra capability of inpainting and controlling th…2024-01
Stable Diffusion Inpaintingsd1.5-inpaintingStable Diffusion Inpainting is a latent text-to-image diffusion model capable of generating photo-realistic images given any text input, wi…2023-09
SDXL Inpaintsdxl-inpaintThis model is capable of generating photo-realistic images given any text input, with the extra capability of inpainting the pictures by us…2023-11
Try-On Diffusiontry-on-diffusionOutfitting Fusion based Latent Diffusion for Controllable Virtual Try-on2024-03

Language models (LLMs)

ModelSlugDescriptionAdded
Claude 4 Sonnetclaude-4-sonnetAdvanced coding and multi-step agentic reasoning model.2025-10
Claude 4.5 Sonnetclaude-4.5-sonnetClaude Sonnet 4.5 empowers developers with advanced coding and reasoning for complex software solutions.2025-09
Claude Opus 4.7claude-opus-4.7Anthropic's most capable AI model excelling at agentic coding, complex reasoning, and high-resolution vision with a 1M-token context window.2026-04
DeepSeek Chatdeepseek-chatDeepSeek V3 combines cutting-edge AI technology with practical usability. Featuring a 671B parameter architecture, enhanced reasoning capab…2025-01
DeepSeek R1deepseek-reasonerDeepSeek-R1 is a cutting-edge AI reasoning model that combines reinforcement learning with supervised fine-tuning. Excels in complex proble…2025-01
Gemini 2.5 Flashgemini-2.5-flashMultimodal AI with transparent reasoning, fast and affordable.2025-10
Gemini 2.5 Flash Litegemini-2.5-flash-liteFastest Gemini 2.5 model for high-volume text and vision tasks.2026-05
Gemini 2.5 PROgemini-2.5-proComplex multimodal reasoning across diverse inputs and formats.2025-10
Gemini 3 Flashgemini-3-flashFrontier-class reasoning and multimodal AI at scale.2026-05
Gemini 3 Progemini-3-proAutonomous multimodal AI for complex reasoning and coding.2025-11
Gemini 3.1 Flash Litegemini-3.1-flash-liteUltra-fast, affordable LLM for high-volume AI pipelines.2026-05
Gemini 3.1 Progemini-3.1-proFrontier reasoning across text, images, video, and code.2026-05
Gemini 3.7 Flashgemini-3.7-flashFast multimodal LLM for coding, agents, and long-document analysis.2026-08
GLM 5.2glm-5.21M-token open-weight LLM for long-horizon coding.2026-07
GPT 4gpt-4GPT-4 outperforms both previous large language models and as of 2023, most state-of-the-art systems (which often have benchmark-specific tr…2024-05
GPT 4 turbogpt-4-turboGPT-4 outperforms both previous large language models and as of 2023, most state-of-the-art systems (which often have benchmark-specific tr…2024-05
GPT 4ogpt-4oGPT-4o (“o” for “omni”) is our most advanced model. It is multimodal (accepting text or image inputs and outputting text), and it has the s…2024-05
GPT 5gpt-5GPT-5 automates complex coding tasks with integrated tools for seamless software development and deployment.2025-08
GPT 5 Minigpt-5-miniRapid high-quality AI across text, images, and files.2025-10
GPT 5 Nanogpt-5-nanoUltra-fast LLM responses for real-time AI applications.2025-10
GPT 5.1gpt-5.1Precise code review and developer workflow assistant.2025-12
GPT 5.2gpt-5.2Advanced reasoning with multimodal input for precise tasks.2025-12
GPT 5.4gpt-5.4Most powerful GPT for frontier reasoning and multimodal tasks.2026-03
GPT 5.4 Minigpt-5.4-miniFastest efficient model for coding and computer-use tasks.2026-03
GPT 5.4 Nanogpt-5.4-nanoFlagship-class AI for classification and extraction tasks.2026-03
GPT 5.5gpt-5.5Frontier reasoning and coding with 1M-token context window.2026-05
Kimi K2 Instruct 0905kimi-k2-instruct-0905Deep contextual understanding and complex code generation.2025-11
Llama 3 8bllama-v3-8b-instructMeta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned gener…2024-05
Llama 3.1 70bllama-v3p1-70b-instructMeta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned gener…2024-07
Llama 3.1 8bllama-v3p1-8b-instructMeta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned gener…2024-07
Llama 4 Maverick Instruct Basicllama4-maverick-instruct-basicLlama 4 Maverick Instruct Basic is a 400B parameter powerhouse with 128 experts for unparalleled text and image understanding.2025-04
Llama 4 Scout Instruct Basicllama4-scout-instruct-basicUnlock powerful multimodal AI with Llama 4 Scout basic, a 17 billion active parameters model offering leading text & image understanding.2025-04
MiniMax M3minimax-m3Reason over 1M-token context for coding and agents.2026-07
Mixtral 8x22bmixtral-8x22b-instructMistral MoE 8x22B Instruct v0.1 model with Sparse Mixture of Experts. Fine tuned for instruction following.2024-05
Nemotron 3 Ultranemotron-3-ultra1M-token reasoning for coding agents and deep research.2026-07
OpenAI o3o3Frontier reasoning model for complex coding, math, and science.2026-03
OpenAI o3 Minio3-miniCost-efficient reasoning model for coding, math, and science.2026-03
O4 Minio4-miniOpenAI o4-mini enhances decision-making by processing text and images with advanced reasoning capabilities.2025-07
QVQ Maxqvq-maxChain-of-thought visual reasoning for math, charts, and diagrams.2026-03
Qwen3.8 Maxqwen-3.8-maxMultimodal reasoning and agentic coding with 1M-token context.2026-08
Qwen Flashqwen-flashFastest low-cost LLM with 1M context for high-volume tasks.2026-03
Qwen Plusqwen-plusMid-tier 1M context LLM for summarization and content tasks.2026-03
Qwen2 VL 72B Instructqwen2-vl-72b-instructQwen2-VL-72B-Instruct is a state-of-the-art multimodal model excelling in image and video understanding, with advanced capabilities for tex…2025-02
Qwen 3 Coder Flashqwen3-coder-flashHigh-volume code generation with 1M token context window.2026-03
Qwen 3 Coder Plusqwen3-coder-plusGenerates, debugs, and refactors entire codebases efficiently.2026-03
Qwen 3 Maxqwen3-max1T-parameter LLM with hybrid reasoning and 262K context.2026-03
Qwen 3 VL Flashqwen3-vl-flashFast, affordable vision-language model with 262K context OCR.2026-03
Qwen 3 VL Plusqwen3-vl-plusPowerful visual QA and document analysis from images.2026-03
Qwen 3.5 Flashqwen3.5-flashFast multimodal AI processing text, images, and video affordably.2026-03
Qwen 3.5 Plusqwen3.5-plusMultimodal 1M context AI for image, video, and text.2026-03
QwQ Plusqwq-plusDeep chain-of-thought reasoning for math, code, and logic.2026-03

Text to embed

ModelSlugDescriptionAdded
Gemini Embedding 001gemini-embedding-001MTEB #1 text embeddings for RAG, search, and clustering.2026-05
Gemini Embedding 2gemini-embedding-2Natively multimodal embeddings — text, image, audio, video and PDF mapped into one vector space, with 8 task-specific modes.2026-05
Text Embedding 3 Largetext-embedding-3-largeText-embedding-3-large is a robust language model by OpenAI designed for generating high-dimensional text embeddings for a wide range of na…2024-08
Text Embedding 3 Smalltext-embedding-3-smallText-embedding-3-small is a compact and efficient model developed for generating high-quality text embeddings. These embeddings are numeric…2024-08

Transcription

ModelSlugDescriptionAdded
Elevenlabs Transcripteleven-labs-transcriptTranscribe audio to accurate text in 99 languages with speaker diarization and word-level timestamps.2025-03
Elevenlabs Dialogue With Timingelevenlabs-dialogue-with-timestampsMulti-speaker dialogue with expressive timestamps included.2025-11
Elevenlabs Forced Alignmentelevenlabs-forced-alignmentPrecise audio-text synchronization with word-level timestamps.2025-11
Elevenlabs Voice Cloningelevenlabs-voice-cloneHyper-realistic voice cloning from short audio samples.2025-11
Elevenlabs Voice Designelevenlabs-voice-designGenerate unique synthetic voices without audio samples.2025-11
TTS Elevenlabs With Timingtts-elevenlabs-with-timestampsEmotionally expressive TTS with word-level timestamp output.2025-11
Whisper Large V3whisper-large-v3Transcribe speech-to-text in 99 languages with timestamps.2026-08

Video editing

ModelSlugDescriptionAdded
Bria Video Eraserbria-erase-videoRemove unwanted objects from videos while preserving audio.2025-12
Bria Increase Video Resolutionbria-increase-video-resolutionTransform your videos with AI-powered upscaling and seamless background removal for professional quality.2025-09
Bria Video Background Removal 3.0bria-video-background-removal-3.0Remove video backgrounds with flicker-free, transparent alpha output.2026-06
Esrgan Video Upscaleresrgan-video-upscalerESRGAN Video Upscaler: Experience sharper, clearer 4k videos with ESRGAN. This AI-powered video upscaler boosts resolution and reduces arti…2024-09
FLUX 3 Draft Enhanceflux-3-draft-enhanceUpscale AI video drafts to Full-HD with native audio.2026-08
FLUX 3 Extend Videoflux-3-extend-videoExtend clips into seamless video continuations with synchronized audio.2026-08
Gemini Omni 1.1 Video Editgemini-omni-1.1-video-editEdit videos with a text prompt, subject preserved.2026-08
Gemini Omni 1.1 Video Extendgemini-omni-1.1-video-extendExtend short video clips into longer seamless scenes.2026-08
Heygen Video Translateheygen-video-translateTranslate videos to multiple languages with natural lip-sync.2025-10
Kling 2.6 Pro Motion Controlkling-2.6-pro-motion-controlTransfer motion from videos to animate custom characters.2025-12
Kling 2.6 Standard Motion Controlkling-2.6-standard-motion-controlPrecise motion transfer from reference videos to characters.2025-12
Kling O1 Video 2 Video Editkling-o1-video-to-video-editEdit any video with precise natural language commands.2026-01
Kling O1 Video 2 Video Referencekling-o1-video-to-video-referenceVideo style transfer using reference character images.2026-01
Kling O3 Video To Video Editkling-o3-video2video-editText-based video editor — swap backgrounds, characters, restyle scenes.2026-03
Kling O3 Video To Video Referencekling-o3-video2video-referenceSwap characters and restyle videos using reference images.2026-03
LTX Retake Videoltx-retake-videoPrecise segment-level video edits maintaining full scene continuity.2025-12
Multi Video Mergemulti-video-mergeMerge multiple videos into a single combined output.2025-10
OpusClip - Clips From Videoopus-clips-from-videoTurn long videos into captioned vertical shorts.2026-07
Pixverse Lipsyncpixverse-lipsyncPixVerse Lipsync expertly synchronizes lip movements to audio for flawless video content creation.2025-07
Runway Gen4 Alephrunway-gen4-alephRunway Aleph revolutionizes video editing with intelligent automation for seamless object and environment manipulation.2025-08
Sam V2 Videosam-v2-videoSAM v2 Video by Meta AI, allows promptable segmentation of objects in videos.2024-08
Sam3 Videosam3-videoReal-time video segmentation and multi-object tracking.2025-11
Sonilo Video to Videosonilo-video-to-videoAdd frame-synced AI music and sound effects to video.2026-08
Sync.so Lipsync 2 Prosync.so-lipsync-2-proLipsync-2-Pro seamlessly synchronizes lips in videos for instant, high-quality multilingual content creation.2025-09
Sync.so React 1sync.so-react-1Edit video actors' emotions with realistic re-expression.2025-12
Topaz Labs Video Upscaletopaz-video-upscaleTopaz Video AI upscales, enhances, denoises, stabilizes, and increases frame rates in video footage, transforming low-quality or standard-d…2025-04
VEED Lipsync v2veed-2-lipsyncDub talking-head videos with emotion-matched lip-sync.2026-07
VEED Lipsyncveed-lipsyncRe-syncs the lips of any talking-head video to a new speech audio track for realistic dubbing and localization.2026-07
VEED Subtitlesveed-subtitlesAutomatically transcribes and burns styled, translated subtitles into any video with 30 presets and a single API call.2026-07
VEED Video Background Removalveed-video-background-removalRemove any video's background with no green screen, or cleanly key chroma footage, using AI matting.2026-07
Video Audio Mergevideo-audio-mergeEffortlessly merge audio and video with our intuitive Video Audio Merge model. Create stunning multimedia content with precise timing, fade…2024-10
Video Captionervideo-captionerWith Video Captioner create accurate, customizable subtitles for your videos effortlessly.2024-10
Video Concatenatevideo-concatenateMerge videos with custom layouts, spacing, and audio.2025-12
Video Editor Agentvideo-editor-agentGeneral-purpose AI media agent: describe the transformation in natural language and it runs ffmpeg in a sandbox (transcode, resize, extract…2026-07
Video Loopvideo-loopEffortlessly loop videos for engaging social media & storytelling with our Video Loop.2025-03
Video Slicervideo-slicerVideo Slicer2025-04
Video Splitvideo-splitUtility node: Video Split. 1->N split; returns videos[] array2026-07
Wan 2.7 Video Editingwan2.7-videoeditEdit existing videos precisely using natural language text instructions.2026-04

Video generation

ModelSlugDescriptionAdded
Cog Video X 5Bcog-video-5b-t2vCogVideo is a groundbreaking AI model that turns text into high-quality videos. Create realistic scenes, animations, and more with ease. Id…2024-09
FLUX 3 Text to Videoflux-3-text-to-videoCinematic text-to-video with native lip-synced audio, up to 20s.2026-08
Gemini Omni 1.1gemini-omni-1.1Text-to-video with synchronized native audio, up to 4K.2026-08
Gemini Omni Flashgemini-omni-flashText-to-video and image-to-video with synchronized native audio.2026-06
Grok Imagine Videogrok-imagine-videoText-to-video and image-to-video with native synchronized audio.2026-06
Grok Imagine Video 1.5 Text to Videogrok-imagine-video-1.5-text-to-videoText-to-video clips up to 1080p with native synchronized audio.2026-08
HeyGen Avatar Vheygen-avatar-vStudio-quality talking-avatar videos from text or audio.2026-05
Kling AI 1.6 Text to Videokling-1.6-text2videoKling AI 1.6 Text-to-Video is a cutting-edge AI tool that transforms text into stunning, lifelike videos. Create professional-quality conte…2025-01
Kling 3.0 Pro Text-to-Videokling-3-pro-text2videoCinematic 1080p videos with realistic audio from text.2026-02
Kling 3.0 Standard Text-to-Videokling-3-standard-text2videoStunning 1080p cinematic videos from simple text prompts.2026-02
Kling O3 Text-to-Videokling-o3-text2video15-second cinematic AI videos with native audio.2026-03
Kling AI Text to Videokling-text2videoKling AI Text-to-Video is a cutting-edge AI tool that transforms text into stunning, lifelike videos. Create professional-quality content e…2024-10
LTX-2-19B I2Vltx-2-19b-i2vSynchronized 4K audio-video generation from images, fast.2026-01
LTX-2-19B T2Vltx-2-19b-t2vSynchronized video and audio from text, multiple input types.2026-01
Minimax AI Directorminimax-ai-directorMinimax video-01-director: Create high-quality videos with control camera movements precisely using text prompts.2025-02
MiniMax Hailuo H3 Text to Videominimax-h3-text-to-videoText-to-video: cinematic 2K clips with native audio.2026-07
Pixverse Text to Videopixverse-text2videoEffortlessly create captivating videos from text with Pixverse text to video AI! Customize style, duration, and more.2025-04
TimelinetimelineUtility node: Timeline. declarative multi-track video compositor2026-07
VEED Avatarsveed-avatarsGenerate UGC-style talking avatar videos from text or audio using 28 stock presenters with realistic lip-sync.2026-07
Google Veo 2veo-2Create stunning, realistic videos with Veo 2, Google's state-of-the-art AI video generation model. Experience enhanced quality & cinematic…2025-03
Google Veo 3veo-3Veo 3 revolutionizes video creation with advanced text-to-video generation and realistic audio synthesis for cinematic content.2025-06
Veo 3 Fastveo-3-fastVeo 3 Fast rapidly creates high-quality, 8-second videos with synchronized audio for diverse content needs.2025-07
Wan 2.2 Text to Video Fastwan-2.2-t2v-fastWan2.2 transforms text and images into high-quality video clips with cinematic flair.2025-08
Wan 2.5 Text to Videowan-2.5-t2vWan2.5-Preview generates synchronized multimedia content, merging text, image, video, and audio seamlessly.2025-09
Wan_2.1 Text to Videowan2.1-t2vCreate visually impressive and feature varied, lifelike motion videos with Wan2.1 using text prompts.2025-03
Wan 2.7 Text to Videowan2.7-t2v1080P cinematic videos with audio sync and multi-shot control.2026-04

Video to audio

ModelSlugDescriptionAdded
Sonilo Video to Audiosonilo-video-to-audioGenerate video-synced music and sound effects from footage.2026-08

Video to image

ModelSlugDescriptionAdded
Start & End Frame Extractorstart-end-frame-extractorExtract first and last frames from any video.2025-10

Video to text

ModelSlugDescriptionAdded
HeyGen Avatar V — Create Avatarheygen-avatar-v-createTrain a Digital Twin avatar from reference video.2026-05

Voice

ModelSlugDescriptionAdded
Kling Create Voicekling-create-voiceClone any voice from a single audio sample.2026-02

On this page