Model Catalog
Every model currently on the Segmind AI Gateway — 672 models with their API slugs, grouped by task. Regenerated automatically from the live catalog.
Every model on the AI Gateway, with the slug you pass
to https://api.segmind.com/v1/<slug> (or /v2 for async) and to
segmind.run() in the Python SDK — 672
models, regenerated automatically from the live catalog.
Checking availability programmatically? This page is a snapshot; the
authoritative source is the live catalog itself. Call
segmind.models.list() or GET
https://api.spotprod.segmind.com/inference-model-information/list — each
entry carries slug, is_depreciated, and the full parameters_schema.
An AI assistant answering "does Segmind support model X?" should prefer that
call over any static list, this one included.
Browsing as a human? segmind.com/models is the searchable catalog with previews, parameters, and pricing.
Recently added
Model slugs for recently added on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
sarvam-bulbul-v3-tts — Text-to-speech in 11 Indian languages with 37 voices. (added 2026-09)
p-video-edit — Edit video from a text prompt, keep original motion. (added 2026-09)
gemini-omni-1.1 — Text-to-video with synchronized native audio, up to 4K. (added 2026-08)
gemini-omni-1.1-video-extend — Extend short video clips into longer seamless scenes. (added 2026-08)
gemini-omni-1.1-video-edit — Edit videos with a text prompt, subject preserved. (added 2026-08)
lyria-3-pro — Full-length text-to-music songs with vocals and lyrics. (added 2026-08)
lyria-3 — Generate 30-second songs with vocals from text or images. (added 2026-08)
gemini-3.7-flash — Fast multimodal LLM for coding, agents, and long-document analysis. (added 2026-08)
kokoro-82m — Text-to-speech with 54 multilingual voices. (added 2026-08)
whisper-large-v3 — Transcribe speech-to-text in 99 languages with timestamps. (added 2026-08)
wan3.0-video — Generate 30-second 1080p video with native audio. (added 2026-08)
wan2.6-i2v-flash — Animate photos into 15-second 1080p video with native audio. (added 2026-08)
grok-imagine-image-2 — Text-to-image and image editing with crisp, legible text. (added 2026-08)
ltx-2.5-pro — Generate 1080p video with native audio and multi-shot scenes. (added 2026-08)
ltx-2.5-fast — Text-to-video and image-to-video with native audio, up to 4K. (added 2026-08)
qwen-image-3 — Generate and edit legible in-image text, up to 2K. (added 2026-08)
seedream-5-pro-layer-decomposition — Split any image into editable transparent PNG layers. (added 2026-08)
bria-extract-object — Extract any named object into a transparent PNG cutout. (added 2026-08)
seedance-2.5 — Generate cinematic multi-shot AI videos up to 30 seconds with synchronized native audio from text, images, or… (added 2026-08)
flux-3-draft-enhance — Upscale AI video drafts to Full-HD with native audio. (added 2026-08)
flux-3-extend-video — Extend clips into seamless video continuations with synchronized audio. (added 2026-08)
flux-3-image-to-video — Animate images into 20-second clips with synchronized native audio. (added 2026-08)
flux-3-text-to-video — Cinematic text-to-video with native lip-synced audio, up to 20s. (added 2026-08)
qwen-3.8-max — Multimodal reasoning and agentic coding with 1M-token context. (added 2026-08)
grok-imagine-video-1.5-reference-to-video — Character-consistent video from up to 7 reference images. (added 2026-08)this is first creation
Model slugs for this is first creation on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
Image-shakil — This is Demo creationAudio & speech synthesis
Model slugs for audio & speech synthesis on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
ace-step-music — ACE-Step generates high-quality music rapidly, enhancing the creative process for developers and artists worl…
chatterbox-tts — Chatterbox transforms text into rich, natural speech with adjustable emotional expressiveness for diverse app…
chatterbox-turbo-tts — Ultra-fast, human-quality TTS with emotional expression.
dia — Dia by Nari Labs is an advanced open-weights TTS model that brings scripts to life with natural speech, emoti…
dubbing — Instantly dubs audio and video into 29 languages while preserving each speaker's original voice.
elevenlabs-dialogue — Immersive, emotionally expressive multi-speaker audio dialogue.
gemini-2.5-flash-tts — Fast, lifelike text-to-speech with expressive emotional tones.
gemini-2.5-pro-tts — Human-like speech synthesis with rich expressive emotional depth.
gemini-3.1-flash-tts — Expressive, controllable TTS with 70+ language support.
grok-tts — Convert text to speech in 20 languages with five voices.
kokoro-82m — Text-to-speech with 54 multilingual voices.
lyria-2 — Lyria 2 by Google DeepMind is an advanced model that generates high-fidelity 48kHz stereo instrumental music…
lyria-3 — Generate 30-second songs with vocals from text or images.
lyria-3-pro — Full-length text-to-music songs with vocals and lyrics.
meta-musicgen-medium — MusicGen: Transform text into music with AI. Create unique, high-quality audio from simple descriptions. Expe…
myshell-tts — MyShell's Voice Cloning and Text to Speech - Transform your audio content with realistic, personalized voices…
openvoice — OpenVoice is a versatile voice cloning model that supports multiple languages and offers precise tone replica…
orpheus-3b-0.1 — Orpheus TTS is an open-source text-to-speech (TTS) system powered by the Llama 3B language model, designed fo…
sam-audio-large — Isolate any described sound from mixed audio tracks.
sarvam-bulbul-v3-tts — Text-to-speech in 11 Indian languages with 37 voices.
seed-audio-1.0 — Generate full audio scenes: dialogue, music, effects, voice cloning.
seed-speech-tts — Natural multilingual text-to-speech and voiceovers from text.
sonilo-text-to-audio — Commercial-safe music and sound effects from text prompts.
sound-generation — Eleven Labs' Sound Generation API provides a robust development tool for programmatically generating audio co…
tts-eleven-labs — ElevenLabs TTS transforms text into captivating, human-like speech for diverse applications.
veena-max-tts — VeenaMAX transforms text into expressive, real-time speech across multiple Indian languages for seamless comm…
veena-tts — Veena transforms text into high-fidelity, expressive speech in Hindi and English for real-time applications.Audio to audio
Model slugs for audio to audio on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
elevenlabs-audio-isolation — Extract clear speech from noisy audio and video.
sts-eleven-labs — Eleven Labs Speech-to-Speech offers AI-powered voice conversion for content creators, media professionals, an…Image editing & transformation
Model slugs for image editing & transformation on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
ai-product-photo-editor — AI Product Photo Editor leverages advanced image-based ML techniques to generate high-quality product visuals…
ai-product-photography — Elevate your product imagery with our AI-powered photography model. Create stunning, professional-quality pho…
alle-v2 — IDM+Faceswap
aura-flow — Largest completely open sourced flow-based generation model that is capable of text-to-image generation
automatic-mask-generator — Automatic Mask Generator is a powerful tool that automates the creation of precise masks for inpainting
become-image — Turn any image of a face into artwork using Stable Diffusion Controlnet and IPAdapter
bg-removal — This model removes the background image from any image
bg-removal-v2 — This model removes the background image from any image
bria-blur-background — Bria AI Image Editing API v2 enables precise and context-aware image manipulation for stunning visual outcome…
bria-enhance-image — Bria AI creates precise, high-quality image enhancements and manipulations for diverse creative applications.
bria-erase-foreground — Seamlessly removes foreground subjects and regenerates backgrounds for flawless image editing.
bria-eraser — AI object removal with seamless context-aware inpainting.
bria-expand-image — Bria Expand enables precise image manipulation and enhancement with generative AI, trained exclusively on lic…
bria-extract-object — Extract any named object into a transparent PNG cutout.
bria-fibo-image-edit — Edit images with structured JSON, masks, and multi-image references.
bria-gen-fill — Bria AI enables precise generative image editing for seamless creative enhancements and transformations.
bria-increase-resolution — Seamlessly upscale and manipulate images while preserving the highest fidelity and safety standards.
bria-lifestyle-shot-by-image — Transforms ordinary product images into stunning, marketing-ready visuals for eCommerce success.
bria-lifestyle-shot-by-text — Transform isolated product images into dynamic lifestyle scenes with AI-driven contextual realism.
bria-product-cutout — Automates precise product cutouts and background removal for professional eCommerce imagery at scale.
bria-product-packshot — Transform product photos into professional, market-ready images with intelligent enhancements and background…
bria-product-shadow — Bria Product Shadow enhances product images with realistic shadows for professional eCommerce presentations.
bria-remove-background — Effortlessly extract backgrounds with unmatched precision, powered by models trained exclusively on licensed…
bria-replace-background — Transform images through advanced background editing and generative content creation for diverse applications.
caricature-style — Transform everyday photos into lively, whimsical caricature illustrations that highlight individual features…
clarity-upscaler — High resolution creative image Upscaler and Enhancer. A free Magnific alternative.
clarityai-creative-upscaler — Creative image upscaling with fine detail enhancement.
clarityai-crystal-upscaler — Upscale images up to 200x with enhanced detail and vibrancy.
clarityai-flux-upscaler — Transform low-resolution images into stunning high-quality visuals.
codeformer — CodeFormer is a robust face restoration algorithm for old photos or AI-generated faces.
consistent-character — Create images of a given character in different poses
consistent-character-with-pose — Create images of a given character in different poses
esrgan — ERGAN is an Image Super-Resolution (upscaler) model that enhances images with stunning, high-quality upscalin…
expression-editor — Expression Editor uses reference images to accurately generate new images with desired expressions. Perfect f…
face-detailer — Restore characters' faces to their original glory with Face Detailer. Enhance facial details, eliminate disto…
face-to-many — Turn a face into 3D, emoji, pixel art, video game, claymation or toy
face-to-sticker — Turn a face into a sticker
faceswap-comic — FaceSwap Comic v1 is an AI-powered face swapping model designed to blend real faces into illustrated or carto…
faceswap-moonfrog-v3 — Take a picture/gif and replace the face in it with a face of your choice. You only need one image of the desi…
faceswap-v2 — Take a picture/gif and replace the face in it with a face of your choice. You only need one image of the desi…
faceswap-v3 — Face Swap V3 is a cutting-edge tool that empowers you to seamlessly swap faces in images. With customizable f…
faceswap-v3-multifaceswap — Faceswap V3 Multifaceswap enables realistic face swapping in images, preserving lighting and expressions for…
faceswap-v4 — Segmind FaceSwap v4 enables fast and precise face or head swapping between images with customizable options f…
faceswap-v5 — Ultra-fast face and head swapping in images.
flux-2-flex — Consistent-style photorealistic images using reference inputs.
flux-2-klein-4b — Sub-second photorealistic image generation and editing.
flux-2-klein-9b — Ultra-fast photorealistic image generation on consumer GPUs.
flux-2-max — Photorealistic images with maximum consistency and fine detail.
flux-2-pro — High-quality photorealistic images with cross-output consistency.
flux-canny-dev — Open-weight edge-guided image generation. Control structure and composition using Canny edge detection.
flux-canny-pro — Professional edge-guided image generation. Control structure and composition using Canny edge detection
flux-controlnet — Flux ControlNets is a collection of models that gives you precise control over image generation. By integrati…
flux-depth-dev — Open-weight depth-aware image generation. Edit images while preserving spatial relationships.
flux-depth-pro — Professional depth-aware image generation. Edit images while preserving spatial relationships.
flux-fill-dev — Open-weight inpainting model for editing and extending images. Guidance-distilled from FLUX.1 Fill Dev
flux-fill-pro — Professional inpainting and outpainting model with state-of-the-art performance. Edit or extend images with n…
flux-img2img — Flux Image-To-Image model by Black Forest Labs is an advanced deep learning tool designed for transforming im…
flux-inpaint — Flux Inpainting is a powerful image editing tool designed to effortlessly edit and enhance your images. It's…
flux-ipadapter — Flux IP Adapter is a cutting-edge AI model that lets you to create stunning, customized images. With its adva…
flux-kontext-dev — FLUX.1 Kontext [dev] creates coherent and editable images by integrating text and visual cues for iterative d…
flux-kontext-max — FLUX.1 Kontext [max] transforms textual descriptions into stunning, high-fidelity images with seamless typogr…
flux-kontext-pro — FLUX.1 Kontext Pro transforms text prompts into high-quality, customized images with remarkable efficiency an…
flux-krea-dev — FLUX.1 Krea generates stunning, photorealistic images with fine-tuned aesthetic control for diverse creative…
flux-pulid — Flux PuLID: Customize AI-generated images with your unique identity. Seamlessly integrate faces into text-to-…
flux-redux-dev — Open-weight image variation model. Create new versions while preserving key elements of your original.
flux-redux-schnell — Fast, efficient image variation model for rapid iteration and experimentation.
focus-outpaint — Fooocus Outpainting transforms ordinary images into extraordinary works of art by seamlessly expanding their…
fooocus — Fooocus enables high-quality image generation effortlessly, combining the best of Stable Diffusion and Midjou…
gpt-image-1-edit — Edit and compose images using natural language with GPT Image 1 Edit, OpenAI’s powerful inpainting and multi-…
gpt-image-1-edit-mini — Affordable text-driven image generation and editing.
gpt-image-1.5-edit — Precise image editing via natural language instructions.
grok-imagine-image — Text-to-image generation and editing, up to 2K resolution.
grok-imagine-image-2 — Text-to-image and image editing with crisp, legible text.
heygen-generate-look — Change avatar outfits and backgrounds while keeping the same face.
hidream-l1-fast — HiDream-I1 is a next-generation, open-source image generative foundation model designed for text-to-image syn…
higgsfield-soul-2 — Generate fashion-editorial photorealistic photos from text or reference.
higgsfield-text2image-soul — SOUL AI transforms text into stunning, customizable visuals with unparalleled style control and precision.
hyperswap-image-faceswap-by-facefusion-labs — High-quality face swapping built for real production workflows.
ic-light — Prompts to auto-magically relight your images.
icon-overlay — icon overlay
ideogram-2a-img-2-img — Ideogram Image to Image: Transform your images with ease! Enhance, modify, or create entirely new visuals usi…
ideogram-3-reframe — Ideogram 3.0's Reframe effortlessly adapts images to diverse formats, enhancing visual content creation for a…
ideogram-3-remix — Ideogram 3 Remix enables versatile image transformation, enhancing creativity through customizable design ite…
ideogram-3-replace-background — Effortlessly replace backgrounds in images, enhancing visual storytelling and creativity with precision and s…
ideogram-character — Achieve perfect character consistency across multiple generations from a single reference image.
ideogram-img-2-img — Ideogram Image to Image: Transform your images with ease! Enhance, modify, or create entirely new visuals usi…
ideogram-reframe — Transform your images with Ideogram Reframe! Easily reframe square images to your chosen resolution.
ideogram-turbo-img-2-img — Transform images instantly with Ideogram Turbo Image to Image! Fast AI for quick edits & creative remixes.
ideogram-v4-remix — Restyle any image into posters with legible in-image text.
idm-vton — Best-in-class clothing virtual try on in the wild
illusion-diffusion-hq — Monster Labs QrCode ControlNet on top of SD Realistic Vision v5.1
image-01 — Generate high-fidelity images from text with precise control & stunning quality with Minimax Image-01.
infinite-you — InfiniteYou generates high-fidelity portraits preserving identity while aligning with creative text prompts.
inpaint-mask-maker — Real-Time Open-Vocabulary Object Detection
insta-depth — InstantID aims to generate customized images with various poses or styles from only a single reference ID ima…
instantid — InstantID aims to generate customized images with various poses or styles from only a single reference ID ima…
ip-sdxl-depth — IP Adapter Depth XL is built on the SDXL framework. This model integrates the IP Adapter and Depth preprocess…
kling-3-image2image — Transform images into photorealistic, production-ready visuals.
kling-o1 — Text-to-video creation with precise AI-driven motion control.
kolors — Kolors is a cutting-edge text-to-image model that bridges language and visual art. Transform your textual ide…
luma-uni-1 — Reasoning-first text-to-image and natural-language image editing.
luma-uni-1-max — Generate and edit images from plain-text instructions.
magic-eraser — LaMA Object Removal- AI Magic Eraser
material-transfer — Transfer a material from an image to a subject
multi-image-kontext-max — FLUX.1 Kontext [max] creates stunning, photorealistic images from text prompts and input images seamlessly.
multi-image-kontext-pro — Transform text into stunning, professional-grade images with precise editing capabilities.
nano-banana-2 — Fast photorealistic images — ideal for marketing and ads.
nano-banana-2-lite — Generate and edit 1K images in about four seconds.
nano-banana-pro — High-fidelity images with accurate multilingual text rendering.
nomos-upscaler — This upscaling model is ideal for enhancing amateur to professional photos, excelling with subjects like cats…
ominicontrol — OminiControl is an innovative framework that optimizes Diffusion Transformer models for versatile image gener…
omni-zero — Omni-Zero: A diffusion pipeline for zero-shot stylized portrait creation.
p-image-edit — Multi-image editing with AI-guided precision and control.
p-image-try-on — Dress photos in multiple garments with photorealistic virtual try-on.
pulid-base — Novel tuning-free ID customization method for text-to-image generation.
qwen-image-edit — Transform images effortlessly through semantic context and pixel-perfect appearance changes.
qwen-image-edit-fast — Qwen-Image-Edit enables precise bilingual image editing for seamless localization and professional content cr…
qwen-image-edit-plus — Multi-image editing with precise text-guided transformations.
qwen-image-edit-plus-add-people — Generate realistic multi-character scenes with natural interactions.
qwen-image-edit-plus-blend-it — Product placement into backgrounds with precise lighting match.
qwen-image-edit-plus-eigen-banana — Precise text-guided image transformation and creative editing.
qwen-image-edit-plus-eraser — Remove unwanted objects while preserving realistic backgrounds.
qwen-image-edit-plus-face-to-portrait — Cropped face into full identity-preserving portrait photo.
qwen-image-edit-plus-group-photo — Merge individual portraits into realistic group photos.
qwen-image-edit-plus-multi-lora — Multi-image editing with superior identity and style control.
qwen-image-edit-plus-multiple-angle — Transform image perspective with natural language prompts.
qwen-image-edit-plus-next-scene — Create cinematic sequences with seamless visual continuity.
qwen-image-edit-plus-product-photography — Transform white-background products into immersive lifestyle scenes.
qwen-image-edit-plus-relight — Advanced image relighting using natural language prompts.
qwen-image-edit-plus-remove-lighting — Remove artificial lighting effects and restore natural tones.
qwen-image-edit-plus-texture-apply — Apply precise textures to images using natural language.
qwen-image-edit-plus-texture-extract — Extract seamless, tileable textures from photographs.
runway-gen4-image — Runway's Gen-4 Image API enables precise, multimodal image generation for innovative creative and technical a…
sam-img2img — The Segment Anything Model (SAM) produces high quality object masks from input prompts such as points or boxe…
sam-v2-image — SAM v2, the next-gen segmentation model from Meta AI, revolutionizes computer vision. Building on SAM's succe…
sam3-image — Precise object segmentation and tracking in images.
sd1.5-controlnet-canny — This model corresponds to the ControlNet conditioned on Canny edges.
sd1.5-controlnet-depth — This model corresponds to the ControlNet conditioned on Depth estimation.
sd1.5-controlnet-openpose — This model corresponds to the ControlNet conditioned on Human Pose Estimation.
sd1.5-controlnet-scribble — This model corresponds to the ControlNet conditioned on Scribble images.
sd1.5-controlnet-softedge — This model corresponds to the ControlNet conditioned on Soft Edge.
sd1.5-img2img — This model uses diffusion-denoising mechanism as first proposed by SDEdit, Stable Diffusion is used for text-…
sd1.5-outpaint — Stable Diffusion Outpainting can extend any image in any direction
sd2.1-faceswapper — Take a picture/gif and replace the face in it with a face of your choice. You only need one image of the desi…
sd3-med-canny — Stable Diffusion 3 (SD3) Medium Canny ControlNet uses Canny edge detection to provide fine-grained control ov…
sd3-med-pose — Stable Diffusion 3 (SD3) Pose ControlNet is a large generative image model tailored for generating images bas…
sd3-med-tile — SD3 Medium Tile ControlNet is a large generative image model designed for generating detailed images based on…
sdxl-controlnet — SDXL ControlNet gives unprecedented control over text-to-image generation. SDXL ControlNet models Introduces…
sdxl-img2img — SDXL Img2Img is used for text-guided image-to-image translation. This model uses the weights from Stable Diff…
sdxl-openpose — This model leverages SDXL to generate the images with ControlNet conditioned on Human Pose Estimation.
seedream-4 — Seedream 4.0 generates high-resolution, professional-grade visuals with superior text rendering for impactful…
seedream-4.5 — Photorealistic image generation with precise text understanding.
seedream-5-pro-layer-decomposition — Split any image into editable transparent PNG layers.
seedream-v5-lite-image-to-image — Transform images intelligently with detailed text prompts.
seg-swap — Swap Objects Instantly. The Segmind SegSwap v0.1 model enables dynamic and precise image editing by allowing…
segfit-v1.1 — Segmind's Fashion and Immersive Try-on model. SegFIT offers effortless AI virtual try-on from just a product…
segfit-v1.2 — SegFit v1.2 creates hyper-realistic virtual try-on images, transforming fashion retail engagement and convers…
segfit-v1.3 — SegFit v1.3 enables hyper-realistic virtual try-ons, enhancing online fashion retail experiences without phys…
segmind-relighting — Prompts to auto-magically relight your images.
segmind-relighting-v2 — Transform images with customizable, photorealistic lighting for unparalleled visual creativity and authentici…
segmind-scenecraft-v01 — SceneCraft transforms plain or existing product images into visually rich, photorealistic scenes. Whether sta…
skin-contrast-upscaler — Enhances skin detail in images while preserving background quality for professional photography and art.
smart-banner-resizer — Recompose one image into multiple ad and banner sizes.
ssd-depth — This model leverages SSD-1B to generate the images with ControlNet conditioned on Depth Estimation
ssd-img2img — This model uses SSD-1B to generate images by passing a text prompt and an initial image to condition the gene…
storydiffusion — Story Diffusion turns your written narratives into stunning image sequences.
style-transfer — Style & Composition Transfer with Stable Diffusion IP Adapter
superimpose — Superimpose model lets you to create captivating visuals by seamlessly overlaying one image on top of another…
superimpose-v2 — Superimpose V2 elevates image editing! Seamlessly layer images with background removal, precise positioning,…
supir — SUPIR restores and enhances images to stunning, photo-realistic quality with advanced AI techniques.
text-overlay — Elevate your visuals withText Overlay Model. Easily add customized text to any image, perfect for social medi…
topaz-image-upscale — Topaz Labs image upscale is an industry-leading AI photo upscaler designed to increase the resolution of phot…
transparent-background-maker — Transform your images with Transparent Background Maker. Quickly remove backgrounds using AI technology, supp…
utility-image-mask — Build / refine binary masks from JSON-described shapes (rect/polygon/ellipse). Optional dilate/erode/feather/…
utility-image-transform — Apply an ordered pipeline of resize / crop / rotate / flip in a single call. Replaces the four separate tools.
w2imgsd1.5-img2img — Create beautifully designed words using Segmind’s word to image for your marketing purposesImage generation
Model slugs for image generation on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
675ed494b9-thejagstudio-LordSwaminaryan —
9c54110684-frank-chieng-sdxl_lora_architecture_siheyuan —
acn-0effee052a — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
aio-fg-full-fc1d87eafb — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Aladdin5k-701ba0d580 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
alexV2-2a5661e8ef — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
AlexV2-655afc6166 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
alexV2FastFlux-e99e4b5aa1 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
alice-in-wonderland-ab7255451b — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
ANIME-GEN-V4-9dae59678b — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
anlora-e8d79af2af — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
ap-fc0862b464 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
arlora-9ceba116a8 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
armp-13752d05df — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
artificialguybr-ClayAnimationRedmond — Clay Animation Redmond based on SDXL 1.0, excels at creating mesmerizing clay animation images with unparalle…
Asherflex-c25117ed67 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
awf-66afd16a65 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
awmcn-617fbb4446 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
awp-29909cf89b — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Ayalora-5ff7f8f2fd — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
b0445a2335-ostris-watercolor_style_lora_sdxl —
background-eraser — Background Eraser helps in flawless background removal with exceptional accuracy.
baiqiang-8812bdefb9 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Balloonblowing-d8d4c47115 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
bed-and-chair-combined-8dd2e74338 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
bizhen-3683a8e947 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
bria-fibo-generate — Generate photorealistic images with structured JSON prompt control.
bria-text-to-image — Bria 3.2 AI transforms natural language into stunning visuals for diverse creative applications — with Base,…
bria-text-to-vector-graphics — Bria Vision enables high-quality text-to-image and text-to-vector graphic generation for versatile commercial…
c42bf9a66d-joachimsallstrom-aether-bubbles-foam-lora-for-sdxl —
cf-6ccbc88846 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
cfoly-c65493430d — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
chillpixel-blacklight-makeup-sdxl-lora — Blacklight Makeup SDXL LoRA is fine-tuned to generate makeup designs that are not only visually striking but…
chroma — Chroma is an open-source, 8.9B parameter text-to-image model (based on FLUX.1-schnell) designed for diverse a…
cloth-finetune-flux-91b5cc1278 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
crazy-8eae969efe — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
ctmn-806488193b — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Delibrate_V2-8dbc3ba5a9 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
dianying-549c4441a6 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Dinaone-ExteriorFloor-c404fe7cc8 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
djwx-635a3531b4 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
dminhk-dog-example-sdxl-lora — Dog Example SDXL LoRA, a specialized AI model within the Stable Diffusion XL framework, uniquely trained to e…
doctor-dolittle-0018c78851 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
dslora-b6e9a6e529 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
dw_bedroom_1_5k-71b62c19fb — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
dw_bedroom_2k-559dc9f020 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
egyptian-d9f4de37d4 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
fast-flux-schnell — Fast Flux.1 Schnell by Segmind is an optimized text-to-image model designed for developers needing faster ima…
fb4ab6b705-naphatmanu-sdxl-lora-index-modern-luxury-1 —
felora-94021b7888 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
fg_v2-5a5c3da2b3 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
fg-individual-15-5b442665e9 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
fg15k-4879e3ca32 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
fg5kFull-2fffd1e2b0 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
fgResize5k-b2981fb7b8 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
flux-1.1-pro — Flux Pro 1.1 is a cutting-edge image generation tool offering exceptional speed, quality, and customization.…
flux-1.1-pro-ultra — Create stunning visuals effortlessly with Flux 1.1 Pro Ultra. Experience unparalleled image quality and speed.
flux-dev — Flux Dev is a 12 billion parameter rectified flow transformer capable of generating images from text descript…
flux-dev-finetuned — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
flux-hanurama-22f3d043c0 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
flux-pixar-27334bad28 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
flux-pro — Flux Pro is a state-of-the-art image generation with top of the line prompt following, visual quality, image…
flux-realism-lora — Flux Realism Lora with upscale, developed by XLabs AI is a cutting-edge model designed to generate realistic…
flux-schnell — Flux Schnell is a state-of-the-art text-to-image generation model engineered for speed and efficiency.
Flux-Turbo-1b4d78be3f — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Flux-Turbo-37e8b8f594 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
fluxanimals-05320abd47 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
fluxdbb-a972e916ef — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
FluxPixar-30c83df57d — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
FluxTurbo-f28792fc23 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
fulora-f3df2e804d — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
fylora-fb68c65109 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Gen-v2-5c9cdcda8e — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Ghibil-27c901962a — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
goofyai-cyborg_style_xl — Cyborg Style SDXL specializes in generating cyborg-themed artwork based on science fiction and futuristic aes…
gpt-image-1 — Create high-quality AI-generated images from text prompts using OpenAI's GPT Image 1 model. Ideal for product…
gpt-image-1-mini — High-quality image generation from text, fast and affordable.
gpt-image-1.5 — Stunning photorealistic images with exceptional instruction-following.
gpt-image-2 — Generate photorealistic images with legible multilingual text and 2K output.
gvrizzo-lora-sdxl-notion-illustration —
haosc-610e769fa8 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
ideogram-2a-txt-2-img — Create captivating designs, realistic images & innovative logos with Ideogram 2a text-to-image.
ideogram-3 — Ideogram 3.0 revolutionizes content creation with photorealistic text-to-image generation and diverse aesthet…
ideogram-4 — Generate 2K posters and logos with accurate text rendering.
ideogram-turbo-txt-2-img — Create stunning images in seconds with Ideogram Turbo Text to Image. Fast AI model for quick ideation & text…
ideogram-txt-2-img — Ideogram Text to Image: Turn your ideas into stunning visuals instantly with this powerful AI tool. Create ca…
ideogram-v4-fast — Generate posters and logos with accurate in-image text.
imagen — Imagen 3 is Google DeepMind's highest quality text-to-image model. Generates detailed images with enhanced li…
imagen-4 — Imagen 4 is Google’s most advanced AI image generation model, creating detailed, photorealistic or abstract i…
imagen-4-fast — Fast photorealistic image generation for bulk and iteration.
imagen-4-ultra — Photorealistic images with native 2K resolution and precise text.
jayashri710-sdxl-khuze-nocrop-1e-4-1200 —
jayashri710-sdxl-lora-khuze-1e-4-1200-512x512images —
jdq-396bb5a145 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Jenny1-6546baecae — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
juggernaut-lightning-flux — Juggernaut Lightning Flux: Blazing fast (<300ms!) & powerful inference with enhanced visuals.
juggernaut-pro-flux — Juggernaut Pro FLUX: Create stunningly realistic AI images with unprecedented detail and sharpness.
jzzs-a0fc97d5df — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
jzzsd-ef9fd76525 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
KappaNeuro-punk-collage — Punk Collage Model offers a unique way to create digital collages that resonate with the punk culture's raw e…
kchoi-lora-sdxl-watercolor —
KellyTest-66fbd687a2 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
kling-3-text2image — Photorealistic, print-ready images from text prompts.
liangnv-b2e39c7ec4 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
M-EdenRock-bed-2d45decb6a — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
M-HubbaArmChair-and-M-RicochetFabric-bed-18fed2c922 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
M-HubbaArmChair-fff004d426 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
M-RicochetFabric-bed-9180ccb8f7 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Marimekko-mrk3903-1e8cab847c — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
meizhuang-2ae6bc0009 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
minimaxir-sdxl-ugly-sonic-lora — SDXL Ugly Sonic LoRA excels at generating quirky and iconic version of one of the most beloved movie characte…
minimaxir-sdxl-wrong-lora — SDXL Wrong LoRA is engineered with a focus on delivering images of higher detail, color saturation and vibran…
mnsh-86fb5b8ef4 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
mnzs-b5e1ea1473 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
multi-lora — SDXL Model with multiple LoRa loading support.
nano-banana — Gemini Image Editor preserves authentic subject identity while enabling seamless image editing and manipulati…
Narinder-dfac11ecfb — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
nav_007_Krishna-6fce6aad5b — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
nerijs-lego-minifig-xl — LEGO Minifig XL is designed to generate LEGO images and excels in creating detailed and accurate representati…
Norod78-SDXL-StickerSheet-Lora — SDXL StickerSheet LoRA is expertly fine-tuned on a comprehensive collection of sticker images, enabling it to…
ostris-crayon_style_lora_sdxl — Crayon Style - SDXL LoRA is a unique model designed to convert any text prompt into a vibrant, crayon-style d…
ostris-ikea-instructions-lora-sdxl — Ikea Instructions LoRA SDXL model is fine-tuned on IKEA diagrams and specializes in generating clear, concise…
ostris-stained-glass-style-sdxl — Stained Glass Style SDXL is trained extensively on diverse stained glass images and can replicate the essence…
p-image — p-image generates high-quality images from text prompts in seconds, optimizing for speed and fidelity.
p-image-ideogram — Sub-second text-to-image with legible in-image text.
pixar-3d-5ab56ddd76 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
playzippyIramayanakids-7f162e1446 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
prabhas-5f173e43a4 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
quanshen-7d0d5c4bdb — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
qwen-image — Qwen-Image revolutionizes image generation and editing with seamless multilingual text integration and photor…
qwen-image-2512 — Photorealistic image generation with precise text description following.
qwen-image-3 — Generate and edit legible in-image text, up to 2K.
qwen-image-fast — Qwen-Image expertly generates stunning images with complex text integration, especially for Chinese typograph…
ra100-sdxl-lora-lower-decks-aesthetic — SDXL LoRA Lower Decks Aesthetic model, inspired by the unique style of “Star Trek: Lower Decks.” generates ar…
Rama-558043a138 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
ramayanaMulti5k-69c1cc4ae9 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
ramsrigouthamg-lora-dog-SSD-1B — LoRA Dog SSD-1B specializes in generating photorealistic images of dogs
recraft-v3 — Recraft V3, the latest iteration of Recraft AI, offers a significant advancement in AI-driven image generatio…
recraft-v3-svg — Recraft V3 SVG generates high-quality, customizable vector graphics with precision and ease. Perfect for logo…
reve-2 — Generate and edit 4K images with sharp in-image text.
rxmx-324bd4c1e2 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
SanaAI-542b8dbe01 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
SBIC-Stone-fbcc5912c5 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
sclora-274da0a25f — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
sd1.5-cyberrealistic — The most versatile photorealistic model that blends various models to achieve the amazing realistic images.
sd1.5-edgeofrealism — This model corresponds to the Stable Diffusion Edge of Realism checkpoint for detailed images at the cost of…
sd1.5-epicrealism — This model corresponds to the Stable Diffusion Epic Realism checkpoint for detailed images at the cost of a s…
sd1.5-juggernaut — The most versatile photorealistic model that blends various models to achieve the amazing realistic images.
sd1.5-realisticvision — This model corresponds to the Stable Diffusion Realistic Vision checkpoint for detailed images at the cost of…
sd1.5-reliberate — This model corresponds to the Stable Diffusion Reliberate checkpoint for detailed images at the cost of a sup…
sdxl1.0-colossus-lightning — Colossus Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024…
sdxl1.0-dreamshaper — The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software.
sdxl1.0-dreamshaper-lightning — DreamShaper Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1…
sdxl1.0-dyanvis-lightning — Dynavis Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024p…
sdxl1.0-juggernaut-lightning — Juggernaut Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 10…
sdxl1.0-newreality-lightning — NewReality Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 10…
sdxl1.0-nightvis-lightning — NightVis Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024…
sdxl1.0-protovis-lightning — ProtoVision Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1…
sdxl1.0-realdream-lightning — RealDream is a sophisticated image generation model utilizing SDXL Lightning architecture. It creates incredi…
sdxl1.0-realdream-pony-v9 — Real Dream Pony V9 is an advanced image generation model based on the Stable Diffusion XL (SDXL) architecture…
sdxl1.0-realism-lightning — Realism Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024p…
sdxl1.0-realvis — The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software.
sdxl1.0-realvis-lightning — Realvis Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024p…
sdxl1.0-samaritan-3d — Samaritan 3D XL leverages the robust capabilities of the SDXL framework, ensuring high-quality, detailed 3D c…
sdxl1.0-samaritan-lightning — Samaritan Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 102…
sdxl1.0-timeless — The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software.
sdxl1.0-txt2img — The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software
sdxl1.0-wildcard-lightning — WildCard Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024…
sdxl1.0-zavychroma — The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software.
seedream-5-pro — Region-precise image editing with native multilingual text.
seedream-v5-lite-text-to-image — Fast, affordable instruction-following image generation.
segmind-vega — The Segmind-Vega Model is a distilled version of the Stable Diffusion XL (SDXL), offering a remarkable 70% re…
segmind-vega-rt-v1 — Segmind-VegaRT a distilled consistency adapter for Segmind-Vega that allows to reduce the number of inference…
Selfie-Aesthetic-f9c0f409c5 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Selfie-Aesthetic1-9e21b5a57f — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
sf-31e1b1ee3c — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
sheb-b899ceaa9f — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
sheb-b9150b5b83 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Shravya_ai-f44dd1e647 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Simple_Vector_Flux — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
siria-e3fb0ee32f — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Sita-b382eb5c53 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
softpasty-flux-dev-a4edbd379c — Flux Schnell is a state-of-the-art text-to-image generation model engineered for speed and efficiency.
SpeedyMary-b361f2fd1d — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
ssd-1b — SSD-1B efficiently generates high-quality, diverse images from text prompts in real-time.
stable-diffusion-3-medium-txt2img — Stable Diffusion is a type of latent diffusion model that can generate images from text. It was created by a…
stable-diffusion-3.5-large-txt2img — Stable Diffusion 3.5 Large offers exceptional customizability, efficient performance on consumer hardware, an…
stable-diffusion-3.5-turbo-txt2img — Stable Diffusion 3.5 Turbo offers exceptional customizability, efficient performance on consumer hardware, an…
Test 001 -4b5d788e6c — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
thumbnail-bc5f00ef3c — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
tianmei-e09a3dfa92 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
v6lora-1745486a5e — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
veo-dd69bf026e — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
wan2.7-image — 2K image generation with precise multilingual text rendering.
wan2.7-image-pro — 4K images with chain-of-thought reasoning and multilingual text.
whnm-339c18f951 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
winnie-the-pooh-c3ccfd99b1 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
wxns-9cf0075aee — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
xse-1f865b3596 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
xz-c445d830b2 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
yujia-0d97d92ca5 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
yunxi-aac5d8354b — Flux Schnell is a state-of-the-art text-to-image generation model engineered for speed and efficiency.
z-image-turbo — Photorealistic images in under one second, bilingual text.
zh-66dcb61456 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
zscj-7c64578707 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
zsjz-7ca8eff909 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
zsnh-7634404088 — Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptionsImage t o image
Model slugs for image t o image on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
sd3-med-img2img — Stable Diffusion 3 Medium image-to-image is a cutting-edge AI tool that uses advanced image-to-image technolo…Image to data
Model slugs for image to data on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
utility-image-metadata — Read image metadata: dimensions, format, EXIF (with GPS decoded to decimal), ICC profile name, raw XMP. Retur…Image to video
Model slugs for image to video on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
ai-face-swap — AI Face Swap: Effortlessly replace faces online. Fine-tune swaps with advanced controls for age, gender, and…
cog-video-5b-i2v — CogVideoX image-to-video is a cutting-edge AI model that converts static images into dynamic, high-quality vi…
flux-3-image-to-video — Animate images into 20-second clips with synchronized native audio.
grok-imagine-video-1.5-image-to-video — Animate a still image into 1080p video with synced audio.
grok-imagine-video-1.5-preview — Image-to-video with native synchronized audio, up to 720p.
grok-imagine-video-1.5-reference-to-video — Character-consistent video from up to 7 reference images.
hailuo-02-fast — Transform any static image into a captivating, high-quality video clip effortlessly.
hailuo-2.3 — Hyper-realistic videos from text with fluid character motion.
hailuo-2.3-fast — Professional-quality videos from text and images at speed.
hallo — Hallo lets you create portrait videos from single images.
happyhorse — Cinematic 1080p text-to-video with native audio and lip-sync.
happyhorse-1.1 — Generate cinematic video with synchronized native audio and multilingual lip-sync from text, an image, or ref…
heygen-avatar-iv — Single photo into a lifelike talking avatar video.
higgsfield-image2video — Transform static images into dynamic, motion-rich videos with unparalleled control and creative depth.
higgsfield-speech2video — Transform images and audio into dynamic, lip-synced videos for engaging digital content.
infinite-talk — Full-body animation from images synchronized perfectly to audio.
kling-1.6-image2video — Kling AI 1.6 Image-to-Video is a powerful AI tool that transforms static images into captivating, animated vi…
kling-2 — Kling 2.0 is an advanced AI video generator (5 and 10 seconds) that creates cinematic, dynamic videos from te…
kling-2.1 — Kling 2.1 offers hyper-realistic video generation with improved motion, sharper 1080p visuals, and instant re…
kling-2.5-turbo — Kling AI 2.5 Turbo generates fluid, cinematic videos from text and images, enhancing content creation and sto…
kling-2.6 — Still images into immersive cinematic videos with synchronized audio.
kling-3-pro-image2video — Animated 1080p videos from images with dynamic motion.
kling-3-standard-image2video — Controlled cinematic 1080p videos from starting images.
kling-bloombloom — Kling AI transforms text and images into dynamic, high-quality video content with realistic motion and sound.
kling-dizzydizzy — Kling DizzyDizzy transforms static content into dynamic, high-resolution videos, enhancing engagement and sto…
kling-expansion — Unleash dynamic visuals with Kling Expansion! Effortlessly inflate and stretch elements for surreal and capti…
kling-fuzzyfuzzy — Transform your photos instantly into adorable, plush-toy-like visuals with Kling fuzzyfuzzy effect.
kling-heart-gesture — Express affection visually with Kling AI's heart gesture effect! Input two portraits and instantly create hea…
kling-hug — Create heartwarming videos instantly with Kling hug effect! Generate tender embracing animations.
kling-image2video — Kling AI Image-to-Video is a powerful AI tool that transforms static images into captivating, animated videos…
kling-kiss — Create a heartfelt video in seconds with Kling kiss effect! Input two portraits and instantly generate a kiss…
kling-o1-image-to-video — Physics-driven animations from images for creative storytelling.
kling-o1-reference-image-to-video — Identity-preserving videos from static images with character reference.
kling-o3-image2video — Images to cinematic videos with precise motion control.
kling-squish — Transform your visuals with Kling AI squish effect! Easily compress and distort images/videos for playful, ex…
kling-v1-pro-ai-avatar — Dynamic AI avatars with synchronized speech from image.
kling-v1-standard-ai-avatar — Lifelike AI avatars with precise lip-sync for presentations.
kling-v2-pro-avatar — Talking avatar videos from image and audio, high quality.
kling-v2-standard-avatar — Lifelike video avatars with precise lip synchronization.
live-portrait — Live Portrait animates static images using a reference driving video through implicit key point based framewo…
live-portrait-video-to-video — Experience the magic of Live Portrait’s Video-to-Video Model! Transform your static images into dynamic video…
ltx-2-fast — Fast, high-quality text-to-video generation by Lightricks.
ltx-2-pro — High-quality video generation with advanced motion control.
ltx-2.5-fast — Text-to-video and image-to-video with native audio, up to 4K.
ltx-2.5-pro — Generate 1080p video with native audio and multi-shot scenes.
ltx-video — LTX-Video is the first DiT-based video generation model capable of generating high-quality videos in real-tim…
luma-ray-3-2 — Cinematic text-to-video and image-to-video clips up to 1080p.
minimax-ai — With Video-01 by MiniMax, create high-definition videos at 720p resolution and 25fps, featuring cinematic cam…
minimax-h3-image-to-video — Animate a still image into 2K video up to 15s.
minimax-h3-reference-to-video — Keep characters and products consistent in 2K reference-to-video.
minimax-hailuo-2 — Generate breathtaking 1080P cinematic videos from text or images with ultra-realistic motion and physics.
modify-video — Transform videos seamlessly with high-fidelity generative edits while preserving original actor performances.
motionctrl-svd — Motion Control SVD is an innovative deep learning framework that breathes life into static images. By intelli…
muscle-surge — Instantly add muscle and strength to your videos with Pixverse Muscle Surge effect!
p-video-animate — Transfer video motion and audio onto any still image.
p-video-avatar — Animate any portrait into a lip-synced talking avatar.
p-video-replace — Swap on-screen video characters while preserving motion and audio.
pixverse-4.5-effects — PixVerse 4.5 transforms photos and text into stunning animated videos for impactful storytelling and marketin…
pixverse-4.5-video — Pixverse 4.5 transforms static images and text into dynamic, engaging videos for captivating social media con…
pixverse-5-extend — Seamlessly extend and continue AI-generated videos.
pixverse-5-transition — Seamless AI-generated video transitions between scenes.
pixverse-5-video — Cinematic videos from text and images with photorealism.
pixverse-image2video — Animate your photos effortlessly with Pixverse Image to Video AI! Upload, add motion prompts and styles.
pixverse-mimic — Transfer motion from reference videos onto still images.
pixverse-v6 — 15-second AI videos with native audio and cinematic controls.
ray-flash-2-720p — Generate stunning 720p videos from text with the Luma ray-flash-2-720p model. Faster & cheaper than Ray 2, of…
runway-gen3-alphaturbo — Runway Gen-3 AlphaTurbo is a cutting-edge AI tool that transforms static images into dynamic videos with exce…
runway-gen4-turbo — Generate videos faster and cheaper with Runway Gen-4 Turbo! Create high-quality text, image, and combined vid…
sadtalker — Audio-based Lip Synchronization for Talking Head Video
scail — Professional character animations from reference images.
seedance-1.0-pro-fast — Cinematic videos from text and images at ultra speed.
seedance-1.5-pro — Synchronized video and audio generation for dynamic storytelling.
seedance-2.0 — Cinematic AI videos with native audio and multi-shot narratives.
seedance-2.0-fast — Professional-grade video creation model with native audio, similar to SeeDance 2.0 but faster and cheaper.
seedance-2.0-mini — Fast text-to-video and image-to-video with synchronized audio.
seedance-2.5 — Generate cinematic multi-shot AI videos up to 30 seconds with synchronized native audio from text, images, or…
seedance-pro — Seedance Pro transforms text and images into engaging 720p dynamic videos with cinematic storytelling.
sora-2 — Stunning dynamic videos from detailed text descriptions.
sora-2-pro — Cinematic-quality videos from text with temporal consistency.
svd — Takes image as input and returns a video.
tooncrafter — Create videos from illustrated input images
v-express — V-Express lets you create portrait videos from single images.
veed-fabric-1.0 — Animate any image into a realistic talking video, lip-synced to your audio or generated from a text script.
veo-2-image2video — Discover Google Veo 2, an AI-powered image-to-video model with 4K resolution, realistic motion, and cinematic…
veo-3.1 — Static images into high-quality videos with synchronized audio.
veo-3.1-fast — Transforms static images into dynamic 1080p videos with synchronized audio and natural motion.
video-effects — Transform your videos with diverse video effects. Start creating captivating videos today.
video-faceswap-by-facefusion-labs — Realistic face swapping in videos from a single image.
video-frame-interpolation — FILM synthesizes smooth, high-quality intermediate frames for fluid motion in videos with significant movemen…
video-stitch — Revolutionize your video editing with the Video Stitch Model. Seamlessly stitch clips, add captivating audio,…
video-tryon — Video Tryon is Segmind’s next-generation AI video model for instant virtual try-on, allowing users to visuali…
video-tryon-v2 — Video Tryon is Segmind’s next-generation AI video model for instant virtual try-on, allowing users to visuali…
video-watermark-remover — Remove watermarks from any video instantly with AI.
videofaceswap — Video Faceswap is a powerful tool for creators, filmmakers, and meme enthusiasts. With this innovative techno…
wan-2.2-i2v-fast — Transforms simple text prompts into breathtaking cinematic-quality videos in minutes.
wan-2.2-i2v-flash — Convert a single image into a coherent dynamic video.
wan-2.5-i2v — Wan2.5-Preview creates stunning, high-resolution videos with flawless audio synchronization from multiple inp…
wan-2.6-i2v — Transform images into high-quality videos with audio sync.
wan-animate — Animate characters and replace video subjects seamlessly.
wan2.1-i2v-720p — Create high-quality 720p videos with excellent visual quality and a broad spectrum of motion from static imag…
wan2.6-i2v-flash — Animate photos into 15-second 1080p video with native audio.
wan2.7-i2v — Animate any image into cinematic 1080P video with audio.
wan2.7-r2v — Character-consistent multi-subject videos from reference images.
wan3.0-video — Generate 30-second 1080p video with native audio.
warmth-of-jesus — Experience the viral "Warmth of Jesus" effect on PixVerse! Transform your images into heartwarming videos of…Image to3d
Model slugs for image to3d on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
hunyuan-3d-2 — Hunyuan3D 2.0 enables the creation of high-quality 3D models with intricate details. Produce assets that are…
hunyuan3d-2.1 — Transform 2D images into photorealistic, high-fidelity 3D assets effortlessly.
hunyuan3d-2mv — Hunyuan3D-2mv is finetuned from Hunyuan3D-2 to support multiview controlled shape generation.
sam-3d-body — Reconstruct 3D human body meshes from a single photo.
sam-3d-objects — Single 2D image into detailed 3D object models.Image understanding
Model slugs for image understanding on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
bria-fibo-generate-structured-prompt — Convert complex inputs into structured JSON prompts for generation.
bria-mask-generator — Bria AI Get Masks automatically generates accurate object masks for advanced image editing and enhancement.
bria-prompt-enhancer — Bria AI generates high-quality, commercially safe images tailored to diverse creative needs.
google-translate — Translate effortlessly with the powerful Google Translation AI model.
ideogram-describe — Ideogram describe can effortlessly generate detailed prompts from images. Perfect for refining creations or r…
image-converter — Convert images between formats instantly.
image-resizer — Resize images to any dimension quickly and precisely.
llava-v1.6 — LLaVa translates images into text descriptions & captions.
nsfw-checker — Detect NSFW and other inappropriate content in images. Returns a boolean has_nsfw_concepts flag, an overall N…
sam-v21-hiera-large — Meta's next-gen segmentation model for images and video.
video-speed-change — Speed up or slow down any video precisely.Inpainting
Model slugs for inpainting on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
focus-inpaint — Fooocus Inpainting is a powerful image generation model that allows you to selectively edit and enhance image…
inpaint-auto — This model is capable of generating photo-realistic images given any text input, with the extra capability of…
sd1.5-inpainting — Stable Diffusion Inpainting is a latent text-to-image diffusion model capable of generating photo-realistic i…
sdxl-inpaint — This model is capable of generating photo-realistic images given any text input, with the extra capability of…
try-on-diffusion — Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-onLanguage models (LLMs)
Model slugs for language models (llms) on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
claude-4-sonnet — Advanced coding and multi-step agentic reasoning model.
claude-4.5-sonnet — Claude Sonnet 4.5 empowers developers with advanced coding and reasoning for complex software solutions.
claude-opus-4.7 — Anthropic's most capable AI model excelling at agentic coding, complex reasoning, and high-resolution vision…
deepseek-chat — DeepSeek V3 combines cutting-edge AI technology with practical usability. Featuring a 671B parameter architec…
deepseek-reasoner — DeepSeek-R1 is a cutting-edge AI reasoning model that combines reinforcement learning with supervised fine-tu…
gemini-2.5-flash — Multimodal AI with transparent reasoning, fast and affordable.
gemini-2.5-flash-lite — Fastest Gemini 2.5 model for high-volume text and vision tasks.
gemini-2.5-pro — Complex multimodal reasoning across diverse inputs and formats.
gemini-3-flash — Frontier-class reasoning and multimodal AI at scale.
gemini-3-pro — Autonomous multimodal AI for complex reasoning and coding.
gemini-3.1-flash-lite — Ultra-fast, affordable LLM for high-volume AI pipelines.
gemini-3.1-pro — Frontier reasoning across text, images, video, and code.
gemini-3.7-flash — Fast multimodal LLM for coding, agents, and long-document analysis.
glm-5.2 — 1M-token open-weight LLM for long-horizon coding.
gpt-4 — GPT-4 outperforms both previous large language models and as of 2023, most state-of-the-art systems (which of…
gpt-4-turbo — GPT-4 outperforms both previous large language models and as of 2023, most state-of-the-art systems (which of…
gpt-4o — GPT-4o (“o” for “omni”) is our most advanced model. It is multimodal (accepting text or image inputs and outp…
gpt-5 — GPT-5 automates complex coding tasks with integrated tools for seamless software development and deployment.
gpt-5-mini — Rapid high-quality AI across text, images, and files.
gpt-5-nano — Ultra-fast LLM responses for real-time AI applications.
gpt-5.1 — Precise code review and developer workflow assistant.
gpt-5.2 — Advanced reasoning with multimodal input for precise tasks.
gpt-5.4 — Most powerful GPT for frontier reasoning and multimodal tasks.
gpt-5.4-mini — Fastest efficient model for coding and computer-use tasks.
gpt-5.4-nano — Flagship-class AI for classification and extraction tasks.
gpt-5.5 — Frontier reasoning and coding with 1M-token context window.
kimi-k2-instruct-0905 — Deep contextual understanding and complex code generation.
llama-v3-8b-instruct — Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrain…
llama-v3p1-70b-instruct — Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrain…
llama-v3p1-8b-instruct — Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrain…
llama4-maverick-instruct-basic — Llama 4 Maverick Instruct Basic is a 400B parameter powerhouse with 128 experts for unparalleled text and ima…
llama4-scout-instruct-basic — Unlock powerful multimodal AI with Llama 4 Scout basic, a 17 billion active parameters model offering leading…
minimax-m3 — Reason over 1M-token context for coding and agents.
mixtral-8x22b-instruct — Mistral MoE 8x22B Instruct v0.1 model with Sparse Mixture of Experts. Fine tuned for instruction following.
nemotron-3-ultra — 1M-token reasoning for coding agents and deep research.
o3 — Frontier reasoning model for complex coding, math, and science.
o3-mini — Cost-efficient reasoning model for coding, math, and science.
o4-mini — OpenAI o4-mini enhances decision-making by processing text and images with advanced reasoning capabilities.
qvq-max — Chain-of-thought visual reasoning for math, charts, and diagrams.
qwen-3.8-max — Multimodal reasoning and agentic coding with 1M-token context.
qwen-flash — Fastest low-cost LLM with 1M context for high-volume tasks.
qwen-plus — Mid-tier 1M context LLM for summarization and content tasks.
qwen2-vl-72b-instruct — Qwen2-VL-72B-Instruct is a state-of-the-art multimodal model excelling in image and video understanding, with…
qwen3-coder-flash — High-volume code generation with 1M token context window.
qwen3-coder-plus — Generates, debugs, and refactors entire codebases efficiently.
qwen3-max — 1T-parameter LLM with hybrid reasoning and 262K context.
qwen3-vl-flash — Fast, affordable vision-language model with 262K context OCR.
qwen3-vl-plus — Powerful visual QA and document analysis from images.
qwen3.5-flash — Fast multimodal AI processing text, images, and video affordably.
qwen3.5-plus — Multimodal 1M context AI for image, video, and text.
qwq-plus — Deep chain-of-thought reasoning for math, code, and logic.Text to embed
Model slugs for text to embed on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
gemini-embedding-001 — MTEB #1 text embeddings for RAG, search, and clustering.
gemini-embedding-2 — Natively multimodal embeddings — text, image, audio, video and PDF mapped into one vector space, with 8 task-…
text-embedding-3-large — Text-embedding-3-large is a robust language model by OpenAI designed for generating high-dimensional text emb…
text-embedding-3-small — Text-embedding-3-small is a compact and efficient model developed for generating high-quality text embeddings…Transcription
Model slugs for transcription on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
eleven-labs-transcript — Transcribe audio to accurate text in 99 languages with speaker diarization and word-level timestamps.
elevenlabs-dialogue-with-timestamps — Multi-speaker dialogue with expressive timestamps included.
elevenlabs-forced-alignment — Precise audio-text synchronization with word-level timestamps.
elevenlabs-voice-clone — Hyper-realistic voice cloning from short audio samples.
elevenlabs-voice-design — Generate unique synthetic voices without audio samples.
tts-elevenlabs-with-timestamps — Emotionally expressive TTS with word-level timestamp output.
whisper-large-v3 — Transcribe speech-to-text in 99 languages with timestamps.Video editing
Model slugs for video editing on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
bria-erase-video — Remove unwanted objects from videos while preserving audio.
bria-increase-video-resolution — Transform your videos with AI-powered upscaling and seamless background removal for professional quality.
bria-video-background-removal-3.0 — Remove video backgrounds with flicker-free, transparent alpha output.
esrgan-video-upscaler — ESRGAN Video Upscaler: Experience sharper, clearer 4k videos with ESRGAN. This AI-powered video upscaler boos…
flux-3-draft-enhance — Upscale AI video drafts to Full-HD with native audio.
flux-3-extend-video — Extend clips into seamless video continuations with synchronized audio.
gemini-omni-1.1-video-edit — Edit videos with a text prompt, subject preserved.
gemini-omni-1.1-video-extend — Extend short video clips into longer seamless scenes.
heygen-video-translate — Translate videos to multiple languages with natural lip-sync.
kling-2.6-pro-motion-control — Transfer motion from videos to animate custom characters.
kling-2.6-standard-motion-control — Precise motion transfer from reference videos to characters.
kling-o1-video-to-video-edit — Edit any video with precise natural language commands.
kling-o1-video-to-video-reference — Video style transfer using reference character images.
kling-o3-video2video-edit — Text-based video editor — swap backgrounds, characters, restyle scenes.
kling-o3-video2video-reference — Swap characters and restyle videos using reference images.
ltx-retake-video — Precise segment-level video edits maintaining full scene continuity.
multi-video-merge — Merge multiple videos into a single combined output.
opus-clips-from-video — Turn long videos into captioned vertical shorts.
p-video-edit — Edit video from a text prompt, keep original motion.
pixverse-lipsync — PixVerse Lipsync expertly synchronizes lip movements to audio for flawless video content creation.
runway-gen4-aleph — Runway Aleph revolutionizes video editing with intelligent automation for seamless object and environment man…
sam-v2-video — SAM v2 Video by Meta AI, allows promptable segmentation of objects in videos.
sam3-video — Real-time video segmentation and multi-object tracking.
sonilo-video-to-video — Add frame-synced AI music and sound effects to video.
sync.so-lipsync-2-pro — Lipsync-2-Pro seamlessly synchronizes lips in videos for instant, high-quality multilingual content creation.
sync.so-react-1 — Edit video actors' emotions with realistic re-expression.
topaz-video-upscale — Topaz Video AI upscales, enhances, denoises, stabilizes, and increases frame rates in video footage, transfor…
veed-2-lipsync — Dub talking-head videos with emotion-matched lip-sync.
veed-lipsync — Re-syncs the lips of any talking-head video to a new speech audio track for realistic dubbing and localizatio…
veed-subtitles — Automatically transcribes and burns styled, translated subtitles into any video with 30 presets and a single…
veed-video-background-removal — Remove any video's background with no green screen, or cleanly key chroma footage, using AI matting.
video-audio-merge — Effortlessly merge audio and video with our intuitive Video Audio Merge model. Create stunning multimedia con…
video-captioner — With Video Captioner create accurate, customizable subtitles for your videos effortlessly.
video-concatenate — Merge videos with custom layouts, spacing, and audio.
video-editor-agent — General-purpose AI media agent: describe the transformation in natural language and it runs ffmpeg in a sandb…
video-loop — Effortlessly loop videos for engaging social media & storytelling with our Video Loop.
video-slicer — Video Slicer
video-split — Utility node: Video Split. 1->N split; returns videos[] array
wan2.7-videoedit — Edit existing videos precisely using natural language text instructions.Video generation
Model slugs for video generation on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
cog-video-5b-t2v — CogVideo is a groundbreaking AI model that turns text into high-quality videos. Create realistic scenes, anim…
flux-3-text-to-video — Cinematic text-to-video with native lip-synced audio, up to 20s.
gemini-omni-1.1 — Text-to-video with synchronized native audio, up to 4K.
gemini-omni-flash — Text-to-video and image-to-video with synchronized native audio.
grok-imagine-video — Text-to-video and image-to-video with native synchronized audio.
grok-imagine-video-1.5-text-to-video — Text-to-video clips up to 1080p with native synchronized audio.
heygen-avatar-v — Studio-quality talking-avatar videos from text or audio.
kling-1.6-text2video — Kling AI 1.6 Text-to-Video is a cutting-edge AI tool that transforms text into stunning, lifelike videos. Cre…
kling-3-pro-text2video — Cinematic 1080p videos with realistic audio from text.
kling-3-standard-text2video — Stunning 1080p cinematic videos from simple text prompts.
kling-o3-text2video — 15-second cinematic AI videos with native audio.
kling-text2video — Kling AI Text-to-Video is a cutting-edge AI tool that transforms text into stunning, lifelike videos. Create…
ltx-2-19b-i2v — Synchronized 4K audio-video generation from images, fast.
ltx-2-19b-t2v — Synchronized video and audio from text, multiple input types.
minimax-ai-director — Minimax video-01-director: Create high-quality videos with control camera movements precisely using text prom…
minimax-h3-text-to-video — Text-to-video: cinematic 2K clips with native audio.
pixverse-text2video — Effortlessly create captivating videos from text with Pixverse text to video AI! Customize style, duration, a…
timeline — Utility node: Timeline. declarative multi-track video compositor
veed-avatars — Generate UGC-style talking avatar videos from text or audio using 28 stock presenters with realistic lip-sync.
veo-2 — Create stunning, realistic videos with Veo 2, Google's state-of-the-art AI video generation model. Experience…
veo-3 — Veo 3 revolutionizes video creation with advanced text-to-video generation and realistic audio synthesis for…
veo-3-fast — Veo 3 Fast rapidly creates high-quality, 8-second videos with synchronized audio for diverse content needs.
wan-2.2-t2v-fast — Wan2.2 transforms text and images into high-quality video clips with cinematic flair.
wan-2.5-t2v — Wan2.5-Preview generates synchronized multimedia content, merging text, image, video, and audio seamlessly.
wan2.1-t2v — Create visually impressive and feature varied, lifelike motion videos with Wan2.1 using text prompts.
wan2.7-t2v — 1080P cinematic videos with audio sync and multi-shot control.Video to audio
Model slugs for video to audio on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
sonilo-video-to-audio — Generate video-synced music and sound effects from footage.Video to image
Model slugs for video to image on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
start-end-frame-extractor — Extract first and last frames from any video.Video to text
Model slugs for video to text on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
heygen-avatar-v-create — Train a Digital Twin avatar from reference video.Voice
Model slugs for voice on the Segmind API — pass one to https://api.segmind.com/v1/<slug> or segmind.run("<slug>", ...):
kling-create-voice — Clone any voice from a single audio sample.