Model Catalog
Explore our comprehensive collection of AI models for image, video, audio generation, and LLM.
Featured Models
Top picks from each category - the best models for getting started
FLUX.1 Schnell
Black Forest Labs
Ultra-fast image generation model optimized for speed. Generates high-quality images in just 1-2 seconds, perfect for real-time applications and rapid prototyping.
MiniMax Hailuo 2.3
MiniMax
Realistic human motion video generation with advanced character consistency and natural movement.
MiniMax Speech-02-Turbo
MiniMax
Low-latency text-to-speech model with multilingual support, emotional voice control, and 300+ voice options.
GPT-4o
OpenAI
OpenAI's flagship multimodal model. Industry-leading performance in reasoning, coding, and creative tasks with native vision capabilities and structured output support.
How to Choose the Right Model
Need Speed?
flux-schnell, Gemini Flash
Need Quality?
FLUX Pro, Kling Pro, Claude
Budget-Friendly?
flux-schnell, MiniMax, Gemini
Most Versatile?
GPT-4o, Claude, FLUX Dev
Image Generation Models
Generate stunning images with FLUX, Stable Diffusion, and more
Background Remover
Featured851 Labs
The most-used background remover on Replicate (27M+ runs). Removes backgrounds with soft alpha or hard segmentation, supports reverse mode (remove foreground), custom background types, and transparent PNG output — all at a very low price.
BiRefNet
BiRefNet
State-of-the-art open-source background removal. BiRefNet's high-fidelity dichotomous image segmentation delivers excellent edge quality on hair, fur, and fine details.
BLIP
Salesforce
Salesforce BLIP (173M+ runs) — image captioning, visual question answering, and image-text matching in one model. The classic choice for bulk captioning at 1 credit per image.
Change Haircut
Black Forest Labs
Change anyone's hairstyle and hair color from a single photo, powered by FLUX.1 Kontext [pro]. Choose from 90+ hairstyles and 30 hair colors — or let 'Random' surprise you — while keeping the face untouched.
Clarity Upscaler
Clarity
The famous creative upscaler (30M+ runs). Instead of just enlarging, it re-imagines detail while upscaling — with controllable creativity, resemblance, prompt guidance, and tiled diffusion for high scale factors.
CLIP Features
CLIP
CLIP ViT-L/14 embeddings for text AND images (163M+ runs) — puts both in the same vector space for cross-modal search, image dedup, and zero-shot classification. 1 credit per run.
CodeFormer
CodeFormer
Robust face restoration for old photos and AI-generated faces (54M+ runs). Its signature fidelity dial balances restoration quality against staying true to the original face, with Real-ESRGAN background enhancement built in.
ControlNet Scribble
ControlNet
The classic sketch-to-image model (38M+ runs) — turn any scribble or line drawing into a detailed image guided by your prompt. Draw the composition, describe the content.
Nano Banana (Edit)
Dedicated edit endpoint for Nano Banana, Google's Gemini 2.5 Flash-based image model. Pass input image URLs to perform conversational editing with character consistency and multi-image fusion.
Nano Banana 2 (Edit)
Edit endpoint for Nano Banana 2, built on Gemini 3.1 Flash Image. Adds resolution control (1K/2K/4K), Google Search grounding, and thinking mode while preserving conversational editing and multi-image fusion.
Nano Banana Pro (Edit)
Edit endpoint for Nano Banana Pro built on Gemini 3 Pro. Professional-grade controls, legible multilingual typography, real-time grounding via Google Search, and resolution up to 2K for editing.
Face to Many
fofr
Turn a face photo into 6 fun styles (15M+ runs) — 3D, Emoji, Video game, Pixels, Clay, or Toy. The viral avatar generator behind countless profile-picture apps.
Face to Sticker
fofr
Turn any face photo into a fun die-cut sticker (1.6M+ runs). InstantID keeps the likeness while IP-Adapter controls take the artwork from faithful caricature to loose cartoon — with an optional 2x upscale for print quality.
Florence-2 Large
Microsoft
Microsoft's Florence-2 all-in-one vision model — captioning, object detection, phrase grounding, OCR, and segmentation in a single API. Pick a task, optionally add text input, done.
FLUX.2 Klein 4B
Black Forest Labs
Very fast image generation and editing model. 4-step distilled, sub-second inference for production and near real-time applications.
FLUX 2 Flex
Black Forest Labs
Maximum-quality FLUX model supporting up to 10 reference images and advanced typography. The most capable model for complex, multi-reference creative projects.
FLUX 2 Max
Black Forest Labs
The highest fidelity image model from Black Forest Labs. Best-in-class prompt following and the most consistent editing in the FLUX.2 lineup — preserves colors, lighting, faces, text, and objects across edits with up to 8 reference images.
FLUX 2 Pro
Black Forest Labs
Professional-grade FLUX 2 with high-quality editing and up to 8 reference image support. Excellent balance of quality, speed, and creative control.
FLUX.2 Dev
Black Forest Labs
Development version of FLUX.2 with image editing capabilities and reference image support. Ideal for iterative design workflows and experimentation.
FLUX 1.1 Pro
Black Forest Labs
Fast high-quality image generation, an upgrade to FLUX.1 Pro with faster speed and improved quality. Perfect for production workloads requiring both speed and fidelity.
FLUX 1.1 Pro Ultra
Black Forest Labs
FLUX1.1 [pro] in ultra and raw modes. Images are up to 4 megapixels — the highest-resolution tier of the FLUX 1.1 Pro family. Use raw mode for realism.
FLUX Dev
Black Forest Labs
A 12 billion parameter rectified flow transformer capable of generating images from text descriptions, tuned for open, high-quality experimentation.
FLUX Kontext Max
Black Forest Labs
A premium text-based image editing model that delivers maximum performance and improved typography generation for transforming images through natural language prompts.
FLUX Kontext Pro
Black Forest Labs
State-of-the-art text-based image editing model that transforms images through natural language. Excellent for style transfer, object modification, text replacement, background changes, and character consistency.
FLUX.1 Krea [dev]
Krea AI
Photorealistic image generation that specifically avoids the 'AI look', producing natural-looking images indistinguishable from real photographs.
FLUX PuLID
ByteDance
PuLID identity customization on FLUX-dev: generate photorealistic portraits of a specific person from one face photo, with markedly higher fidelity than SDXL-based variants. Tune id_weight and start_step to balance likeness against prompt editability.
FLUX.1 Schnell
Black Forest Labs
Ultra-fast image generation model optimized for speed. Generates high-quality images in just 1-2 seconds, perfect for real-time applications and rapid prototyping.
Runway Gen-4 Image
Runway
Runway's Gen-4 Image model with references: combine up to 3 reference images with @tag mentions in your prompt to keep characters, objects, and locations consistent across every angle and scene, at 720p or 1080p.
Runway Gen-4 Image Turbo
Runway
Gen-4 Image Turbo is 2.5x faster and cheaper than Gen-4 Image, with the same reference-driven API: use 1 to 3 reference images with @tag mentions for consistent characters and objects, at a flat price regardless of resolution.
GFPGAN
Tencent ARC
Tencent ARC's face restoration model — 115M+ runs, the most-run model in the restoration collection. Restores old/blurry photos and fixes faces in AI-generated images.
GPT Image 2
OpenAI
OpenAI's state-of-the-art image generation and editing model with strong instruction following, sharp text rendering, and detailed editing. Quality-based pricing lets you trade off cost vs. fidelity.
GPT Image 1.5
OpenAI
OpenAI's latest image generation model with better instruction following and adherence to prompts, including sharp text rendering and detailed editing.
Grok Imagine Image
xAI
Generate images using xAI's Grok Imagine model. Sibling to Grok Imagine Video, sharing the same underlying Grok Imagine architecture for fast text-to-image generation.
Grounding DINO
Grounding DINO
Text-prompted object detection (39M+ runs) — describe what to find in natural language ('red car, person wearing a hat') and get bounding boxes with confidence scores plus an annotated image.
Hunyuan 3D 3.1
Tencent
Tencent's flagship 3D generation — create high-polygon textured 3D models from a text prompt OR an image, with optional PBR (physically based rendering) materials.
Ideogram V4 Balanced
Ideogram
A middle-ground tier in Ideogram's v4 family, balancing generation speed and output quality. Delivers strong typography and photorealism at a lower cost than the Quality tier.
Ideogram V4 Quality
Ideogram
The highest-fidelity tier of Ideogram's v4 model family, tuned for maximum detail, realism, and typography accuracy. Best suited for final production assets where quality matters more than speed.
Ideogram V4 Turbo
Ideogram
The fastest and cheapest model in Ideogram's v4 family, built for rapid iteration while retaining Ideogram's signature text rendering and style consistency.
Ideogram V3 Turbo
Ideogram
The fastest and cheapest Ideogram v3 tier. V3 creates images with stunning realism, creative designs, and consistent styles.
Ideogram Character
Ideogram
Generate consistent characters from a single reference image. Render the same character in many styles — realistic or fiction — insert them into existing photos with mask inpainting, and rely on Ideogram's signature text rendering for legible signs and typography.
MiniMax Image-01
MiniMax
MiniMax's first image generation model with character reference support: provide a single face photo via subject_reference and generate consistent images of that person across prompts, styles, and aspect ratios — up to 9 images per request.
Topaz Image Upscale
Topaz Labs
Professional-grade upscaling from Topaz Labs, the industry standard for photo enhancement. Five specialized enhance models, up to 6x upscale, subject detection, and optional face enhancement.
Imagen 4
Google's Imagen 4 flagship text-to-image model.
Imagen 4 Fast
A fast version of Imagen 4 for when speed and cost are more important than maximum quality.
Bria Increase Resolution
Bria
Bria's commercially-safe image upscaler (130K+ runs). Increase resolution 2x or 4x with a model trained exclusively on licensed data, preserving alpha transparency — built for enterprise pipelines that require full legal liability coverage.
Krea 2 Large
Krea AI
Krea AI's flagship text-to-image model, focused on photorealistic output with strong prompt adherence. Supports style-reference and moodboard-guided generation for consistent visual direction.
Krea 2 Medium
Krea AI
A lower-cost variant of Krea 2 Large, trading some fidelity for faster and cheaper generation while keeping the same photorealistic style focus.
Microsoft MAI-Image 2.5 Pro
Microsoft
Microsoft's highest-fidelity image model for production-grade text-to-image generation via Fal.AI. Built for hero imagery, detailed compositions, precise text rendering, photorealism, stylized illustration, commercial design, and visually rich concept work.
Moondream2
Moondream
Small but capable vision-language model (14M+ runs) — ask free-form questions about any image and get detailed answers. Efficient VQA for tagging, moderation prep, and rich alt-text.
Multilingual E5 Large
E5
Multilingual text embeddings (74M+ runs) — 1024-dimension vectors across 100 languages including Korean. Pairs perfectly with Core.Today customer databases' vector search (knn_vector).
Nano Banana 2 (Gemini 3.1 Flash Image)
Google's fast image generation model built on Gemini 3.1 Flash Image. The high-efficiency counterpart to Nano Banana Pro — combining Pro-level visual quality with Flash-level speed and pricing. Features conversational editing, multi-image fusion, character consistency, accurate text rendering, and Google Search grounding. Supports up to 14 reference images and resolutions up to 4K.
Nano Banana 2 Lite
Google's lightweight Nano Banana 2 variant built on Gemini 3.1 Flash Image, tuned for faster and cheaper generation. Retains conversational editing, multi-image fusion, and character consistency from the full Nano Banana 2.
Nano Banana
Google Gemini 2.5 Flash-based image generation with multimodal editing capabilities. Fast and versatile for both creation and editing tasks.
Nano Banana Pro (Gemini 3 Pro Image)
Google's state-of-the-art image generation and editing model built on Gemini 3 Pro. Creates detailed visuals with legible text in multiple languages, connects to real-time information from Google Search, and provides professional-grade creative controls. Supports up to 14 reference images and resolutions up to 4K.
NSFW Image Detection
Falcons.ai
The standard NSFW image classifier (127M+ runs) — returns 'normal' or 'nsfw' for any image. An essential, ultra-cheap moderation gate for UGC platforms.
Pruna P-Image
Pruna AI
Pruna AI's distilled text-to-image model optimized for extremely low-cost, high-throughput generation. One of the most-run community models on Replicate (15.6M+ runs) thanks to its speed and price.
Pruna P-Image Upscale
Pruna AI
PrunaAI's high-end image upscaler (575K+ runs) reaching up to 128-megapixel output. Target-megapixel or factor-based control, optional detail/realism enhancement passes, and megapixel-tiered pricing that starts at just 15 credits.
PhotoMaker
Tencent ARC
Generate stylized photos of a person from 1-4 reference photos (9M+ runs). Ten styles including Cinematic, Disney Character, and Digital Art — keep the identity, change everything else.
PhotoMaker Style
Tencent ARC
The stylization-focused variant of PhotoMaker: turn 1-4 photos of a person into paintings, comics, 3D art, and more with stronger style transfer. Pairs with the base PhotoMaker — use this one when style matters more than photorealism.
Professional Headshot
Black Forest Labs
Turn any single photo into a polished professional business headshot, powered by FLUX.1 Kontext [pro]. Pick a background — white, black, gray, neutral, or office — and get a LinkedIn-ready portrait in one step.
Proteus v0.2
Proteus
Popular anime-focused image model (12M+ runs) — high-quality anime and illustration styles with img2img and inpainting support. The go-to for anime avatars and webtoon-style art.
PuLID
ByteDance
ByteDance PuLID: tuning-free identity customization on SDXL. Give it one face photo and a prompt to generate portraits in any scene or style — no training, 4-step fast sampling, and even two-identity blending. Extremely cost-effective at 2 credits per image.
Real-ESRGAN
Real-ESRGAN
The classic image upscaler (94M+ runs). Real-ESRGAN super-resolution up to 10x with optional GFPGAN face enhancement — the go-to default for cleanly enlarging photos and AI images.
Recraft V4
Recraft
Recraft's next-generation text-to-image model, improving on V3's typography, style range, and prompt adherence for production-grade brand and design assets.
Recraft V4 SVG
Recraft
Generates vector graphics directly in SVG format using Recraft V4, ideal for logos, icons, and scalable illustrations that need to stay crisp at any size.
Recraft V3
Recraft
Recraft V3 (code-named red_panda) is a text-to-image model with the ability to generate long texts, and images in a wide list of styles. SOTA in image generation per the Artificial Analysis Text-to-Image Benchmark.
Recraft Crisp Upscale
Recraft
Fast, affordable upscaler from Recraft designed for sharp, crisp results — a single image input with no tuning needed. Great default for UI assets, illustrations, and product images.
Remove Background
Bria AI
AI-powered background removal tool for images. Clean, accurate cutouts for any subject with professional-quality edge detection.
Remove BG
lucataco
One of the most-run background removers on Replicate (17M+ runs). A single-parameter API that strips the background and returns a transparent PNG — at just 1 credit per image, the cheapest cutout on the platform.
Stable Diffusion XL
Stability AI
Stability AI's classic SDXL (85M+ runs) — the battle-tested text-to-image model with img2img, inpainting, refiner, and LoRA support. A dependable workhorse with a huge ecosystem.
SDXL Lightning 4-step
ByteDance
ByteDance's 4-step SDXL Lightning — the most-run model on Replicate (1B+ runs). Near-instant 1024px image generation at one of the lowest prices in the catalog.
Seedream 5 Lite
ByteDance
Seedream 5.0 lite: image generation with built-in reasoning, example-based editing, and deep domain knowledge. Supports multi-reference generation with up to 14 images and sequential batch generation.
Seedream 5 Pro
ByteDance
ByteDance's flagship Seedream 5.0 Pro image generation model, with built-in reasoning, precise instruction following, and multi-reference support for up to 10 images. Offers 1K and 2K resolution output with higher fidelity than Seedream 5 Lite.
Seedream 4.5
ByteDance
Upgraded ByteDance image model with stronger spatial understanding and world knowledge. Supports single/multi-reference image-to-image editing and sequential (multi-image) generation.
Seedream 4.0
ByteDance
ByteDance's latest image generation model with exceptional prompt understanding and creative capabilities.
Text Extract OCR
OCR
Simple, massively-used OCR (91M+ runs) — extracts text from an image with a single input and returns plain text. Great default for receipts, screenshots, and scanned documents at just 1 credit.
TRELLIS
Microsoft
Microsoft's TRELLIS image-to-3D (836K+ runs — the most-used 3D model on Replicate). Turns one or more images into a textured GLB 3D asset, with turntable render videos and optional Gaussian PLY.
Google Upscaler
Google's image upscaler (770K+ runs). Upscale images 2x or 4x using generative AI while preserving natural detail, with adjustable output compression — a simple, reliable enhancer at a flat price.
Z-Image Turbo
Pruna AI
Pruna's ultra-fast Z-Image Turbo (48M+ runs). Megapixel-priced image generation up to 2048x2048 — pay exactly for the resolution you generate, with excellent price/performance for high-volume use.
Video Generation Models
Create AI videos with Kling, MiniMax, and cutting-edge models
Add Watermark
FeaturedFullJourney
Add a text watermark to any video — simple, fast brand protection for generated or user content at 2 credits per video.
CogVLM2 Video
CogVLM
Video understanding and captioning — ask free-form questions about a video and get detailed answers about actions, scenes, and content. Great for video search indexing and moderation prep.
Google Gemini Omni Flash
Google Gemini Omni Flash text-to-video via Fal.AI. Generates a video directly from a descriptive text prompt — pacing and audio (dialogue, background music) are controlled in the prompt itself, in 16:9 or 9:16 at 3-10 second durations.
Gen-4.5
Runway
Runway's Gen-4.5 model, offering state-of-the-art video motion quality, prompt adherence, and visual fidelity for text-to-video and image-to-video generation.
Grok Imagine Video 1.5
xAI
xAI's Grok Imagine Video 1.5 preview: image-to-video generation with synchronized audio. An upgraded successor to Grok Imagine Video with flat per-second pricing regardless of resolution.
Grok Imagine Video
xAI
Generate videos using xAI's Grok Imagine Video model. Supports text-to-video, image-to-video, and editing an existing short video clip.
MiniMax Hailuo 2.3
MiniMax
Realistic human motion video generation with advanced character consistency and natural movement.
MiniMax Hailuo 2.3 Fast
MiniMax
Lower-latency version of Hailuo 2.3 optimized for faster generation while maintaining good quality for human motion videos.
Happy Horse 1.1
Alibaba
Alibaba's Happy Horse 1.1 generates videos from text, animates a single image, or builds a video from multiple reference images. Supports 720p and 1080p, 3-15 second durations, and five aspect ratios.
MiniMax Hailuo-03 (H3) Image to Video
MiniMax
MiniMax Hailuo-03 (H3) image-to-video via Fal.AI. 2K video generation from a first-frame image, 5-15 second duration, native audio, with optional first-to-last keyframe control via end_image_url.
Kling v3 Omni Video
Kuaishou
Kling Video 3.0 Omni: a unified multimodal video model that generates and edits video from text, images, reference images, and existing video. Combines text-to-video, image-to-video, reference-based generation, and video editing with native audio and multi-shot control.
Kling v3 Video
Kuaishou
Kling Video 3.0: Kuaishou's flagship text/image-to-video model generating cinematic videos up to 15 seconds with multi-shot control, native audio, start/end frame images, and a dedicated 4K mode.
Kling v2.6
Kuaishou
Kling 2.6 Pro: top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation. Audio generation is enabled by default.
Kling 2.5 Turbo Pro
Kuaishou
Cinematic-grade video generation with enhanced motion and scene coherence. Top-tier Kling model for professional output.
Kling v2.1
Kuaishou
Kling v2.1 with 720p/1080p support and frame transition capabilities for smooth, high-quality video generation.
LatentSync
ByteDance
ByteDance's open-source lipsync — re-syncs a video's mouth movements to any audio track using latent diffusion. State-of-the-art open lipsync quality for dubbing and localization.
MMAudio
MMAudio
Add AI-generated sound to any video (5M+ runs) — synthesizes synchronized audio (ambience, effects, foley) from the video content and an optional text prompt. The perfect finisher for silent AI-generated clips.
Pruna P-Video
Pruna AI
PrunaAI's fast video generator with a built-in draft mode for rapid creative iteration. Text-to-video, image-to-video, and audio-conditioned generation in a single endpoint, with clips up to 20 seconds — one of the longest durations on the platform.
Pruna P-Video Avatar
Pruna AI
Pruna AI's talking-avatar video model — animates a single input image to speak provided text (voice_script) or an uploaded audio track, with selectable Gemini-family preset voices, language, and visual delivery prompt.
PixVerse V6
PixVerse
PixVerse's flagship video generation model. Generates cinematic videos with synchronized audio, multi-shot sequences with scene transitions, astonishing physics, and precise camera control at up to 1080p.
PixVerse V5
PixVerse
Advanced video generation with special effects capabilities and anime-optimized output, supporting multiple visual styles.
Luma Ray 3.2
Luma
Luma's flagship Ray video model. Text-to-video and keyframe (start/end image) generation with optional HDR-encoded output and professional EXR export.
Luma Ray Flash 2 720p
Luma
Luma's Ray Flash 2 generates 5 or 9 second 720p videos faster and cheaper than Ray 2. Supports keyframe control via start and end images, seamless loops, and a rich library of camera motion concepts — the first Luma model on the platform.
Real-ESRGAN Video
Real-ESRGAN
Video upscaling with Real-ESRGAN — enhance videos to FHD, 2K, or 4K frame by frame. The go-to open-source video upscaler for old footage and AI-generated clips.
SadTalker
SadTalker
Talking-head video from a single photo and an audio track — animates the face with natural head motion and eye blinks, with optional GFPGAN face enhancement.
Seedance 2.0
ByteDance
ByteDance's next-generation multimodal video model with native synchronized audio. Combines up to 9 reference images, 3 videos, and 3 audio files in a single generation for character-consistent, lip-synced video creation, editing, and extension.
Seedance 2.0 Fast
ByteDance
A faster, cheaper variant of Seedance 2.0 for quicker video generation with multimodal reference inputs (up to 9 images, 3 videos, 3 audios) and native audio, at 480p or 720p.
Seedance 2.0 Mini
ByteDance
Lighter, cheaper variant of ByteDance's Seedance 2.0. Native audio, multimodal reference inputs (images/videos/audio), text-to-video and image-to-video, capped at 720p (no 1080p/4K tier).
Seedance 1 Lite
ByteDance
A lightweight ByteDance video generation model offering text-to-video and image-to-video support for 4-12 second videos at 480p, 720p, or 1080p resolution.
Seedance 1 Pro
ByteDance
A pro version of Seedance that offers text-to-video and image-to-video support for 2-12 second videos, at 480p, 720p, and 1080p resolution.
Seedance 1 Pro Fast
ByteDance
ByteDance's cinematic video generation model with fast generation speed and professional output quality.
OpenAI Sora 2
OpenAI
OpenAI's video generation model with realistic physics simulation and audio generation capabilities, producing highly coherent videos.
OpenAI Sora 2 Pro
OpenAI
OpenAI's most advanced synced-audio video generation model. The premium tier of Sora 2 with higher fidelity, up to 1024p resolution, and image-to-video via an input reference frame.
MiniMax Hailuo-03 (H3) Text to Video
MiniMax
MiniMax Hailuo-03 (H3) text-to-video via Fal.AI. State-of-the-art 2K video generation from a text prompt, 5-15 second duration, native audio, and wide aspect-ratio support (21:9 through 9:16).
Veed Lipsync v2
Veed
Veed Lipsync v2 via Fal.AI. Replaces a source video's mouth movements to articulate a new audio track — takes any source video + a new audio track and produces a lip-synced output video.
Google Veo 3.1
Google's state-of-the-art video generation model with built-in audio generation, producing cinematic-quality videos with synchronized sound.
Google Veo 3.1 Fast
Fast version of Veo 3.1 with audio generation, optimized for speed while maintaining high quality output.
Video Utils
FFmpeg
FFmpeg-powered video utilities (20M+ runs) — convert to mp4/gif, extract audio as mp3, or dump zipped frames, all with one task parameter. The Swiss-army knife for media pipelines.
Wan 2.7 I2V
Alibaba
Alibaba Wan 2.7 image-to-video model. Animates a first frame (and optional last frame or continuation clip) with audio synchronization, up to 15 seconds, at 720p or 1080p.
Wan 2.7 T2V
Alibaba
Alibaba Wan 2.7 text-to-video model. Supports up to 15 seconds, audio synchronization for voice/music, multilingual prompts, and prompt expansion, at 720p or 1080p.
Wan 2.5 I2V
Alibaba
Image-to-video model with lip sync support, animating still images into realistic videos with natural motion.
Wan 2.5 I2V Fast
Alibaba
Fast image-to-video variant of Wan 2.5, optimized for rapid generation of animated videos from still images.
Wan 2.5 T2V
Alibaba
Text-to-video model with audio synchronization support, producing high-quality videos from text prompts with natural motion.
Wan 2.5 T2V Fast
Alibaba
Fast text-to-video generation variant of Wan 2.5, optimized for speed with good quality output.
Wan 2.2 I2V Fast
Alibaba
A very fast and cheap PrunaAI-optimized version of Alibaba's Wan 2.2 A14B image-to-video model. Turns a single still image plus a prompt into a short animated clip at 480p or 720p, with an optional frame-interpolation pass for smoother motion.
Wan 2.2 T2V Fast
Alibaba
A very fast and cheap PrunaAI-optimized version of Alibaba's Wan 2.2 A14B text-to-video model. Generates short clips at 480p or 720p directly from a text prompt, with 30 FPS frame interpolation enabled by default — the text-to-video sibling of Wan 2.2 I2V Fast.
Audio & TTS Models
Text-to-speech, voice cloning, and audio generation
Bark
FeaturedSuno
Suno's text-prompted generative audio model. Produces speech with nonverbal sounds like [laughs] and [sighs], plus music and sound effects, in 100+ speaker presets across 13 languages. Returns audio plus an optional .npz history file for voice continuity.
Chatterbox
Resemble AI
Resemble AI's production-grade open-source TTS with unique emotion exaggeration control and instant voice cloning from a short reference audio. MIT-licensed and benchmarked against leading closed-source systems.
Chatterbox Multilingual
Resemble AI
Chatterbox open-source TTS in 23 languages with instant voice cloning and emotion exaggeration control. Max 300 characters per request.
Chatterbox Turbo
Resemble AI
Resemble AI's fastest open-source TTS without sacrificing quality. 20 pre-made voices, paralinguistic tags like [sigh] and [chuckle], and optional instant voice cloning from 5s+ reference audio. Max 500 characters per request.
Gemini 3.1 Flash TTS
Google Gemini 3.1 Flash native-audio TTS via Replicate. Style instructions (separate from the spoken text) plus 30 preset voices and 40+ language/locale codes, with inline markup tags for expressive delivery ([sigh], [whispering], [shouting], etc.).
Incredibly Fast Whisper
Whisper
Whisper large-v3 optimized for speed (38M+ runs) — transcribes roughly 150 minutes of audio in under 100 seconds using batched inference. Chunk-level or word-level timestamps.
Suno Music V5.5
Suno
Suno's latest model, with the highest audio quality and richest arrangements. Creates two complete songs from a text description, or from exact lyrics with style and title in custom mode. Returns audio, cover image, and metadata for each track.
Suno Music V5
Suno
Suno V5 music generation with improved audio quality and musicality. Creates two complete songs from a text description, or from exact lyrics with style and title in custom mode. Returns audio, cover image, and metadata for each track.
Suno Music V4.5
Suno
Suno V4.5 music generation. Creates two complete songs from a text description, or from exact lyrics with style and title in custom mode. Returns audio, cover image, and metadata for each track.
MiniMax Music 2.6
MiniMax
Generate full-length songs or instrumentals (up to ~6 minutes) from a text prompt. 99%+ accurate BPM/key control, 14+ structure tags, auto-generated lyrics, and instrumental-only mode.
MiniMax Music 2.5
MiniMax
Generate full-length songs (up to ~5 minutes) with vocals, lyrics, and rich instrumentation. 14 structure tags, style-aware mixing, and an expanded instrument library including orchestral and traditional instruments.
MiniMax Music Cover
MiniMax
Reimagine any song in a different style. The model extracts the melodic structure from the input audio and regenerates the track — the melody and duration stay the same, but voice, instruments, genre, and arrangement can all change.
Qwen3 TTS
Qwen
Alibaba's unified TTS with three modes: preset speakers (custom_voice), instant voice cloning from reference audio (voice_clone), and creating a brand-new voice from a text description (voice_design).
Inworld Realtime TTS 2
Inworld
Inworld's most expressive TTS with natural-language steering — place bracketed instructions like [speak quickly] before the text they apply to. Real-time latency and 15+ language support.
Suno SFX V5
Suno
Suno V5 sound effect generation. Creates two short sound effect variants from a text description, with optional loop mode, tempo, and musical key controls. Ideal for UI sounds, game audio, and video foley.
MiniMax Speech 2.8 HD
MiniMax
Ranked #1 on Artificial Analysis Speech Arena and Hugging Face TTS Arena. Broadcast-quality TTS with autoregressive Transformer + Flow-VAE decoder, 32+ languages, voice cloning, natural interjections, and emotion control.
MiniMax Speech 2.8 Turbo
MiniMax
Low-latency MiniMax Speech 2.8 Turbo with under 250ms latency, 40+ languages, voice cloning, natural interjections, and real-time pricing. Ideal for interactive and real-time applications.
MiniMax Speech 2.6 HD
MiniMax
Studio-quality multilingual text-to-speech with nuanced prosody, emotion control, and premium voices for professional applications.
MiniMax Speech 2.6 Turbo
MiniMax
Fast multilingual text-to-speech with emotional control, optimized for real-time applications with low latency.
MiniMax Speech-02-Turbo
MiniMax
Low-latency text-to-speech model with multilingual support, emotional voice control, and 300+ voice options.
Sonilo v1.1 Text to Music
Sonilo
Sonilo v1.1 text-to-music via Fal.AI. Generates up to 3 distinct music tracks (up to 600 seconds each) from a text description of the desired sound.
Inworld TTS 1.5 Max
Inworld
Inworld's highest-quality realtime TTS with under 200ms latency. Supports SSML break tags for pauses, emotion markups like [happy], and 15 languages.
Clova Voice TTS Premium
NCP Clova
NAVER Clova Voice Premium TTS with 108 voices across 6 languages. High-quality Korean voice synthesis with emotion control, Pro voices, and bilingual support.
ElevenLabs Turbo v2.5
ElevenLabs
High-quality, low-latency ElevenLabs text-to-speech in 32 languages. The same 26 premium voices as v3 at half the price, optimized for real-time and high-volume use.
ElevenLabs v3
ElevenLabs
ElevenLabs' most expressive text-to-speech model. 26 premium voices, inline audio tags like [laughs] and [whispers], fine-grained style and stability controls, and 70+ language support.
ElevenLabs v2 Multilingual
ElevenLabs
ElevenLabs Multilingual v2 text-to-speech in over 30 languages. Stable, proven voice quality with the same premium voice lineup and fine-grained voice settings.
Sonilo v1.1 Video to Sound Effects
Sonilo
Sonilo v1.1 video-to-sound-effects via Fal.AI. Adds AI-generated sound (ambience, effects, foley) to an input video — auto-captions the scene if no prompt is given, or accepts per-segment sound descriptions for finer control.
Suno Vocal Separation
Suno
Separate a Suno-generated track into clean vocal and instrumental stems. Takes the taskId and audioId from a previous Suno music generation and returns separated vocalUrl/instrumentalUrl audio files.
MiniMax Voice Cloning
MiniMax
Clone any voice from a 10-second to 5-minute audio sample. Returns a custom voice_id you can pass to MiniMax speech models, plus a preview clip synthesized with the cloned voice.
Whisper
OpenAI
OpenAI's Whisper large-v3 speech recognition (144M+ runs) — the standard for transcription. Automatic language detection across ~100 languages, English translation, and plain text / SRT / VTT output formats.
Whisper Diarization
Whisper
Whisper transcription with speaker diarization (8M+ runs) — returns who said what, with per-segment speaker labels and timestamps. The go-to for meetings and interviews.
XTTS-v2
Coqui
Coqui XTTS-v2 multilingual voice-cloning TTS. Clone a voice from a single short audio sample and speak in 16 languages — one of the most popular open-source voice cloning models.
LLM Models
GPT-4o, Claude, Gemini - OpenAI-compatible chat API
Claude Haiku 4.5
FeaturedAnthropic
Fast, cost-effective model for everyday tasks. Great balance of speed, intelligence, and cost for high-volume applications.
Claude Opus 5
Anthropic
Anthropic's latest flagship Opus model, with a 1M-token context window by default and 128K max output tokens. Same pricing as Opus 4.5–4.8 ($5/$25 per M tokens) with prompt caching (read $0.50/M, write $6.25/M) and web search. Compatible with the Anthropic Messages and OpenAI Chat Completions formats.
Claude Opus 4.8
Anthropic
Anthropic's most capable Opus-tier model, with a 1M-token context window (200K on some surfaces), 128K max output tokens, and knowledge cutoff to January 2026. Builds on Opus 4.7 with stronger long-horizon agentic coding, better tool triggering, and adaptive thinking that reasons only when a turn needs it. Compatible with the Anthropic Messages and OpenAI Chat Completions formats.
Claude Opus 4.7
Anthropic
Anthropic's latest flagship model with reliable knowledge cutoff to January 2026 and 128K max output tokens. Builds on Opus 4.6 with improved reasoning, coding, and instruction-following while staying compatible with the Anthropic Messages and OpenAI Chat Completions formats.
Claude Opus 4.6
Anthropic
Anthropic's most capable model. Delivers breakthrough performance in reasoning, coding, and complex analysis with enhanced safety and instruction following.
Claude Opus 4.5
Anthropic
Anthropic's most powerful model for highly complex tasks. Exceptional at research, analysis, and creative projects requiring deep expertise.
Claude Sonnet 5
Anthropic
Anthropic's newest Sonnet model, tuned for the best balance of speed, cost, and intelligence. Features a 1M-token context window, 64K max output tokens, and knowledge cutoff to August 2025. Supports adaptive thinking that reasons only when a turn needs it, vision, and tool use. Compatible with the Anthropic Messages and OpenAI Chat Completions formats.
Claude Sonnet 4.5
Anthropic
Anthropic's most intelligent and capable Sonnet model. Best-in-class for complex reasoning, nuanced understanding, and coding tasks with exceptional instruction following.
Claude Sonnet 4
Anthropic
Balanced Sonnet 4 model offering strong reasoning and coding abilities at an efficient price point. Ideal for everyday production workloads that need a good mix of speed and intelligence.
Gemini 3.5 Flash
Google's latest Flash-tier workhorse model. Flat pricing across all input modalities (text, image, video, audio, PDF) at $1.50/$9 per million tokens with no long-context premium, a 1M token context window, and 65,536 max output tokens. Cached inputs get a 90% discount.
Gemini 3.1 Flash Image Preview
Gemini 3.1 Flash with native image generation capabilities. Can generate images directly in chat responses alongside text. Features separate pricing for text and image output tokens.
Gemini 3.1 Flash Lite
Stable release of the ultra-lightweight Gemini 3.1 Flash Lite. The most cost-effective Gemini model at $0.25/$1.50 per million tokens, with cached input (90% discount), audio input, and batch (50% discount) support. Ideal for high-throughput, budget-conscious applications.
Gemini 3.1 Flash Lite Preview
Ultra-lightweight variant of Gemini 3.1 Flash. The most cost-effective Gemini model with support for cached input and audio input. Ideal for high-throughput, budget-conscious applications.
Gemini 3.1 Flash Live Preview
Gemini 3.1 Flash optimized for real-time interactions and live streaming scenarios. Features low-latency responses with audio input support at dedicated pricing.
Gemini 3.1 Pro Preview
Google's latest and most capable Gemini model in preview. Features dynamic pricing that adjusts based on context length, with enhanced pricing for inputs over 200K tokens.
Gemini 3 Flash
Google's most advanced reasoning model with state-of-the-art multimodal understanding, PhD-level reasoning, and leading coding performance.
Gemini 3 Pro Image Preview
Google's premium image generation model within the Gemini 3 Pro family. Generates high-quality images directly in chat with the highest fidelity among Gemini image models. Image output tokens are priced at 10x text output tokens.
Gemini 3 Pro Preview
Google's most powerful Gemini model in preview. Features breakthrough reasoning, coding, and multimodal capabilities with the largest context window.
Gemini 2.5 Flash
Google's fast and efficient model with built-in thinking capabilities. Great balance of speed, reasoning, and cost for high-volume applications.
Gemini 2.5 Pro
Google's most capable model with state-of-the-art reasoning and 1M token context. Excels at complex coding, math, and multi-document analysis.
Gemini 2.0 Flash
Google's fastest and most capable model. Features a massive 1M token context window, native multimodal support, and real-time capabilities.
Gemini 2.0 Flash Lite
Ultra-lightweight version of Gemini 2.0 Flash optimized for maximum speed and minimal cost. Perfect for high-volume, latency-sensitive applications.
Gemini Embedding 001
Google's text embedding model for generating vector representations. Optimized for semantic search, clustering, and similarity tasks.
GPT-5.6 Luna
OpenAI
The fast, low-cost tier of OpenAI's GPT-5.6 family (GA July 2026). Luna is built for high-throughput workloads at $1/$6 per million tokens, with a 1M token context window, 128K max output tokens, and cached inputs at a 90% discount.
GPT-5.6 Sol
OpenAI
The flagship tier of OpenAI's GPT-5.6 family (GA July 2026). Sol delivers the strongest reasoning, coding, and multimodal performance of the generation with a 1M token context window, 128K max output tokens, and cached inputs at a 90% discount.
GPT-5.6 Terra
OpenAI
The balanced tier of OpenAI's GPT-5.6 family (GA July 2026). Terra offers near-flagship quality at half the price of Sol, with a 1M token context window, 128K max output tokens, and cached inputs at a 90% discount.
GPT-5.5
OpenAI
OpenAI's newest flagship model with a 1.05M token context window and 128K max output tokens. Supports cached inputs at 10× discount and improved reasoning, coding, and multimodal performance over the GPT-5.4 series.
GPT-5.4
OpenAI
OpenAI's newest flagship model with 1M context window and 128K output tokens. Delivers top-tier reasoning across all domains with adjustable reasoning effort levels from none to xhigh.
GPT-5.4 Mini
OpenAI
Fast and cost-efficient variant of GPT-5.4 with 400K context window and 128K output tokens. Excellent balance of performance and affordability for everyday tasks.
GPT-5.4 Nano
OpenAI
Ultra-lightweight and fastest GPT-5.4 variant with 400K context and 128K output. Designed for high-throughput, low-latency applications at minimal cost. Supports MCP for tool integration.
GPT-5.2
OpenAI
OpenAI's latest and most advanced GPT model. Delivers state-of-the-art performance across reasoning, coding, and creative tasks with enhanced capabilities.
GPT-5.1 (2025-11-13)
OpenAI
Dated snapshot of GPT-5.1 for reproducible results. Supports cached input tokens for cost savings on repeated context. Ideal for production deployments requiring model version pinning.
GPT-5
OpenAI
OpenAI's latest flagship model. Delivers exceptional performance across reasoning, coding, and creative tasks with a massive 1M token context window and 32K output tokens. Supports vision, function calling, and JSON mode.
GPT-5 Mini
OpenAI
Fast and efficient variant of GPT-5. Delivers strong performance across reasoning, coding, and creative tasks with a 1M token context window and 32K output tokens, at a fraction of the cost of GPT-5.
GPT-5 Nano
OpenAI
Ultra-fast and lightweight variant of GPT-5. Designed for high-throughput, low-latency applications with a 1M token context window and 32K output tokens at minimal cost.
GPT-4.1
OpenAI
OpenAI's most capable model for coding and instruction following. Features a 1M token context window, 32K output tokens, and major improvements in coding, complex prompts, and long-context tasks. 20% cheaper than GPT-4o on output.
GPT-4.1 Mini
OpenAI
A significant leap in small model performance. Matches or exceeds GPT-4o in intelligence while reducing latency by nearly half and cost by 83%. Ideal balance of speed, quality, and affordability.
GPT-4.1 Nano
OpenAI
OpenAI's fastest and cheapest model. Optimized for classification, autocompletion, and low-latency tasks. Ultra-affordable at $0.10/1M input tokens.
GPT-4o
OpenAI
OpenAI's flagship multimodal model. Industry-leading performance in reasoning, coding, and creative tasks with native vision capabilities and structured output support.
GPT-4o Mini
OpenAI
Cost-effective, fast model with strong performance. Best for high-volume tasks where speed and cost matter more than absolute capability.
GPT Audio Mini
OpenAI
Lightweight multimodal model with native audio input/output capabilities. Optimized for voice-based interactions and audio processing tasks.
MiniMax M2.7
MiniMax
MiniMax's flagship M2-series language model, served through an OpenAI-compatible API. Strong multilingual capability (notably Chinese and English) at a very low price point ($0.30/$1.20 per million tokens) with prompt cache reads at $0.06/M.
OpenAI o4-mini
OpenAI
Fast, cost-effective reasoning model optimized for coding and STEM tasks. Provides strong reasoning at a fraction of the cost of larger reasoning models.
OpenAI o3-mini
OpenAI
Efficient reasoning model that delivers strong performance at lower cost. Ideal for tasks requiring reasoning without the overhead of larger models.
OpenAI o1
OpenAI
OpenAI's most advanced reasoning model. Uses extended thinking time to solve complex problems in science, coding, and math with exceptional accuracy.
Pricing Comparison
Compare credit costs across all models to find the best fit for your needs
| Category | Model | Credits | Unit | Best For |
|---|---|---|---|---|
| Image | flux-schnell | 7 | per image | Fast generation, real-time apps |
| flux-2-max | 160 | per image at 1 MP — varies by resolution (0.5MP 130, 2MP 230, 4MP/match_input 370, custom 390) | Best quality, commercial | |
| seedream-4.5 | 90 | per image — sequential mode (auto) scales per generated image, up to max_images (15) | Text rendering, logos | |
| Video | seedance-1-pro-fast | 290 | per video | Low-cost fast generation |
| hailuo-2.3 | 650 | per video | High quality, balanced | |
| veo-3.1 | 7,440 | per video | Best quality video | |
| Audio | speech-02-turbo | 142 | per 1000 characters (0.135 credits/char, billed per character) | Real-time TTS, fast response |
| speech-2.8-hd | 237 | per 1000 characters (0.225 credits/char, billed per character) | High quality HD voice | |
| LLM | gpt-4o | 3 | per 1K tokens (avg) | General, code, multimodal |
| claude-sonnet-5 | 4 | per 1K tokens (avg) | Analysis, long-form, code | |
| gemini-2.0-flash | 1 | per 1K tokens (avg) | Ultra-fast, bulk processing |
Ready to Get Started?
Try our models in the Playground or follow the Quickstart guide