Skip to main content
Core.Today
50+ AI Models

Model Catalog

Explore our comprehensive collection of AI models for image, video, audio generation, and LLM.

79
Image Models
46
Video Models
32
Audio Models
45
LLM Models

Featured Models

Top picks from each category - the best models for getting started

How to Choose the Right Model

Need Speed?

flux-schnell, Gemini Flash

Need Quality?

FLUX Pro, Kling Pro, Claude

Budget-Friendly?

flux-schnell, MiniMax, Gemini

Most Versatile?

GPT-4o, Claude, FLUX Dev

Image Generation Models

Generate stunning images with FLUX, Stable Diffusion, and more

Background Remover

Featured

851 Labs

1 credits

The most-used background remover on Replicate (27M+ runs). Removes backgrounds with soft alpha or hard segmentation, supports reverse mode (remove foreground), custom background types, and transparent PNG output — all at a very low price.

Most-used background remover on Replicate
Soft alpha matte or hard segmentation (threshold)
Reverse mode — remove the foreground instead
FastHigh
View Details

BiRefNet

BiRefNet

5 credits

State-of-the-art open-source background removal. BiRefNet's high-fidelity dichotomous image segmentation delivers excellent edge quality on hair, fur, and fine details.

State-of-the-art segmentation quality
Excellent edges on hair, fur, and fine details
Adjustable processing resolution
FastUltra
View Details

BLIP

Salesforce

1 credits

Salesforce BLIP (173M+ runs) — image captioning, visual question answering, and image-text matching in one model. The classic choice for bulk captioning at 1 credit per image.

Three tasks: captioning, VQA, image-text matching
173M+ runs — the classic captioner
1 credit per image — ideal for bulk pipelines
FastStandard
View Details

Change Haircut

Black Forest Labs

93 credits

Change anyone's hairstyle and hair color from a single photo, powered by FLUX.1 Kontext [pro]. Choose from 90+ hairstyles and 30 hair colors — or let 'Random' surprise you — while keeping the face untouched.

90+ hairstyle presets: bobs, braids, updos, fades, waves, and more
30 hair colors from natural shades to rose gold, blue, and silver
Identity-preserving edits — only the hair changes
FastHigh
View Details

Clarity Upscaler

Clarity

60 credits

The famous creative upscaler (30M+ runs). Instead of just enlarging, it re-imagines detail while upscaling — with controllable creativity, resemblance, prompt guidance, and tiled diffusion for high scale factors.

Creative detail re-imagination while upscaling
Creativity and resemblance dials
Prompt-guided enhancement
MediumUltra
View Details

CLIP Features

CLIP

1 credits

CLIP ViT-L/14 embeddings for text AND images (163M+ runs) — puts both in the same vector space for cross-modal search, image dedup, and zero-shot classification. 1 credit per run.

Text AND image embeddings in one space
163M+ runs — the CLIP standard on Replicate
Batch multiple inputs per call (newline-separated)
FastHigh
View Details

CodeFormer

CodeFormer

8 credits

Robust face restoration for old photos and AI-generated faces (54M+ runs). Its signature fidelity dial balances restoration quality against staying true to the original face, with Real-ESRGAN background enhancement built in.

Quality vs fidelity balance dial (codeformer_fidelity)
Robust on heavily degraded faces
Built-in Real-ESRGAN background enhancement
FastHigh
View Details

ControlNet Scribble

ControlNet

120 credits

The classic sketch-to-image model (38M+ runs) — turn any scribble or line drawing into a detailed image guided by your prompt. Draw the composition, describe the content.

Sketch controls composition, prompt controls content
38M+ runs — the scribble standard
1 or 4 samples per run
MediumHigh
View Details

Nano Banana (Edit)

Google

91 credits

Dedicated edit endpoint for Nano Banana, Google's Gemini 2.5 Flash-based image model. Pass input image URLs to perform conversational editing with character consistency and multi-image fusion.

Edit-mode endpoint optimized for image inputs
Multimodal editing with character consistency
Multi-image fusion
FastHigh
View Details

Nano Banana 2 (Edit)

Google

190 credits

Edit endpoint for Nano Banana 2, built on Gemini 3.1 Flash Image. Adds resolution control (1K/2K/4K), Google Search grounding, and thinking mode while preserving conversational editing and multi-image fusion.

Resolution control: 1K, 2K, 4K
Optional web search grounding
Thinking mode (minimal/high)
MediumUltra
View Details

Nano Banana Pro (Edit)

Google

350 credits

Edit endpoint for Nano Banana Pro built on Gemini 3 Pro. Professional-grade controls, legible multilingual typography, real-time grounding via Google Search, and resolution up to 2K for editing.

Gemini 3 Pro reasoning for complex edits
Legible text rendering in multiple languages
Web search grounding
MediumUltra
View Details

Face to Many

fofr

21 credits

Turn a face photo into 6 fun styles (15M+ runs) — 3D, Emoji, Video game, Pixels, Clay, or Toy. The viral avatar generator behind countless profile-picture apps.

6 preset styles: 3D, Emoji, Video game, Pixels, Clay, Toy
Identity preserved via InstantID
Custom LoRA support for brand styles
FastHigh
View Details

Face to Sticker

fofr

22 credits

Turn any face photo into a fun die-cut sticker (1.6M+ runs). InstantID keeps the likeness while IP-Adapter controls take the artwork from faithful caricature to loose cartoon — with an optional 2x upscale for print quality.

One face photo in, sticker-style artwork out
InstantID-based likeness with instant_id_strength control
ip_adapter_weight and prompt_strength to balance likeness vs style
FastHigh
View Details

Florence-2 Large

Microsoft

2 credits

Microsoft's Florence-2 all-in-one vision model — captioning, object detection, phrase grounding, OCR, and segmentation in a single API. Pick a task, optionally add text input, done.

One model, many tasks (task_input selection)
Caption / Detailed Caption / Object Detection
OCR and OCR with region boxes
FastHigh
View Details

FLUX.2 Klein 4B

Black Forest Labs

2 credits

Very fast image generation and editing model. 4-step distilled, sub-second inference for production and near real-time applications.

4-step distilled model — sub-second inference
Supports both text-to-image and image-to-image editing (up to 5 reference images)
Selectable output resolution from 0.25MP to 4MP
FastStandard
View Details

FLUX 2 Flex

Black Forest Labs

280 credits

Maximum-quality FLUX model supporting up to 10 reference images and advanced typography. The most capable model for complex, multi-reference creative projects.

Up to 10 reference images supported
Advanced typography rendering
Ultra-quality output
SlowUltra
View Details

FLUX 2 Max

Black Forest Labs

160 credits

The highest fidelity image model from Black Forest Labs. Best-in-class prompt following and the most consistent editing in the FLUX.2 lineup — preserves colors, lighting, faces, text, and objects across edits with up to 8 reference images.

Highest editing consistency in the FLUX.2 lineup — preserves identity, colors, lighting, and text
Best-in-class prompt following for both short and long prompts
Multi-reference editing with up to 8 input images
FastUltra
View Details

FLUX 2 Pro

Black Forest Labs

70 credits

Professional-grade FLUX 2 with high-quality editing and up to 8 reference image support. Excellent balance of quality, speed, and creative control.

Up to 8 reference images supported
High-quality image editing
Ultra-quality generation
MediumUltra
View Details

FLUX.2 Dev

Black Forest Labs

28 credits

Development version of FLUX.2 with image editing capabilities and reference image support. Ideal for iterative design workflows and experimentation.

Image editing capabilities
Reference image support
Balanced speed and quality
MediumHigh
View Details

FLUX 1.1 Pro

Black Forest Labs

93 credits

Fast high-quality image generation, an upgrade to FLUX.1 Pro with faster speed and improved quality. Perfect for production workloads requiring both speed and fidelity.

Faster than FLUX.1 Pro
Improved image quality
Ultra-quality output
FastUltra
View Details

FLUX 1.1 Pro Ultra

Black Forest Labs

140 credits

FLUX1.1 [pro] in ultra and raw modes. Images are up to 4 megapixels — the highest-resolution tier of the FLUX 1.1 Pro family. Use raw mode for realism.

Up to 4-megapixel ultra-resolution output
Raw mode for less processed, more natural-looking images
Flux Redux image_prompt support for composition guidance
MediumUltra
View Details

FLUX Dev

Black Forest Labs

58 credits

A 12 billion parameter rectified flow transformer capable of generating images from text descriptions, tuned for open, high-quality experimentation.

12B parameter rectified flow transformer
Optional fp8 'go_fast' mode for faster generation
Supports image-to-image via the image parameter
FastHigh
View Details

FLUX Kontext Max

Black Forest Labs

190 credits

A premium text-based image editing model that delivers maximum performance and improved typography generation for transforming images through natural language prompts.

Top-tier Kontext editing performance, above FLUX Kontext Pro
Improved typography and text rendering in edited images
Natural language instruction-based editing
MediumUltra
View Details

FLUX Kontext Pro

Black Forest Labs

93 credits

State-of-the-art text-based image editing model that transforms images through natural language. Excellent for style transfer, object modification, text replacement, background changes, and character consistency.

Text-driven image editing with high fidelity
Style transfer and transformation
Object and text replacement
MediumUltra
View Details

FLUX.1 Krea [dev]

Krea AI

58 credits

Photorealistic image generation that specifically avoids the 'AI look', producing natural-looking images indistinguishable from real photographs.

Photorealistic output without AI artifacts
Natural-looking image generation
Avoids typical AI aesthetic tells
MediumHigh
View Details

FLUX PuLID

ByteDance

47 credits

PuLID identity customization on FLUX-dev: generate photorealistic portraits of a specific person from one face photo, with markedly higher fidelity than SDXL-based variants. Tune id_weight and start_step to balance likeness against prompt editability.

FLUX-dev backbone for photorealistic, high-fidelity portraits
Tuning-free ID preservation from a single face photo
id_weight (0-3) controls how strongly the face is preserved
MediumHigh
View Details

FLUX.1 Schnell

Black Forest Labs

7 credits

Ultra-fast image generation model optimized for speed. Generates high-quality images in just 1-2 seconds, perfect for real-time applications and rapid prototyping.

Ultra-fast generation (1-2 seconds)
High-quality output at 1024x1024
Excellent prompt understanding
FastHigh
View Details

Runway Gen-4 Image

Runway

190 credits

Runway's Gen-4 Image model with references: combine up to 3 reference images with @tag mentions in your prompt to keep characters, objects, and locations consistent across every angle and scene, at 720p or 1080p.

Up to 3 reference images for consistent characters, objects, and locations
@tag_name mentions in the prompt bind each reference to a role
720p and 1080p output resolutions with differential pricing
MediumUltra
View Details

Runway Gen-4 Image Turbo

Runway

70 credits

Gen-4 Image Turbo is 2.5x faster and cheaper than Gen-4 Image, with the same reference-driven API: use 1 to 3 reference images with @tag mentions for consistent characters and objects, at a flat price regardless of resolution.

2.5x faster and cheaper than Gen-4 Image
Flat 90-credit price for both 720p and 1080p
1 to 3 reference images with @tag_name prompt mentions (at least 1 required)
FastHigh
View Details

GFPGAN

Tencent ARC

6 credits

Tencent ARC's face restoration model — 115M+ runs, the most-run model in the restoration collection. Restores old/blurry photos and fixes faces in AI-generated images.

115M+ runs — the standard for face restoration
Restores old, blurry, or damaged photos
Fixes distorted faces in AI-generated images
FastHigh
View Details

GPT Image 2

OpenAI

300 credits

OpenAI's state-of-the-art image generation and editing model with strong instruction following, sharp text rendering, and detailed editing. Quality-based pricing lets you trade off cost vs. fidelity.

Strong prompt and instruction following
Sharp typography and text rendering
Image editing via input_images (composition/edit)
MediumUltra
View Details

GPT Image 1.5

OpenAI

310 credits

OpenAI's latest image generation model with better instruction following and adherence to prompts, including sharp text rendering and detailed editing.

Four quality tiers (low/medium/high/auto) with proportional pricing
Strong instruction following and prompt adherence
Sharp in-image text rendering
MediumUltra
View Details

Grok Imagine Image

xAI

47 credits

Generate images using xAI's Grok Imagine model. Sibling to Grok Imagine Video, sharing the same underlying Grok Imagine architecture for fast text-to-image generation.

Text-to-image generation built on the Grok Imagine architecture
Image editing mode — pass an existing image plus a prompt to edit it
14 aspect ratio presets plus an auto option
FastStandard
View Details

Grounding DINO

Grounding DINO

2 credits

Text-prompted object detection (39M+ runs) — describe what to find in natural language ('red car, person wearing a hat') and get bounding boxes with confidence scores plus an annotated image.

Zero-shot detection from natural language
Bounding boxes + confidence scores (JSON)
Annotated visualization image
FastHigh
View Details

Hunyuan 3D 3.1

Tencent

1160 credits

Tencent's flagship 3D generation — create high-polygon textured 3D models from a text prompt OR an image, with optional PBR (physically based rendering) materials.

Text-to-3D AND image-to-3D in one model
Up to 500K-face high-polygon output
Optional PBR material generation
SlowUltra
View Details

Ideogram V4 Balanced

Ideogram

140 credits

A middle-ground tier in Ideogram's v4 family, balancing generation speed and output quality. Delivers strong typography and photorealism at a lower cost than the Quality tier.

Balanced speed and quality within the Ideogram v4 family
Magic Prompt auto-enhancement when using natural-language prompt
Structured json_prompt input for precise, repeatable control
MediumHigh
View Details

Ideogram V4 Quality

Ideogram

230 credits

The highest-fidelity tier of Ideogram's v4 model family, tuned for maximum detail, realism, and typography accuracy. Best suited for final production assets where quality matters more than speed.

Highest-fidelity tier in the Ideogram v4 family
Magic Prompt auto-enhancement when using natural-language prompt
Structured json_prompt input for precise, repeatable control
SlowUltra
View Details

Ideogram V4 Turbo

Ideogram

70 credits

The fastest and cheapest model in Ideogram's v4 family, built for rapid iteration while retaining Ideogram's signature text rendering and style consistency.

Fastest and lowest-cost tier in the Ideogram v4 family
Magic Prompt auto-enhancement when using natural-language prompt
Structured json_prompt input for precise, repeatable control
FastStandard
View Details

Ideogram V3 Turbo

Ideogram

70 credits

The fastest and cheapest Ideogram v3 tier. V3 creates images with stunning realism, creative designs, and consistent styles.

Fastest and cheapest tier of Ideogram v3
60+ fixed resolution presets in addition to aspect ratio
Inpainting support via image + mask
FastHigh
View Details

Ideogram Character

Ideogram

350 credits

Generate consistent characters from a single reference image. Render the same character in many styles — realistic or fiction — insert them into existing photos with mask inpainting, and rely on Ideogram's signature text rendering for legible signs and typography.

Consistent characters from just one reference image
Auto, Fiction, or Realistic character style types
Mask inpainting to add your character to an existing image
MediumUltra
View Details

MiniMax Image-01

MiniMax

23 credits

MiniMax's first image generation model with character reference support: provide a single face photo via subject_reference and generate consistent images of that person across prompts, styles, and aspect ratios — up to 9 images per request.

Character consistency from a single face photo via subject_reference
Batch generation of 1-9 images per request
8 aspect ratios from square 1:1 to cinematic 21:9
FastHigh
View Details

Topaz Image Upscale

Topaz Labs

190 credits

Professional-grade upscaling from Topaz Labs, the industry standard for photo enhancement. Five specialized enhance models, up to 6x upscale, subject detection, and optional face enhancement.

Industry-standard Topaz enhancement quality
5 enhance models: Standard / Low Resolution / CGI / High Fidelity / Text Refine
Up to 6x upscale
MediumUltra
View Details

Imagen 4

Google

93 credits

Google's Imagen 4 flagship text-to-image model.

Google's flagship Imagen 4 quality tier
1K and 2K output resolution at the same flat price
Configurable safety filter strictness
MediumUltra
View Details

Imagen 4 Fast

Google

46 credits

A fast version of Imagen 4 for when speed and cost are more important than maximum quality.

Lowest-cost tier in the Imagen 4 family
Fast generation optimized for iteration speed
Configurable safety filter strictness
FastStandard
View Details

Bria Increase Resolution

Bria

93 credits

Bria's commercially-safe image upscaler (130K+ runs). Increase resolution 2x or 4x with a model trained exclusively on licensed data, preserving alpha transparency — built for enterprise pipelines that require full legal liability coverage.

Trained exclusively on licensed data — commercially safe
2x or 4x resolution increase (desired_increase)
Preserves alpha channel — ideal after background removal
FastHigh
View Details

Krea 2 Large

Krea AI

140 credits

Krea AI's flagship text-to-image model, focused on photorealistic output with strong prompt adherence. Supports style-reference and moodboard-guided generation for consistent visual direction.

Photorealistic output with strong prompt adherence
Style-reference transfer from up to 10 images with adjustable strength
Moodboard-guided generation via the Krea webapp
MediumHigh
View Details

Krea 2 Medium

Krea AI

70 credits

A lower-cost variant of Krea 2 Large, trading some fidelity for faster and cheaper generation while keeping the same photorealistic style focus.

Cheaper, faster sibling of Krea 2 Large with the same photorealistic style focus
Creativity slider (raw/low/medium/high) controls how far the model deviates from the literal prompt
Style reference images (up to 10) transfer a visual style to the output
FastStandard
View Details

Microsoft MAI-Image 2.5 Pro

Microsoft

400 credits

Microsoft's highest-fidelity image model for production-grade text-to-image generation via Fal.AI. Built for hero imagery, detailed compositions, precise text rendering, photorealism, stylized illustration, commercial design, and visually rich concept work.

High-fidelity output tuned for production and commercial use
Precise in-image text rendering
Generate up to 4 images per call
MediumUltra
View Details

Moondream2

Moondream

5 credits

Small but capable vision-language model (14M+ runs) — ask free-form questions about any image and get detailed answers. Efficient VQA for tagging, moderation prep, and rich alt-text.

Free-form questions about any image
Detailed multi-sentence answers
Small model — fast and cheap (7 credits)
FastHigh
View Details

Multilingual E5 Large

E5

2 credits

Multilingual text embeddings (74M+ runs) — 1024-dimension vectors across 100 languages including Korean. Pairs perfectly with Core.Today customer databases' vector search (knn_vector).

1024-dim embeddings, 100 languages incl. Korean
Batch embedding of multiple texts per call
Normalized embeddings option (cosine-ready)
FastHigh
View Details

Nano Banana 2 (Gemini 3.1 Flash Image)

Google

190 credits

Google's fast image generation model built on Gemini 3.1 Flash Image. The high-efficiency counterpart to Nano Banana Pro — combining Pro-level visual quality with Flash-level speed and pricing. Features conversational editing, multi-image fusion, character consistency, accurate text rendering, and Google Search grounding. Supports up to 14 reference images and resolutions up to 4K.

Gemini 3.1 Flash Image-powered generation with Pro-level quality
Accurate text rendering in multiple languages
Google Web Search and Image Search grounding
FastUltra
View Details

Nano Banana 2 Lite

Google

80 credits

Google's lightweight Nano Banana 2 variant built on Gemini 3.1 Flash Image, tuned for faster and cheaper generation. Retains conversational editing, multi-image fusion, and character consistency from the full Nano Banana 2.

Faster, lower-cost variant of the full Nano Banana 2 model
Built on Gemini 3.1 Flash Image
Conversational, instruction-based editing
FastStandard
View Details

Nano Banana

Google

91 credits

Google Gemini 2.5 Flash-based image generation with multimodal editing capabilities. Fast and versatile for both creation and editing tasks.

Gemini 2.5 Flash-powered generation
Multimodal image editing
Fast generation speed
FastHigh
View Details

Nano Banana Pro (Gemini 3 Pro Image)

Google

350 credits

Google's state-of-the-art image generation and editing model built on Gemini 3 Pro. Creates detailed visuals with legible text in multiple languages, connects to real-time information from Google Search, and provides professional-grade creative controls. Supports up to 14 reference images and resolutions up to 4K.

Gemini 3 Pro-powered generation with advanced reasoning
Legible text rendering in multiple languages
Google Search integration for real-time data visualization
MediumUltra
View Details

NSFW Image Detection

Falcons.ai

1 credits

The standard NSFW image classifier (127M+ runs) — returns 'normal' or 'nsfw' for any image. An essential, ultra-cheap moderation gate for UGC platforms.

127M+ runs — the moderation standard
Simple normal/nsfw output
1 credit per image — run on every upload
FastHigh
View Details

Pruna P-Image

Pruna AI

12 credits

Pruna AI's distilled text-to-image model optimized for extremely low-cost, high-throughput generation. One of the most-run community models on Replicate (15.6M+ runs) thanks to its speed and price.

Distilled architecture tuned for extremely low-cost, high-throughput generation
One of the most-run community models on Replicate (15.6M+ runs)
Custom aspect_ratio=custom mode with explicit width/height (multiples of 16, up to 1440px)
FastStandard
View Details

Pruna P-Image Upscale

Pruna AI

12 credits

PrunaAI's high-end image upscaler (575K+ runs) reaching up to 128-megapixel output. Target-megapixel or factor-based control, optional detail/realism enhancement passes, and megapixel-tiered pricing that starts at just 15 credits.

Up to 128MP output — poster and large-print territory
Two modes: target megapixels (precise billing) or scale factor (1-8x)
Optional enhance_details and enhance_realism passes
MediumUltra
View Details

PhotoMaker

Tencent ARC

10 credits

Generate stylized photos of a person from 1-4 reference photos (9M+ runs). Ten styles including Cinematic, Disney Character, and Digital Art — keep the identity, change everything else.

Identity-preserving generation from 1-4 photos
10 built-in styles
Prompt-controlled scenes and outfits
MediumHigh
View Details

PhotoMaker Style

Tencent ARC

15 credits

The stylization-focused variant of PhotoMaker: turn 1-4 photos of a person into paintings, comics, 3D art, and more with stronger style transfer. Pairs with the base PhotoMaker — use this one when style matters more than photorealism.

Stronger style transfer than the base PhotoMaker
Identity-preserving generation from 1-4 reference photos
10 style presets: Cinematic, Disney Character, Digital Art, Comic book, and more
MediumHigh
View Details

Professional Headshot

Black Forest Labs

93 credits

Turn any single photo into a polished professional business headshot, powered by FLUX.1 Kontext [pro]. Pick a background — white, black, gray, neutral, or office — and get a LinkedIn-ready portrait in one step.

One photo in, professional headshot out — no studio required
FLUX.1 Kontext [pro] app tuned specifically for business portraits
5 background presets: white, black, neutral, gray, office
FastHigh
View Details

Proteus v0.2

Proteus

35 credits

Popular anime-focused image model (12M+ runs) — high-quality anime and illustration styles with img2img and inpainting support. The go-to for anime avatars and webtoon-style art.

Anime and illustration specialization
img2img and inpainting with mask
Up to 4 outputs per run (billed per image)
MediumHigh
View Details

PuLID

ByteDance

6 credits

ByteDance PuLID: tuning-free identity customization on SDXL. Give it one face photo and a prompt to generate portraits in any scene or style — no training, 4-step fast sampling, and even two-identity blending. Extremely cost-effective at 2 credits per image.

Tuning-free identity preservation from a single face photo
Lightning-fast 4-step sampling on SDXL
mix_identities: blend two different faces into one person
FastHigh
View Details

Real-ESRGAN

Real-ESRGAN

5 credits

The classic image upscaler (94M+ runs). Real-ESRGAN super-resolution up to 10x with optional GFPGAN face enhancement — the go-to default for cleanly enlarging photos and AI images.

Battle-tested upscaler (94M+ runs)
Scale factor up to 10x
Optional GFPGAN face enhancement
FastHigh
View Details

Recraft V4

Recraft

93 credits

Recraft's next-generation text-to-image model, improving on V3's typography, style range, and prompt adherence for production-grade brand and design assets.

Improved typography handling over Recraft V3
Wider style range and stronger prompt adherence
Long-form prompts up to 10,000 characters
MediumHigh
View Details

Recraft V4 SVG

Recraft

190 credits

Generates vector graphics directly in SVG format using Recraft V4, ideal for logos, icons, and scalable illustrations that need to stay crisp at any size.

Native SVG output — true vector graphics, not a rasterized image
Built on Recraft V4 for improved typography and prompt adherence
Ideal for assets that must stay crisp at any scale
MediumHigh
View Details

Recraft V3

Recraft

93 credits

Recraft V3 (code-named red_panda) is a text-to-image model with the ability to generate long texts, and images in a wide list of styles. SOTA in image generation per the Artificial Analysis Text-to-Image Benchmark.

State-of-the-art benchmark performance among text-to-image models
Long, accurate in-image text rendering
19 curated styles spanning realistic photography and digital illustration
MediumHigh
View Details

Recraft Crisp Upscale

Recraft

14 credits

Fast, affordable upscaler from Recraft designed for sharp, crisp results — a single image input with no tuning needed. Great default for UI assets, illustrations, and product images.

Sharp, crisp upscaling optimized for clarity
Zero-tuning single-input API
Fast processing
FastHigh
View Details

Remove Background

Bria AI

32 credits

AI-powered background removal tool for images. Clean, accurate cutouts for any subject with professional-quality edge detection.

Precise edge detection
Clean background removal
Works with any subject type
FastHigh
View Details

Remove BG

lucataco

1 credits

One of the most-run background removers on Replicate (17M+ runs). A single-parameter API that strips the background and returns a transparent PNG — at just 1 credit per image, the cheapest cutout on the platform.

17M+ runs — battle-tested at scale
Single parameter: just the image
Transparent PNG output
FastStandard
View Details

Stable Diffusion XL

Stability AI

9 credits

Stability AI's classic SDXL (85M+ runs) — the battle-tested text-to-image model with img2img, inpainting, refiner, and LoRA support. A dependable workhorse with a huge ecosystem.

Battle-tested classic (85M+ runs)
img2img and inpainting with mask
Expert-ensemble refiner support
MediumHigh
View Details

SDXL Lightning 4-step

ByteDance

4 credits

ByteDance's 4-step SDXL Lightning — the most-run model on Replicate (1B+ runs). Near-instant 1024px image generation at one of the lowest prices in the catalog.

1B+ runs — the most-run model on Replicate
4-step generation: near-instant results
Up to 1280px output, up to 4 images per request
FastStandard
View Details

Seedream 5 Lite

ByteDance

82 credits

Seedream 5.0 lite: image generation with built-in reasoning, example-based editing, and deep domain knowledge. Supports multi-reference generation with up to 14 images and sequential batch generation.

Example-based editing — show a before/after pair and the model applies the same transformation
Logical reasoning for spatial relationships, physics, and processes
Deep domain knowledge across architecture, science, health, and design
FastHigh
View Details

Seedream 5 Pro

ByteDance

210 credits

ByteDance's flagship Seedream 5.0 Pro image generation model, with built-in reasoning, precise instruction following, and multi-reference support for up to 10 images. Offers 1K and 2K resolution output with higher fidelity than Seedream 5 Lite.

Built-in reasoning for precise instruction following
Multi-reference generation with up to 10 input images
1K and 2K resolution output
MediumUltra
View Details

Seedream 4.5

ByteDance

90 credits

Upgraded ByteDance image model with stronger spatial understanding and world knowledge. Supports single/multi-reference image-to-image editing and sequential (multi-image) generation.

Strong spatial understanding and world knowledge
Multi-reference image-to-image generation (up to 14 input images)
Sequential generation mode producing up to 15 related images in one call
MediumUltra
View Details

Seedream 4.0

ByteDance

70 credits

ByteDance's latest image generation model with exceptional prompt understanding and creative capabilities.

Advanced language understanding
Creative composition
Excellent text rendering
MediumHigh
View Details

Text Extract OCR

OCR

1 credits

Simple, massively-used OCR (91M+ runs) — extracts text from an image with a single input and returns plain text. Great default for receipts, screenshots, and scanned documents at just 1 credit.

One input, plain-text output
91M+ runs — battle-tested
1 credit per image — ideal for bulk
FastStandard
View Details

TRELLIS

Microsoft

77 credits

Microsoft's TRELLIS image-to-3D (836K+ runs — the most-used 3D model on Replicate). Turns one or more images into a textured GLB 3D asset, with turntable render videos and optional Gaussian PLY.

The most-used 3D model on Replicate
Single or multi-image input
Textured GLB export (generate_model)
MediumHigh
View Details

Google Upscaler

Google

46 credits

Google's image upscaler (770K+ runs). Upscale images 2x or 4x using generative AI while preserving natural detail, with adjustable output compression — a simple, reliable enhancer at a flat price.

2x or 4x upscaling with generative detail preservation
Google-grade quality and consistency
Adjustable output compression quality (1-100)
FastHigh
View Details

Z-Image Turbo

Pruna AI

40 credits

Pruna's ultra-fast Z-Image Turbo (48M+ runs). Megapixel-priced image generation up to 2048x2048 — pay exactly for the resolution you generate, with excellent price/performance for high-volume use.

Ultra-fast turbo generation (8 steps)
Megapixel-based pricing — pay for what you generate
Up to 2048x2048 output
FastHigh
View Details

Video Generation Models

Create AI videos with Kling, MiniMax, and cutting-edge models

Add Watermark

Featured

FullJourney

1 credits

Add a text watermark to any video — simple, fast brand protection for generated or user content at 2 credits per video.

Text watermark on any video
Adjustable size
2 credits per video
FastStandard
View Details

CogVLM2 Video

CogVLM

28 credits

Video understanding and captioning — ask free-form questions about a video and get detailed answers about actions, scenes, and content. Great for video search indexing and moderation prep.

Free-form questions about video content
Action and scene understanding
Detailed multi-sentence answers
MediumHigh
View Details

Google Gemini Omni Flash

Google

2320 credits

Google Gemini Omni Flash text-to-video via Fal.AI. Generates a video directly from a descriptive text prompt — pacing and audio (dialogue, background music) are controlled in the prompt itself, in 16:9 or 9:16 at 3-10 second durations.

Text-to-video generation from a single descriptive prompt
Dialogue and background music controlled directly in the prompt text — no separate audio toggle
16:9 (landscape) or 9:16 (vertical) aspect ratios
FastHigh
View Details

Gen-4.5

Runway

1400 credits

Runway's Gen-4.5 model, offering state-of-the-art video motion quality, prompt adherence, and visual fidelity for text-to-video and image-to-video generation.

State-of-the-art motion quality and prompt adherence
Flat per-second pricing regardless of aspect ratio
Text-to-video or image-to-video from an optional first frame
MediumUltra
View Details

Grok Imagine Video 1.5

xAI

930 credits

xAI's Grok Imagine Video 1.5 preview: image-to-video generation with synchronized audio. An upgraded successor to Grok Imagine Video with flat per-second pricing regardless of resolution.

Image-to-video animation with synchronized audio included
Upgraded 1.5 preview of xAI's Grok Imagine Video
Flexible 1-15 second duration, billed per second
MediumHigh
View Details

Grok Imagine Video

xAI

580 credits

Generate videos using xAI's Grok Imagine Video model. Supports text-to-video, image-to-video, and editing an existing short video clip.

Flat per-second pricing regardless of resolution or aspect ratio
Text-to-video and image-to-video generation
Video editing mode — modify an existing short clip (up to 8.7s) with a prompt
FastHigh
View Details

MiniMax Hailuo 2.3

MiniMax

650 credits

Realistic human motion video generation with advanced character consistency and natural movement.

Realistic human motion
Advanced character consistency
Natural movement generation
SlowUltra
View Details

MiniMax Hailuo 2.3 Fast

MiniMax

450 credits

Lower-latency version of Hailuo 2.3 optimized for faster generation while maintaining good quality for human motion videos.

Faster generation than Hailuo 2.3
Good human motion quality
Character consistency
MediumHigh
View Details

Happy Horse 1.1

Alibaba

2090 credits

Alibaba's Happy Horse 1.1 generates videos from text, animates a single image, or builds a video from multiple reference images. Supports 720p and 1080p, 3-15 second durations, and five aspect ratios.

Three modes in one model: text-to-video, image-to-video, reference-to-video
Reference-to-video with up to 9 images — keeps subjects and scenes consistent
720p and 1080p output at 3-15 second durations
MediumHigh
View Details

MiniMax Hailuo-03 (H3) Image to Video

MiniMax

3020 credits

MiniMax Hailuo-03 (H3) image-to-video via Fal.AI. 2K video generation from a first-frame image, 5-15 second duration, native audio, with optional first-to-last keyframe control via end_image_url.

Generates video directly from a single first-frame image
Up to 2K resolution output
5-15 second adjustable duration
MediumUltra
View Details

Kling v3 Omni Video

Kuaishou

2610 credits

Kling Video 3.0 Omni: a unified multimodal video model that generates and edits video from text, images, reference images, and existing video. Combines text-to-video, image-to-video, reference-based generation, and video editing with native audio and multi-shot control.

Three quality tiers: standard (720p), pro (1080p, default), and 4k
Reference-based generation from up to 7 images for character/style consistency
Video editing mode via reference_video + video_reference_type=base
SlowUltra
View Details

Kling v3 Video

Kuaishou

2610 credits

Kling Video 3.0: Kuaishou's flagship text/image-to-video model generating cinematic videos up to 15 seconds with multi-shot control, native audio, start/end frame images, and a dedicated 4K mode.

Cinematic videos up to 15 seconds (3-15s range)
Three quality tiers: standard (720p), pro (1080p, default), and 4k
Multi-shot control via multi_prompt (up to 6 sequential shots in one generation)
SlowUltra
View Details

Kling v2.6

Kuaishou

1630 credits

Kling 2.6 Pro: top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation. Audio generation is enabled by default.

Cinematic visual quality with fluid, natural motion
Native synchronized audio generation (enabled by default)
5 or 10 second duration options
MediumUltra
View Details

Kling 2.5 Turbo Pro

Kuaishou

810 credits

Cinematic-grade video generation with enhanced motion and scene coherence. Top-tier Kling model for professional output.

Cinematic-grade output quality
Enhanced motion coherence
Superior scene consistency
MediumUltra
View Details

Kling v2.1

Kuaishou

580 credits

Kling v2.1 with 720p/1080p support and frame transition capabilities for smooth, high-quality video generation.

720p and 1080p output support
Smooth frame transitions
Image-to-video capability
SlowUltra
View Details

LatentSync

ByteDance

220 credits

ByteDance's open-source lipsync — re-syncs a video's mouth movements to any audio track using latent diffusion. State-of-the-art open lipsync quality for dubbing and localization.

Video + audio → lip-synced video
State-of-the-art open-source lipsync
Latent diffusion mouth re-generation
SlowHigh
View Details

MMAudio

MMAudio

11 credits

Add AI-generated sound to any video (5M+ runs) — synthesizes synchronized audio (ambience, effects, foley) from the video content and an optional text prompt. The perfect finisher for silent AI-generated clips.

Video-synchronized audio synthesis
Optional prompt to steer the sound
Negative prompt (default excludes music)
MediumHigh
View Details

Pruna P-Video

Pruna AI

240 credits

PrunaAI's fast video generator with a built-in draft mode for rapid creative iteration. Text-to-video, image-to-video, and audio-conditioned generation in a single endpoint, with clips up to 20 seconds — one of the longest durations on the platform.

Built-in draft mode: lower-quality previews at roughly a quarter of the standard price
Text-to-video and image-to-video in one endpoint (image parameter)
Optional audio input (flac/mp3/wav) to condition video generation
FastStandard
View Details

Pruna P-Video Avatar

Pruna AI

870 credits

Pruna AI's talking-avatar video model — animates a single input image to speak provided text (voice_script) or an uploaded audio track, with selectable Gemini-family preset voices, language, and visual delivery prompt.

Talking-avatar animation from a single input image
Text-to-speech via 29 Gemini-family preset voices, or drive speech from an uploaded audio track
10-language voice selection with adjustable delivery style via voice_prompt
MediumHigh
View Details

PixVerse V6

PixVerse

810 credits

PixVerse's flagship video generation model. Generates cinematic videos with synchronized audio, multi-shot sequences with scene transitions, astonishing physics, and precise camera control at up to 1080p.

Synchronized AI audio: BGM, sound effects, and character dialogue
Multi-shot generation with scene transitions via generate_multi_clip_switch
Astonishingly realistic physics and precise camera control
MediumHigh
View Details

PixVerse V5

PixVerse

24 credits

Advanced video generation with special effects capabilities and anime-optimized output, supporting multiple visual styles.

Special effects capabilities
Anime-optimized output
Multiple visual styles (realistic, anime, 3D)
SlowUltra
View Details

Luma Ray 3.2

Luma

1740 credits

Luma's flagship Ray video model. Text-to-video and keyframe (start/end image) generation with optional HDR-encoded output and professional EXR export.

Flagship text-to-video generation with detailed prompt control
Keyframe mode using start/end images to anchor the first and last frame
HDR-encoded MP4 output for wider dynamic range
MediumUltra
View Details

Luma Ray Flash 2 720p

Luma

700 credits

Luma's Ray Flash 2 generates 5 or 9 second 720p videos faster and cheaper than Ray 2. Supports keyframe control via start and end images, seamless loops, and a rich library of camera motion concepts — the first Luma model on the platform.

Faster and cheaper than Ray 2 at 720p output
Keyframe control: anchor the first frame (start_image) and last frame (end_image)
Loop option for seamless, continuously repeating playback
FastHigh
View Details

Real-ESRGAN Video

Real-ESRGAN

1280 credits

Video upscaling with Real-ESRGAN — enhance videos to FHD, 2K, or 4K frame by frame. The go-to open-source video upscaler for old footage and AI-generated clips.

Upscale to FHD / 2K / 4K
Three Real-ESRGAN model variants (incl. anime)
Frame-by-frame enhancement
SlowHigh
View Details

SadTalker

SadTalker

91 credits

Talking-head video from a single photo and an audio track — animates the face with natural head motion and eye blinks, with optional GFPGAN face enhancement.

One photo + audio → talking-head video
Natural head motion and eye blinks
Still mode for subtle motion
SlowHigh
View Details

Seedance 2.0

ByteDance

2090 credits

ByteDance's next-generation multimodal video model with native synchronized audio. Combines up to 9 reference images, 3 videos, and 3 audio files in a single generation for character-consistent, lip-synced video creation, editing, and extension.

Native audio generation — dialogue, sound effects, and background music synced with visuals
Multimodal reference inputs: up to 9 images, 3 videos, 3 audio files in one pass
Character consistency across shots using reference images
MediumUltra
View Details

Seedance 2.0 Fast

ByteDance

1750 credits

A faster, cheaper variant of Seedance 2.0 for quicker video generation with multimodal reference inputs (up to 9 images, 3 videos, 3 audios) and native audio, at 480p or 720p.

Faster, lower-cost alternative to Seedance 2.0 (480p/720p only, no 1080p/4K)
Native audio generation — dialogue, sound effects, and background music
Multimodal reference inputs: up to 9 images, 3 videos, 3 audio files
FastHigh
View Details

Seedance 2.0 Mini

ByteDance

1280 credits

Lighter, cheaper variant of ByteDance's Seedance 2.0. Native audio, multimodal reference inputs (images/videos/audio), text-to-video and image-to-video, capped at 720p (no 1080p/4K tier).

Native synchronized audio generation (dialogue, sound effects, background music)
Multimodal reference inputs — up to 9 reference images, 3 reference videos, 3 reference audio clips
Text-to-video and image-to-video (first/last frame) generation
FastStandard
View Details

Seedance 1 Lite

ByteDance

420 credits

A lightweight ByteDance video generation model offering text-to-video and image-to-video support for 4-12 second videos at 480p, 720p, or 1080p resolution.

Text-to-video and image-to-video generation in one model
4-12 second duration range with per-second billing
480p / 720p / 1080p resolution options
FastStandard
View Details

Seedance 1 Pro

ByteDance

700 credits

A pro version of Seedance that offers text-to-video and image-to-video support for 2-12 second videos, at 480p, 720p, and 1080p resolution.

Higher fidelity than Seedance 1 Lite, up to full 1080p output
Text-to-video and image-to-video generation in one model
2-12 second duration range with per-second billing
MediumUltra
View Details

Seedance 1 Pro Fast

ByteDance

290 credits

ByteDance's cinematic video generation model with fast generation speed and professional output quality.

Fast generation speed
Professional output quality
Cinematic video generation
FastHigh
View Details

OpenAI Sora 2

OpenAI

930 credits

OpenAI's video generation model with realistic physics simulation and audio generation capabilities, producing highly coherent videos.

Realistic physics simulation
Audio generation support
Up to 20-second videos
SlowUltra
View Details

OpenAI Sora 2 Pro

OpenAI

2790 credits

OpenAI's most advanced synced-audio video generation model. The premium tier of Sora 2 with higher fidelity, up to 1024p resolution, and image-to-video via an input reference frame.

OpenAI's most advanced video model with natively synced audio
Premium tier of Sora 2 — higher fidelity and detail
Two resolution tiers: standard (720p) and high (1024p)
SlowUltra
View Details

MiniMax Hailuo-03 (H3) Text to Video

MiniMax

3020 credits

MiniMax Hailuo-03 (H3) text-to-video via Fal.AI. State-of-the-art 2K video generation from a text prompt, 5-15 second duration, native audio, and wide aspect-ratio support (21:9 through 9:16).

State-of-the-art 2K resolution text-to-video generation
Native audio generation alongside video
5-15 second configurable duration
MediumUltra
View Details

Veed Lipsync v2

Veed

2440 credits

Veed Lipsync v2 via Fal.AI. Replaces a source video's mouth movements to articulate a new audio track — takes any source video + a new audio track and produces a lip-synced output video.

Re-syncs mouth movements in a source video to a new audio track
Works with any source video + audio track pairing
Output video length automatically matches the input video/audio
MediumHigh
View Details

Google Veo 3.1

Google

7440 credits

Google's state-of-the-art video generation model with built-in audio generation, producing cinematic-quality videos with synchronized sound.

Built-in audio generation
Cinematic-quality output
Image-to-video with start frame
SlowUltra
View Details

Google Veo 3.1 Fast

Google

2790 credits

Fast version of Veo 3.1 with audio generation, optimized for speed while maintaining high quality output.

Faster generation than Veo 3.1
Built-in audio generation
High-quality output
MediumHigh
View Details

Video Utils

FFmpeg

5 credits

FFmpeg-powered video utilities (20M+ runs) — convert to mp4/gif, extract audio as mp3, or dump zipped frames, all with one task parameter. The Swiss-army knife for media pipelines.

Convert any video to mp4 or gif
Extract audio as mp3
Dump frames as a zip (with fps control)
FastStandard
View Details

Wan 2.7 I2V

Alibaba

1740 credits

Alibaba Wan 2.7 image-to-video model. Animates a first frame (and optional last frame or continuation clip) with audio synchronization, up to 15 seconds, at 720p or 1080p.

Image-to-video animation from a first frame
Optional last-frame targeting or continuation from an existing video clip
Audio synchronization for voice and music (auto-generated or user-supplied)
MediumHigh
View Details

Wan 2.7 T2V

Alibaba

1740 credits

Alibaba Wan 2.7 text-to-video model. Supports up to 15 seconds, audio synchronization for voice/music, multilingual prompts, and prompt expansion, at 720p or 1080p.

Text-to-video generation up to 15 seconds
Audio synchronization for voice and music (auto-generated or user-supplied)
Multilingual prompt support
MediumHigh
View Details

Wan 2.5 I2V

Alibaba

1160 credits

Image-to-video model with lip sync support, animating still images into realistic videos with natural motion.

Image-to-video animation
Lip sync support
Natural motion from still images
SlowUltra
View Details

Wan 2.5 I2V Fast

Alibaba

790 credits

Fast image-to-video variant of Wan 2.5, optimized for rapid generation of animated videos from still images.

Faster generation than Wan 2.5 I2V
Image-to-video animation
Good motion quality
MediumHigh
View Details

Wan 2.5 T2V

Alibaba

2320 credits

Text-to-video model with audio synchronization support, producing high-quality videos from text prompts with natural motion.

High-quality text-to-video generation
Audio synchronization support
Natural motion synthesis
SlowUltra
View Details

Wan 2.5 T2V Fast

Alibaba

790 credits

Fast text-to-video generation variant of Wan 2.5, optimized for speed with good quality output.

Faster generation than Wan 2.5 T2V
Good quality text-to-video
Audio synchronization support
MediumHigh
View Details

Wan 2.2 I2V Fast

Alibaba

260 credits

A very fast and cheap PrunaAI-optimized version of Alibaba's Wan 2.2 A14B image-to-video model. Turns a single still image plus a prompt into a short animated clip at 480p or 720p, with an optional frame-interpolation pass for smoother motion.

PrunaAI-optimized inference — fast and low-cost compared to the base Wan 2.2 model
Image-to-video with optional last-frame conditioning for smoother transitions
Optional 30 FPS frame interpolation (ffmpeg-based) for silkier motion
FastStandard
View Details

Wan 2.2 T2V Fast

Alibaba

240 credits

A very fast and cheap PrunaAI-optimized version of Alibaba's Wan 2.2 A14B text-to-video model. Generates short clips at 480p or 720p directly from a text prompt, with 30 FPS frame interpolation enabled by default — the text-to-video sibling of Wan 2.2 I2V Fast.

PrunaAI-optimized inference — fast and low-cost compared to the base Wan 2.2 model
Pure text-to-video: no reference image required, just a prompt
30 FPS frame interpolation (ffmpeg-based) enabled by default for smoother motion
FastStandard
View Details

Audio & TTS Models

Text-to-speech, voice cloning, and audio generation

Bark

Featured

Suno

190 credits

Suno's text-prompted generative audio model. Produces speech with nonverbal sounds like [laughs] and [sighs], plus music and sound effects, in 100+ speaker presets across 13 languages. Returns audio plus an optional .npz history file for voice continuity.

Speech with nonverbal sounds: [laughs], [sighs], [gasps]
Music notes (♪) and sound effect generation
131 speaker presets across 13 languages
SlowStandard
View Details

Chatterbox

Resemble AI

59 credits

Resemble AI's production-grade open-source TTS with unique emotion exaggeration control and instant voice cloning from a short reference audio. MIT-licensed and benchmarked against leading closed-source systems.

Unique emotion exaggeration control
Instant voice cloning from short audio
Production-grade open-source quality
FastHigh
View Details

Chatterbox Multilingual

Resemble AI

7 credits

Chatterbox open-source TTS in 23 languages with instant voice cloning and emotion exaggeration control. Max 300 characters per request.

23 language support including Korean
Instant voice cloning from short audio
Emotion exaggeration control
FastHigh
View Details

Chatterbox Turbo

Resemble AI

59 credits

Resemble AI's fastest open-source TTS without sacrificing quality. 20 pre-made voices, paralinguistic tags like [sigh] and [chuckle], and optional instant voice cloning from 5s+ reference audio. Max 500 characters per request.

Fastest open-source TTS quality tier
Paralinguistic tags: [sigh], [chuckle], [cough], [clear throat]
20 pre-made voices
FastHigh
View Details

Gemini 3.1 Flash TTS

Google

299 credits

Google Gemini 3.1 Flash native-audio TTS via Replicate. Style instructions (separate from the spoken text) plus 30 preset voices and 40+ language/locale codes, with inline markup tags for expressive delivery ([sigh], [whispering], [shouting], etc.).

Native-audio TTS powered by Gemini 3.1 Flash
Separate style-instruction field (prompt) independent from the spoken text
30 preset voices
FastHigh
View Details

Incredibly Fast Whisper

Whisper

10 credits

Whisper large-v3 optimized for speed (38M+ runs) — transcribes roughly 150 minutes of audio in under 100 seconds using batched inference. Chunk-level or word-level timestamps.

~150 minutes of audio transcribed in under 100 seconds
Whisper large-v3 accuracy
Chunk or word-level timestamps
FastUltra
View Details

Suno Music V5.5

Suno

140 credits

Suno's latest model, with the highest audio quality and richest arrangements. Creates two complete songs from a text description, or from exact lyrics with style and title in custom mode. Returns audio, cover image, and metadata for each track.

Two complete songs per generation
Highest audio quality and richest arrangements in the Suno lineup
Custom mode with exact lyrics, style, and title
MediumUltra
View Details

Suno Music V5

Suno

140 credits

Suno V5 music generation with improved audio quality and musicality. Creates two complete songs from a text description, or from exact lyrics with style and title in custom mode. Returns audio, cover image, and metadata for each track.

Two complete songs per generation
Improved audio quality and musicality over V4.5
Custom mode with exact lyrics, style, and title
MediumHigh
View Details

Suno Music V4.5

Suno

140 credits

Suno V4.5 music generation. Creates two complete songs from a text description, or from exact lyrics with style and title in custom mode. Returns audio, cover image, and metadata for each track.

Two complete songs per generation
Custom mode with exact lyrics, style, and title
Instrumental-only option
MediumHigh
View Details

MiniMax Music 2.6

MiniMax

350 credits

Generate full-length songs or instrumentals (up to ~6 minutes) from a text prompt. 99%+ accurate BPM/key control, 14+ structure tags, auto-generated lyrics, and instrumental-only mode.

Full-length songs up to ~6 minutes (most 2-4 min)
99%+ accurate BPM and key control via prompt
14+ structure tags: [Intro], [Verse], [Chorus], [Hook], [Drop], [Bridge]...
MediumHigh
View Details

MiniMax Music 2.5

MiniMax

350 credits

Generate full-length songs (up to ~5 minutes) with vocals, lyrics, and rich instrumentation. 14 structure tags, style-aware mixing, and an expanded instrument library including orchestral and traditional instruments.

Full-length songs up to ~5 minutes (most 2:30-4:30)
14 structure tags: [Intro], [Verse], [Chorus], [Hook], [Drop], [Bridge]...
Realistic vocals with breathing and pitch transitions
MediumHigh
View Details

MiniMax Music Cover

MiniMax

700 credits

Reimagine any song in a different style. The model extracts the melodic structure from the input audio and regenerates the track — the melody and duration stay the same, but voice, instruments, genre, and arrangement can all change.

Melody and song duration preserved
Voice, instruments, genre, arrangement all changeable
Genre transformation: pop→jazz, rock→bossa nova, ballad→EDM
MediumHigh
View Details

Qwen3 TTS

Qwen

47 credits

Alibaba's unified TTS with three modes: preset speakers (custom_voice), instant voice cloning from reference audio (voice_clone), and creating a brand-new voice from a text description (voice_design).

Three modes: preset voice, voice cloning, voice design
Voice design from natural-language description
9 preset speakers including Korean (Sohee)
FastHigh
View Details

Inworld Realtime TTS 2

Inworld

59 credits

Inworld's most expressive TTS with natural-language steering — place bracketed instructions like [speak quickly] before the text they apply to. Real-time latency and 15+ language support.

Natural-language steering with bracketed instructions
Real-time latency
15+ language support with auto detection
FastHigh
View Details

Suno SFX V5

Suno

29 credits

Suno V5 sound effect generation. Creates two short sound effect variants from a text description, with optional loop mode, tempo, and musical key controls. Ideal for UI sounds, game audio, and video foley.

Two SFX variants per generation
Seamless loop mode for game/ambient audio
Tempo (BPM) and musical key controls
FastHigh
View Details

MiniMax Speech 2.8 HD

MiniMax

237 credits

Ranked #1 on Artificial Analysis Speech Arena and Hugging Face TTS Arena. Broadcast-quality TTS with autoregressive Transformer + Flow-VAE decoder, 32+ languages, voice cloning, natural interjections, and emotion control.

#1 ranked on major TTS benchmarks
Studio-grade broadcast quality audio
Natural interjections (laughs, sighs, coughs, etc.)
MediumUltra
View Details

MiniMax Speech 2.8 Turbo

MiniMax

142 credits

Low-latency MiniMax Speech 2.8 Turbo with under 250ms latency, 40+ languages, voice cloning, natural interjections, and real-time pricing. Ideal for interactive and real-time applications.

Under 250ms latency for real-time use
40+ language support
Natural interjections (laughs, sighs, coughs, etc.)
FastHigh
View Details

MiniMax Speech 2.6 HD

MiniMax

237 credits

Studio-quality multilingual text-to-speech with nuanced prosody, emotion control, and premium voices for professional applications.

Studio-grade audio quality
Nuanced prosody control
Subtitle timestamp export
MediumUltra
View Details

MiniMax Speech 2.6 Turbo

MiniMax

142 credits

Fast multilingual text-to-speech with emotional control, optimized for real-time applications with low latency.

Ultra-low latency synthesis
Real-time streaming support
300+ voice presets
FastHigh
View Details

MiniMax Speech-02-Turbo

MiniMax

142 credits

Low-latency text-to-speech model with multilingual support, emotional voice control, and 300+ voice options.

Low-latency real-time synthesis
300+ voice presets
Emotional expression control
FastHigh
View Details

Sonilo v1.1 Text to Music

Sonilo

520 credits

Sonilo v1.1 text-to-music via Fal.AI. Generates up to 3 distinct music tracks (up to 600 seconds each) from a text description of the desired sound.

Generates full music tracks from a text description
Up to 600 seconds (10 minutes) per track
Generate up to 3 distinct track variations in a single call
FastStandard
View Details

Inworld TTS 1.5 Max

Inworld

24 credits

Inworld's highest-quality realtime TTS with under 200ms latency. Supports SSML break tags for pauses, emotion markups like [happy], and 15 languages.

Under 200ms median latency
Emotion markups: [happy], [sad], [angry]
SSML break tags for precise pauses
FastHigh
View Details

Clova Voice TTS Premium

NCP Clova

153 credits

NAVER Clova Voice Premium TTS with 108 voices across 6 languages. High-quality Korean voice synthesis with emotion control, Pro voices, and bilingual support.

108 voice options across 6 languages
High-quality Korean voice synthesis
Emotion control (neutral, sad, happy, angry)
FastHigh
View Details

ElevenLabs Turbo v2.5

ElevenLabs

119 credits

High-quality, low-latency ElevenLabs text-to-speech in 32 languages. The same 26 premium voices as v3 at half the price, optimized for real-time and high-volume use.

Low latency optimized for real-time apps
32 language support
26 premium preset voices
FastHigh
View Details

ElevenLabs v3

ElevenLabs

237 credits

ElevenLabs' most expressive text-to-speech model. 26 premium voices, inline audio tags like [laughs] and [whispers], fine-grained style and stability controls, and 70+ language support.

Most expressive ElevenLabs model
Inline audio tags: [laughs], [whispers], [sighs], [excited]
26 premium preset voices
MediumUltra
View Details

ElevenLabs v2 Multilingual

ElevenLabs

237 credits

ElevenLabs Multilingual v2 text-to-speech in over 30 languages. Stable, proven voice quality with the same premium voice lineup and fine-grained voice settings.

30+ language support
Stable, production-proven voice quality
26 premium preset voices
MediumHigh
View Details

Sonilo v1.1 Video to Sound Effects

Sonilo

310 credits

Sonilo v1.1 video-to-sound-effects via Fal.AI. Adds AI-generated sound (ambience, effects, foley) to an input video — auto-captions the scene if no prompt is given, or accepts per-segment sound descriptions for finer control.

Adds ambience, sound effects, and foley to silent or under-scored video
Auto-captions the scene when no prompt is given
Per-segment sound descriptions for finer, time-ranged control
FastStandard
View Details

Suno Vocal Separation

Suno

120 credits

Separate a Suno-generated track into clean vocal and instrumental stems. Takes the taskId and audioId from a previous Suno music generation and returns separated vocalUrl/instrumentalUrl audio files.

Clean vocal + instrumental stem separation
Chains directly from Suno music generation output
Both stems mirrored to CDN as downloadable MP3s
FastHigh
View Details

MiniMax Voice Cloning

MiniMax

6980 credits

Clone any voice from a 10-second to 5-minute audio sample. Returns a custom voice_id you can pass to MiniMax speech models, plus a preview clip synthesized with the cloned voice.

Voice cloning from as little as 10 seconds of audio
Cloned voice_id works with MiniMax speech models
Preview audio clip included in the result
MediumHigh
View Details

Whisper

OpenAI

10 credits

OpenAI's Whisper large-v3 speech recognition (144M+ runs) — the standard for transcription. Automatic language detection across ~100 languages, English translation, and plain text / SRT / VTT output formats.

The standard for speech-to-text (144M+ runs)
~100 languages with auto detection (Korean included)
Translate any language to English
MediumUltra
View Details

Whisper Diarization

Whisper

7 credits

Whisper transcription with speaker diarization (8M+ runs) — returns who said what, with per-segment speaker labels and timestamps. The go-to for meetings and interviews.

Speaker-labeled transcription (who said what)
Per-segment timestamps
Auto speaker-count detection (or specify)
FastHigh
View Details

XTTS-v2

Coqui

26 credits

Coqui XTTS-v2 multilingual voice-cloning TTS. Clone a voice from a single short audio sample and speak in 16 languages — one of the most popular open-source voice cloning models.

Voice cloning from one short audio sample
16 language support including Korean
Cross-language cloning (speak any language in the cloned voice)
MediumStandard
View Details

LLM Models

GPT-4o, Claude, Gemini - OpenAI-compatible chat API

Claude Haiku 4.5

Featured

Anthropic

1 credits

Fast, cost-effective model for everyday tasks. Great balance of speed, intelligence, and cost for high-volume applications.

200K context window
8K max output tokens
Vision capabilities
FastHigh
View Details

Claude Opus 5

Anthropic

5 credits

Anthropic's latest flagship Opus model, with a 1M-token context window by default and 128K max output tokens. Same pricing as Opus 4.5–4.8 ($5/$25 per M tokens) with prompt caching (read $0.50/M, write $6.25/M) and web search. Compatible with the Anthropic Messages and OpenAI Chat Completions formats.

Most capable Anthropic model
1M token context window (default)
128K max output tokens
MediumUltra
View Details

Claude Opus 4.8

Anthropic

5 credits

Anthropic's most capable Opus-tier model, with a 1M-token context window (200K on some surfaces), 128K max output tokens, and knowledge cutoff to January 2026. Builds on Opus 4.7 with stronger long-horizon agentic coding, better tool triggering, and adaptive thinking that reasons only when a turn needs it. Compatible with the Anthropic Messages and OpenAI Chat Completions formats.

Most capable Anthropic model
1M token context window
128K max output tokens
MediumUltra
View Details

Claude Opus 4.7

Anthropic

5 credits

Anthropic's latest flagship model with reliable knowledge cutoff to January 2026 and 128K max output tokens. Builds on Opus 4.6 with improved reasoning, coding, and instruction-following while staying compatible with the Anthropic Messages and OpenAI Chat Completions formats.

Most capable Anthropic model
200K context window
128K max output tokens
MediumUltra
View Details

Claude Opus 4.6

Anthropic

5 credits

Anthropic's most capable model. Delivers breakthrough performance in reasoning, coding, and complex analysis with enhanced safety and instruction following.

Most capable Anthropic model
200K context window
Enhanced reasoning and coding
MediumUltra
View Details

Claude Opus 4.5

Anthropic

5 credits

Anthropic's most powerful model for highly complex tasks. Exceptional at research, analysis, and creative projects requiring deep expertise.

200K context window
32K max output tokens
Vision capabilities
MediumUltra
View Details

Claude Sonnet 5

Anthropic

4 credits

Anthropic's newest Sonnet model, tuned for the best balance of speed, cost, and intelligence. Features a 1M-token context window, 64K max output tokens, and knowledge cutoff to August 2025. Supports adaptive thinking that reasons only when a turn needs it, vision, and tool use. Compatible with the Anthropic Messages and OpenAI Chat Completions formats.

Balanced speed, cost, and intelligence
1M token context window
64K max output tokens
FastUltra
View Details

Claude Sonnet 4.5

Anthropic

4 credits

Anthropic's most intelligent and capable Sonnet model. Best-in-class for complex reasoning, nuanced understanding, and coding tasks with exceptional instruction following.

200K context window
16K max output tokens
Vision capabilities
FastUltra
View Details

Claude Sonnet 4

Anthropic

3 credits

Balanced Sonnet 4 model offering strong reasoning and coding abilities at an efficient price point. Ideal for everyday production workloads that need a good mix of speed and intelligence.

200K context window
16K max output tokens
Vision capabilities
FastHigh
View Details

Gemini 3.5 Flash

Google

2 credits

Google's latest Flash-tier workhorse model. Flat pricing across all input modalities (text, image, video, audio, PDF) at $1.50/$9 per million tokens with no long-context premium, a 1M token context window, and 65,536 max output tokens. Cached inputs get a 90% discount.

Flat pricing across all input modalities
No long-context premium
1M token context window
FastHigh
View Details

Gemini 3.1 Flash Image Preview

Google

500 credits

Gemini 3.1 Flash with native image generation capabilities. Can generate images directly in chat responses alongside text. Features separate pricing for text and image output tokens.

Native image generation in chat
131,072 token context window
32,768 max output tokens
FastHigh
View Details

Gemini 3.1 Flash Lite

Google

1 credits

Stable release of the ultra-lightweight Gemini 3.1 Flash Lite. The most cost-effective Gemini model at $0.25/$1.50 per million tokens, with cached input (90% discount), audio input, and batch (50% discount) support. Ideal for high-throughput, budget-conscious applications.

Most cost-effective Gemini model
Stable model ID (no preview suffix)
1,048,576 token context window
FastStandard
View Details

Gemini 3.1 Flash Lite Preview

Google

100 credits

Ultra-lightweight variant of Gemini 3.1 Flash. The most cost-effective Gemini model with support for cached input and audio input. Ideal for high-throughput, budget-conscious applications.

Most cost-effective Gemini model
1,048,576 token context window
65,536 max output tokens
FastStandard
View Details

Gemini 3.1 Flash Live Preview

Google

300 credits

Gemini 3.1 Flash optimized for real-time interactions and live streaming scenarios. Features low-latency responses with audio input support at dedicated pricing.

Live API support (real-time bidirectional)
Low-latency streaming responses
131,072 token context window
FastHigh
View Details

Gemini 3.1 Pro Preview

Google

500 credits

Google's latest and most capable Gemini model in preview. Features dynamic pricing that adjusts based on context length, with enhanced pricing for inputs over 200K tokens.

Dynamic pricing (standard / long-context >200K)
Advanced reasoning and analysis
1,048,576 token context window
MediumUltra
View Details

Gemini 3 Flash

Google

500 credits

Google's most advanced reasoning model with state-of-the-art multimodal understanding, PhD-level reasoning, and leading coding performance.

PhD-level reasoning ability
1,048,576 token context window
65,536 max output tokens
MediumUltra
View Details

Gemini 3 Pro Image Preview

Google

500 credits

Google's premium image generation model within the Gemini 3 Pro family. Generates high-quality images directly in chat with the highest fidelity among Gemini image models. Image output tokens are priced at 10x text output tokens.

Premium image generation quality
65,536 token context window
32,768 max output tokens
MediumUltra
View Details

Gemini 3 Pro Preview

Google

4 credits

Google's most powerful Gemini model in preview. Features breakthrough reasoning, coding, and multimodal capabilities with the largest context window.

1,048,576 token context window
65,536 max output tokens
Multimodal input: text, image, video, audio, PDF
MediumUltra
View Details

Gemini 2.5 Flash

Google

1 credits

Google's fast and efficient model with built-in thinking capabilities. Great balance of speed, reasoning, and cost for high-volume applications.

1,048,576 token context window
65,536 max output tokens
Multimodal input: text, image, video, audio
FastHigh
View Details

Gemini 2.5 Pro

Google

3 credits

Google's most capable model with state-of-the-art reasoning and 1M token context. Excels at complex coding, math, and multi-document analysis.

1,048,576 token context window
65,536 max output tokens
Multimodal input: audio, image, video, text, PDF
MediumUltra
View Details

Gemini 2.0 Flash

Google

1 credits

Google's fastest and most capable model. Features a massive 1M token context window, native multimodal support, and real-time capabilities.

1,048,576 token context window
8,192 max output tokens
Native multimodal (text, image, audio, video)
FastHigh
View Details

Gemini 2.0 Flash Lite

Google

0.5 credits

Ultra-lightweight version of Gemini 2.0 Flash optimized for maximum speed and minimal cost. Perfect for high-volume, latency-sensitive applications.

Ultra-fast inference
Minimal cost per request
1,048,576 token context window
FastStandard
View Details

Gemini Embedding 001

Google

0.1 credits

Google's text embedding model for generating vector representations. Optimized for semantic search, clustering, and similarity tasks.

High-quality text embeddings
2,048 token max input per request
Output dimensions: 128–3,072 (default 3,072; recommended 768 / 1,536 / 3,072)
FastHigh
View Details

GPT-5.6 Luna

OpenAI

1 credits

The fast, low-cost tier of OpenAI's GPT-5.6 family (GA July 2026). Luna is built for high-throughput workloads at $1/$6 per million tokens, with a 1M token context window, 128K max output tokens, and cached inputs at a 90% discount.

Fast, low-cost tier of the GPT-5.6 family
1M token context window
128K max output tokens
FastHigh
View Details

GPT-5.6 Sol

OpenAI

5 credits

The flagship tier of OpenAI's GPT-5.6 family (GA July 2026). Sol delivers the strongest reasoning, coding, and multimodal performance of the generation with a 1M token context window, 128K max output tokens, and cached inputs at a 90% discount.

Flagship tier of the GPT-5.6 family
1M token context window
128K max output tokens
FastUltra
View Details

GPT-5.6 Terra

OpenAI

3 credits

The balanced tier of OpenAI's GPT-5.6 family (GA July 2026). Terra offers near-flagship quality at half the price of Sol, with a 1M token context window, 128K max output tokens, and cached inputs at a 90% discount.

Balanced tier of the GPT-5.6 family
1M token context window
128K max output tokens
FastHigh
View Details

GPT-5.5

OpenAI

5 credits

OpenAI's newest flagship model with a 1.05M token context window and 128K max output tokens. Supports cached inputs at 10× discount and improved reasoning, coding, and multimodal performance over the GPT-5.4 series.

1.05M token context window
128K max output tokens
Cached input pricing (10× cheaper)
FastUltra
View Details

GPT-5.4

OpenAI

5 credits

OpenAI's newest flagship model with 1M context window and 128K output tokens. Delivers top-tier reasoning across all domains with adjustable reasoning effort levels from none to xhigh.

1M token context window
128K max output tokens
Adjustable reasoning (none/low/medium/high/xhigh)
FastUltra
View Details

GPT-5.4 Mini

OpenAI

2 credits

Fast and cost-efficient variant of GPT-5.4 with 400K context window and 128K output tokens. Excellent balance of performance and affordability for everyday tasks.

400K context window
128K max output tokens
Fast inference speed
FastHigh
View Details

GPT-5.4 Nano

OpenAI

1 credits

Ultra-lightweight and fastest GPT-5.4 variant with 400K context and 128K output. Designed for high-throughput, low-latency applications at minimal cost. Supports MCP for tool integration.

400K context window
128K max output tokens
Ultra-fast inference
FastStandard
View Details

GPT-5.2

OpenAI

4 credits

OpenAI's latest and most advanced GPT model. Delivers state-of-the-art performance across reasoning, coding, and creative tasks with enhanced capabilities.

Latest GPT architecture
256K context window
32K max output tokens
FastUltra
View Details

GPT-5.1 (2025-11-13)

OpenAI

3 credits

Dated snapshot of GPT-5.1 for reproducible results. Supports cached input tokens for cost savings on repeated context. Ideal for production deployments requiring model version pinning.

Fixed model snapshot for reproducibility
Cached input tokens for cost savings
Strong reasoning and coding performance
FastHigh
View Details

GPT-5

OpenAI

3 credits

OpenAI's latest flagship model. Delivers exceptional performance across reasoning, coding, and creative tasks with a massive 1M token context window and 32K output tokens. Supports vision, function calling, and JSON mode.

1M token context window
32K max output tokens
Advanced reasoning capabilities
FastUltra
View Details

GPT-5 Mini

OpenAI

1 credits

Fast and efficient variant of GPT-5. Delivers strong performance across reasoning, coding, and creative tasks with a 1M token context window and 32K output tokens, at a fraction of the cost of GPT-5.

1M token context window
32K max output tokens
Fast inference speed
FastHigh
View Details

GPT-5 Nano

OpenAI

1 credits

Ultra-fast and lightweight variant of GPT-5. Designed for high-throughput, low-latency applications with a 1M token context window and 32K output tokens at minimal cost.

1M token context window
32K max output tokens
Ultra-fast inference
FastHigh
View Details

GPT-4.1

OpenAI

3 credits

OpenAI's most capable model for coding and instruction following. Features a 1M token context window, 32K output tokens, and major improvements in coding, complex prompts, and long-context tasks. 20% cheaper than GPT-4o on output.

1M token context window
32K max output tokens
Best-in-class coding performance
FastUltra
View Details

GPT-4.1 Mini

OpenAI

1 credits

A significant leap in small model performance. Matches or exceeds GPT-4o in intelligence while reducing latency by nearly half and cost by 83%. Ideal balance of speed, quality, and affordability.

1M token context window
32K max output tokens
Matches GPT-4o intelligence at fraction of cost
FastHigh
View Details

GPT-4.1 Nano

OpenAI

1 credits

OpenAI's fastest and cheapest model. Optimized for classification, autocompletion, and low-latency tasks. Ultra-affordable at $0.10/1M input tokens.

1M token context window
32K max output tokens
Ultra-low latency
FastStandard
View Details

GPT-4o

OpenAI

3 credits

OpenAI's flagship multimodal model. Industry-leading performance in reasoning, coding, and creative tasks with native vision capabilities and structured output support.

Native multimodal (text + vision + audio)
128K context window
16K max output tokens
FastUltra
View Details

GPT-4o Mini

OpenAI

1 credits

Cost-effective, fast model with strong performance. Best for high-volume tasks where speed and cost matter more than absolute capability.

128K context window
16K max output tokens
Vision capabilities
FastHigh
View Details

GPT Audio Mini

OpenAI

1 credits

Lightweight multimodal model with native audio input/output capabilities. Optimized for voice-based interactions and audio processing tasks.

Native audio input/output
Voice-based interactions
Cost-effective multimodal
FastHigh
View Details

MiniMax M2.7

MiniMax

1 credits

MiniMax's flagship M2-series language model, served through an OpenAI-compatible API. Strong multilingual capability (notably Chinese and English) at a very low price point ($0.30/$1.20 per million tokens) with prompt cache reads at $0.06/M.

OpenAI-compatible API (SDK drop-in)
Very low price point ($0.30/$1.20 per M tokens)
Prompt cache read discount ($0.06/M)
FastHigh
View Details

OpenAI o4-mini

OpenAI

2 credits

Fast, cost-effective reasoning model optimized for coding and STEM tasks. Provides strong reasoning at a fraction of the cost of larger reasoning models.

Optimized for coding/STEM
200K context window
100K max output tokens
FastUltra
View Details

OpenAI o3-mini

OpenAI

2 credits

Efficient reasoning model that delivers strong performance at lower cost. Ideal for tasks requiring reasoning without the overhead of larger models.

Efficient reasoning capabilities
200K context window
100K max output tokens
FastHigh
View Details

OpenAI o1

OpenAI

15 credits

OpenAI's most advanced reasoning model. Uses extended thinking time to solve complex problems in science, coding, and math with exceptional accuracy.

Advanced chain-of-thought reasoning
200K context window
100K max output tokens
SlowUltra
View Details

Pricing Comparison

Compare credit costs across all models to find the best fit for your needs

CategoryModelCreditsUnitBest For
Imageflux-schnell7per imageFast generation, real-time apps
flux-2-max160per image at 1 MP — varies by resolution (0.5MP 130, 2MP 230, 4MP/match_input 370, custom 390)Best quality, commercial
seedream-4.590per image — sequential mode (auto) scales per generated image, up to max_images (15)Text rendering, logos
Videoseedance-1-pro-fast290per videoLow-cost fast generation
hailuo-2.3650per videoHigh quality, balanced
veo-3.17,440per videoBest quality video
Audiospeech-02-turbo142per 1000 characters (0.135 credits/char, billed per character)Real-time TTS, fast response
speech-2.8-hd237per 1000 characters (0.225 credits/char, billed per character)High quality HD voice
LLMgpt-4o3per 1K tokens (avg)General, code, multimodal
claude-sonnet-54per 1K tokens (avg)Analysis, long-form, code
gemini-2.0-flash1per 1K tokens (avg)Ultra-fast, bulk processing

Ready to Get Started?

Try our models in the Playground or follow the Quickstart guide