Vision & Utilities
음성 인식(Whisper), OCR·이미지 캡셔닝, 임베딩, 3D 생성, 립싱크·토킹헤드, 비디오 향상까지 — 생성 외의 모든 AI 유틸리티를 하나의 API로 사용하세요.
모델 비교표
| 모델 | 크레딧 | 유형 | 특징 |
|---|---|---|---|
| salesforce/blip | 1 | OCR/캡션 | Salesforce, 빠른 처리 |
| andreasjansson/clip-features | 1 | 임베딩 | CLIP, 빠른 처리 |
| falcons-ai/nsfw_image_detection | 1 | 안전 필터 | Falcons.ai, 빠른 처리 |
| abiruyt/text-extract-ocr | 1 | OCR/캡션 | OCR, 빠른 처리 |
| charlesmccarthy/addwatermark | 1 | 비디오 유틸 | FullJourney, 빠른 처리 |
| lucataco/florence-2-large | 2 | 객체 감지 | Microsoft, 빠른 처리 |
| adirik/grounding-dino | 2 | 객체 감지 | Grounding DINO, 빠른 처리 |
| beautyyuyanli/multilingual-e5-large | 2 | 임베딩 | E5, 빠른 처리 |
| lucataco/moondream2 | 5 | OCR/캡션 | Moondream, 빠른 처리 |
| nicolascoutureau/video-utils | 5 | 비디오 유틸 | FFmpeg, 빠른 처리 |
| thomasmol/whisper-diarization | 7 | 음성 인식(STT) | Whisper, 빠른 처리 |
| vaibhavs10/incredibly-fast-whisper | 10 | 음성 인식(STT) | Whisper, 빠른 처리 |
| openai/whisper | 10 | 음성 인식(STT) | OpenAI, 최고 품질 |
| zsxkib/mmaudio | 11 | 비디오 유틸 | MMAudio, 고품질 |
| chenxwh/cogvlm2-video | 28 | 비디오 유틸 | CogVLM, 고품질 |
| firtoz/trellis | 77 | 3D 생성 | Microsoft, 고품질 |
| cjwbw/sadtalker | 91 | 립싱크/아바타 | SadTalker, 고품질 |
| bytedance/latentsync | 220 | 립싱크/아바타 | ByteDance, 고품질 |
| tencent/hunyuan-3d-3.1 | 1,160 | 3D 생성 | Tencent, 최고 품질 |
| lucataco/real-esrgan-video | 1,280 | 비디오 향상 | Real-ESRGAN, 고품질 |
| veed/lipsync/v2 | 2,440 | 립싱크/아바타 | Veed, 고품질 |
오디오·비디오·이미지 파일은 POST /v1/files/upload-url로 업로드한 뒤 해당 URL을 입력으로 사용하세요. 처리 시간이 파일 길이에 비례하는 모델은 상세 문서에 표기되어 있습니다.
모델 상세 정보
각 모델의 상세한 파라미터, 예제, 활용 팁은 개별 문서에서 확인하세요.
Incredibly Fast Whisper
Whisper
Whisper large-v3 optimized for speed (38M+ runs) — transcribes roughly 150 minutes of audio in under 100 seconds using batched inference. Chunk-level or word-level timestamps.
Whisper
OpenAI
OpenAI's Whisper large-v3 speech recognition (144M+ runs) — the standard for transcription. Automatic language detection across ~100 languages, English translation, and plain text / SRT / VTT output formats.
Whisper Diarization
Whisper
Whisper transcription with speaker diarization (8M+ runs) — returns who said what, with per-segment speaker labels and timestamps. The go-to for meetings and interviews.
BLIP
Salesforce
Salesforce BLIP (173M+ runs) — image captioning, visual question answering, and image-text matching in one model. The classic choice for bulk captioning at 1 credit per image.
CLIP Features
CLIP
CLIP ViT-L/14 embeddings for text AND images (163M+ runs) — puts both in the same vector space for cross-modal search, image dedup, and zero-shot classification. 1 credit per run.
Florence-2 Large
Microsoft
Microsoft's Florence-2 all-in-one vision model — captioning, object detection, phrase grounding, OCR, and segmentation in a single API. Pick a task, optionally add text input, done.
Grounding DINO
Grounding DINO
Text-prompted object detection (39M+ runs) — describe what to find in natural language ('red car, person wearing a hat') and get bounding boxes with confidence scores plus an annotated image.
Hunyuan 3D 3.1
Tencent
Tencent's flagship 3D generation — create high-polygon textured 3D models from a text prompt OR an image, with optional PBR (physically based rendering) materials.
Moondream2
Moondream
Small but capable vision-language model (14M+ runs) — ask free-form questions about any image and get detailed answers. Efficient VQA for tagging, moderation prep, and rich alt-text.
Multilingual E5 Large
E5
Multilingual text embeddings (74M+ runs) — 1024-dimension vectors across 100 languages including Korean. Pairs perfectly with Core.Today customer databases' vector search (knn_vector).
NSFW Image Detection
Falcons.ai
The standard NSFW image classifier (127M+ runs) — returns 'normal' or 'nsfw' for any image. An essential, ultra-cheap moderation gate for UGC platforms.
Text Extract OCR
OCR
Simple, massively-used OCR (91M+ runs) — extracts text from an image with a single input and returns plain text. Great default for receipts, screenshots, and scanned documents at just 1 credit.
TRELLIS
Microsoft
Microsoft's TRELLIS image-to-3D (836K+ runs — the most-used 3D model on Replicate). Turns one or more images into a textured GLB 3D asset, with turntable render videos and optional Gaussian PLY.
Add Watermark
FullJourney
Add a text watermark to any video — simple, fast brand protection for generated or user content at 2 credits per video.
CogVLM2 Video
CogVLM
Video understanding and captioning — ask free-form questions about a video and get detailed answers about actions, scenes, and content. Great for video search indexing and moderation prep.
LatentSync
ByteDance
ByteDance's open-source lipsync — re-syncs a video's mouth movements to any audio track using latent diffusion. State-of-the-art open lipsync quality for dubbing and localization.
MMAudio
MMAudio
Add AI-generated sound to any video (5M+ runs) — synthesizes synchronized audio (ambience, effects, foley) from the video content and an optional text prompt. The perfect finisher for silent AI-generated clips.
Real-ESRGAN Video
Real-ESRGAN
Video upscaling with Real-ESRGAN — enhance videos to FHD, 2K, or 4K frame by frame. The go-to open-source video upscaler for old footage and AI-generated clips.
SadTalker
SadTalker
Talking-head video from a single photo and an audio track — animates the face with natural head motion and eye blinks, with optional GFPGAN face enhancement.
Veed Lipsync v2
Veed
Veed Lipsync v2 via Fal.AI. Replaces a source video's mouth movements to articulate a new audio track — takes any source video + a new audio track and produces a lip-synced output video.
Video Utils
FFmpeg
FFmpeg-powered video utilities (20M+ runs) — convert to mp4/gif, extract audio as mp3, or dump zipped frames, all with one task parameter. The Swiss-army knife for media pipelines.
빠른 시작 — 음성 인식 (Whisper)
오디오 URL 하나로 100개 언어 자동 감지 전사를 받을 수 있습니다 (13크레딧):
curl -X POST https://api.core.today/v1/predictions \
-H "Content-Type: application/json" \
-H "X-API-Key: cdt_your_api_key" \
-d '{
"model": "openai/whisper",
"input": {
"audio": "https://example.com/meeting.mp3",
"language": "auto"
}
}'이미지에서 텍스트 추출(OCR)은 단 1크레딧입니다:
curl -X POST https://api.core.today/v1/predictions \
-H "Content-Type: application/json" \
-H "X-API-Key: cdt_your_api_key" \
-d '{
"model": "abiruyt/text-extract-ocr",
"input": { "image": "https://example.com/receipt.jpg" }
}'모델 선택 가이드
- 음성 전사 — openai/whisper (13크레딧, 표준). 긴 오디오를 빠르게는 incredibly-fast-whisper (13크레딧), 회의록처럼 화자 구분이 필요하면 whisper-diarization (9크레딧)
- OCR·캡션 — 텍스트 추출은 text-extract-ocr (1크레딧), 이미지 캡션·VQA는 blip (1크레딧), 자유 질문은 moondream2 (7크레딧)
- 임베딩 — 다국어 텍스트는 multilingual-e5-large (3크레딧, 한국어 지원 + 고객 DB 벡터 검색 조합), 이미지·텍스트 교차 검색은 clip-features (1크레딧)
- 3D 생성 — 이미지→3D는 trellis (100크레딧, GLB 출력), 프롬프트/이미지 겸용 최고 품질은 hunyuan-3d-3.1 (1,500크레딧, PBR 지원)
- 립싱크·토킹헤드 — 영상 립싱크는 latentsync (285크레딧), 사진 한 장으로 토킹헤드는 sadtalker (120크레딧)
- 콘텐츠 모더레이션 — nsfw_image_detection (1크레딧, UGC 필수 게이트)
- 객체 감지 — 자연어로 찾는 grounding-dino (3크레딧), 감지+캡션+OCR 올인원 florence-2-large (3크레딧)
- 비디오 업스케일·유틸 — real-esrgan-video (1,650크레딧, FHD/2K/4K), 비디오에 AI 사운드는 mmaudio (15크레딧), 변환·추출은 video-utils (7크레딧), 비디오 내용 질문은 cogvlm2-video (36크레딧)