Skip to main content
Core.Today
Model APIs

Vision & Utilities

음성 인식(Whisper), OCR·이미지 캡셔닝, 임베딩, 3D 생성, 립싱크·토킹헤드, 비디오 향상까지 — 생성 외의 모든 AI 유틸리티를 하나의 API로 사용하세요.

Give this page to your AI — all model specs as LLM-friendly text
llms.txt ↗

모델 비교표

모델크레딧유형특징
salesforce/blip1OCR/캡션Salesforce, 빠른 처리
andreasjansson/clip-features1임베딩CLIP, 빠른 처리
falcons-ai/nsfw_image_detection1안전 필터Falcons.ai, 빠른 처리
abiruyt/text-extract-ocr1OCR/캡션OCR, 빠른 처리
charlesmccarthy/addwatermark1비디오 유틸FullJourney, 빠른 처리
lucataco/florence-2-large2객체 감지Microsoft, 빠른 처리
adirik/grounding-dino2객체 감지Grounding DINO, 빠른 처리
beautyyuyanli/multilingual-e5-large2임베딩E5, 빠른 처리
lucataco/moondream25OCR/캡션Moondream, 빠른 처리
nicolascoutureau/video-utils5비디오 유틸FFmpeg, 빠른 처리
thomasmol/whisper-diarization7음성 인식(STT)Whisper, 빠른 처리
vaibhavs10/incredibly-fast-whisper10음성 인식(STT)Whisper, 빠른 처리
openai/whisper10음성 인식(STT)OpenAI, 최고 품질
zsxkib/mmaudio11비디오 유틸MMAudio, 고품질
chenxwh/cogvlm2-video28비디오 유틸CogVLM, 고품질
firtoz/trellis773D 생성Microsoft, 고품질
cjwbw/sadtalker91립싱크/아바타SadTalker, 고품질
bytedance/latentsync220립싱크/아바타ByteDance, 고품질
tencent/hunyuan-3d-3.11,1603D 생성Tencent, 최고 품질
lucataco/real-esrgan-video1,280비디오 향상Real-ESRGAN, 고품질
veed/lipsync/v22,440립싱크/아바타Veed, 고품질

오디오·비디오·이미지 파일은 POST /v1/files/upload-url로 업로드한 뒤 해당 URL을 입력으로 사용하세요. 처리 시간이 파일 길이에 비례하는 모델은 상세 문서에 표기되어 있습니다.

모델 상세 정보

각 모델의 상세한 파라미터, 예제, 활용 팁은 개별 문서에서 확인하세요.

21 models

Incredibly Fast Whisper

Whisper

10 credits

Whisper large-v3 optimized for speed (38M+ runs) — transcribes roughly 150 minutes of audio in under 100 seconds using batched inference. Chunk-level or word-level timestamps.

FastUltra
View details

Whisper

OpenAI

10 credits

OpenAI's Whisper large-v3 speech recognition (144M+ runs) — the standard for transcription. Automatic language detection across ~100 languages, English translation, and plain text / SRT / VTT output formats.

MediumUltra
View details

Whisper Diarization

Whisper

7 credits

Whisper transcription with speaker diarization (8M+ runs) — returns who said what, with per-segment speaker labels and timestamps. The go-to for meetings and interviews.

FastHigh
View details

BLIP

Salesforce

1 credits

Salesforce BLIP (173M+ runs) — image captioning, visual question answering, and image-text matching in one model. The classic choice for bulk captioning at 1 credit per image.

FastHigh
View details

CLIP Features

CLIP

1 credits

CLIP ViT-L/14 embeddings for text AND images (163M+ runs) — puts both in the same vector space for cross-modal search, image dedup, and zero-shot classification. 1 credit per run.

FastHigh
View details

Florence-2 Large

Microsoft

2 credits

Microsoft's Florence-2 all-in-one vision model — captioning, object detection, phrase grounding, OCR, and segmentation in a single API. Pick a task, optionally add text input, done.

FastHigh
View details

Grounding DINO

Grounding DINO

2 credits

Text-prompted object detection (39M+ runs) — describe what to find in natural language ('red car, person wearing a hat') and get bounding boxes with confidence scores plus an annotated image.

FastHigh
View details

Hunyuan 3D 3.1

Tencent

1,160 credits

Tencent's flagship 3D generation — create high-polygon textured 3D models from a text prompt OR an image, with optional PBR (physically based rendering) materials.

SlowUltra
View details

Moondream2

Moondream

5 credits

Small but capable vision-language model (14M+ runs) — ask free-form questions about any image and get detailed answers. Efficient VQA for tagging, moderation prep, and rich alt-text.

FastHigh
View details

Multilingual E5 Large

E5

2 credits

Multilingual text embeddings (74M+ runs) — 1024-dimension vectors across 100 languages including Korean. Pairs perfectly with Core.Today customer databases' vector search (knn_vector).

FastHigh
View details

NSFW Image Detection

Falcons.ai

1 credits

The standard NSFW image classifier (127M+ runs) — returns 'normal' or 'nsfw' for any image. An essential, ultra-cheap moderation gate for UGC platforms.

FastHigh
View details

Text Extract OCR

OCR

1 credits

Simple, massively-used OCR (91M+ runs) — extracts text from an image with a single input and returns plain text. Great default for receipts, screenshots, and scanned documents at just 1 credit.

FastHigh
View details

TRELLIS

Microsoft

77 credits

Microsoft's TRELLIS image-to-3D (836K+ runs — the most-used 3D model on Replicate). Turns one or more images into a textured GLB 3D asset, with turntable render videos and optional Gaussian PLY.

MediumHigh
View details

Add Watermark

FullJourney

1 credits

Add a text watermark to any video — simple, fast brand protection for generated or user content at 2 credits per video.

FastHigh
View details

CogVLM2 Video

CogVLM

28 credits

Video understanding and captioning — ask free-form questions about a video and get detailed answers about actions, scenes, and content. Great for video search indexing and moderation prep.

MediumHigh
View details

LatentSync

ByteDance

220 credits

ByteDance's open-source lipsync — re-syncs a video's mouth movements to any audio track using latent diffusion. State-of-the-art open lipsync quality for dubbing and localization.

SlowHigh
View details

MMAudio

MMAudio

11 credits

Add AI-generated sound to any video (5M+ runs) — synthesizes synchronized audio (ambience, effects, foley) from the video content and an optional text prompt. The perfect finisher for silent AI-generated clips.

MediumHigh
View details

Real-ESRGAN Video

Real-ESRGAN

1,280 credits

Video upscaling with Real-ESRGAN — enhance videos to FHD, 2K, or 4K frame by frame. The go-to open-source video upscaler for old footage and AI-generated clips.

SlowHigh
View details

SadTalker

SadTalker

91 credits

Talking-head video from a single photo and an audio track — animates the face with natural head motion and eye blinks, with optional GFPGAN face enhancement.

SlowHigh
View details

Veed Lipsync v2

Veed

2,440 credits

Veed Lipsync v2 via Fal.AI. Replaces a source video's mouth movements to articulate a new audio track — takes any source video + a new audio track and produces a lip-synced output video.

MediumHigh
View details

Video Utils

FFmpeg

5 credits

FFmpeg-powered video utilities (20M+ runs) — convert to mp4/gif, extract audio as mp3, or dump zipped frames, all with one task parameter. The Swiss-army knife for media pipelines.

FastHigh
View details

빠른 시작 — 음성 인식 (Whisper)

오디오 URL 하나로 100개 언어 자동 감지 전사를 받을 수 있습니다 (13크레딧):

curl -X POST https://api.core.today/v1/predictions \
  -H "Content-Type: application/json" \
  -H "X-API-Key: cdt_your_api_key" \
  -d '{
    "model": "openai/whisper",
    "input": {
      "audio": "https://example.com/meeting.mp3",
      "language": "auto"
    }
  }'

이미지에서 텍스트 추출(OCR)은 단 1크레딧입니다:

curl -X POST https://api.core.today/v1/predictions \
  -H "Content-Type: application/json" \
  -H "X-API-Key: cdt_your_api_key" \
  -d '{
    "model": "abiruyt/text-extract-ocr",
    "input": { "image": "https://example.com/receipt.jpg" }
  }'

모델 선택 가이드

  • 음성 전사 — openai/whisper (13크레딧, 표준). 긴 오디오를 빠르게는 incredibly-fast-whisper (13크레딧), 회의록처럼 화자 구분이 필요하면 whisper-diarization (9크레딧)
  • OCR·캡션 — 텍스트 추출은 text-extract-ocr (1크레딧), 이미지 캡션·VQA는 blip (1크레딧), 자유 질문은 moondream2 (7크레딧)
  • 임베딩 — 다국어 텍스트는 multilingual-e5-large (3크레딧, 한국어 지원 + 고객 DB 벡터 검색 조합), 이미지·텍스트 교차 검색은 clip-features (1크레딧)
  • 3D 생성 — 이미지→3D는 trellis (100크레딧, GLB 출력), 프롬프트/이미지 겸용 최고 품질은 hunyuan-3d-3.1 (1,500크레딧, PBR 지원)
  • 립싱크·토킹헤드 — 영상 립싱크는 latentsync (285크레딧), 사진 한 장으로 토킹헤드는 sadtalker (120크레딧)
  • 콘텐츠 모더레이션 — nsfw_image_detection (1크레딧, UGC 필수 게이트)
  • 객체 감지 — 자연어로 찾는 grounding-dino (3크레딧), 감지+캡션+OCR 올인원 florence-2-large (3크레딧)
  • 비디오 업스케일·유틸 — real-esrgan-video (1,650크레딧, FHD/2K/4K), 비디오에 AI 사운드는 mmaudio (15크레딧), 변환·추출은 video-utils (7크레딧), 비디오 내용 질문은 cogvlm2-video (36크레딧)