Skip to main content
Core.Today
Model APIs

Audio & TTS

MiniMax Speech, ElevenLabs, Inworld, Chatterbox, NAVER Clova Voice 등 다양한 TTS 모델과 음성 복제 모델을 지원합니다. 음악·효과음 생성과 보컬 분리는 Music Generation 문서에서 다룹니다.

Give this page to your AI — all model specs as LLM-friendly text
llms.txt ↗

모델 비교표

모델크레딧유형특징
chatterbox-multilingual7TTSResemble AI, 빠른 생성
tts-1.5-max24TTSInworld, 빠른 생성
xtts-v226TTSCoqui, 고품질
qwen3-tts47TTSQwen, 빠른 생성
chatterbox59TTSResemble AI, 빠른 생성
chatterbox-turbo59TTSResemble AI, 빠른 생성
realtime-tts-259TTSInworld, 빠른 생성
turbo-v2.5119TTSElevenLabs, 빠른 생성
speech-2.8-turbo142TTSMiniMax, 빠른 생성
speech-2.6-turbo142TTSMiniMax, 빠른 생성
speech-02-turbo142TTSMiniMax, 빠른 생성
tts-premium153TTSNCP Clova, 빠른 생성
bark190TTSSuno, 고품질
speech-2.8-hd237TTSMiniMax, 최고 품질
speech-2.6-hd237TTSMiniMax, 최고 품질
v3237TTSElevenLabs, 최고 품질
v2-multilingual237TTSElevenLabs, 고품질
gemini-3.1-flash-tts299TTSGoogle, 빠른 생성
v1.1/video-to-sound-effects310TTSSonilo, 빠른 생성
voice-cloning6,980TTSMiniMax, 고품질

모델 상세 정보

각 모델의 상세한 파라미터, 예제 코드, 활용 팁을 확인하세요.

20 models

Bark

Suno

190 credits

Suno's text-prompted generative audio model. Produces speech with nonverbal sounds like [laughs] and [sighs], plus music and sound effects, in 100+ speaker presets across 13 languages. Returns audio plus an optional .npz history file for voice continuity.

SlowHigh
View details

Chatterbox

Resemble AI

59 credits

Resemble AI's production-grade open-source TTS with unique emotion exaggeration control and instant voice cloning from a short reference audio. MIT-licensed and benchmarked against leading closed-source systems.

FastHigh
View details

Chatterbox Multilingual

Resemble AI

7 credits

Chatterbox open-source TTS in 23 languages with instant voice cloning and emotion exaggeration control. Max 300 characters per request.

FastHigh
View details

Chatterbox Turbo

Resemble AI

59 credits

Resemble AI's fastest open-source TTS without sacrificing quality. 20 pre-made voices, paralinguistic tags like [sigh] and [chuckle], and optional instant voice cloning from 5s+ reference audio. Max 500 characters per request.

FastHigh
View details

Gemini 3.1 Flash TTS

Google

299 credits

Google Gemini 3.1 Flash native-audio TTS via Replicate. Style instructions (separate from the spoken text) plus 30 preset voices and 40+ language/locale codes, with inline markup tags for expressive delivery ([sigh], [whispering], [shouting], etc.).

FastHigh
View details

Qwen3 TTS

Qwen

47 credits

Alibaba's unified TTS with three modes: preset speakers (custom_voice), instant voice cloning from reference audio (voice_clone), and creating a brand-new voice from a text description (voice_design).

FastHigh
View details

Inworld Realtime TTS 2

Inworld

59 credits

Inworld's most expressive TTS with natural-language steering — place bracketed instructions like [speak quickly] before the text they apply to. Real-time latency and 15+ language support.

FastHigh
View details

MiniMax Speech 2.8 HD

MiniMax

237 credits

Ranked #1 on Artificial Analysis Speech Arena and Hugging Face TTS Arena. Broadcast-quality TTS with autoregressive Transformer + Flow-VAE decoder, 32+ languages, voice cloning, natural interjections, and emotion control.

MediumUltra
View details

MiniMax Speech 2.8 Turbo

MiniMax

142 credits

Low-latency MiniMax Speech 2.8 Turbo with under 250ms latency, 40+ languages, voice cloning, natural interjections, and real-time pricing. Ideal for interactive and real-time applications.

FastHigh
View details

MiniMax Speech 2.6 HD

MiniMax

237 credits

Studio-quality multilingual text-to-speech with nuanced prosody, emotion control, and premium voices for professional applications.

MediumUltra
View details

MiniMax Speech 2.6 Turbo

MiniMax

142 credits

Fast multilingual text-to-speech with emotional control, optimized for real-time applications with low latency.

FastHigh
View details

MiniMax Speech-02-Turbo

MiniMax

142 credits

Low-latency text-to-speech model with multilingual support, emotional voice control, and 300+ voice options.

FastHigh
View details

Inworld TTS 1.5 Max

Inworld

24 credits

Inworld's highest-quality realtime TTS with under 200ms latency. Supports SSML break tags for pauses, emotion markups like [happy], and 15 languages.

FastHigh
View details

Clova Voice TTS Premium

NCP Clova

153 credits

NAVER Clova Voice Premium TTS with 108 voices across 6 languages. High-quality Korean voice synthesis with emotion control, Pro voices, and bilingual support.

FastHigh
View details

ElevenLabs Turbo v2.5

ElevenLabs

119 credits

High-quality, low-latency ElevenLabs text-to-speech in 32 languages. The same 26 premium voices as v3 at half the price, optimized for real-time and high-volume use.

FastHigh
View details

ElevenLabs v3

ElevenLabs

237 credits

ElevenLabs' most expressive text-to-speech model. 26 premium voices, inline audio tags like [laughs] and [whispers], fine-grained style and stability controls, and 70+ language support.

MediumUltra
View details

ElevenLabs v2 Multilingual

ElevenLabs

237 credits

ElevenLabs Multilingual v2 text-to-speech in over 30 languages. Stable, proven voice quality with the same premium voice lineup and fine-grained voice settings.

MediumHigh
View details

Sonilo v1.1 Video to Sound Effects

Sonilo

310 credits

Sonilo v1.1 video-to-sound-effects via Fal.AI. Adds AI-generated sound (ambience, effects, foley) to an input video — auto-captions the scene if no prompt is given, or accepts per-segment sound descriptions for finer control.

FastHigh
View details

MiniMax Voice Cloning

MiniMax

6,980 credits

Clone any voice from a 10-second to 5-minute audio sample. Returns a custom voice_id you can pass to MiniMax speech models, plus a preview clip synthesized with the cloned voice.

MediumHigh
View details

XTTS-v2

Coqui

26 credits

Coqui XTTS-v2 multilingual voice-cloning TTS. Clone a voice from a single short audio sample and speak in 16 languages — one of the most popular open-source voice cloning models.

MediumHigh
View details
NEW

Clova Voice TTS Premium

108개 음성, 6개 언어를 지원하는 고품질 한국어 특화 TTS.

curl -X POST https://api.core.today/v1/predictions \
  -H "Content-Type: application/json" \
  -H "X-API-Key: cdt_your_api_key" \
  -d '{
    "model": "ncp-clova/tts-premium",
    "input": {
      "text": "안녕하세요. 코어닷 AI API 게이트웨이의 음성합성 서비스입니다.",
      "speaker": "nara",
      "emotion": 2,
      "speed": 0,
      "format": "mp3"
    }
  }'

주요 파라미터

speaker (108개 음성)
  • nara - 아라 (여, 한국어)
  • nminsang - 민상 (남, 한국어)
  • vara - 아라 Pro (여, 한국어)
  • clara - 클라라 (여, 영어)
emotion / speed / pitch
  • emotion - 0(중립), 1(슬픔), 2(기쁨), 3(분노)
  • speed - -5 ~ 10 (기본 0)
  • pitch - -5 ~ 5 (기본 0)
  • volume - -5 ~ 5 (기본 0)

음성 목록 (언어별)

한국어 (72개 + Pro 9개)
ID이름성별
nara아라
nara_call아라(상담원)
dara_ang아라(화남)
nminsang민상
nminseo민서
njinho진호
nbora보라
ndaeseong대성
ndain다인아동여
ndonghyun동현
neunseo은서
neunwoo은우
neunyoung은영
ngaram가람아동여
ngoeun고은
ngyeongjun경준
nhajun하준아동남
nheera희라
nian이안
nihyun이현
njaewook재욱
njangj드림
njihun지훈
njihwan지환
njiwon지원
njiyun지윤
njonghyeok종혁
njonghyun종현
njooahn주안
njoonyoung준영
nkitae기태
nkyunglee경리
nkyungtae경태
nkyuwon규원
nmammon악마 마몬
nmeow야옹이아동여
nmijin미진
nminjeong민정
nminyoung민영
nmovie최무비
noyj봄달
nraewon래원
nreview박리뷰
nsabina마녀 사비나
nsangdo상도
nseonghoon성훈
nseungpyo승표
nshasha샤샤
nsinu신우
nsiyoon시윤
nsujin수진
nsunhee선희
nsunkyung선경
ntaejin태진
ntiffany기서
nwontak원탁
nwoof멍멍이아동남
nwoosik우식
nyeji예지
nyejin예진
nyounghwa정영화
nyoungil영일
nyoungmi영미
nyujin유진
nyuna유나
jinho진호
mijin미진
napple늘봄
nes_c_hyeri혜리
nes_c_kihyo기효
nes_c_mikyung미경
nes_c_sohyun소현

Pro 음성 (고품질):

vara아라 Pro
vdaeseong대성 Pro
vdain다인 Pro
vdonghyun동현 Pro
vgoeun고은 Pro
vhyeri혜리 Pro
vian이안 Pro
vmikyung미경 Pro
vyuna유나 Pro
일본어 (15개)
ID이름성별
shinji신지
dayumu아유무
ddaiki다이키
deriko에리코
dhajime하지메
dmio미오
dnaomi나오미
dnaomi_formal나오미(뉴스)
dnaomi_joyful나오미(기쁨)
driko리코
dsayuri사유리
dtomoko토모코
nnaomi나오미
nsayuri사유리
ntomoko토모코
영어 (4개)
ID이름성별
clara클라라
danna안나
djoey조이
matt매트
기타 (한국어+영어 2개, 중국어 2개, 스페인어 2개, 대만어 2개)
ID이름언어성별
dara-danna아라&안나한국어+영어
dsinu-matt신우&매트한국어+영어
liangliang량량중국어
meimei메이메이중국어
carmen카르멘스페인어
jose호세스페인어
chiahua차화대만어
kuanlin관린대만어
추천

MiniMax Speech-02-Turbo

실시간 TTS에 최적화된 고속 음성 합성 모델.

curl -X POST https://api.core.today/v1/predictions \
  -H "Content-Type: application/json" \
  -H "X-API-Key: cdt_your_api_key" \
  -d '{
    "model": "minimax/speech-02-turbo",
    "input": {
      "text": "안녕하세요! Core.Today AI API에 오신 것을 환영합니다.",
      "voice_id": "male-qn-qingse",
      "speed": 1.0
    }
  }'

주요 파라미터

voice_id
  • male-qn-qingse - 남성 (청아한)
  • female-shaonv - 여성 (소녀)
  • male-qn-jingying - 남성 (정영)
  • female-yujie - 여성 (우아한)
speed
  • 0.5 - 느리게
  • 1.0 - 기본 속도
  • 1.5 - 빠르게
  • 2.0 - 매우 빠르게

모델 선택 가이드

한국어 특화 TTS

clova-tts-premium (153크레딧, 108개 음성, 감정 표현)

실시간 TTS

speech-02-turbo (142크레딧, 빠른 응답)

고품질 음성

speech-2.8-hd (237크레딧, HD 품질)

범용 TTS

speech-2.8-turbo (142크레딧, 균형 잡힌 성능)