Audio & TTS
MiniMax Speech, ElevenLabs, Inworld, Chatterbox, NAVER Clova Voice 등 다양한 TTS 모델과 음성 복제 모델을 지원합니다. 음악·효과음 생성과 보컬 분리는 Music Generation 문서에서 다룹니다.
모델 비교표
| 모델 | 크레딧 | 유형 | 특징 |
|---|---|---|---|
| chatterbox-multilingual | 7 | TTS | Resemble AI, 빠른 생성 |
| tts-1.5-max | 24 | TTS | Inworld, 빠른 생성 |
| xtts-v2 | 26 | TTS | Coqui, 고품질 |
| qwen3-tts | 47 | TTS | Qwen, 빠른 생성 |
| chatterbox | 59 | TTS | Resemble AI, 빠른 생성 |
| chatterbox-turbo | 59 | TTS | Resemble AI, 빠른 생성 |
| realtime-tts-2 | 59 | TTS | Inworld, 빠른 생성 |
| turbo-v2.5 | 119 | TTS | ElevenLabs, 빠른 생성 |
| speech-2.8-turbo | 142 | TTS | MiniMax, 빠른 생성 |
| speech-2.6-turbo | 142 | TTS | MiniMax, 빠른 생성 |
| speech-02-turbo | 142 | TTS | MiniMax, 빠른 생성 |
| tts-premium | 153 | TTS | NCP Clova, 빠른 생성 |
| bark | 190 | TTS | Suno, 고품질 |
| speech-2.8-hd | 237 | TTS | MiniMax, 최고 품질 |
| speech-2.6-hd | 237 | TTS | MiniMax, 최고 품질 |
| v3 | 237 | TTS | ElevenLabs, 최고 품질 |
| v2-multilingual | 237 | TTS | ElevenLabs, 고품질 |
| gemini-3.1-flash-tts | 299 | TTS | Google, 빠른 생성 |
| v1.1/video-to-sound-effects | 310 | TTS | Sonilo, 빠른 생성 |
| voice-cloning | 6,980 | TTS | MiniMax, 고품질 |
모델 상세 정보
각 모델의 상세한 파라미터, 예제 코드, 활용 팁을 확인하세요.
Bark
Suno
Suno's text-prompted generative audio model. Produces speech with nonverbal sounds like [laughs] and [sighs], plus music and sound effects, in 100+ speaker presets across 13 languages. Returns audio plus an optional .npz history file for voice continuity.
Chatterbox
Resemble AI
Resemble AI's production-grade open-source TTS with unique emotion exaggeration control and instant voice cloning from a short reference audio. MIT-licensed and benchmarked against leading closed-source systems.
Chatterbox Multilingual
Resemble AI
Chatterbox open-source TTS in 23 languages with instant voice cloning and emotion exaggeration control. Max 300 characters per request.
Chatterbox Turbo
Resemble AI
Resemble AI's fastest open-source TTS without sacrificing quality. 20 pre-made voices, paralinguistic tags like [sigh] and [chuckle], and optional instant voice cloning from 5s+ reference audio. Max 500 characters per request.
Gemini 3.1 Flash TTS
Google Gemini 3.1 Flash native-audio TTS via Replicate. Style instructions (separate from the spoken text) plus 30 preset voices and 40+ language/locale codes, with inline markup tags for expressive delivery ([sigh], [whispering], [shouting], etc.).
Qwen3 TTS
Qwen
Alibaba's unified TTS with three modes: preset speakers (custom_voice), instant voice cloning from reference audio (voice_clone), and creating a brand-new voice from a text description (voice_design).
Inworld Realtime TTS 2
Inworld
Inworld's most expressive TTS with natural-language steering — place bracketed instructions like [speak quickly] before the text they apply to. Real-time latency and 15+ language support.
MiniMax Speech 2.8 HD
MiniMax
Ranked #1 on Artificial Analysis Speech Arena and Hugging Face TTS Arena. Broadcast-quality TTS with autoregressive Transformer + Flow-VAE decoder, 32+ languages, voice cloning, natural interjections, and emotion control.
MiniMax Speech 2.8 Turbo
MiniMax
Low-latency MiniMax Speech 2.8 Turbo with under 250ms latency, 40+ languages, voice cloning, natural interjections, and real-time pricing. Ideal for interactive and real-time applications.
MiniMax Speech 2.6 HD
MiniMax
Studio-quality multilingual text-to-speech with nuanced prosody, emotion control, and premium voices for professional applications.
MiniMax Speech 2.6 Turbo
MiniMax
Fast multilingual text-to-speech with emotional control, optimized for real-time applications with low latency.
MiniMax Speech-02-Turbo
MiniMax
Low-latency text-to-speech model with multilingual support, emotional voice control, and 300+ voice options.
Inworld TTS 1.5 Max
Inworld
Inworld's highest-quality realtime TTS with under 200ms latency. Supports SSML break tags for pauses, emotion markups like [happy], and 15 languages.
Clova Voice TTS Premium
NCP Clova
NAVER Clova Voice Premium TTS with 108 voices across 6 languages. High-quality Korean voice synthesis with emotion control, Pro voices, and bilingual support.
ElevenLabs Turbo v2.5
ElevenLabs
High-quality, low-latency ElevenLabs text-to-speech in 32 languages. The same 26 premium voices as v3 at half the price, optimized for real-time and high-volume use.
ElevenLabs v3
ElevenLabs
ElevenLabs' most expressive text-to-speech model. 26 premium voices, inline audio tags like [laughs] and [whispers], fine-grained style and stability controls, and 70+ language support.
ElevenLabs v2 Multilingual
ElevenLabs
ElevenLabs Multilingual v2 text-to-speech in over 30 languages. Stable, proven voice quality with the same premium voice lineup and fine-grained voice settings.
Sonilo v1.1 Video to Sound Effects
Sonilo
Sonilo v1.1 video-to-sound-effects via Fal.AI. Adds AI-generated sound (ambience, effects, foley) to an input video — auto-captions the scene if no prompt is given, or accepts per-segment sound descriptions for finer control.
MiniMax Voice Cloning
MiniMax
Clone any voice from a 10-second to 5-minute audio sample. Returns a custom voice_id you can pass to MiniMax speech models, plus a preview clip synthesized with the cloned voice.
XTTS-v2
Coqui
Coqui XTTS-v2 multilingual voice-cloning TTS. Clone a voice from a single short audio sample and speak in 16 languages — one of the most popular open-source voice cloning models.
Clova Voice TTS Premium
108개 음성, 6개 언어를 지원하는 고품질 한국어 특화 TTS.
curl -X POST https://api.core.today/v1/predictions \
-H "Content-Type: application/json" \
-H "X-API-Key: cdt_your_api_key" \
-d '{
"model": "ncp-clova/tts-premium",
"input": {
"text": "안녕하세요. 코어닷 AI API 게이트웨이의 음성합성 서비스입니다.",
"speaker": "nara",
"emotion": 2,
"speed": 0,
"format": "mp3"
}
}'주요 파라미터
speaker (108개 음성)
nara- 아라 (여, 한국어)nminsang- 민상 (남, 한국어)vara- 아라 Pro (여, 한국어)clara- 클라라 (여, 영어)
emotion / speed / pitch
emotion- 0(중립), 1(슬픔), 2(기쁨), 3(분노)speed- -5 ~ 10 (기본 0)pitch- -5 ~ 5 (기본 0)volume- -5 ~ 5 (기본 0)
음성 목록 (언어별)
한국어 (72개 + Pro 9개)
| ID | 이름 | 성별 |
|---|---|---|
| nara | 아라 | 여 |
| nara_call | 아라(상담원) | 여 |
| dara_ang | 아라(화남) | 여 |
| nminsang | 민상 | 남 |
| nminseo | 민서 | 여 |
| njinho | 진호 | 남 |
| nbora | 보라 | 여 |
| ndaeseong | 대성 | 남 |
| ndain | 다인 | 아동여 |
| ndonghyun | 동현 | 남 |
| neunseo | 은서 | 여 |
| neunwoo | 은우 | 남 |
| neunyoung | 은영 | 여 |
| ngaram | 가람 | 아동여 |
| ngoeun | 고은 | 여 |
| ngyeongjun | 경준 | 남 |
| nhajun | 하준 | 아동남 |
| nheera | 희라 | 여 |
| nian | 이안 | 남 |
| nihyun | 이현 | 여 |
| njaewook | 재욱 | 남 |
| njangj | 드림 | 여 |
| njihun | 지훈 | 남 |
| njihwan | 지환 | 남 |
| njiwon | 지원 | 여 |
| njiyun | 지윤 | 여 |
| njonghyeok | 종혁 | 남 |
| njonghyun | 종현 | 남 |
| njooahn | 주안 | 남 |
| njoonyoung | 준영 | 남 |
| nkitae | 기태 | 남 |
| nkyunglee | 경리 | 여 |
| nkyungtae | 경태 | 남 |
| nkyuwon | 규원 | 남 |
| nmammon | 악마 마몬 | 남 |
| nmeow | 야옹이 | 아동여 |
| nmijin | 미진 | 여 |
| nminjeong | 민정 | 여 |
| nminyoung | 민영 | 여 |
| nmovie | 최무비 | 남 |
| noyj | 봄달 | 여 |
| nraewon | 래원 | 남 |
| nreview | 박리뷰 | 남 |
| nsabina | 마녀 사비나 | 여 |
| nsangdo | 상도 | 남 |
| nseonghoon | 성훈 | 남 |
| nseungpyo | 승표 | 남 |
| nshasha | 샤샤 | 여 |
| nsinu | 신우 | 남 |
| nsiyoon | 시윤 | 남 |
| nsujin | 수진 | 여 |
| nsunhee | 선희 | 여 |
| nsunkyung | 선경 | 여 |
| ntaejin | 태진 | 남 |
| ntiffany | 기서 | 여 |
| nwontak | 원탁 | 남 |
| nwoof | 멍멍이 | 아동남 |
| nwoosik | 우식 | 남 |
| nyeji | 예지 | 여 |
| nyejin | 예진 | 여 |
| nyounghwa | 정영화 | 여 |
| nyoungil | 영일 | 남 |
| nyoungmi | 영미 | 여 |
| nyujin | 유진 | 여 |
| nyuna | 유나 | 여 |
| jinho | 진호 | 남 |
| mijin | 미진 | 여 |
| napple | 늘봄 | 여 |
| nes_c_hyeri | 혜리 | 여 |
| nes_c_kihyo | 기효 | 남 |
| nes_c_mikyung | 미경 | 여 |
| nes_c_sohyun | 소현 | 여 |
Pro 음성 (고품질):
| vara | 아라 Pro | 여 |
| vdaeseong | 대성 Pro | 남 |
| vdain | 다인 Pro | 여 |
| vdonghyun | 동현 Pro | 남 |
| vgoeun | 고은 Pro | 여 |
| vhyeri | 혜리 Pro | 여 |
| vian | 이안 Pro | 남 |
| vmikyung | 미경 Pro | 여 |
| vyuna | 유나 Pro | 여 |
일본어 (15개)
| ID | 이름 | 성별 |
|---|---|---|
| shinji | 신지 | 남 |
| dayumu | 아유무 | 남 |
| ddaiki | 다이키 | 남 |
| deriko | 에리코 | 여 |
| dhajime | 하지메 | 남 |
| dmio | 미오 | 여 |
| dnaomi | 나오미 | 여 |
| dnaomi_formal | 나오미(뉴스) | 여 |
| dnaomi_joyful | 나오미(기쁨) | 여 |
| driko | 리코 | 여 |
| dsayuri | 사유리 | 여 |
| dtomoko | 토모코 | 여 |
| nnaomi | 나오미 | 여 |
| nsayuri | 사유리 | 여 |
| ntomoko | 토모코 | 여 |
영어 (4개)
| ID | 이름 | 성별 |
|---|---|---|
| clara | 클라라 | 여 |
| danna | 안나 | 여 |
| djoey | 조이 | 여 |
| matt | 매트 | 남 |
기타 (한국어+영어 2개, 중국어 2개, 스페인어 2개, 대만어 2개)
| ID | 이름 | 언어 | 성별 |
|---|---|---|---|
| dara-danna | 아라&안나 | 한국어+영어 | 여 |
| dsinu-matt | 신우&매트 | 한국어+영어 | 남 |
| liangliang | 량량 | 중국어 | 남 |
| meimei | 메이메이 | 중국어 | 여 |
| carmen | 카르멘 | 스페인어 | 여 |
| jose | 호세 | 스페인어 | 남 |
| chiahua | 차화 | 대만어 | 여 |
| kuanlin | 관린 | 대만어 | 남 |
MiniMax Speech-02-Turbo
실시간 TTS에 최적화된 고속 음성 합성 모델.
curl -X POST https://api.core.today/v1/predictions \
-H "Content-Type: application/json" \
-H "X-API-Key: cdt_your_api_key" \
-d '{
"model": "minimax/speech-02-turbo",
"input": {
"text": "안녕하세요! Core.Today AI API에 오신 것을 환영합니다.",
"voice_id": "male-qn-qingse",
"speed": 1.0
}
}'주요 파라미터
voice_id
male-qn-qingse- 남성 (청아한)female-shaonv- 여성 (소녀)male-qn-jingying- 남성 (정영)female-yujie- 여성 (우아한)
speed
0.5- 느리게1.0- 기본 속도1.5- 빠르게2.0- 매우 빠르게
모델 선택 가이드
한국어 특화 TTS
clova-tts-premium (153크레딧, 108개 음성, 감정 표현)
실시간 TTS
speech-02-turbo (142크레딧, 빠른 응답)
고품질 음성
speech-2.8-hd (237크레딧, HD 품질)
범용 TTS
speech-2.8-turbo (142크레딧, 균형 잡힌 성능)