Skip to main content
Core.Today
|
MiniMaxFastHigh

MiniMax Speech 2.8 Turbo

Low-latency MiniMax Speech 2.8 Turbo with under 250ms latency, 40+ languages, voice cloning, natural interjections, and real-time pricing. Ideal for interactive and real-time applications.

142 credits
per 1000 characters (0.135 credits/char, billed per character)
Under 250ms latency for real-time use
40+ language support
Natural interjections (laughs, sighs, coughs, etc.)
Voice cloning from short audio samples
Emotional expression control
Pause markers and pronunciation control

Run it right now

Test this model instantly in the Console Playground โ€” no code required

Sign in to try

Use with AI Assistant

Copy usage instructions for Claude, ChatGPT, or other AI

Quick Start

curl -X POST "https://api.core.today/v1/predictions" \
  -H "Content-Type: application/json" \
  -H "X-API-Key: cdt_your_api_key" \
  -d '{
  "model": "minimax/speech-2.8-turbo",
  "input": {
    "text": "์•ˆ๋…•ํ•˜์„ธ์š”! (laughs) ์ฝ”์–ด๋‹ทํˆฌ๋ฐ์ด์— ์˜ค์‹  ๊ฒƒ์„ ํ™˜์˜ํ•ฉ๋‹ˆ๋‹ค. <#0.3#> ์ตœ๊ณ ์˜ AI ์Œ์„ฑ์„ ๋น ๋ฅด๊ฒŒ ๊ฒฝํ—˜ํ•ด ๋ณด์„ธ์š”!",
    "voice_id": "Korean_CheerfulLittleSister",
    "emotion": "happy",
    "speed": 1,
    "language_boost": "Korean"
  }
}'

Voice Gallery

Showing 73 of 73

Every voice reads the same line. Play to compare timbre, then click an ID to copy it into reference_id.

Sample line (Korean): โ€œ์•ˆ๋…•ํ•˜์„ธ์š”? ์ฝ”์–ด๋‹ทํˆฌ๋ฐ์ด์— ์˜ค์‹  ๊ฒƒ์„ ํ™˜์˜ํ•ฉ๋‹ˆ๋‹คโ€โ€” โ€œHello? Welcome to Core.Todayโ€

์—‰๋šฑํ•œ ์†Œ๋…€F
ํ™œ๋ฐœํ•œ ์†Œ๋…€F
์šฉ๊ฐํ•œ ์—ฌ์ „์‚ฌF
์ฐจ๋ถ„ํ•œ ์ˆ™๋…€F
๋”ฐ๋œปํ•œ ์—ฌ์„ฑF
๋งค๋ ฅ์ ์ธ ์–ธ๋‹ˆF
๋งค๋ ฅ์ ์ธ ์—ฌ๋™์ƒF
๋ฐœ๋ž„ํ•œ ์—ฌ๋™์ƒF
์†Œ๊ฟ‰์นœ๊ตฌ ์†Œ๋…€F
์ฐจ๊ฐ€์šด ์†Œ๋…€F
์นด๋ฆฌ์Šค๋งˆ ์—ฌ์™•F
์šฐ์•„ํ•œ ๊ณต์ฃผF
๋งคํ˜น์ ์ธ ์–ธ๋‹ˆF
๋‹ค์ •ํ•œ ์–ธ๋‹ˆF
๋ถ€๋“œ๋Ÿฌ์šด ์—ฌ์„ฑF
๋„๋„ํ•œ ์ˆ™๋…€F
์„ฑ์ˆ™ํ•œ ์—ฌ์„ฑF
์‹ ๋น„๋กœ์šด ์†Œ๋…€F
๊ฐœ์„ฑ ์žˆ๋Š” ์†Œ๋…€F
๋“ฌ์งํ•œ ์–ธ๋‹ˆF
๋‹น๋‹นํ•œ ์†Œ๋…€F
์ˆ˜์ค์€ ์†Œ๋…€F
ํŽธ์•ˆํ•œ ์—ฌ์„ฑF
๋‹ฌ์ฝคํ•œ ์†Œ๋…€F
์‚ฌ๋ ค ๊นŠ์€ ์—ฌ์„ฑF
ํ˜„๋ช…ํ•œ ์—˜ํ”„F
์šด๋™ํ•˜๋Š” ํ•™์ƒM
์šฉ๊ฐํ•œ ๋ชจํ—˜๊ฐ€M
์šฉ๊ฐํ•œ ์ฒญ๋…„M
์ฐจ๋ถ„ํ•œ ์‹ ์‚ฌM
์พŒํ™œํ•œ ๋‚จ์ž์นœ๊ตฌM
์พŒํ™œํ•œ ํ›„๋ฐฐM
๊ฑด๋ฐฉ์ง„ ๋‚จ์žM
์ฐจ๊ฐ€์šด ์ฒญ๋…„M
์ž์‹ ๊ฐ ์žˆ๋Š” ๋ณด์ŠคM
๋ฐฐ๋ คํ•˜๋Š” ์„ ๋ฐฐM
๊ฐ•์ธํ•œ ๋‚จ์„ฑM
์—ด์ •์ ์ธ ์‹ญ๋Œ€M
์˜จํ™”ํ•œ ๋ณด์ŠคM
์ˆœ์ˆ˜ํ•œ ์†Œ๋…„M
์ง€์ ์ธ ๋‚จ์„ฑM
์ง€์ ์ธ ์„ ๋ฐฐM
๊ณ ๋…ํ•œ ์ „์‚ฌM
๊ธ์ •์ ์ธ ์ฒญ๋…„M
๋งค๋ ฅ์ ์ธ ๋‚จ์žM
์†Œ์œ ์š• ๊ฐ•ํ•œ ๋‚จ์žM
๋“ฌ์งํ•œ ์ฒญ๋…„M
์—„๊ฒฉํ•œ ๋ณด์ŠคM
์ง€ํ˜œ๋กœ์šด ์„ ์ƒ๋‹˜M
Wise Lady (Default)F
Expressive NarratorM
Calm WomanF
Deep Voice ManM
Friendly PersonN
Captivating StorytellerM
Kind LadyF
Gentle ButlerM
๋‹ค์ •ํ•œ ์‚ฌ๋žŒN
์˜๊ฐ์„ ์ฃผ๋Š” ์†Œ๋…€F
๊นŠ์€ ๋ชฉ์†Œ๋ฆฌ ๋‚จ์„ฑM
์ฐจ๋ถ„ํ•œ ์—ฌ์„ฑF
์บ์ฃผ์–ผํ•œ ๋‚จ์žM
ํ™œ๋ฐœํ•œ ์†Œ๋…€F
์ธ๋‚ด์‹ฌ ์žˆ๋Š” ๋‚จ์„ฑM
์ Š์€ ๊ธฐ์‚ฌM
๊ฒฐ๋‹จ๋ ฅ ์žˆ๋Š” ๋‚จ์„ฑM
์‚ฌ๋ž‘์Šค๋Ÿฌ์šด ์†Œ๋…€F
์˜ˆ์˜ ๋ฐ”๋ฅธ ์†Œ๋…„M
์œ„์—„ ์žˆ๋Š” ํƒœ๋„M
์šฐ์•„ํ•œ ๋‚จ์„ฑM
์ˆ˜๋…€์›์žฅF
๋‹ฌ์ฝคํ•œ ์†Œ๋…€ 2F
ํ™œ๊ธฐ์ฐฌ ์†Œ๋…€F

Parameters

ParameterTypeRequiredDefaultDescription
textstringYes-Text to narrate (max 10,000 characters). Use <#0.5#> to insert pauses. Supports interjections: (laughs), (sighs), (coughs), (gasps), (humming), (whistles), (sneezes), etc.
voice_idstringNoEnglish_WiseladyVoice preset or cloned voice ID. 17+ built-in voices available.
speednumberNo1Speech speed multiplier (0.5-2.0)
volumenumberNo1Relative loudness. 1.0 is default MiniMax gain. Range 0โ€“10.
pitchintegerNo0Semitone offset (-12 to +12)
emotionstringNoautoDelivery style
autohappysadangryfearfuldisgustedsurprisedcalmfluentneutral
english_normalizationbooleanNofalseImprove number/date reading for English text (adds a small amount of latency).
sample_rateintegerNo32000Audio sample rate in Hz.
80001600022050240003200044100
bitrateintegerNo128000MP3 bitrate in bits per second. Only used when audio_format is mp3.
3200064000128000256000
audio_formatstringNomp3File format for the generated audio. Choose mp3 for general use, wav/flac for lossless, or pcm for raw bytes.
mp3wavflacpcm
channelstringNomonomono for 1 channel (default), stereo for 2 channels.
monostereo
subtitle_enablebooleanNofalseReserved: the upstream currently returns audio only โ€” no subtitle data is included even when enabled
language_booststringNoNoneLanguage hint for better pronunciation
NoneAutomaticChineseChinese,YueCantoneseEnglishArabicRussianSpanishFrenchPortugueseGermanTurkishDutchUkrainianVietnameseIndonesianJapaneseItalianKoreanThaiPolishRomanianGreekCzechFinnishHindiBulgarianDanishHebrewMalayPersianSlovakSwedishCroatianFilipinoHungarianNorwegianSlovenianCatalanNynorskTamilAfrikaans

Common Parameters

Common parameters used when calling POST /v1/predictions.

ParameterTypeRequiredDefaultDescription
modelstringYes-Model identifier
inputobjectYes-Object containing the model-specific parameters from the table above
output_folderstringNo-Folder path for output files (max 256 chars, '..' not allowed)
webhook_urlstringNo-Webhook URL to call on completion
is_publicbooleanNofalseIf true, output files are also available via permanent public URLs

Examples

Interactive Voice Agent

Generate low-latency voice response with natural interjections

curl -X POST "https://api.core.today/v1/predictions" \
  -H "Content-Type: application/json" \
  -H "X-API-Key: cdt_your_api_key" \
  -d '{
  "model": "minimax/speech-2.8-turbo",
  "input": {
    "text": "์•ˆ๋…•ํ•˜์„ธ์š”! (laughs) ์ฝ”์–ด๋‹ทํˆฌ๋ฐ์ด์— ์˜ค์‹  ๊ฒƒ์„ ํ™˜์˜ํ•ฉ๋‹ˆ๋‹ค. <#0.3#> ์ตœ๊ณ ์˜ AI ์Œ์„ฑ์„ ๋น ๋ฅด๊ฒŒ ๊ฒฝํ—˜ํ•ด ๋ณด์„ธ์š”!",
    "voice_id": "Korean_CheerfulLittleSister",
    "emotion": "happy",
    "speed": 1,
    "language_boost": "Korean"
  }
}'

Tips & Best Practices

1Best choice for real-time applications where latency under 250ms matters
2Use interjection tags like (laughs), (sighs) for more natural delivery
3Insert pauses with <#x#> markers for better pacing (0.01-99.99 seconds)
4Set speed to 1.1-1.2 for chatbot responses to feel more responsive
5Use language_boost for non-English text to improve pronunciation accuracy
6For production-quality output, consider Speech 2.8 HD instead
7subtitle_enable is accepted but the upstream does not return subtitle data yet โ€” responses contain audio only

Use Cases

Real-time voice assistants and agents
Live streaming and interactive content
Customer service chatbots
Game character voices
Multilingual content localization