Skip to main content
Core.Today
|
Billing note Voice gallery samples on this page were rendered with S2.1 Pro (same voice ids). Same price as S2.1 Pro โ€” choose this model only for prompts already tuned on it.
Fish AudioFastHigh

Fish Audio S2 Pro

Fish Audio S2-generation TTS: 80+ languages, bracketed natural-language emotion and tone tags, Voice Library voices. Same price as S2.1 Pro; kept for pipelines tuned on S2. Billed per UTF-8 byte.

105 credits
per 1,000 Korean characters (0.034875 credits per UTF-8 byte; ~35 credits per 1,000 English characters)
80+ languages from a single model
Bracketed natural-language emotion and tone tags ([sad], [shouting], [sighing])
Voice Library voices via reference_id, plus cloned voices
Same price as S2.1 Pro โ€” keep it for pipelines tuned on S2
mp3 / wav / opus output, speed and volume prosody control

Run it right now

Test this model instantly in the Console Playground โ€” no code required

Sign in to try

Use with AI Assistant

Copy usage instructions for Claude, ChatGPT, or other AI

Quick Start

curl -X POST "https://api.core.today/v1/predictions" \
  -H "Content-Type: application/json" \
  -H "X-API-Key: cdt_your_api_key" \
  -d '{
  "model": "fishaudio/s2-pro",
  "input": {
    "text": "์•ˆ๋…•ํ•˜์„ธ์š”, ์ฝ”์–ด๋‹ทํˆฌ๋ฐ์ด์ž…๋‹ˆ๋‹ค. ํ•˜๋‚˜์˜ API ํ‚ค๋กœ ์ด๋ฏธ์ง€, ์˜์ƒ, ์Œ์„ฑ ์ƒ์„ฑ ๋ชจ๋ธ์„ ๋ชจ๋‘ ์—ฐ๊ฒฐํ•ด ๋ณด์„ธ์š”.",
    "reference_id": "29da56534ac84ccd81092be4359a1639",
    "format": "mp3"
  }
}'

Voice Gallery

Showing 37 of 37

Every voice reads the same line. Play to compare timbre, then click an ID to copy it into reference_id.

Sample line (Korean): โ€œ์•ˆ๋…•ํ•˜์„ธ์š”, ์ฝ”์–ด๋‹ทํˆฌ๋ฐ์ด์ž…๋‹ˆ๋‹ค. ํ•˜๋‚˜์˜ API ํ‚ค๋กœ ์ด๋ฏธ์ง€, ์˜์ƒ, ์Œ์„ฑ ์ƒ์„ฑ ๋ชจ๋ธ์„ ๋ชจ๋‘ ์—ฐ๊ฒฐํ•ด ๋ณด์„ธ์š”.โ€โ€” โ€œHi, this is Core.Today. One API key connects you to image, video and speech models.โ€

Fish Official Korean (15)

Designed by Fish rather than cloned from a person, so the rights position is the simplest here. Their registered descriptions are identical apart from gender and city, and the city is a label on the entry โ€” nobody measured the accent โ€” so pick these by ear.

์ง€์šฐF

Seoul female ยท chatbot replies and voice prompts

conversationalnatural
ํ˜œ์ง„F

Seoul female ยท chatbot replies and voice prompts

conversationalnatural
์„œ์œคF

Seoul female ยท chatbot replies and voice prompts

conversationalnatural
์„œ์—ฐF

Seoul female ยท chatbot replies and voice prompts

conversationalnatural
ํ•˜์€F

Seoul female ยท chatbot replies and voice prompts

conversationalnatural
์ˆ˜์•„F

Registered as a Busan female conversational voice

conversationalregional
์˜ˆ๋ฆฐF

Registered as a Gyeongju female conversational voice

conversationalregional
์œ ๋‚˜F

Crisp young female ยท product and business narration

professionalclear
๋ฏผ์ค€M

Seoul male ยท chatbot replies and voice prompts

conversationalnatural
์‹œ์šฐM

Seoul male ยท chatbot replies and voice prompts

conversationalnatural
์ค€์„œM

Seoul male ยท chatbot replies and voice prompts

conversationalnatural
๋„์œค (Doyoon)M

Seoul male ยท a different voice from Doyun below

conversationalnatural
๋„์œค (Doyun)M

Seoul male ยท a different voice from Doyoon above

conversationalnatural
์žฌ๋ฏผM

Registered as a Busan male conversational voice

conversationalregional
ํƒœ์˜คM

Registered as a Gyeongju male conversational voice

conversationalregional

Korean community (18)

Uploaded by Fish users, and the place to look for a specific job โ€” narration, shorts, advertising reads, interview hosting. Voices cloned from identifiable people are not listed here.

์ฐจ๋ถ„ํ•œ ๋‚ด๋ ˆ์ด์…˜F

Documentary and reflective narration

calmclearprofessional
๊ธฐ๋ณธ ๋‚˜๋ ˆ์ด์…˜ (๋‚จ)M

News reads and educational narration

steadymeasuredauthoritative
๋ถ€๋“œ๋Ÿฌ์šด ๋‚ด๋ ˆ์ดํ„ฐM

Warm storytelling and essay reads

warmgentleempathetic
๋‹คํ ๋‚˜๋ ˆ์ด์…˜F

Documentary and factual narration

calmprofessionalclear
๊ธ์ • ์•„๋‚˜์šด์„œF

Explainer videos and course narration

articulateprofessionalfriendly
๋”ฐ๋œปํ•œ ์—ฌ์„ฑF

Conversation and storytelling

warmrelaxedexpressive
20๋Œ€ ์—ฌ์„ฑ (๊ด‘๊ณ )F

Product ads and informational reads

smoothcalmpersuasive
๋‚˜๊ธ‹๋‚˜๊ธ‹ ์‡ผ์ธ F

Shorts and social clips

brightenergeticfriendly
ํšจ์ • (ํŒŸ์บ์ŠคํŠธ)F

Podcasts and conversational content

smoothfriendlynatural
์• ๋ค (๊ด‘๊ณ )M

Advertising and promo reads

energeticconfidentpersuasive
์œ ๋ผF

Composed female narration

calmsmoothprofessional
์ง„์šฐM

Storytelling and long-form reads

calmsmoothnarrative
๋ณด์ด์Šค1M

Entertainment and character lines

brightenergeticenthusiastic
30๋Œ€ ๋‚จ์ž ์ธํ„ฐ๋ทฐ์–ดM

Interview hosting and course content

calmclearfriendly
์ •๋ณด์„ฑ ์œ ํŠœ๋ฒ„ (๋‚จ)M

Fast-paced explainer videos

energeticfastconfident
20๋Œ€ ์—ฌ์„ฑ ์‡ผ์ธ F

Beauty, tutorials and social

brightenergeticfriendly
ํ˜ธ๊ธฐ์š”M

Educational content and narration

calmclearfriendly
์„ฑํ˜„M

Intimate conversation and ASMR

warmgentleintimate

Robot and character (4)

Registered as English, but measured on Korean: our gateway synthesized a Korean line, our own ASR transcribed it back, and the accuracy score is how closely that transcript matched (1.0 = exact). Voices that failed are not listed โ€” one Spanish-registered robot scored 0.0 and produced no Korean at all.

์žฅ๋‚œ๊พธ๋Ÿฌ๊ธฐ ๋กœ๋ด‡N

Playful robot character ยท Korean accuracy 0.97

high-pitchedenergeticmechanical
ํ™œ๊ธฐ์ฐฌ ๋กœ๋ด‡N

Game and animation robot ยท Korean accuracy 0.97

expressiveenergetichigh-pitched
๋ ˆํŠธ๋กœ ์ปดํ“จํ„ฐM

Retro computer voice ยท Korean accuracy 0.94

monotoneelectronicretro
์•ˆ๋‚ด์šฉ ๋กœ๋ด‡M

Announcements and instructions ยท Korean accuracy 0.85

monotoneclearneutral

Parameters

ParameterTypeRequiredDefaultDescription
textstringYes-Text to synthesize (UTF-8). Billed by UTF-8 bytes โ€” one Korean character is 3 bytes. Emotion, tone and effect cues go in brackets before the text they steer, e.g. [sad] I missed you, [shouting], [sighing], [long-break].
reference_idstringNo2f06c7a428e9431fadd60af4dfe91763Voice ID (32-hex id from the Fish Audio Voice Library). 37 curated voices (15 Fish Official Korean, 18 Korean community, 4 robot/character) can be previewed and copied from the Voice Gallery on this model's Core.Today Docs page (docs/models/fishaudio/โ€ฆ); any other public or cloned voice id works too โ€” you are responsible for the rights to the voice you pass.
formatstringNomp3Output audio format.
mp3wavopus
mp3_bitrateintegerNo128MP3 bitrate in kbps (mp3 only).
64128192
speednumberNo1.0Speaking rate multiplier, 0.5โ€“2.0 (1.0 = default).
volumenumberNo0Volume adjustment in dB, -20 to 20 (0 = default).
latencystringNobalancedbalanced favors low latency (default); normal favors stability โ€” a good choice for batch generation.
balancednormal
normalizebooleanNotrueExpand numbers, dates and units into spoken form.
temperaturenumberNo0.7Sampling temperature, 0.1โ€“1.0. Lower is more consistent, higher more expressive.
top_pnumberNo0.7Nucleus sampling probability mass, 0.1โ€“1.0.
chunk_lengthintegerNo200Internal chunk length, 100โ€“300. Affects phrasing of long passages.

Common Parameters

Common parameters used when calling POST /v1/predictions.

ParameterTypeRequiredDefaultDescription
modelstringYes-Model identifier
inputobjectYes-Object containing the model-specific parameters from the table above
output_folderstringNo-Folder path for output files (max 256 chars, '..' not allowed)
webhook_urlstringNo-Webhook URL to call on completion
is_publicbooleanNofalseIf true, output files are also available via permanent public URLs

Examples

Core.Today intro

Same script as the voice gallery, rendered by S2 Pro.

curl -X POST "https://api.core.today/v1/predictions" \
  -H "Content-Type: application/json" \
  -H "X-API-Key: cdt_your_api_key" \
  -d '{
  "model": "fishaudio/s2-pro",
  "input": {
    "text": "์•ˆ๋…•ํ•˜์„ธ์š”, ์ฝ”์–ด๋‹ทํˆฌ๋ฐ์ด์ž…๋‹ˆ๋‹ค. ํ•˜๋‚˜์˜ API ํ‚ค๋กœ ์ด๋ฏธ์ง€, ์˜์ƒ, ์Œ์„ฑ ์ƒ์„ฑ ๋ชจ๋ธ์„ ๋ชจ๋‘ ์—ฐ๊ฒฐํ•ด ๋ณด์„ธ์š”.",
    "reference_id": "29da56534ac84ccd81092be4359a1639",
    "format": "mp3"
  }
}'

Emotion tags

Bracketed natural-language cues, same syntax as S2.1 Pro.

curl -X POST "https://api.core.today/v1/predictions" \
  -H "Content-Type: application/json" \
  -H "X-API-Key: cdt_your_api_key" \
  -d '{
  "model": "fishaudio/s2-pro",
  "input": {
    "text": "[calm] ์˜ค๋Š˜์˜ ์‚ฌ์šฉ๋Ÿ‰ ๋ฆฌํฌํŠธ๋ฅผ ์ฝ์–ด ๋“œ๋ฆด๊ฒŒ์š”. [emphasis] ํฌ๋ ˆ๋”ง์ด 20% ๋‚จ์•˜์Šต๋‹ˆ๋‹ค.",
    "reference_id": "d74f023d1525420797aed41b5d421c05"
  }
}'

Tips & Best Practices

1The 15 Fish Official Korean voices (์ง€์šฐ, ๋ฏผ์ค€, โ€ฆ) are designed by Fish rather than cloned from a person, which makes their rights position the simplest โ€” a good baseline to compare against
2The 4 robot voices are registered as English but were measured on Korean: our gateway synthesized a Korean line and our own ASR transcribed it back at 0.85โ€“0.97 similarity. Voices that failed that test are not in the gallery
3Voices cloned from named characters or identifiable voice actors are deliberately absent. You can still pass any reference_id to the API โ€” the rights to that voice are yours to clear
4Prefer S2.1 Pro for new work โ€” same price, better quality and latency; S2 Pro is here for pipelines already tuned on it
5Bracket tags work the same as S2.1 Pro: [happy], [whispering], [break]
6Korean is 3 UTF-8 bytes per character โ€” 1,000 Korean characters cost about 105 credits
7The docs Voice Gallery samples were rendered with S2.1 Pro; the voice ids are identical, only the model differs

Use Cases

Korean narration for shorts, ads and explainer videos
Multilingual product voice-overs from one model
Audiobooks and long-form storytelling with emotion tags
Announcements and IVR prompts
Podcast intros and character voices