Skip to main content
Core.Today
|
ElevenLabsMediumUltra

ElevenLabs Eleven v4

ElevenLabs Eleven v4 text-to-speech via Fal.AI. Expressive speech with inline audio tags ([whispering], [excited]), IPA pronunciation, stability/similarity controls, up to 5,000 characters per request and optional character-level timestamps.

186 credits
per 1000 characters (0.186 credits/char, billed per character)
Expressive delivery steered by inline audio tags like [whispering] and [excited]
IPA pronunciation control enclosed in forward slashes
Stability and similarity controls to shape delivery
Up to 5,000 characters per request with optional character-level timestamps

Run it right now

Test this model instantly in the Console Playground โ€” no code required

Sign in to try

Use with AI Assistant

Copy usage instructions for Claude, ChatGPT, or other AI

Quick Start

curl -X POST "https://api.core.today/v1/predictions" \
  -H "Content-Type: application/json" \
  -H "X-API-Key: cdt_your_api_key" \
  -d '{
  "model": "elevenlabs/tts/eleven-v4",
  "input": {
    "text": "[excited] Hello! Welcome to Eleven v4.",
    "voice": "Aria"
  }
}'

Parameters

ParameterTypeRequiredDefaultDescription
textstringYes-The text to convert to speech (max 5,000 characters, billed per character). Supports audio tags such as [whispering] or [excited] and IPA pronunciation enclosed in forward slashes.
voicestringNoRachelThe voice to use โ€” an ElevenLabs premade voice name (e.g. Rachel, Aria, Roger, Sarah, George) or a voice ID.
stabilitynumberNo0.5Voice stability (0-1). Lower values allow more expressive delivery; higher values make delivery more consistent.
similarity_boostnumberNo0.75How closely the output follows the reference voice (0-1). Higher values increase similarity but may reduce naturalness.
language_codestringNo-Language code (ISO 639-1, e.g. en, ko, ja) for speech generation and text normalization. Auto-detected when omitted.
apply_text_normalizationstringNoautoWhether to normalize text such as numbers and dates before generation.
autoonoff
output_formatstringNomp3_44100_128Output audio format, formatted as codec_sample_rate_bitrate (MP3 variants).
mp3_22050_32mp3_44100_32mp3_44100_64mp3_44100_96mp3_44100_128mp3_44100_192
timestampsbooleanNofalseWhether to return character-level timing information with the generated audio.
seedintegerNo-Seed for best-effort reproducibility. Identical output is not guaranteed.

Common Parameters

Common parameters used when calling POST /v1/predictions.

ParameterTypeRequiredDefaultDescription
modelstringYes-Model identifier
inputobjectYes-Object containing the model-specific parameters from the table above
output_folderstringNo-Folder path for output files (max 256 chars, '..' not allowed)
webhook_urlstringNo-Webhook URL to call on completion
is_publicbooleanNofalseIf true, output files are also available via permanent public URLs

Examples

Excited greeting

A short line using an audio tag to set an excited delivery.

curl -X POST "https://api.core.today/v1/predictions" \
  -H "Content-Type: application/json" \
  -H "X-API-Key: cdt_your_api_key" \
  -d '{
  "model": "elevenlabs/tts/eleven-v4",
  "input": {
    "text": "[excited] Hello! Welcome to Eleven v4.",
    "voice": "Aria"
  }
}'

Whispered Korean narration

Korean narration with a whispered delivery and an explicit language code.

curl -X POST "https://api.core.today/v1/predictions" \
  -H "Content-Type: application/json" \
  -H "X-API-Key: cdt_your_api_key" \
  -d '{
  "model": "elevenlabs/tts/eleven-v4",
  "input": {
    "text": "[whispering] ์กฐ์šฉํžˆ ํ•ด ๋ด์š”. ์ˆฒ์ด ์ž ๋“ค๊ณ  ์žˆ์–ด์š”.",
    "voice": "Rachel",
    "language_code": "ko",
    "stability": 0.4
  }
}'

Tips & Best Practices

1Billing counts every character of text, including audio tags โ€” keep tags purposeful.
2Lower stability for more expressive, varied delivery; raise it for consistent narration.
3Set language_code when the text mixes languages or when auto-detection picks the wrong one.
4Use Eleven v4 Turbo for the same controls at half the per-character rate.

Use Cases

Narration and voiceover for videos, ads and explainers
Audiobook and long-form reading with emotional nuance
Character dialogue for games and interactive content
Subtitle or lip-sync alignment using character-level timestamps