All skills
elevenlabs avatar

/text-to-speech

@279173d official
by elevenlabselevenlabs/skills462 stars
74

Convert text to speech using ElevenLabs voice AI. Use when generating audio from text, creating voiceovers, building voice apps, or synthesizing speech in 90+ languages.

Use this Skill: https://skilld.dev/gh/elevenlabs/skills/text-to-speech

This session only. Nothing lands on disk.

SKILL.md

≈46 tokens always: the name and description. ≈1.9k when used: this file. ≈3.6k more on demand in 3 files.

ElevenLabs Text-to-Speech

Generate natural speech from text - supports 90+ languages, multiple models for quality vs latency tradeoffs.

Setup: See Installation Guide. For JavaScript, use @elevenlabs/* packages only.

Quick Start

Python

from elevenlabs import ElevenLabs

client = ElevenLabs()

audio = client.text_to_speech.convert(
    text="Hello, welcome to ElevenLabs!",
    voice_id="JBFqnCBsd6RMkjVDRZzb",  # George
    model_id="eleven_multilingual_v2"
)

with open("output.mp3", "wb") as f:
    for chunk in audio:
        f.write(chunk)

JavaScript

import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createWriteStream } from "fs";
import { Readable } from "stream";

const client = new ElevenLabsClient();
const audio = await client.textToSpeech.convert("JBFqnCBsd6RMkjVDRZzb", {
  text: "Hello, welcome to ElevenLabs!",
  modelId: "eleven_multilingual_v2",
});
// convert() returns a web ReadableStream — bridge it to a Node stream to write to disk
Readable.fromWeb(audio).pipe(createWriteStream("output.mp3"));

CLI

Use say to play text immediately with the default voice and eleven_v3 model:

elevenlabs say "Hello!"

Pipe text into say when another command produces the input:

echo "The build finished successfully." | elevenlabs say

Use the API command when you need to set request parameters directly:

elevenlabs text-to-speech convert --voice-id JBFqnCBsd6RMkjVDRZzb \
  --text "Hello!" --model-id eleven_multilingual_v2 --output output.mp3

The CLI reads ELEVENLABS_API_KEY from the environment automatically.

Models

Model ID Languages Latency Best For
eleven_v4 90+ Standard Highest quality, expressive content and dialogue
eleven_v4_turbo 90+ ~100ms Expressive real-time dialogue through the Text to Dialogue WebSocket
eleven_v3 70+ Standard Highest quality, emotional range
eleven_multilingual_v2 29 Standard High quality, long-form content
eleven_flash_v2_5 32 ~75ms Ultra-low latency, real-time
eleven_flash_v2 English ~75ms English-only, fastest
eleven_turbo_v2_5 32 ~250-300ms Balanced quality/speed
eleven_turbo_v2 English ~250-300ms English-only, balanced

Voice IDs

Use pre-made voices or create custom voices in the dashboard.

Popular voices:

  • JBFqnCBsd6RMkjVDRZzb - George (male, narrative)
  • EXAVITQu4vr4xnSDxMaL - Sarah (female, soft)
  • onwK4e9ZLuTAKqWW03F9 - Daniel (male, authoritative)
  • XB0fDUnXU5powFXDhCwa - Charlotte (female, conversational)
voices = client.voices.get_all()
for voice in voices.voices:
    print(f"{voice.voice_id}: {voice.name}")

Voice Settings

Fine-tune how the voice sounds:

  • Stability: How consistent the voice stays. Lower values = more emotional range and variation, but can sound unstable. Higher = steady, predictable delivery.
  • Similarity boost: How closely to match the original voice sample. Higher values sound more like the original but may amplify audio artifacts.
  • Style: Exaggerates the voice's unique style characteristics. It is not available for Eleven v4 models.
  • Speed: Adjusts speech rate on supported models. It is not available for Eleven v4 models.
  • Speaker boost: Post-processing that enhances clarity and voice similarity.
from elevenlabs import VoiceSettings

audio = client.text_to_speech.convert(
    text="Customize my voice settings.",
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    voice_settings=VoiceSettings(
        stability=0.5,
        similarity_boost=0.75,
        style=0.5,
        speed=1.0,             # 0.25 to 4.0 (default 1.0)
        use_speaker_boost=True
    )
)

Language Selection

Use language_code with models that support language enforcement to guide pronunciation and text normalization. Unsupported language codes are ignored, and language_code is not supported on eleven_multilingual_v2.

audio = client.text_to_speech.convert(
    text="Bonjour, comment allez-vous?",
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    model_id="eleven_v3",
    language_code="fr"  # ISO 639-1 code
)

Text Normalization

Controls how numbers, dates, and abbreviations are converted to spoken words. For example, "01/15/2026" becomes "January fifteenth, twenty twenty-six":

  • "auto" (default): Model decides based on context
  • "on": Always normalize (use when you want natural speech)
  • "off": Speak literally (use when you want "zero one slash one five...")
audio = client.text_to_speech.convert(
    text="Call 1-800-555-0123 on 01/15/2026",
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    apply_text_normalization="on"
)

Request Stitching

When generating long audio in multiple requests, the audio can have pops, unnatural pauses, or tone shifts at the boundaries. Request stitching solves this by letting each request know what comes before/after it:

# First request
audio1 = client.text_to_speech.convert(
    text="This is the first part.",
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    next_text="And this continues the story."
)

# Second request using previous context
audio2 = client.text_to_speech.convert(
    text="And this continues the story.",
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    previous_text="This is the first part."
)

Output Formats

Format Description
mp3_44100_128 MP3 44.1kHz 128kbps (default) - compressed, good for web/apps
mp3_44100_192 MP3 44.1kHz 192kbps (Creator+) - higher quality compressed
mp3_44100_64 MP3 44.1kHz 64kbps - lower quality, smaller files
mp3_22050_32 MP3 22.05kHz 32kbps - smallest MP3 files
pcm_16000 Raw PCM 16kHz - use for real-time processing
pcm_22050 Raw PCM 22.05kHz
pcm_24000 Raw PCM 24kHz - good balance for streaming
pcm_44100 Raw PCM 44.1kHz (Pro+) - CD quality
pcm_48000 Raw PCM 48kHz (Pro+) - highest quality
ulaw_8000 μ-law 8kHz - standard for phone systems (Twilio, telephony)
alaw_8000 A-law 8kHz - telephony (alternative to μ-law)
opus_48000_64 Opus 48kHz 64kbps - efficient streaming codec
wav_44100 WAV 44.1kHz - uncompressed with headers

Streaming

For real-time applications, use the stream method (returns audio chunks as they're generated):

audio_stream = client.text_to_speech.stream(
    text="This text will be streamed as audio.",
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    model_id="eleven_flash_v2_5"  # Ultra-low latency
)

for chunk in audio_stream:
    play_audio(chunk)

See references/streaming.md for WebSocket streaming.

Error Handling

try:
    audio = client.text_to_speech.convert(
        text="Generate speech",
        voice_id="invalid-voice-id"
    )
except Exception as e:
    print(f"API error: {e}")

Common errors:

  • 401: Invalid API key
  • 422: Invalid parameters (check voice_id, model_id)
  • 429: Rate limit exceeded

Tracking Costs

Monitor character usage via response headers (x-character-count, request-id):

response = client.text_to_speech.convert.with_raw_response(
    text="Hello!", voice_id="JBFqnCBsd6RMkjVDRZzb", model_id="eleven_multilingual_v2"
)
audio = response.parse()
print(f"Characters used: {response.headers.get('x-character-count')}")

References

Source: SKILL.md on GitHub

2 warnings1d5 checks · Risk SAFE
  • Gen Agent Trust Hub1d

    This skill provides standard documentation and code examples for using the ElevenLabs Text-to-Speech API. It follows best practices for secret management and uses official vendor libraries and tools.

  • Socket1d

    No alerts

  • Snyk1d

    Risk: MEDIUM · 1 issue

  • Runlayer7mo

    4/4 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 279173d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 19 hours ago.

Activeupdated 2 days ago
compatibility
Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
Other metadata
metadata
{
  "openclaw": {
    "requires": {
      "env": [
        "ELEVENLABS_API_KEY"
      ]
    },
    "primaryEnv": "ELEVENLABS_API_KEY"
  }
}
  • Python
  • API
  • text-to-speech
  • elevenlabs
  • audio
  • voice-synthesis
  • multilingual
  • streaming
  • javascript

README badge

README badge for elevenlabs/skills/text-to-speech

Converts text to speech using the ElevenLabs API with support for 70+ languages, multiple quality/latency models, and voice customization. Use this skill when building voice apps, generating voiceovers, or synthesizing speech in real-time applications; it requires an ELEVENLABS_API_KEY environment variable.

Generated from the current SKILL.md.

Does this skill support all languages?
The skill supports 70+ languages depending on the model. eleven_v3 supports 70+, eleven_multilingual_v2 supports 29, and flash/turbo variants support 32. You can enforce a specific language with the language_code parameter.
What latency should I expect?
Latency varies by model: eleven_flash_v2_5 and eleven_flash_v2 offer ~75ms, turbo variants ~250-300ms, and v3/multilingual_v2 use standard latency. Choose eleven_flash for real-time applications.
Can I use custom voices?
Yes. The skill includes pre-made voice IDs like George and Sarah, but you can also create and use custom voices via the ElevenLabs dashboard.
Does this require an API key?
Yes. The skill requires an ElevenLabs API key set in the ELEVENLABS_API_KEY environment variable.
What output formats are supported?
The skill supports MP3 (multiple bitrates), PCM (multiple sample rates), Opus, WAV, and telephony codecs (ulaw/alaw). Default is MP3 44.1kHz 128kbps.

Generated from the current SKILL.md. These answers refresh after source changes.