All skills
elevenlabs avatar

/text-to-speech

@279173d official
by elevenlabselevenlabs/skills462 stars
74

Convert text to speech using ElevenLabs voice AI. Use when generating audio from text, creating voiceovers, building voice apps, or synthesizing speech in 90+ languages.

Use this Skill: https://skilld.dev/gh/elevenlabs/skills/text-to-speech

This session only. Nothing lands on disk.

referencesvoice-settings.md

≈872 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Voice Settings

Fine-tune voice characteristics for your use case.

Parameters

Parameter Range Default Description
stability 0.0 - 1.0 0.5 How consistent the voice sounds across the generation. Lower = more emotional variation and expressiveness (but can sound erratic). Higher = steady, predictable tone.
similarity_boost 0.0 - 1.0 0.75 How closely to match the original voice sample. Higher sounds more like the source voice but may amplify audio artifacts or background noise from the original recording.
style 0.0 - 1.0 0.0 Exaggerates the unique characteristics of the voice's speaking style. Higher values make the voice more "characterful" but can reduce stability. Not available for Eleven v4 models.
speed 0.25 - 4.0 1.0 Speech speed multiplier. 1.0 = normal speed. Range is 0.25-4.0 for the REST API; the agents platform restricts to 0.7-1.2. Not available for Eleven v4 models.
use_speaker_boost boolean true Post-processing that enhances voice clarity and similarity to the original. Generally leave this on unless you're experiencing artifacts.

Python Example

from elevenlabs import ElevenLabs
from elevenlabs import VoiceSettings

client = ElevenLabs()

audio = client.text_to_speech.convert(
    text="Testing different voice settings.",
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    model_id="eleven_v3",
    voice_settings=VoiceSettings(
        stability=0.5,
        similarity_boost=0.75,
        style=0.0,
        use_speaker_boost=True
    )
)

JavaScript Example

const audio = await client.textToSpeech.convert("JBFqnCBsd6RMkjVDRZzb", {
  text: "Testing different voice settings.",
  modelId: "eleven_v3",
  voiceSettings: {
    stability: 0.5,
    similarityBoost: 0.75,
    style: 0.0,
    useSpeakerBoost: true,
  },
});

CLI Example

elevenlabs text-to-speech convert \
  --voice-id JBFqnCBsd6RMkjVDRZzb \
  --text "Testing different voice settings." \
  --model-id eleven_v3 \
  --params '{"voice_settings": {"stability": 0.5, "similarity_boost": 0.75, "style": 0.0, "use_speaker_boost": true}}' \
  --output output.mp3

The CLI reads ELEVENLABS_API_KEY from the environment automatically. Pass nested JSON objects like voice_settings via --params.

Use Case Recommendations

Audiobooks / Narration

voice_settings=VoiceSettings(
    stability=0.7,        # Consistent tone
    similarity_boost=0.5, # Natural variation
    style=0.0
)

Conversational / Chatbots

voice_settings=VoiceSettings(
    stability=0.4,        # More expressive
    similarity_boost=0.75,
    style=0.3             # Slight style emphasis
)

News / Professional

voice_settings=VoiceSettings(
    stability=0.8,        # Very consistent
    similarity_boost=0.6,
    style=0.0
)

Character Voices / Drama

voice_settings=VoiceSettings(
    stability=0.3,        # Highly expressive
    similarity_boost=0.8,
    style=0.5             # Strong style
)

Tips

  • Start with defaults and adjust incrementally
  • Lower stability if voice sounds monotonous
  • Reduce similarity_boost if you hear audio artifacts
  • Style is unavailable for Eleven v4 models
  • Speed is unavailable for Eleven v4 models
  • Test with representative text from your actual use case
  • Flash models ignore some voice settings for speed

Source: SKILL.md on GitHub

2 warnings2d5 checks · Risk SAFE
  • Gen Agent Trust Hub2d

    This skill provides standard documentation and code examples for using the ElevenLabs Text-to-Speech API. It follows best practices for secret management and uses official vendor libraries and tools.

  • Socket2d

    No alerts

  • Snyk2d

    Risk: MEDIUM · 1 issue

  • Runlayer7mo

    4/4 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 279173d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 days ago
compatibility
Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
Other metadata
metadata
{
  "openclaw": {
    "requires": {
      "env": [
        "ELEVENLABS_API_KEY"
      ]
    },
    "primaryEnv": "ELEVENLABS_API_KEY"
  }
}
  • Python
  • API
  • text-to-speech
  • elevenlabs
  • audio
  • voice-synthesis
  • multilingual
  • streaming
  • javascript

README badge

README badge for elevenlabs/skills/text-to-speech

Converts text to speech using the ElevenLabs API with support for 70+ languages, multiple quality/latency models, and voice customization. Use this skill when building voice apps, generating voiceovers, or synthesizing speech in real-time applications; it requires an ELEVENLABS_API_KEY environment variable.

Generated from the current SKILL.md.

Does this skill support all languages?
The skill supports 70+ languages depending on the model. eleven_v3 supports 70+, eleven_multilingual_v2 supports 29, and flash/turbo variants support 32. You can enforce a specific language with the language_code parameter.
What latency should I expect?
Latency varies by model: eleven_flash_v2_5 and eleven_flash_v2 offer ~75ms, turbo variants ~250-300ms, and v3/multilingual_v2 use standard latency. Choose eleven_flash for real-time applications.
Can I use custom voices?
Yes. The skill includes pre-made voice IDs like George and Sarah, but you can also create and use custom voices via the ElevenLabs dashboard.
Does this require an API key?
Yes. The skill requires an ElevenLabs API key set in the ELEVENLABS_API_KEY environment variable.
What output formats are supported?
The skill supports MP3 (multiple bitrates), PCM (multiple sample rates), Opus, WAV, and telephony codecs (ulaw/alaw). Default is MP3 44.1kHz 128kbps.

Generated from the current SKILL.md. These answers refresh after source changes.