All skills
elevenlabs avatar

/speech-engine

@44a05ea official
by elevenlabselevenlabs/skills462 stars
74

Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine. Use when building Speech Engine servers, WebSocket handlers, WebRTC browser clients, conversation token endpoints, interruption-aware streaming responses, or voice-enabled chat agents that connect developer-owned server logic to ElevenLabs speech-to-text and text-to-speech.

Use this Skill: https://skilld.dev/gh/elevenlabs/skills/speech-engine

This session only. Nothing lands on disk.

referencespython-sdk-reference.md

≈1.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Python SDK Reference

Use the async ElevenLabs SDK for Speech Engine servers.

from elevenlabs import AsyncElevenLabs

elevenlabs = AsyncElevenLabs()

Resource Methods

Create

Only speech_engine.ws_url is required. Use a secure WebSocket URL such as wss://example.com/ws. Add optional config blocks when the Speech Engine needs custom voice, speech recognition, turn-taking, request headers, client-side first-message overrides, or privacy behavior.

engine = await elevenlabs.speech_engine.create(
    name="My Speech Engine",
    speech_engine={
        "ws_url": "wss://example.com/ws",
        "request_headers": {
            "x-agent-runtime": "openclaw",
        },
    },
    overrides={
        "first_message": True,
    },
    tts={
        "model_id": "eleven_flash_v2_5",
        "voice_id": "cjVigY5qzO86Huf0OWal",
        "optimize_streaming_latency": "2",
    },
    asr={
        "provider": "scribe_realtime",
        "keywords": ["OpenClaw", "Acme Cloud"],
    },
    turn={
        "turn_eagerness": "normal",
        "speculative_turn": True,
    },
    privacy={
        "record_voice": False,
    },
)

print(engine.engine_id)

Enable overrides.first_message before using overrides.agent.firstMessage when starting a browser session.

Get

engine = await elevenlabs.speech_engine.get("seng_...")

The returned resource has an engine ID plus helpers for serving Speech Engine traffic from a trusted Python process.

Serve

Run a Speech Engine server on the configured WebSocket path. Keep response generation behind a validation boundary so raw speech-recognition text does not directly control responses, tools, secrets, or privileged actions.

engine = await elevenlabs.speech_engine.get(os.environ["ELEVENLABS_SPEECH_ENGINE_ID"])
await engine.serve(port=3001, path="/ws", debug=True, callbacks=validated_callbacks)

Key parameters:

Parameter Default Purpose
port 3001 Port to listen on
path None Restrict WebSocket connections to one path
debug False Log protocol details while developing
disable_auth False Skip JWT verification. Dangerous — see below

Common callback keys include on_init, on_transcript, on_close, on_disconnect, and on_error. Use on_close for clean disconnects from ElevenLabs and on_disconnect when the WebSocket drops unexpectedly.

disable_auth (dangerous)

engine.serve() and SpeechEngineServer verify the X-Elevenlabs-Speech-Engine-Authorization JWT on every incoming connection by default. Passing disable_auth=True turns that check off. When it is off, the server accepts any client that can reach it — an attacker who finds the URL can open unlimited conversations, drain your ElevenLabs and downstream LLM quota, and inject arbitrary transcripts into your response pipeline.

Only recommend this option when the developer has already implemented at least one compensating control:

  • an IP allowlist restricting inbound traffic to ElevenLabs' egress ranges, or
  • a custom shared-secret header — configured via speech_engine.request_headers on the Speech Engine resource at create time — validated by an upstream proxy or middleware before the request reaches the SDK.

If neither is in place, do not disable auth. When it is enabled, the SDK emits a UserWarning at startup to make the state visible in logs. api_key is not required in this mode, since it is only used for JWT verification.

verify_request

Use only when managing WebSocket upgrades manually:

is_valid = engine.verify_request(headers)

It checks X-Elevenlabs-Speech-Engine-Authorization against a JWT signed with the SHA-256 hash of the ElevenLabs API key.

Session API

Each Speech Engine session represents one conversation.

Member Purpose
conversation_id Assigned after initialization
is_open Whether the WebSocket is open
send_response(response) Send response text or a text stream back for TTS
run() Run the receive loop for manual sessions
close() Close the WebSocket

send_response() accepts a string or async iterable of response text.

Safety

Speech-recognition text is untrusted user-controlled data. Validate intent with deterministic checks, allowlists, or explicit confirmation before it affects response generation, tool calls, secrets, or privileged workflows.

Wire Protocol

The SDK handles protocol details automatically. Outgoing messages from your server are response text chunks and connection keep-alives.

Source: SKILL.md on GitHub

1 warning17d3 checks · Risk SAFE
  • Gen Agent Trust Hub17d

    This skill provides a secure framework for building voice-enabled agents using ElevenLabs' Speech Engine. It correctly implements security best practices, including mandatory JWT authentication for WebSockets and explicit warnings against treating user speech transcripts as trusted input.

  • Socket17d

    No alerts

  • Snyk17d

    Risk: MEDIUM · 1 issue

Signed by skilld at 44a05ea. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last month
compatibility
Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
Other metadata
metadata
{
  "openclaw": {
    "requires": {
      "env": [
        "ELEVENLABS_API_KEY"
      ]
    },
    "primaryEnv": "ELEVENLABS_API_KEY"
  }
}
  • elevenlabs
  • speech-to-text
  • text-to-speech
  • websocket
  • voice
  • real-time
  • webrtc
  • conversational-ai
  • agent-runtime
  • streaming

README badge

README badge for elevenlabs/skills/speech-engine

Builds real-time voice conversations by connecting a custom server to ElevenLabs speech-to-text and text-to-speech over WebSocket. Use this skill to add voice interfaces to chat apps or agent runtimes while keeping application logic server-side, with built-in support for interruption handling and turn-taking.

Generated from the current SKILL.md.

Does this skill work with hosted ElevenLabs Conversational AI agents?
No. This skill is for custom agent runtimes with developer-owned server logic. Use the agents skill instead if you're creating or configuring a hosted ElevenLabs agent with platform-managed prompts and tools.
What programming languages does this skill support?
Python and TypeScript/JavaScript. Use @elevenlabs/* packages for JavaScript and the Python SDK for async implementations.
Do I need to expose my server to the internet?
Yes. The Speech Engine WebSocket URL must be publicly accessible. For local development, use ngrok or similar; for production, deploy to a public HTTPS endpoint.
How should I handle security with speech-to-text input?
Treat all speech-recognition text as untrusted input. Validate it against allowlisted intents or confirm with the user before it affects responses, tools, or privileged actions.
Can the browser directly access the ElevenLabs API key?
No. Keep ELEVENLABS_API_KEY on the server only. The browser requests a conversation token from a server endpoint before starting a voice session.

Generated from the current SKILL.md. These answers refresh after source changes.