All skills
elevenlabs avatar

/speech-to-text

@279173d official
by elevenlabselevenlabs/skills462 stars
74

Transcribe audio to text using ElevenLabs Scribe v2. Use when converting audio/video to text, generating subtitles, transcribing meetings, or processing spoken content.

Use this Skill: https://skilld.dev/gh/elevenlabs/skills/speech-to-text

This session only. Nothing lands on disk.

referencesrealtime-events.md

≈2.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Real-Time Event Reference

Complete reference for events in real-time speech-to-text streaming.

Sent Events (Client → Server)

input_audio_chunk

Send audio data for transcription.

{
  "message_type": "input_audio_chunk",
  "audio_base_64": "<base64-encoded-pcm-audio>",
  "commit": false,
  "sample_rate": 16000
}
Field Type Required Description
message_type string Yes Always "input_audio_chunk"
audio_base_64 string Yes Base64-encoded PCM audio data
commit boolean Yes Whether to commit after this chunk
sample_rate number No Sample rate in Hz (8000-48000)
previous_text string No Context from prior transcript (first chunk only, max 50 chars)

commit

Finalize the current transcript segment.

{
  "message_type": "commit"
}

Received Events (Server → Client)

All received events use message_type as the discriminator field.

session_started

Connection established successfully.

{
  "message_type": "session_started",
  "session_id": "0b0a72b57fd743ebbed6555d44836cf2",
  "config": {
    "sample_rate": 16000,
    "audio_format": "pcm_16000",
    "language_code": "en",
    "model_id": "scribe_v2_realtime",
    "commit_strategy": "manual",
    "enable_logging": true,
    "include_timestamps": true
  }
}

partial_transcript

Interim transcription results, updates frequently as audio is processed.

{
  "message_type": "partial_transcript",
  "text": "Hello, how are"
}
Field Type Description
message_type string "partial_transcript"
text string Current partial transcription

final_transcript

A stable result for a segment after speech settles but before the segment is committed.

{
  "message_type": "final_transcript",
  "text": "Hello, how are you today?"
}

final_transcript_with_timestamps

A delayed final result sent when timestamps or language detection are enabled.

{
  "message_type": "final_transcript_with_timestamps",
  "text": "Hello, how are you today?",
  "language_code": "en",
  "words": [
    {"text": "Hello", "start": 0.0, "end": 0.32, "type": "word", "logprob": -0.03}
  ]
}

language_code and words are optional and depend on the enabled session options.

committed_transcript

Final transcription after commit.

{
  "message_type": "committed_transcript",
  "text": "Hello, how are you today?"
}
Field Type Description
message_type string "committed_transcript"
text string Finalized transcription

committed_transcript_with_timestamps

Final transcription with word-level timing. Sent after committed_transcript when include_timestamps=true.

{
  "message_type": "committed_transcript_with_timestamps",
  "text": "Hello, how are you today?",
  "language_code": "en",
  "words": [
    {"text": "Hello", "start": 0.0, "end": 0.32, "type": "word"},
    {"text": " ", "start": 0.32, "end": 0.35, "type": "spacing"},
    {"text": "how", "start": 0.40, "end": 0.55, "type": "word"}
  ]
}
Field Type Description
message_type string "committed_transcript_with_timestamps"
text string Full transcription text
language_code string Detected language code
words array Word-level timing data
words[].text string The word or token
words[].start number Start time in seconds
words[].end number End time in seconds
words[].type string "word", "spacing", or "audio_event"
words[].speaker_id string Speaker identifier (if diarization enabled)

edited_transcript

Edited version of a committed segment. Sent when the connection includes transcript_edit.

{
  "message_type": "edited_transcript",
  "text": "our next meeting is on the twelfth of July twenty twenty-six",
  "edited_text": "our next meeting is on 2026-07-12"
}
Field Type Description
message_type string "edited_transcript"
text string Committed transcript text that the instruction was applied to
edited_text string Edited text; identical to text when no edits were made

committed_transcript_entities

Entities detected in a committed segment when the connection includes entity_detection.

{
  "message_type": "committed_transcript_entities",
  "text": "My name is Alice.",
  "entities": [
    {
      "text": "Alice",
      "entity_type": "name_given",
      "start_char": 11,
      "end_char": 16
    }
  ]
}
Field Type Description
message_type string "committed_transcript_entities"
text string Committed transcript segment scanned for entities
entities array Detected entities; empty when none are found
entities[].text string Text identified as an entity
entities[].entity_type string Detected entity type
entities[].start_char integer Start character offset in text
entities[].end_char integer End character offset in text

Error Events

error

Sent when an error occurs.

{
  "message_type": "error",
  "error": "input_error"
}

Error Codes

Code Description
auth_error Invalid API key or token
quota_exceeded Usage limit reached
input_error Unsupported audio format or invalid input
rate_limited Too many requests
commit_throttled Commits sent too frequently
invalid_request Connection parameters were rejected; the session closes
session_time_limit_exceeded Session exceeded max duration
unaccepted_terms Terms not accepted in dashboard
resource_exhausted Server capacity reached
queue_overflow Server queue capacity reached
chunk_size_exceeded Audio chunk too large
insufficient_audio_activity Not enough speech detected
transcriber_error Internal processing error

Connection Events

open

WebSocket connection established (standard WebSocket event, not a JSON message).

close

WebSocket connection closed (standard WebSocket close frame with code and reason).

Event Handling Examples

Python

The Python SDK abstracts the wire protocol. You can use event.type (not message_type) when using the SDK's event objects:

async for event in connection:
    if event.type == "session_started":
        print(f"Session: {event.session_id}")
    elif event.type == "partial_transcript":
        print(f"Partial: {event.text}")
    elif event.type == "final_transcript":
        print(f"Stable segment: {event.text}")
    elif event.type == "committed_transcript":
        print(f"Final: {event.text}")
    elif event.type == "committed_transcript_with_timestamps":
        for word in event.words:
            print(f"  {word.text}: {word.start}s - {word.end}s")
    elif event.type == "edited_transcript":
        print(f"Edited: {event.edited_text}")
    elif event.type == "error":
        print(f"Error: {event.error}")
    elif event.type == "invalid_request":
        print(f"Connection rejected: {event.error}")

JavaScript

The JavaScript SDK uses event names matching the message_type values:

connection.on("session_started", (data) => {
  console.log("Session:", data.sessionId);
});

connection.on("partial_transcript", (data) => {
  console.log("Partial:", data.text);
});

connection.on("final_transcript", (data) => {
  console.log("Stable segment:", data.text);
});

connection.on("committed_transcript", (data) => {
  console.log("Final:", data.text);
});

connection.on("committed_transcript_with_timestamps", (data) => {
  for (const word of data.words) {
    console.log(`  ${word.text}: ${word.start}s - ${word.end}s`);
  }
});

connection.on("edited_transcript", (data) => {
  console.log("Edited:", data.edited_text);
});

connection.on("error", (error) => {
  console.error("Error:", error);
});

connection.on("invalid_request", (error) => {
  console.error("Connection rejected:", error.error);
});

Source: SKILL.md on GitHub

2 warnings1d5 checks · Risk SAFE
  • Gen Agent Trust Hub1d

    The skill facilitates audio and video transcription through ElevenLabs Scribe v2. Security analysis identifies a standard indirect prompt injection surface where instructions contained within the audio/video input could influence agent behavior after transcription. All external tools and scripts originate from official vendor repositories.

  • Socket1d

    No alerts

  • Snyk1d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    4/7 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 279173d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 days ago
compatibility
Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
Other metadata
metadata
{
  "openclaw": {
    "requires": {
      "env": [
        "ELEVENLABS_API_KEY"
      ]
    },
    "primaryEnv": "ELEVENLABS_API_KEY"
  }
}
  • API
  • Python
  • elevenlabs
  • speech-to-text
  • transcription
  • audio
  • diarization
  • timestamps
  • javascript

README badge

README badge for elevenlabs/skills/speech-to-text

Transcribe audio and video to text using ElevenLabs Scribe v2, with support for 90+ languages, speaker diarization, word-level timestamps, and real-time streaming (~150ms latency). Covers batch transcription for meetings and subtitles, live voice agent input, and keyterm prompting for domain-specific vocabulary.

Generated from the current SKILL.md.

Does this skill support real-time transcription?
Yes. Use `scribe_v2_realtime` model for live transcription with ~150ms latency. It supports Voice Activity Detection (VAD) auto-commit or manual commit strategies, and works with both server-side streaming and client-side React components.
What languages does this support?
The skill supports 90+ languages with automatic detection. You can optionally provide a language hint using ISO 639-1 or ISO 639-3 codes.
Does this provide speaker identification?
Yes. Enable `diarize=true` to label each word with a speaker ID. For call recordings, you can also set `detect_speaker_roles=true` to label speakers as `agent` and `customer`.
What file formats and sizes does this handle?
Supports MP3, WAV, M4A, FLAC, OGG, WebM, AAC, AIFF, Opus for audio and MP4, AVI, MKV, MOV, WMV, FLV, WebM, MPEG, 3GPP for video. Maximum file size is 5.0GB with a 10-hour duration limit.
Can I get word-level timing information?
Yes. Set `timestamps_granularity="word"` to get start/end times for each word, plus type classification (word, spacing, audio_event) and speaker identification.

Generated from the current SKILL.md. These answers refresh after source changes.