All skills
elevenlabs avatar

/speech-to-text

@279173d official
by elevenlabselevenlabs/skills462 stars
74

Transcribe audio to text using ElevenLabs Scribe v2. Use when converting audio/video to text, generating subtitles, transcribing meetings, or processing spoken content.

Use this Skill: https://skilld.dev/gh/elevenlabs/skills/speech-to-text

This session only. Nothing lands on disk.

SKILL.md

≈46 tokens always: the name and description. ≈2.4k when used: this file. ≈11k more on demand in 6 files.

ElevenLabs Speech-to-Text

Transcribe audio to text with Scribe v2 - supports 90+ languages, speaker diarization, and word-level timestamps.

Setup: See Installation Guide. For JavaScript, use @elevenlabs/* packages only.

Quick Start

Python

from elevenlabs import ElevenLabs

client = ElevenLabs()

with open("audio.mp3", "rb") as audio_file:
    result = client.speech_to_text.convert(file=audio_file, model_id="scribe_v2")

print(result.text)

JavaScript

import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createReadStream } from "fs";

const client = new ElevenLabsClient();
const result = await client.speechToText.convert({
  file: createReadStream("audio.mp3"),
  modelId: "scribe_v2",
});
console.log(result.text);

CLI

elevenlabs speech-to-text convert --file audio.mp3 --model-id scribe_v2

Models

Model ID Description Best For
scribe_v2 State-of-the-art accuracy, 90+ languages Batch transcription, subtitles, long-form audio
scribe_v2_medical Specialized recognition for medical and clinical audio, 90+ languages Clinical documentation, medical dictation, patient calls
scribe_v2_realtime Low latency (~150ms) Live transcription, voice agents
scribe_v2_realtime_turbo Realtime transcription variant Live transcription
scribe_v2_realtime_lite Realtime transcription variant Live transcription

Transcription with Timestamps

Word-level timestamps include type classification and speaker identification:

result = client.speech_to_text.convert(
    file=audio_file, model_id="scribe_v2", timestamps_granularity="word"
)

for word in result.words:
    print(f"{word.text}: {word.start}s - {word.end}s (type: {word.type})")

Speaker Diarization

Identify WHO said WHAT - the model labels each word with a speaker ID, useful for meetings, interviews, or any multi-speaker audio:

result = client.speech_to_text.convert(
    file=audio_file,
    model_id="scribe_v2",
    diarize=True
)

for word in result.words:
    print(f"[{word.speaker_id}] {word.text}")

For call recordings, the batch API can label diarized speakers as agent and customer by setting detect_speaker_roles=true alongside diarize=true. This option is not compatible with use_multi_channel=true.

If your workspace has registered speaker profiles, set use_speaker_library=true with diarize=true to match detected speakers against the speaker library.

elevenlabs speech-to-text convert \
  --file call.mp3 \
  --model-id scribe_v2 \
  --diarize true \
  --detect-speaker-roles true \
  --use-speaker-library true

Multichannel Audio

Use use_multi_channel=true when each speaker is isolated on a separate audio channel. By default, the API returns one transcript per channel under transcripts; set multichannel_output_style="combined" to receive one transcript merged by timestamp, with channel_index on each word.

result = client.speech_to_text.convert(
    file=audio_file,
    model_id="scribe_v2",
    use_multi_channel=True,
    multichannel_output_style="combined",
)

Keyterm Prompting

Help the model recognize specific words it might otherwise mishear - product names, technical jargon, or unusual spellings (up to 100 terms):

result = client.speech_to_text.convert(
    file=audio_file,
    model_id="scribe_v2",
    keyterms=["ElevenLabs", "Scribe", "API"]
)

Language Detection

Automatic detection with optional language hint:

result = client.speech_to_text.convert(
    file=audio_file,
    model_id="scribe_v2",
    language_code="eng"  # ISO 639-1 or ISO 639-3 code
)

print(f"Detected: {result.language_code} ({result.language_probability:.0%})")

Supported Formats

Audio: MP3, WAV, M4A, FLAC, OGG, WebM, AAC, AIFF, Opus Video: MP4, AVI, MKV, MOV, WMV, FLV, WebM, MPEG, 3GPP

Limits: Up to 5.0GB file size, 10 hours duration

Response Format

{
  "text": "The full transcription text",
  "language_code": "eng",
  "language_probability": 0.98,
  "words": [
    {"text": "The", "start": 0.0, "end": 0.15, "type": "word", "speaker_id": "speaker_0"},
    {"text": " ", "start": 0.15, "end": 0.16, "type": "spacing", "speaker_id": "speaker_0"}
  ]
}

Word types:

  • word - An actual spoken word
  • spacing - Whitespace between words (useful for precise timing)
  • audio_event - Non-speech sounds the model detected (laughter, applause, music, etc.)

Error Handling

try:
    result = client.speech_to_text.convert(file=audio_file, model_id="scribe_v2")
except Exception as e:
    print(f"Transcription failed: {e}")

Common errors:

  • 401: Invalid API key
  • 422: Invalid parameters
  • 429: Rate limit exceeded

Tracking Costs

Monitor usage via request-id response header:

response = client.speech_to_text.with_raw_response.convert(file=audio_file, model_id="scribe_v2")
result = response.data
print(f"Request ID: {response.headers.get('request-id')}")

Real-Time Streaming

For live transcription with ultra-low latency (~150ms), use the real-time API. The real-time API produces two types of transcripts:

  • Partial transcripts: Interim results that update frequently as audio is processed - use these for live feedback (e.g., showing text as the user speaks)
  • Committed transcripts: Final, stable results after you "commit" - use these as the source of truth for your application

A "commit" tells the model to finalize the current segment. You can commit manually (e.g., when the user pauses) or use Voice Activity Detection (VAD) to auto-commit on silence.

Python (Server-Side)

import asyncio
from elevenlabs import ElevenLabs

client = ElevenLabs()

async def transcribe_realtime():
    async with client.speech_to_text.realtime.connect(
        model_id="scribe_v2_realtime",
        include_timestamps=True,
        keyterms=["ElevenLabs", "Scribe"],
        no_verbatim=True,
    ) as connection:
        await connection.stream_url("https://example.com/audio.mp3")

        async for event in connection:
            if event.type == "partial_transcript":
                print(f"Partial: {event.text}")
            elif event.type == "committed_transcript":
                print(f"Final: {event.text}")

asyncio.run(transcribe_realtime())

JavaScript (Client-Side with React)

import { useScribe, CommitStrategy } from "@elevenlabs/react";

function TranscriptionComponent() {
  const [transcript, setTranscript] = useState("");

  const scribe = useScribe({
    modelId: "scribe_v2_realtime",
    commitStrategy: CommitStrategy.VAD, // Auto-commit on silence for mic input
    keyterms: ["ElevenLabs", "Scribe"],
    noVerbatim: true,
    includeLanguageDetection: true,
    onPartialTranscript: (data) => console.log("Partial:", data.text),
    onCommittedTranscript: (data) => setTranscript((prev) => prev + data.text),
  });

  const start = async () => {
    // Get token from your backend (never expose API key to client)
    const { token } = await fetch("/scribe-token").then((r) => r.json());

    await scribe.connect({
      token,
      microphone: { echoCancellation: true, noiseSuppression: true },
    });
  };

  return <button onClick={start}>Start Recording</button>;
}

Commit Strategies

Strategy Description
Manual You call commit() when ready - use for file processing or when you control the audio segments
VAD Voice Activity Detection auto-commits when silence is detected - use for live microphone input

Set includeLanguageDetection: true to receive the detected language code in delayed final transcript events.

// React: set commitStrategy on the hook (recommended for mic input)
import { useScribe, CommitStrategy } from "@elevenlabs/react";

const scribe = useScribe({
  modelId: "scribe_v2_realtime",
  commitStrategy: CommitStrategy.VAD,
  keyterms: ["ElevenLabs", "Scribe"],
  noVerbatim: true,
  // Optional VAD tuning:
  vadSilenceThresholdSecs: 1.5,
  vadThreshold: 0.4,
});
// JavaScript client: pass vad config on connect
const connection = await client.speechToText.realtime.connect({
  modelId: "scribe_v2_realtime",
  keyterms: ["ElevenLabs", "Scribe"],
  noVerbatim: true,
  vad: {
    silenceThresholdSecs: 1.5,
    threshold: 0.4,
  },
});

Event Types

Event Description
partial_transcript Live interim results
final_transcript Stable segment result sent before the segment is committed
final_transcript_with_timestamps Delayed final result with timestamps and/or detected language
committed_transcript Final results after commit
committed_transcript_with_timestamps Final with word timing
committed_transcript_entities Entities detected in a committed segment
edited_transcript Original and edited text for a committed segment when transcript editing is enabled
invalid_request Connection parameters were rejected and the session closes
error Error occurred

See real-time references for complete documentation.

References

Source: SKILL.md on GitHub

2 warnings1d5 checks · Risk SAFE
  • Gen Agent Trust Hub1d

    The skill facilitates audio and video transcription through ElevenLabs Scribe v2. Security analysis identifies a standard indirect prompt injection surface where instructions contained within the audio/video input could influence agent behavior after transcription. All external tools and scripts originate from official vendor repositories.

  • Socket1d

    No alerts

  • Snyk1d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    4/7 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 279173d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 19 hours ago.

Activeupdated 2 days ago
compatibility
Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
Other metadata
metadata
{
  "openclaw": {
    "requires": {
      "env": [
        "ELEVENLABS_API_KEY"
      ]
    },
    "primaryEnv": "ELEVENLABS_API_KEY"
  }
}
  • API
  • Python
  • elevenlabs
  • speech-to-text
  • transcription
  • audio
  • diarization
  • timestamps
  • javascript

README badge

README badge for elevenlabs/skills/speech-to-text

Transcribe audio and video to text using ElevenLabs Scribe v2, with support for 90+ languages, speaker diarization, word-level timestamps, and real-time streaming (~150ms latency). Covers batch transcription for meetings and subtitles, live voice agent input, and keyterm prompting for domain-specific vocabulary.

Generated from the current SKILL.md.

Does this skill support real-time transcription?
Yes. Use `scribe_v2_realtime` model for live transcription with ~150ms latency. It supports Voice Activity Detection (VAD) auto-commit or manual commit strategies, and works with both server-side streaming and client-side React components.
What languages does this support?
The skill supports 90+ languages with automatic detection. You can optionally provide a language hint using ISO 639-1 or ISO 639-3 codes.
Does this provide speaker identification?
Yes. Enable `diarize=true` to label each word with a speaker ID. For call recordings, you can also set `detect_speaker_roles=true` to label speakers as `agent` and `customer`.
What file formats and sizes does this handle?
Supports MP3, WAV, M4A, FLAC, OGG, WebM, AAC, AIFF, Opus for audio and MP4, AVI, MKV, MOV, WMV, FLV, WebM, MPEG, 3GPP for video. Maximum file size is 5.0GB with a 10-hour duration limit.
Can I get word-level timing information?
Yes. Set `timestamps_granularity="word"` to get start/end times for each word, plus type classification (word, spacing, audio_event) and speaker identification.

Generated from the current SKILL.md. These answers refresh after source changes.