All skills
elevenlabs avatar

/speech-to-text

@279173d official
by elevenlabselevenlabs/skills462 stars
74

Transcribe audio to text using ElevenLabs Scribe v2. Use when converting audio/video to text, generating subtitles, transcribing meetings, or processing spoken content.

Use this Skill: https://skilld.dev/gh/elevenlabs/skills/speech-to-text

This session only. Nothing lands on disk.

referencesrealtime-commit-strategies.md

≈1.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Transcripts and Commit Strategies

Control when and how transcripts are finalized in real-time streaming.

Why Commits Matter

In real-time transcription, the model continuously refines its understanding as more audio arrives. A word that sounds like "their" might become "there" or "they're" once more context is heard. The commit mechanism lets you decide when to "lock in" the transcript.

Transcript Types

Type Description
Partial Interim "best guess" results that update frequently as audio is processed. Use for live feedback (showing text as the user speaks), but don't save these - they may change.
Committed Final, stable results after a commit occurs. Use these as the source of truth for your application - they won't change.
Committed with Timestamps Same as committed, but includes word-level timing data for subtitles, karaoke, or lip-sync.

Manual Commit (Default)

You explicitly control when transcript segments finalize.

Python

async with client.speech_to_text.realtime.connect(
    model_id="scribe_v2_realtime",
) as connection:
    # Send audio
    await connection.send({
        "audio_base_64": audio_base_64,
        "sample_rate": 16000,
    })

    # Commit when ready (e.g., pause in speech, end of sentence)
    await connection.commit()

JavaScript

const connection = await client.speechToText.realtime.connect({
  modelId: "scribe_v2_realtime",
});

// Send audio
connection.send({
  audioBase64: audioBase64,
  sampleRate: 16000,
});

// Commit when ready
connection.commit();

Best Practices

  • Commit every 20-30 seconds for optimal performance
  • Commit during silence or logical breaks (end of sentence, speaker change)
  • Auto-commit at 90 seconds if no manual commit is sent

Providing Context

Send previous text with the first audio chunk to help the model:

await connection.send({
    "audio_base_64": first_chunk,
    "sample_rate": 16000,
    "previous_text": "So as I was saying,"  # Keep under 50 characters
})

This helps with:

  • Continuing conversations after reconnection
  • Providing context for better accuracy
  • Handling sentence fragments

Voice Activity Detection (VAD)

VAD listens for silence and automatically commits when the speaker pauses. This creates natural transcript segments that match how people actually speak - pausing between sentences and thoughts. Recommended for live microphone input.

Configuration

React (useScribe)
import { useScribe, CommitStrategy } from "@elevenlabs/react";

const scribe = useScribe({
  modelId: "scribe_v2_realtime",
  commitStrategy: CommitStrategy.VAD,
  // Optional VAD tuning:
  vadSilenceThresholdSecs: 1.5,    // Silence duration before commit
  vadThreshold: 0.4,               // Speech detection sensitivity (0-1)
  minSpeechDurationMs: 100,        // Minimum speech length required
  minSilenceDurationMs: 100,       // Minimum silence length required
});

Important: The default is CommitStrategy.MANUAL. For microphone input, always set CommitStrategy.VAD — without it, committed transcripts will never fire and the connection may drop.

JavaScript client
const connection = await client.speechToText.realtime.connect({
  modelId: "scribe_v2_realtime",
  vad: {
    silenceThresholdSecs: 1.5,    // Silence duration before commit
    threshold: 0.4,               // Speech detection sensitivity (0-1)
    minSpeechDurationMs: 100,     // Minimum speech length required
    minSilenceDurationMs: 100,    // Minimum silence length required
  },
});

Parameters

Parameter Description Default
silenceThresholdSecs Seconds of silence before auto-commit 1.5
threshold Speech detection sensitivity (lower = more sensitive) 0.4
minSpeechDurationMs Ignore speech shorter than this 100
minSilenceDurationMs Ignore silence shorter than this 100

When to Use VAD

  • Live microphone input
  • Conversational applications
  • When natural speech boundaries are preferred
  • Client-side implementations

When to Use Manual Commit

  • Processing audio files
  • Known segment boundaries
  • Maximum control over timing
  • Server-side batch processing

Supported Audio Formats

Format Sample Rate Notes
PCM 16-bit 16kHz Recommended, best balance
PCM 16-bit 8kHz - 48kHz Supported range
μ-law 8-bit 8kHz Telephony compatibility

Source: SKILL.md on GitHub

2 warnings1d5 checks · Risk SAFE
  • Gen Agent Trust Hub1d

    The skill facilitates audio and video transcription through ElevenLabs Scribe v2. Security analysis identifies a standard indirect prompt injection surface where instructions contained within the audio/video input could influence agent behavior after transcription. All external tools and scripts originate from official vendor repositories.

  • Socket1d

    No alerts

  • Snyk1d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    4/7 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 279173d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 days ago
compatibility
Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
Other metadata
metadata
{
  "openclaw": {
    "requires": {
      "env": [
        "ELEVENLABS_API_KEY"
      ]
    },
    "primaryEnv": "ELEVENLABS_API_KEY"
  }
}
  • API
  • Python
  • elevenlabs
  • speech-to-text
  • transcription
  • audio
  • diarization
  • timestamps
  • javascript

README badge

README badge for elevenlabs/skills/speech-to-text

Transcribe audio and video to text using ElevenLabs Scribe v2, with support for 90+ languages, speaker diarization, word-level timestamps, and real-time streaming (~150ms latency). Covers batch transcription for meetings and subtitles, live voice agent input, and keyterm prompting for domain-specific vocabulary.

Generated from the current SKILL.md.

Does this skill support real-time transcription?
Yes. Use `scribe_v2_realtime` model for live transcription with ~150ms latency. It supports Voice Activity Detection (VAD) auto-commit or manual commit strategies, and works with both server-side streaming and client-side React components.
What languages does this support?
The skill supports 90+ languages with automatic detection. You can optionally provide a language hint using ISO 639-1 or ISO 639-3 codes.
Does this provide speaker identification?
Yes. Enable `diarize=true` to label each word with a speaker ID. For call recordings, you can also set `detect_speaker_roles=true` to label speakers as `agent` and `customer`.
What file formats and sizes does this handle?
Supports MP3, WAV, M4A, FLAC, OGG, WebM, AAC, AIFF, Opus for audio and MP4, AVI, MKV, MOV, WMV, FLV, WebM, MPEG, 3GPP for video. Maximum file size is 5.0GB with a 10-hour duration limit.
Can I get word-level timing information?
Yes. Set `timestamps_granularity="word"` to get start/end times for each word, plus type classification (word, spacing, audio_event) and speaker identification.

Generated from the current SKILL.md. These answers refresh after source changes.