All skills
elevenlabs avatar

/speech-to-text

@279173d official
by elevenlabselevenlabs/skills462 stars
74

Transcribe audio to text using ElevenLabs Scribe v2. Use when converting audio/video to text, generating subtitles, transcribing meetings, or processing spoken content.

Use this Skill: https://skilld.dev/gh/elevenlabs/skills/speech-to-text

This session only. Nothing lands on disk.

referencesrealtime-client-side.md

≈1.7k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Client-Side Real-Time Streaming

Stream audio from the browser directly to ElevenLabs for real-time transcription.

Installation

# React
npm install @elevenlabs/react@latest @elevenlabs/elevenlabs-js@latest

# JavaScript
npm install @elevenlabs/client@latest @elevenlabs/elevenlabs-js@latest

Warning: Always use the @elevenlabs/* namespace for client-side packages.

Token Generation

Client-side streaming requires a single-use token to protect your API key. Generate tokens on your backend:

import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";

const elevenlabs = new ElevenLabsClient({
  apiKey: process.env.ELEVENLABS_API_KEY,
});

app.get("/scribe-token", yourAuthMiddleware, async (req, res) => {
  const token = await elevenlabs.tokens.singleUse.create("realtime_scribe");
  res.json(token);
});

Note: Single-use tokens expire after 15 minutes.

React Implementation

import { useScribe, CommitStrategy } from "@elevenlabs/react";

function TranscriptionComponent() {
  const [transcript, setTranscript] = useState("");

  const scribe = useScribe({
    modelId: "scribe_v2_realtime",
    commitStrategy: CommitStrategy.VAD, // Auto-commit on silence for mic input
    includeLanguageDetection: true,
    onPartialTranscript: (data) => {
      // Show live feedback as user speaks
      console.log("Partial:", data.text);
    },
    onCommittedTranscript: (data) => {
      // Final transcript for this segment
      setTranscript((prev) => prev + data.text);
    },
  });

  const startRecording = async () => {
    const tokenResponse = await fetch("/scribe-token");
    const { token } = await tokenResponse.json();

    await scribe.connect({
      token,
      microphone: {
        echoCancellation: true,
        noiseSuppression: true,
        autoGainControl: true,
      },
    });
  };

  const stopRecording = () => {
    scribe.disconnect();
  };

  return (
    <div>
      <div>Status: {scribe.status}</div>
      <button onClick={startRecording}>Start</button>
      <button onClick={stopRecording}>Stop</button>
      <p>{transcript}</p>
    </div>
  );
}

Important: The default commit strategy is CommitStrategy.MANUAL, which requires you to call scribe.commit() explicitly. For microphone input, always set CommitStrategy.VAD so the server auto-commits when silence is detected. Without this, committed transcripts will never fire and the connection may drop.

scribe.status Values

Status Meaning
"disconnected" No active connection
"connecting" Connection is being established
"connected" Connected and ready to receive audio
"transcribing" Actively processing speech (transitions from "connected" when audio is detected or VAD commits)
"error" An error occurred

Important: When checking if the session is active, always check for both "connected" and "transcribing". The status transitions to "transcribing" during speech processing, so checking only "connected" will cause UI elements (buttons, waveforms, indicators) to incorrectly reset mid-session.

// Correct - handles both active states
const isListening = scribe.status === "connected" || scribe.status === "transcribing";

// Wrong - will flicker/reset when VAD commits
const isListening = scribe.status === "connected";

JavaScript Implementation

import { Scribe, RealtimeEvents } from "@elevenlabs/client";

async function startTranscription() {
  const tokenResponse = await fetch("/scribe-token");
  const { token } = await tokenResponse.json();

  const connection = Scribe.connect({
    token,
    modelId: "scribe_v2_realtime",
    includeTimestamps: true,
    includeLanguageDetection: true,
    keyterms: ["ElevenLabs", "Scribe"],
    noVerbatim: true,
    microphone: {
      echoCancellation: true,
      noiseSuppression: true,
      autoGainControl: true,
    },
  });

  connection.on(RealtimeEvents.OPEN, () => {
    console.log("Connected");
  });

  connection.on(RealtimeEvents.PARTIAL_TRANSCRIPT, (data) => {
    console.log("Partial:", data.text);
  });

  connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT, (data) => {
    console.log("Committed:", data.text);
  });

  connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT_WITH_TIMESTAMPS, (data) => {
    for (const word of data.words) {
      console.log(`${word.text}: ${word.start}s - ${word.end}s`);
    }
  });

  connection.on(RealtimeEvents.ERROR, (error) => {
    console.error("Error:", error);
  });

  connection.on(RealtimeEvents.CLOSE, () => {
    console.log("Disconnected");
  });

  return connection;
}

keyterms biases realtime recognition toward important terms. noVerbatim removes filler words, false starts, and disfluencies from committed transcripts. includeLanguageDetection returns the detected language code in a delayed final transcript event.

Both Scribe.connect and useScribe accept secondaryLanguages for expected additional languages, entityDetection for entity events, and filterBackgroundAudio to reduce false activation from background speech and ambient noise. Do not combine filterBackgroundAudio with includeTimestamps. Enterprise zero-retention sessions can set enableLogging: false.

Set transcriptEdit on Scribe.connect to apply a natural-language instruction of up to 2,000 characters to each committed segment. Listen for RealtimeEvents.EDITED_TRANSCRIPT; each event contains the original text and the edited_text. Transcript editing cannot be combined with entityDetection and adds a 30% premium, billed for at least 10 seconds of audio per committed segment.

Manual Audio Chunking

For file uploads or custom audio sources, encode to PCM-16 and send in chunks:

const chunkSize = 4096;

for (let offset = 0; offset < pcmData.length; offset += chunkSize) {
  const chunk = pcmData.slice(offset, offset + chunkSize);
  const bytes = new Uint8Array(chunk.buffer);
  const base64 = btoa(String.fromCharCode(...bytes));

  scribe.sendAudio(base64);

  // Simulate real-time streaming
  await new Promise((resolve) => setTimeout(resolve, 50));
}

// Finalize transcription
scribe.commit();

Microphone Options

Option Description
echoCancellation Remove echo from speakers
noiseSuppression Filter background noise
autoGainControl Normalize volume levels

Security

  • Never expose your API key to the client
  • Always generate single-use tokens on your backend
  • Use authentication middleware to protect token endpoints
  • For enterprise zero-retention sessions, set enableLogging: false in Scribe.connect or useScribe; this disables history features for the session

Source: SKILL.md on GitHub

2 warnings1d5 checks · Risk SAFE
  • Gen Agent Trust Hub1d

    The skill facilitates audio and video transcription through ElevenLabs Scribe v2. Security analysis identifies a standard indirect prompt injection surface where instructions contained within the audio/video input could influence agent behavior after transcription. All external tools and scripts originate from official vendor repositories.

  • Socket1d

    No alerts

  • Snyk1d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    4/7 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 279173d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 days ago
compatibility
Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
Other metadata
metadata
{
  "openclaw": {
    "requires": {
      "env": [
        "ELEVENLABS_API_KEY"
      ]
    },
    "primaryEnv": "ELEVENLABS_API_KEY"
  }
}
  • API
  • Python
  • elevenlabs
  • speech-to-text
  • transcription
  • audio
  • diarization
  • timestamps
  • javascript

README badge

README badge for elevenlabs/skills/speech-to-text

Transcribe audio and video to text using ElevenLabs Scribe v2, with support for 90+ languages, speaker diarization, word-level timestamps, and real-time streaming (~150ms latency). Covers batch transcription for meetings and subtitles, live voice agent input, and keyterm prompting for domain-specific vocabulary.

Generated from the current SKILL.md.

Does this skill support real-time transcription?
Yes. Use `scribe_v2_realtime` model for live transcription with ~150ms latency. It supports Voice Activity Detection (VAD) auto-commit or manual commit strategies, and works with both server-side streaming and client-side React components.
What languages does this support?
The skill supports 90+ languages with automatic detection. You can optionally provide a language hint using ISO 639-1 or ISO 639-3 codes.
Does this provide speaker identification?
Yes. Enable `diarize=true` to label each word with a speaker ID. For call recordings, you can also set `detect_speaker_roles=true` to label speakers as `agent` and `customer`.
What file formats and sizes does this handle?
Supports MP3, WAV, M4A, FLAC, OGG, WebM, AAC, AIFF, Opus for audio and MP4, AVI, MKV, MOV, WMV, FLV, WebM, MPEG, 3GPP for video. Maximum file size is 5.0GB with a 10-hour duration limit.
Can I get word-level timing information?
Yes. Set `timestamps_granularity="word"` to get start/end times for each word, plus type classification (word, spacing, audio_event) and speaker identification.

Generated from the current SKILL.md. These answers refresh after source changes.