All skills
elevenlabs avatar

/voice-isolator

@62eb4fc official
by elevenlabselevenlabs/skills462 stars
74

Remove background noise and isolate vocals/speech from audio using ElevenLabs Voice Isolator (audio isolation) API. Use when cleaning up noisy recordings, removing music or background ambience from dialogue, isolating speech from field recordings, preparing audio for transcription, extracting vocals, or any "denoise / clean up / isolate voice" task.

Use this Skill: https://skilld.dev/gh/elevenlabs/skills/voice-isolator

This session only. Nothing lands on disk.

SKILL.md

≈92 tokens always: the name and description. ≈775 when used: this file. ≈495 more on demand in 1 file.

ElevenLabs Voice Isolator

Removes background noise from audio and isolates vocals/speech — useful for cleaning up noisy recordings, prepping audio for transcription, or pulling dialogue out of a mixed track.

Setup: See Installation Guide. For JavaScript, use @elevenlabs/* packages only.

Quick Start

Python

from elevenlabs import ElevenLabs

client = ElevenLabs()

with open("noisy.mp3", "rb") as audio_file:
    audio_stream = client.audio_isolation.convert(audio=audio_file)

with open("clean.mp3", "wb") as f:
    for chunk in audio_stream:
        f.write(chunk)

JavaScript

import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createReadStream, createWriteStream } from "fs";

const client = new ElevenLabsClient();

const audioStream = await client.audioIsolation.convert({
  audio: createReadStream("noisy.mp3"),
});

audioStream.pipe(createWriteStream("clean.mp3"));

CLI

elevenlabs audio-isolation convert --audio noisy.mp3 --output clean.mp3

Parameters

Parameter Type Default Description
audio file (required) — Audio file with vocals/speech to isolate
file_format string other other for any encoded audio, or pcm_s16le_16 for 16-bit PCM mono @ 16kHz little-endian (lower latency)

Isolating from a URL

import requests
from io import BytesIO
from elevenlabs import ElevenLabs

client = ElevenLabs()

audio_url = "https://example.com/noisy.mp3"
response = requests.get(audio_url)
audio_data = BytesIO(response.content)

audio_stream = client.audio_isolation.convert(audio=audio_data)

with open("clean.mp3", "wb") as f:
    for chunk in audio_stream:
        f.write(chunk)

Low-Latency PCM Input

If you already have raw 16-bit PCM mono @ 16kHz, passing file_format="pcm_s16le_16" skips decoding and reduces latency:

audio_stream = client.audio_isolation.convert(
    audio=pcm_bytes,
    file_format="pcm_s16le_16",
)

Supported Formats

Any common encoded audio/video container works as input (MP3, WAV, M4A, FLAC, OGG, WebM, MP4, etc.). Response is a streamed MP3 by default.

Common Workflows

  • Clean up interview/podcast recordings — strip room tone, HVAC, traffic before editing.
  • Prep noisy audio for Speech-to-Text — isolate voice first, then pass through speech_to_text.convert() for better transcription accuracy.
  • Extract dialogue from mixed tracks — pull vocals out of a track with music/SFX.
  • Pre-processing for Voice Changer — isolate the source voice before applying voice transformation.

Error Handling

try:
    audio_stream = client.audio_isolation.convert(audio=audio_file)
except Exception as e:
    print(f"Voice isolation failed: {e}")

Common errors:

  • 401: Invalid API key
  • 422: Invalid parameters (e.g. wrong file_format for the supplied audio)
  • 429: Rate limit exceeded

References

Source: SKILL.md on GitHub

No alerts16d3 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    The skill is safe. It provides standard integration with the ElevenLabs Voice Isolator API for audio cleaning. All installation methods, dependencies, and external URLs trace back to the official vendor's infrastructure.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

Signed by skilld at 62eb4fc. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last month
compatibility
Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
Other metadata
metadata
{
  "openclaw": {
    "requires": {
      "env": [
        "ELEVENLABS_API_KEY"
      ]
    },
    "primaryEnv": "ELEVENLABS_API_KEY"
  }
}
  • API
  • Python
  • elevenlabs
  • audio-isolation
  • noise-removal
  • speech-extraction
  • audio-processing
  • transcription
  • javascript

README badge

README badge for elevenlabs/skills/voice-isolator

Removes background noise and isolates speech or vocals from audio files using the ElevenLabs Voice Isolator API. Useful for cleaning up noisy recordings, extracting dialogue from mixed tracks, and preparing audio for transcription or downstream processing.

Generated from the current SKILL.md.

What audio formats does this skill accept?
Any common encoded audio or video container works as input: MP3, WAV, M4A, FLAC, OGG, WebM, MP4, etc. The response is streamed as MP3 by default.
Do I need an ElevenLabs API key?
Yes. The skill requires an ELEVENLABS_API_KEY environment variable to authenticate with the ElevenLabs API.
What programming languages are supported?
The skill includes examples for Python (via the ElevenLabs SDK), JavaScript (via @elevenlabs/elevenlabs-js), and cURL.
Can I reduce latency when processing audio?
Yes. If you have raw 16-bit PCM mono audio at 16kHz, pass file_format='pcm_s16le_16' to skip decoding and lower latency.
What happens if the API call fails?
Common errors include 401 (invalid API key), 422 (invalid parameters like wrong file format), and 429 (rate limit exceeded). The skill includes error handling patterns.

Generated from the current SKILL.md. These answers refresh after source changes.