All skills
elevenlabs avatar

/music

@68fff59 official
by elevenlabselevenlabs/skills462 stars
74

Generate music using ElevenLabs Music API. Use when creating instrumental tracks, songs with lyrics, background music, jingles, or any AI-generated music composition. Supports prompt-based generation, composition plans for granular control, and detailed output with metadata.

Use this Skill: https://skilld.dev/gh/elevenlabs/skills/music

This session only. Nothing lands on disk.

SKILL.md

≈71 tokens always: the name and description. ≈3.3k when used: this file. ≈5k more on demand in 2 files.

ElevenLabs Music Generation

Generate music from text prompts - supports instrumental tracks, songs with lyrics, and fine-grained control via composition plans.

Setup: See Installation Guide. For JavaScript, use @elevenlabs/* packages only.

All examples below use music_v2_5, the most advanced generation model. Pass music_v2 or music_v1 only when an older model is explicitly requested.

Quick Start

Python

from elevenlabs import ElevenLabs

client = ElevenLabs()

audio = client.music.compose(
    prompt="A chill lo-fi hip hop beat with jazzy piano chords",
    music_length_ms=30000,
    model_id="music_v2_5",
)

with open("output.mp3", "wb") as f:
    for chunk in audio:
        f.write(chunk)

TypeScript

import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createWriteStream } from "fs";

const client = new ElevenLabsClient();
const audio = await client.music.compose({
  prompt: "A chill lo-fi hip hop beat with jazzy piano chords",
  musicLengthMs: 30000,
  modelId: "music_v2_5",
});
audio.pipe(createWriteStream("output.mp3"));

CLI

elevenlabs music compose \
  --prompt "A chill lo-fi beat" \
  --music-length-ms 30000 \
  --model-id music_v2_5 \
  --output output.mp3

Methods

Method Description
music.compose Generate audio from a prompt or composition plan
music.stream Stream audio chunks as they are generated (paid plans)
music.composition_plan.create Generate a structured plan for fine-grained control
music.compose_detailed Generate audio + composition plan + metadata; pass store_for_inpainting=True to enable inpainting
music.compose_detailed_stream Stream audio plus composition plan, metadata, and optional word timestamps as Server-Sent Events
music.video_to_music Generate background music from one or more uploaded video files
music.upload Upload an audio file for later inpainting workflows, optionally extracting its composition plan or word-level timestamps
music.finetunes.list List accessible music finetunes
music.finetunes.create Train a music finetune from uploaded audio
music.finetunes.get Retrieve finetune status and metadata
music.finetunes.update Update finetune metadata or visibility
music.finetunes.delete Delete a music finetune

See API Reference for full parameter details.

music.upload is available to enterprise clients with access to the inpainting feature.

Music Finetunes

Create a finetune from training audio with POST /v1/music/finetunes, then poll the get endpoint until its status is completed. Pass the returned id as finetune_id when composing music.

Use the list, update, and delete endpoints to manage accessible finetunes.

Video to Music

Generate background music from uploaded video clips via POST /v1/music/video-to-music (client.music.video_to_music). This is separate from prompt-based music.compose (POST /v1/music).

The API combines videos in order, accepts an optional natural-language description, and lets you steer style with up to 10 tags such as upbeat or cinematic. This endpoint still defaults to music_v1; pass model_id="music_v2_5" to use the most advanced model.

Python

from elevenlabs import ElevenLabs

client = ElevenLabs()

audio = client.music.video_to_music(
    videos=["trailer.mp4"],
    description="Build suspense, then resolve with a warm cinematic finish.",
    tags=["cinematic", "suspenseful", "uplifting"],
    model_id="music_v2_5",
)

with open("video-score.mp3", "wb") as f:
    for chunk in audio:
        f.write(chunk)

TypeScript

import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createReadStream, createWriteStream } from "fs";

const client = new ElevenLabsClient();

const audio = await client.music.videoToMusic({
  videos: [createReadStream("trailer.mp4")],
  description: "Build suspense, then resolve with a warm cinematic finish.",
  tags: ["cinematic", "suspenseful", "uplifting"],
  modelId: "music_v2_5",
});

audio.pipe(createWriteStream("video-score.mp3"));

CLI

elevenlabs music video_to_music \
  --videos trailer.mp4 \
  --description "Build suspense, then resolve with a warm cinematic finish." \
  --tags cinematic \
  --model-id music_v2_5 \
  --output video-score.mp3

The CLI currently accepts one --videos file and one --tags value per request; use the Python or TypeScript SDK to send multiple videos or tags.

Constraints from the current API schema:

  • Upload 1-10 video files per request
  • Keep total combined upload size at or below 200 MB
  • Keep total combined video duration at or below 600 seconds
  • Use description for high-level musical direction and tags for concise style cues

Composition Plans

music_v2_5 composition plans are an ordered list of chunks. Each chunk specifies its own text (section label, lyrics, inline cues), duration_ms, positive_styles, negative_styles, and context_adherence (low, medium, or high, default high). Up to 30 chunks per plan, each 3,000–120,000 ms, total length 3 s to 10 minutes. Each chunk's text supports up to 6,132 characters, with up to 30 lines of 200 characters each.

Generate a plan first, edit it, then compose:

plan = client.music.composition_plan.create(
    prompt="An epic orchestral piece building to a climax",
    music_length_ms=60000,
    model_id="music_v2_5",
)

# Edit chunks in place
plan["chunks"][0]["text"] = "[Intro]\nQuiet strings rising"

audio = client.music.compose(
    composition_plan=plan,
    model_id="music_v2_5",
)
const plan = await client.music.compositionPlan.create({
  prompt: "An epic orchestral piece building to a climax",
  musicLengthMs: 60000,
  modelId: "music_v2_5",
});

plan.chunks[0].text = "[Intro]\nQuiet strings rising";

const audio = await client.music.compose({
  compositionPlan: plan,
  modelId: "music_v2_5",
});

Or hand-build a plan to control lyrics and style per section:

composition_plan = {
    "chunks": [
        {
            "text": "[Verse]\nWalking down an empty street",
            "duration_ms": 15000,
            "positive_styles": ["pop", "upbeat", "female vocals", "acoustic guitar"],
            "negative_styles": ["dark", "slow"],
            "context_adherence": "high",
        },
        {
            "text": "[Chorus]\nThis is my moment",
            "duration_ms": 15000,
            "positive_styles": ["powerful vocals", "full band"],
            "negative_styles": [],
            "context_adherence": "high",
        },
    ]
}

audio = client.music.compose(composition_plan=composition_plan, model_id="music_v2_5")
const compositionPlan = {
  chunks: [
    {
      text: "[Verse]\nWalking down an empty street",
      durationMs: 15000,
      positiveStyles: ["pop", "upbeat", "female vocals", "acoustic guitar"],
      negativeStyles: ["dark", "slow"],
      contextAdherence: "high",
    },
    {
      text: "[Chorus]\nThis is my moment",
      durationMs: 15000,
      positiveStyles: ["powerful vocals", "full band"],
      negativeStyles: [],
      contextAdherence: "high",
    },
  ],
};

const audio = await client.music.compose({
  compositionPlan,
  modelId: "music_v2_5",
});

Put broader characteristics (genre, instrumentation, vocal style) in positive_styles, not in text. The first chunk's styles set the overall tone — include 6–7 styles there.

Output Formats

Use the output_format query parameter on compose, detailed compose, or stream requests to select the generated audio format. auto chooses a model-appropriate MP3 format; for music_v2 and music_v2_5, it selects mp3_48000_192. Higher-bitrate MP3 options include mp3_48000_240 and mp3_48000_320.

Streaming

For paid plans, stream audio chunks as they are generated instead of waiting for the full file:

from io import BytesIO

stream = client.music.stream(
    prompt="A driving synthwave track with arpeggiated leads",
    music_length_ms=30000,
    model_id="music_v2_5",
)

buffer = BytesIO()
for chunk in stream:
    if chunk:
        buffer.write(chunk)
const stream = await client.music.stream({
  prompt: "A driving synthwave track with arpeggiated leads",
  musicLengthMs: 30000,
  modelId: "music_v2_5",
});

const chunks: Buffer[] = [];
for await (const chunk of stream) {
  chunks.push(chunk);
}

Detailed streaming

Use detailed streaming when the application needs generated music metadata while audio is still arriving. POST /v1/music/detailed/stream accepts the same prompt or composition-plan body as detailed compose, streams text/event-stream, and can include word timestamps with with_timestamps.

elevenlabs music compose_detailed_stream \
  --prompt "A bright indie pop hook with warm guitars" \
  --music-length-ms 30000 \
  --model-id music_v2_5 \
  --with-timestamps true \
  --output-format auto

Inpainting

Inpainting edits or extends a stored song by mixing audio reference chunks (unchanged slices of a stored song) with new generation chunks in a single composition plan.

Step 1 — get a song_id, either by storing a fresh generation or uploading existing audio:

# Option A: keep a generation for later editing
result = client.music.compose_detailed(
    prompt="An upbeat pop song with verse and chorus",
    music_length_ms=60000,
    model_id="music_v2_5",
    store_for_inpainting=True,
)
song_id = result.song_id

# Option B: upload an existing track and extract its plan
uploaded = client.music.upload(
    file=open("my-song.mp3", "rb"),
    extract_composition_plan="music_v2_5",
)
song_id = uploaded.song_id
composition_plan = uploaded.composition_plan
import { createReadStream } from "fs";

// Option A: keep a generation for later editing
const result = await client.music.composeDetailed({
  prompt: "An upbeat pop song with verse and chorus",
  musicLengthMs: 60000,
  modelId: "music_v2_5",
  storeForInpainting: true,
});
let songId = result.songId;

// Option B: upload an existing track and extract its plan
const uploaded = await client.music.upload({
  file: createReadStream("my-song.mp3"),
  extractCompositionPlan: "music_v2_5",
});
songId = uploaded.songId;
const compositionPlan = uploaded.compositionPlan;

Step 2 — compose a plan that references the stored audio and regenerates the part you want to change:

plan = {
    "chunks": [
        {"song_id": song_id, "range": {"start_ms": 0, "end_ms": 30000}},
        {
            "text": "[Chorus]\nWe're rising up tonight",
            "duration_ms": 30000,
            "positive_styles": ["bigger drums", "layered vocals", "anthemic"],
            "negative_styles": ["sparse"],
            "context_adherence": "high",
        },
    ]
}

audio = client.music.compose(composition_plan=plan, model_id="music_v2_5")
const plan = {
  chunks: [
    { songId, range: { startMs: 0, endMs: 30000 } },
    {
      text: "[Chorus]\nWe're rising up tonight",
      durationMs: 30000,
      positiveStyles: ["bigger drums", "layered vocals", "anthemic"],
      negativeStyles: ["sparse"],
      contextAdherence: "high",
    },
  ],
};

const audio = await client.music.compose({
  compositionPlan: plan,
  modelId: "music_v2_5",
});

To match the feel of a stored slice without copying it, attach a conditioning_ref (up to 30,000 ms) plus a condition_strength of low, medium, high, or xhigh to a generation chunk. Conditioning placed on the first chunk influences every later chunk.

See API Reference for the full inpainting parameter list.

Content Restrictions

  • Cannot reference specific artists, bands, or copyrighted lyrics
  • bad_prompt errors include a prompt_suggestion with alternative phrasing
  • bad_composition_plan errors include a composition_plan_suggestion

Error Handling

try:
    audio = client.music.compose(prompt="...", music_length_ms=30000)
except Exception as e:
    print(f"API error: {e}")
try {
  const audio = await client.music.compose({
    prompt: "...",
    musicLengthMs: 30000,
  });
} catch (err) {
  console.error("API error:", err);
}

Common errors: 401 (invalid key), 422 (invalid params), 429 (rate limit).

References

Source: SKILL.md on GitHub

1 warning5d5 checks · Risk SAFE
  • Gen Agent Trust Hub5d

    The skill provides instructions for generating music via the ElevenLabs Music API using official SDKs and CLI tools. All external resources and installation scripts originate from the official vendor and follow secure practices for secret management and tool installation.

  • Socket5d

    No alerts

  • Snyk5d

    Risk: LOW · No issues

  • Runlayer6mo

    2/3 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 68fff59. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated last week
compatibility
Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
Other metadata
metadata
{
  "openclaw": {
    "requires": {
      "env": [
        "ELEVENLABS_API_KEY"
      ]
    },
    "primaryEnv": "ELEVENLABS_API_KEY"
  }
}

README badge

README badge for elevenlabs/skills/music

Generates instrumental tracks, songs with lyrics, and background music using the ElevenLabs Music API. Supports prompt-based composition, fine-grained control via composition plans with per-section styling, video-to-music generation, and inpainting to edit or extend stored tracks. Requires an ElevenLabs API key and internet access.

Generated from the current SKILL.md.

Does this skill require an API key?
Yes. You must set the ELEVENLABS_API_KEY environment variable to authenticate with the ElevenLabs Music API.
Can I generate music with lyrics?
Yes. Use composition plans to specify lyrics per section (e.g. [Verse], [Chorus]), or pass a prompt that describes a song with vocals. The API will generate audio matching your text.
Does the skill support editing or extending existing audio?
Yes, via inpainting. You can store a generated track or upload existing audio, then create a composition plan that references unchanged slices and regenerates only the parts you want to change.
What models are available?
The skill defaults to music_v2, the current generation model. You can pass model_id='music_v1' if explicitly needed, though music_v2 is recommended.
Can I generate music from video?
Yes. The video_to_music method accepts 1-10 video files (up to 200 MB combined, 600 seconds duration) and generates background music based on a description and optional style tags.

Generated from the current SKILL.md. These answers refresh after source changes.