All skills
heygen-com avatar

/media-use

@ff6e210
by HeyGenheygen-com/hyperframes55k stars
4,998

Agent Media OS, the single skill for every media need in a HyperFrames project. Resolve BGM, SFX, image, icon, brand logo, voice, color grade, or LUT into a frozen local file or paste-ready block + ledger record (one verb, `resolve`); generate via TTS / music / image models when the catalog misses; produce voiceover, transcription, captions, and background removal through one shared audio engine; operate on media (cut / reframe / transform); and reuse assets across projects. Also use for vague feedback that real footage looks dark, flat, boring, should feel retro/camcorder/print/ASCII, needs privacy, or needs a media reveal.

Use this Skill: https://skilld.dev/gh/heygen-com/hyperframes/media-use

This session only. Nothing lands on disk.

audioreferencesrequirements.md

≈1.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Requirements & Caches

Credential & key priority

Run npx hyperframes auth status to see what's configured and which engines a workflow will use (see the skill's Preflight section). Keys resolve in this order — first match wins:

Provider Resolution order (first non-empty wins) Local deps when used
HeyGen (TTS + BGM/SFX retrieval) $HEYGEN_API_KEY → $HYPERFRAMES_API_KEY → ~/.heygen/credentials (shared with heygen-cli; $HEYGEN_CONFIG_DIR overrides the dir; written by hyperframes auth login) none (REST)
ElevenLabs (TTS fallback) $ELEVENLABS_API_KEY pip install elevenlabs
Lyria (BGM fallback) $GEMINI_API_KEY → $GOOGLE_API_KEY pip install google-genai
Kokoro (TTS, no key) always — final voice fallback pip install kokoro-onnx soundfile
MusicGen (BGM, no key) always — final music fallback pip install transformers torch soundfile numpy

hyperframes auth login (browser OAuth) is the recommended setup: one sign-in, every project, no per-repo .env. An OAuth login is sent as Authorization: Bearer; an API key as X-Api-Key; both are tagged with X-HeyGen-Source: cli. OAuth CLI users can consume the web-plan free allowance for HeyGen TTS (10 min/month); API keys follow the normal API billing path. With no HeyGen credential, voice/BGM run fully locally (Kokoro / MusicGen) — hyperframes auth status and hyperframes doctor both report whether those local deps are installed.

Model caches & system dependencies

Each command downloads its own model on first run and caches it under ~/.cache/hyperframes/:

  • TTS (HeyGen) — no local deps; needs a HeyGen credential + ffmpeg on PATH (to transcode the mp3 response to .wav). Credential resolves like the CLI: $HEYGEN_API_KEY → $HYPERFRAMES_API_KEY → ~/.heygen/credentials (shared with heygen-cli; run npx hyperframes auth login). An OAuth login is sent as Authorization: Bearer; an API key as X-Api-Key; both include X-HeyGen-Source: cli so the backend can apply CLI OAuth free usage.
  • TTS (ElevenLabs) — same as HeyGen: API key + ffmpeg.
  • TTS (Kokoro) — Kokoro-82M (~311 MB) + voices (~27 MB) in tts/. Requires Python 3.8+ with kokoro-onnx and soundfile (pip install kokoro-onnx soundfile). Non-English text also needs espeak-ng system-wide.
  • BGM (Lyria) — needs $GEMINI_API_KEY or $GOOGLE_API_KEY + pip install google-genai. No local model cache.
  • BGM (MusicGen) — pip install transformers torch soundfile. facebook/musicgen-small (~300 MB) cached under ~/.cache/huggingface/ on first run.
  • Transcribe — Whisper model size depending on choice (75 MB – 3.1 GB) in whisper/, downloaded from HuggingFace on first use. whisper.cpp itself is NOT bundled: the CLI resolves it from PATH, installs via Homebrew (macOS), or builds it from source with git+cmake on first use ($HYPERFRAMES_WHISPER_PATH overrides).
  • Remove-background — u2net_human_seg (~168 MB ONNX) in background-removal/models/. Peak inference RAM ~1.5 GB.

Run npx hyperframes doctor if a command fails because of a missing dependency.

Source: SKILL.md on GitHub

No alerts2d3 checks · Risk SAFE
  • Gen Agent Trust Hub2d

    The skill is a comprehensive media management system for the HyperFrames platform, authored by heygen-com. It provides functionality to resolve, generate, and process audio, images, icons, and video. It integrates with reputable AI providers including HeyGen, Google Gemini, OpenAI, and ElevenLabs. The skill follows secure practices such as shell-less command execution to prevent injection, uses standard environment-based secret management, and includes a documented telemetry system for usage tracking with built-in opt-out mechanisms.

  • Socket2d

    No alerts

  • Snyk2d

    Risk: LOW · No issues

Signed by skilld at ff6e210. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 10 hours ago.

Activeupdated 4 days ago

README badge

README badge for heygen-com/hyperframes/media-use