All skills
heygen-com avatar

/media-use

@ff6e210
by HeyGenheygen-com/hyperframes55k stars
4,998

Agent Media OS, the single skill for every media need in a HyperFrames project. Resolve BGM, SFX, image, icon, brand logo, voice, color grade, or LUT into a frozen local file or paste-ready block + ledger record (one verb, `resolve`); generate via TTS / music / image models when the catalog misses; produce voiceover, transcription, captions, and background removal through one shared audio engine; operate on media (cut / reframe / transform); and reuse assets across projects. Also use for vague feedback that real footage looks dark, flat, boring, should feel retro/camcorder/print/ASCII, needs privacy, or needs a media reveal.

Use this Skill: https://skilld.dev/gh/heygen-com/hyperframes/media-use

This session only. Nothing lands on disk.

audioreferencestts-to-captions.md

≈439 tokens on demand. Your agent reads this file only when SKILL.md points to it.

TTS → Captions

When no recorded voiceover exists, generate one and obtain word-level caption timing. Two paths depending on which TTS provider is in use:

Path A — HeyGen (single call, no Whisper)

HeyGen returns word timestamps in the same response as the audio. Use the bundled REST helper (the hyperframes tts command is Kokoro-only):

node skills/media-use/audio/scripts/heygen-tts.mjs \
  script.txt --output narration.wav --words narration.words.json

narration.words.json is already in the [{ id, text, start, end }] shape the captions pipeline consumes — no separate transcribe pass.

Path B — Gemini / ElevenLabs / Kokoro (TTS → transcription)

These adapters supply audio without word data. The shared audio engine runs transcription automatically when timings are absent. For Gemini, use the request in Text to speech, then consume audio_meta.json → voices[].words.

For a standalone local Kokoro generation, generate the audio, then transcribe:

npx hyperframes tts script.txt --voice af_heart --output narration.wav
npx hyperframes transcribe narration.wav --model small.en   # voice af_heart is American English

Whisper extracts precise word boundaries from the generated audio, so caption timing matches delivery without hand-tuning. Match --model to the voice's language (use small.en for a/b prefixes, small --language <code> otherwise). Then consume transcript.json via the caption references in captions/.

For Gemini, verify that transcription preserved the script, especially names, numbers, and delivery pauses. If words is empty, resolve the transcription failure before captioning. Generate and align again after changing the read.

Source: SKILL.md on GitHub

No alerts2d3 checks · Risk SAFE
  • Gen Agent Trust Hub2d

    The skill is a comprehensive media management system for the HyperFrames platform, authored by heygen-com. It provides functionality to resolve, generate, and process audio, images, icons, and video. It integrates with reputable AI providers including HeyGen, Google Gemini, OpenAI, and ElevenLabs. The skill follows secure practices such as shell-less command execution to prevent injection, uses standard environment-based secret management, and includes a documented telemetry system for usage tracking with built-in opt-out mechanisms.

  • Socket2d

    No alerts

  • Snyk2d

    Risk: LOW · No issues

Signed by skilld at ff6e210. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 10 hours ago.

Activeupdated 4 days ago

README badge

README badge for heygen-com/hyperframes/media-use