All skills
hoodini avatar

/video-edit

@d87beb7

Edit any video into a captioned showcase — transcribe (any language, defaults to large-v3), present a transcript_review.txt for the user to fix mishears BEFORE rendering, then build a HyperFrames composition with liquid-glass caption pills, liquid blob background, liquid morph wipes, optional behind-subject text via background removal, and render the final video. Use whenever the user provides a video file and asks to edit it, caption it, add subtitles, fix existing captions, make a reel/promo/captioned tutorial, or "do the same" pattern as a prior captioned video. Supports English, Hebrew, and any Whisper-supported language. **Renders both 16:9 (YouTube / horizontal) and 9:16 (TikTok / Instagram Reels / YouTube Shorts) from the SAME 16:9 source** — vertical mode uses a centered footage strip with a blurred backdrop + liquid blobs and a vertical-tuned caption pill, no need to re-shoot. THE PIPELINE PAUSES FOR USER APPROVAL on the transcript before final render — this is the support mechanism for getting captions perfect (especially Hebrew). Pairs with hyperframes, hyperframes-cli, hyperframes-registry, and yuv-design-system skills.

Use this Skill: https://skilld.dev/gh/hoodini/ai-agents-skills/video-edit

This session only. Nothing lands on disk.

referencessetup.md

≈675 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Setup — Prerequisites for the video-edit skill

This skill orchestrates several external tools. A fresh machine needs the following one-time installs. The skill itself just calls these CLIs; nothing else is bundled.

Required

Tool Version Purpose
Node.js ≥ 22 npx hyperframes … (scaffold, lint, render, background-remove)
Python ≥ 3.10 Whisper transcription + caption generator + review scripts
FFmpeg any recent Audio extraction, frame extraction, footage re-encoding
faster-whisper latest Word-level transcription (Python package)

Install commands

Windows (PowerShell, with winget)

winget install OpenJS.NodeJS.LTS
winget install Python.Python.3.12
winget install Gyan.FFmpeg
pip install faster-whisper

macOS (with Homebrew)

brew install node@22 python@3.12 ffmpeg
pip3 install faster-whisper

Linux (Debian / Ubuntu)

sudo apt update
sudo apt install -y python3 python3-pip ffmpeg
curl -fsSL https://deb.nodesource.com/setup_22.x | sudo bash -
sudo apt install -y nodejs
pip3 install faster-whisper

Verify

node --version          # v22+ expected
python --version        # 3.10+ expected
ffmpeg -version         # any recent build
python -c "import faster_whisper; print(faster_whisper.__version__)"
npx hyperframes doctor  # checks Chrome / FFmpeg / memory for renders

Optional but recommended

GPU acceleration for Whisper

faster-whisper can run on CUDA, but on Windows it usually crashes mid-decode because cuDNN is not on the PATH. The bundled transcribe.py is hard-coded to CPU int8 for that reason — fast enough on a modern CPU (~5 min for a 3-minute clip with the large-v3 model). If you have cuDNN properly installed on Linux, swap device="cpu" for device="cuda" and compute_type="int8" for compute_type="float16".

GPU for background removal (npx hyperframes remove-background)

CoreML (macOS), CUDA (Linux with proper drivers) or DirectML (Windows) accelerate the u2net mask model. Without GPU it falls back to CPU — a 10-second 1440p clip takes ~3-8 minutes.

npx hyperframes remove-background --info  # lists available execution providers

Pre-cache Whisper model

The first transcribe downloads the large-v3 model (~3 GB) into ~/.cache/huggingface/hub/. Subsequent runs are instant to start.

Where files live

The skill expects to operate inside a HyperFrames project directory created by npx hyperframes init. Reference assets in this skill (references/) are copied into that project as part of step 4–6 in the main SKILL.md workflow.

Source: SKILL.md on GitHub

1 alert3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    The video-edit skill provides a legitimate and well-documented pipeline for automated video captioning and editing. It employs industry-standard tools like ffmpeg and faster-whisper. While the skill includes powerful features like a local web server for transcript review and remote script execution for installation, these are implemented using author-owned resources and restricted to the local loopback interface, aligning with the skill's primary purpose.

  • Socket3mo

    1 alert: gptAnomaly

  • Snyk3mo

    Risk: CRITICAL · 2 issues

Signed by skilld at d87beb7. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Steadyupdated 4 months ago

README badge

README badge for hoodini/ai-agents-skills/video-edit