All skills
aktsmm avatar

/context-to-video

@848fd9b
by yamapanaktsmm/agent-skills26 stars
4

Turn any context (blog URL, pasted article, PR diff, meeting notes, release notes, raw prompt) into a narrated explainer mp4 with slides, subtitles, and optionally a talking-head avatar. Uses a free local stack (Pillow + edge-tts + ffmpeg) with optional SadTalker for the avatar. Use when the user says "動画化", "explainer video", "解説動画作って", "ブログを動画に", "記事を動画化", "PRを動画で説明", "議事録を動画化", "context to video", or wants an mp4 from arbitrary text/source. Output is mp4 + srt under the user's chosen project workspace.

Use this Skill: https://skilld.dev/gh/aktsmm/agent-skills/context-to-video

This session only. Nothing lands on disk.

referencesproduction-upgrade.md

≈673 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Production Upgrade

無料 edge-tts の代替で顧客配布・常時運用に耐える構成に上げる手順。

なぜ差し替えるか

項目 edge-tts Azure Speech 正規
認証 なし (匿名) サブスクキー or Entra ID
SLA なし 99.9%
商用利用 グレー 明示OK
安定性 エンドポイント変更で突然停止リスク 公式
機能 Neural voices + HD voices / Custom Neural Voice / TTS Avatar

Azure Speech (TTS) に差し替え

pip install azure-cognitiveservices-speech
import azure.cognitiveservices.speech as speechsdk

def tts_azure(text: str, out_path: str, voice: str = "ja-JP-NanamiNeural"):
    cfg = speechsdk.SpeechConfig(subscription=KEY, region=REGION)
    cfg.speech_synthesis_voice_name = voice
    cfg.set_speech_synthesis_output_format(
        speechsdk.SpeechSynthesisOutputFormat.Audio48Khz192KBitRateMonoMp3
    )
    audio_out = speechsdk.audio.AudioOutputConfig(filename=out_path)
    synth = speechsdk.SpeechSynthesizer(speech_config=cfg, audio_config=audio_out)
    synth.speak_text_async(text).get()

scripts/build_video.py の tts() をこれに差し替えるだけ。声名は edge-tts と同じものが多くそのまま使える。

参考: https://learn.microsoft.com/azure/ai-services/speech-service/

Azure TTS Avatar (アバター動画)

スライド+アバター解説をやるなら Azure Speech の Text to speech avatar API。

  • バッチ合成 → mp4 が返ってくる
  • 標準アバター (Lisa 他) は申請不要、Custom Avatar は審査あり
  • 既存スライド mp4 と ffmpeg overlay で右下ワイプ合成すれば「スライド+喋るLisa」になる
ffmpeg -i slides_video.mp4 -i avatar.mp4 -filter_complex \
  "[1:v]scale=480:-1[av];[0:v][av]overlay=W-w-40:H-h-40" \
  -map 0:a -c:a copy out.mp4

公式: https://learn.microsoft.com/azure/ai-services/speech-service/text-to-speech-avatar/what-is-text-to-speech-avatar

Entra ID 認証 (会社サブで disableLocalAuth=true の場合)

会社サブの Azure AI Services は API key 無効化ポリシーが当たっていることがある。その場合は DefaultAzureCredential で Bearer token を取得。

from azure.identity import DefaultAzureCredential
cred = DefaultAzureCredential()
token = cred.get_token("https://cognitiveservices.azure.com/.default").token
cfg = speechsdk.SpeechConfig(auth_token=f"aad#{RESOURCE_ID}#{token}", region=REGION)

RESOURCE_ID は Speech リソースの完全な ARM ID (/subscriptions/.../providers/Microsoft.CognitiveServices/accounts/<name>)。

Source: SKILL.md on GitHub

1 warning1mo3 checks · Risk SAFE
  • Gen Agent Trust Hub1mo

    The skill is functionally safe but possesses a surface for indirect prompt injection because it processes external content (URLs, PR diffs, and notes) using media generation tools without explicit content sanitization.

  • Socket1mo

    No alerts

  • Snyk1mo

    Risk: MEDIUM · 2 issues

Signed by skilld at 848fd9b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 18 hours ago.

Activeupdated 3 months ago
argument-hint
入力(URL/テキスト/PR/メモ)、言語(JP/EN)、長さ目安、アバター要否、出力先パス
user-invocable
true
metadata
{
  "author": "yamapan"
}

README badge

README badge for aktsmm/agent-skills/context-to-video