All skills
jwynia avatar

/document-to-narration

@6ac6c30
by J Wyniajwynia/agent-skills160 stars
20

Convert written documents to narrated video scripts with TTS audio and word-level timing. Use when preparing essays, blog posts, or articles for video narration. Outputs scene files, audio, and VTT with precise word timestamps. Keywords: narration, voiceover, TTS, scenes, audio, timing, video script, spoken.

Use this Skill: https://skilld.dev/gh/jwynia/agent-skills/document-to-narration

This session only. Nothing lands on disk.

referencesspoken-adaptation-guide.md

≈908 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Spoken Adaptation Guide

Principle: Sound Natural When Read Aloud

The test for spoken adaptation: read the text aloud. Does it flow? Do you stumble? Would a listener understand on first hearing?

When to Adapt

Adapt text when:

  • Parenthetical asides break the flow
  • Sentences are too long or nested
  • Abbreviations would sound awkward spoken
  • Punctuation creates unnatural pauses
  • Academic hedging obscures the point

Do NOT adapt when:

  • The author's voice would be lost
  • Rhetorical devices are intentional
  • The original phrasing is memorable
  • Simplification would change meaning

Adaptation Categories

1. Structural Simplification

Written: "The tool — which, it should be noted, was available to both participants — produced dramatically different results."

Spoken: "The tool was available to both participants. But it produced dramatically different results."

2. Parenthetical Handling

Written: "They found the floor (where anyone can get results) but never discovered the ceiling."

Spoken Options:

  • "They found the floor — where anyone can get results — but never discovered the ceiling."
  • "They found the floor, where anyone can get results. But they never discovered the ceiling."

3. Academic Hedging

Written: "It could be argued that the distinction between..."

Spoken (if author's voice permits): "The distinction between..."

Written: "Research suggests that..."

Spoken: "Research shows that..." (if the research is definitive)

4. List Handling

Written: "The elements include: clarity, brevity, and impact."

Spoken: "The elements include clarity, brevity, and impact."

For longer lists: Written: "1. First item 2. Second item 3. Third item"

Spoken: "First: [item]. Second: [item]. And third: [item]."

5. Abbreviation Expansion

Written Spoken
e.g. for example
i.e. that is
etc. and so on
vs. versus
w/ with

6. Emphasis Preservation

Italics in writing often indicate emphasis, but TTS needs context cues.

Written: "The microwave doesn't have a 'more skill' setting."

Spoken: "The microwave doesn't have a 'more skill' setting." (Natural stress from sentence structure)

If emphasis is critical and might be lost: Spoken: "The microwave... it simply doesn't have a 'more skill' setting."

What NOT to Change

  • Author's distinctive phrasing
  • Rhetorical questions (these work great in audio)
  • Parallel structure ("Some... others...")
  • Deliberate fragment sentences
  • Characteristic word choices
  • The argument's logic or claims

Scene Transition Language

When a section break becomes a scene break, sometimes you need to add spoken transitions:

Context Transition Type Example
New major section Introduction "Now..." or "Here's where..."
Returning to main thread Callback "Back to our question:"
Contrast or pivot Pivot "But here's the thing:"
Building on previous Connection "And this connects to..."
Conclusion Summary "So what does this mean?"

Use sparingly. Only add transitions when the audio would feel jarring without them.

Quality Checklist

Before finalizing adapted text:

  • Read it aloud - does it flow naturally?
  • Check for tongue-twisters or awkward consonant clusters
  • Verify the author's voice is preserved
  • Confirm no meaning was changed
  • Ensure transitions feel organic, not mechanical
  • Verify abbreviations are expanded
  • Check sentence length (aim for 15-25 words average)

Source: SKILL.md on GitHub

1 warning16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    The document-to-narration skill is a utility for converting markdown into video scripts with synchronized audio and captions. It utilizes Deno, Python, and external tools like ffmpeg and Whisper. The skill is well-structured and uses safe practices for command execution. A low-risk surface for indirect prompt injection is identified because the skill processes user-supplied markdown documents to influence scene splitting and narration adaptation.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer7mo

    14/14 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 6ac6c30. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 months ago.

Dormantupdated 8 months ago
compatibility
Requires Deno, Python 3.12 with venv, ffmpeg, whisper-cpp
Other metadata
metadata
{
  "author": "jwynia",
  "version": "1.0",
  "domain": "video-production",
  "type": "generator",
  "mode": "generative"
}

README badge

README badge for jwynia/agent-skills/document-to-narration