All skills
hoodini avatar

/director

@6c79770

Award-winning director's brain for making films/videos with AI. Runs a gated, interrogation-first process: lock the IDEA and the SCRIPT (the story) in words before generating a single pixel, then plan shots/camera/light/sound, then generate and edit. Use whenever creating a video, film, teaser, trailer, promo, ad, short, reel, documentary, or narrative piece AND the storytelling must actually work — or any time the user says the story/script keeps failing. Hebrew triggers: סרטון, טיזר, פרומו, פרסומת, ריל, סרט, תסריט, סטוריבורד. Handles development, story, and shot-planning; composes with cinematic-ai-video and yuv-fomo-teaser (style/manipulation), hyperframes (render), nano-banana-2/ElevenLabs (assets). Backed by a 28-chapter Director's Bible in references/.

Use this Skill: https://skilld.dev/gh/hoodini/ai-agents-skills/director

This session only. Nothing lands on disk.

references23-dialogue-subtext-voice.md

≈6.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Dialogue, Subtext & Voice

This is line-level craft — the actual words that come out of mouths (or VO) and crawl across the screen. The chapters before this built the bones: structure (01-story-structure.md), scene-as-value-change (03-character-and-scene-craft.md), the hook engine (04-engagement-psychology-hooks.md). This chapter is the muscle and skin. It is also where AI-generated work most reliably sounds fake — not because the model can't form sentences, but because it writes people saying what they mean, which real people almost never do. Fixing that is the single highest-leverage line-level skill you can learn.

The thesis of the whole chapter, in one sentence: dialogue is not communication, it is combat conducted by other means. Hold that, and the rest follows.


1. Dialogue Is Action (the core reframe)

The amateur model of dialogue is information transfer: Character A knows a thing, says the thing, Character B now knows the thing. That is how email works. It is not how drama works.

Robert McKee's book Dialogue puts it precisely: tradition defines dialogue as talk between characters, but "all talk responds to a need, engages a purpose, and performs an action. To say something is to do something." Hence the book's subtitle — The Art of Verbal Action. Beneath every line, the writer creates a desire, an intent, and an action; the action expressed through words is what we call a line of dialogue.

The intuition: when you say "Is it cold in here?" you are rarely reporting a temperature. You are getting someone to close the window, or testing whether they care about your comfort, or stalling, or flirting. The literal content is a vehicle. The thing underneath — what you are trying to do to the other person — is the line's real payload.

This is why the practical question for every line is not "what does the character say?" but "what is the character trying to get, and who are they getting it from?" A line is a tactic. Characters switch tactics when one stops working — they charge, then they wheedle, then they threaten, then they go quiet. A scene is a sequence of escalating tactics in pursuit of a goal (this connects directly to scene-as-value-shift in 03-character-and-scene-craft.md: the tactics are how the value flips).

The "talking vs. doing-with-words" test

Take any line and ask: could a character accomplish this by handing the other person a note? If yes — if the line only conveys facts — it is talking, not doing. Rewrite it so the line tries to change the other person's state: make them afraid, ashamed, hopeful, defensive, complicit.

→ SHORT-FORM: In a 15–90s teaser you have room for maybe 1–6 spoken lines. Every single one must be a power move. A short-form line that merely describes (the product, the feature, the setting) is dead weight you cannot afford. The single line of dialogue in a great ad is usually a dare, a confession, or a threat — never a caption read aloud.

→ AI APPLICATION: Prompt the model with the goal beneath the line, not the line: instead of "write dialogue where she tells him she's leaving," prompt "she wants him to beg her to stay but refuses to ask for it — write the scene." The model writes far better dialogue when given a tactic and an obstacle than when given a message to deliver.


2. Subtext — the Iceberg

Subtext is everything a line means that it does not say. The text is the spoken words; the subtext is the want, fear, and tactic underneath. In good drama the gap between the two is where all the heat lives.

Harold Pinter, in his 1962 essay/speech "Writing for the Theatre," gave the canonical statement: "The speech we hear is an indication of that which we don't hear… The speech we hear is an indication of that which we don't hear. It is a necessary avoidance, a violent, sly, anguished or mocking smoke screen which keeps the other in its true place." Speech, for Pinter, is the cover over the real transaction — which is why the famous Pinter pause carries so much weight: the silence is where the unsaid thing presses hardest against the surface.

The practical heuristic: write the "no" that means "yes."

On the nose (text = subtext) With subtext (text ≠ subtext)
"I'm still in love with you." "You left your umbrella. …It's by the door."
"I'm terrified of dying alone." "You should get a dog. People with dogs live longer."
"I resent that you got the promotion." "Big office. They give you the plant, or did you buy that yourself?"

In every right-hand example the audience does the work of decoding — and that act of decoding is the engagement. The viewer leans in, feels clever, becomes a participant. This is the dopamine-as-prediction loop from 05-neuroscience-honest.md operating at the level of a single line: a tiny gap, a tiny resolution, a tiny reward.

On-the-nose dialogue — the cardinal sin

"On the nose" means the line states its own subtext, leaving nothing to decode. It is the most common dialogue failure and the loudest AI tell. The cure is not cleverness; it is obliqueness: characters approach the real subject sideways, talk about the wrong thing on purpose, or fight about the dishes when they're really fighting about the marriage.

The Mamet / Pinter school is the purest expression of this discipline. David Mamet, in his (in)famous memo to the writers of The Unit, hammered the rule in all caps: any time a character says "AS YOU KNOW" — i.e. tells another character a fact the writer needs the audience to know — "THE SCENE IS A CROCK OF SHIT." Mamet's positive formulation was three questions every scene must answer: "WHO WANTS WHAT? WHAT HAPPENS IF HE DON'T GET IT? WHY NOW?" — the verbal-action principle compressed to a memo. (Mamet writes clipped, overlapping, evasive lines; Pinter writes menacing silences and non-sequiturs. Both refuse to let people say what they mean.)

→ SHORT-FORM: Subtext is expensive in time — it needs the viewer to know enough to decode. In a 30s spot you usually cannot build the context for a fully oblique line. The short-form compromise: let the image carry the subtext while one spare line of VO plays against it. Visual of a man eating dinner alone, perfectly composed; VO: "He finally has the house to himself." The irony (loneliness vs. "finally") is the subtext, delivered in one beat.

→ AI APPLICATION: This is the prompt that most improves AI dialogue: "Write the subtext, not the statement. No character may say what they actually want. Show it through what they choose to talk about instead." Then adversarially review: scan the output for any line where a character announces their own emotion ("I'm so angry," "I've always loved you," "I'm scared") and force a rewrite. Over-explaining is the model's default failure mode; you must counter-instruct against it explicitly.


3. Character Voice / Idiolect

Idiolect is an individual's unique way of using language — vocabulary, rhythm, syntax, fillers, what they won't say. The gold standard, attributed to many screenwriting teachers and worth treating as a hard test: cover the character names and you should still know who is speaking. If every character sounds like the writer, you have one character wearing different hats.

Voice is built from a small number of dials:

  • Vocabulary / register — does she say "intoxicated," "drunk," or "hammered"? A cop, a priest, and a teenager describe the same corpse in three different lexicons.
  • Sentence length & rhythm — short and clipped (Mamet's tough guys) vs. long and qualified (an anxious over-explainer). Rhythm is audible; it's the fastest voice cue.
  • What they avoid — the most distinctive voices are defined by omission. A man who never says "I love you" but says "drive safe." A character who deflects every serious question with a joke. The negative space of speech is character.
  • Tics and verbal habits — a catchphrase, a repeated hedge ("I mean…"), grammatical quirks. Use sparingly; one or two per character, or it becomes a cartoon.
  • Pronoun/topic gravity — what they keep steering the conversation toward (always money? always status? always the past?).

The deeper principle: voice is the surface of psychology. How a person talks encodes how they think and what they fear. The over-explainer is anxious about being misunderstood; the man of few words is protecting something. Build the voice from the wound and it will be consistent without a glossary.

→ SHORT-FORM: With one or two lines you can't establish a full idiolect — so pick one dominant dial and commit hard. A single line in heavy dialect, or one piece of jargon, or one telling omission, does more characterization in a teaser than a paragraph would. The casting of the voice (timbre, accent, age) does the rest; see 18-ai-audio-vo-music-sfx-2026.md.

→ AI APPLICATION: LLMs collapse toward a generic "smart, balanced, well-spoken" register — every character sounds like a helpful assistant. Counter it with a per-character voice card: 4–6 bullets (vocabulary level, sentence length, one tic, one thing they never say, what they steer toward). Pass the card with every generation. Then run the cover-the-names test on the output and regenerate any line that's swappable between speakers.


4. Exposition Without Dumping

Exposition is the necessary background — who these people are, what's at stake, the rules of the world. The audience needs it; the trick is delivering it without the characters obviously narrating it for the camera's benefit.

The failure has names. "As you know, Bob" dialogue — characters telling each other things they both already know — is also called "maid and butler dialogue" (the term attributed to Algis Budrys) and "Rod and Don dialogue" (attributed to Damon Knight), both catalogued in the Turkey City Lexicon, the science-fiction workshop glossary. The image: two servants discussing their employer's affairs solely so the audience overhears the plot.

Fixes, roughly in order of power:

  1. Make the listener not know — and not want to be told. Exposition lands cleanly when one character needs the information and the other is withholding it. Now the fact-delivery is also a power struggle (verbal action again). The classic device is the newcomer/ingénue: someone genuinely ignorant whose questions are real.
  2. Exposition under conflict (the "argument" trick). Bury facts inside a fight. When two people hurl backstory at each other as ammunition — "You've been late with rent every month since your sister moved in" — the audience absorbs the facts while watching the duel. McKee and others call this delivering exposition "on a need-to-know basis," as ammunition rather than information.
  3. Show, then don't tell. Dramatize the fact instead of stating it. We don't need "I'm a recovering alcoholic"; we need him hesitating at the open bar, then asking for soda water. (Mamet's extreme version of this advice in the memo: pretend the characters can't speak and you'd be forced to tell the story in pictures — "great drama.")
  4. Reveal exposition late and as a turn. Information withheld and dropped at the worst possible moment becomes a reversal (Aristotle's peripeteia) instead of a briefing. The fact does double duty: it informs and it detonates.

→ SHORT-FORM: Short content has almost no room for exposition — and this is a feature, not a bug. The discipline is start in the middle (in medias res) and let the audience infer the world from a single charged image plus, at most, one line. A teaser that spends its opening seconds explaining is already lost. If a fact isn't load-bearing in the next five seconds, cut it.

→ AI APPLICATION: Models love to front-load exposition — they'll open a scene with a character reciting the premise. Instruct: "Reveal nothing the audience can infer from action. Deliver any necessary fact as a weapon inside a conflict, never as a statement." When generating a VO script, cap the setup at one sentence and force the rest to emerge from images.


5. The Four Functions (every line earns its place)

A line of dialogue should do at least one of these — ideally two or three at once. Lines that do none are cuttable:

  1. Reveal character — show who's speaking through how they say it (idiolect, §3).
  2. Advance the plot — change the situation; move the want closer or further.
  3. Create or escalate conflict — apply pressure, raise a stake, force a choice.
  4. Set tone — establish genre, mood, the rules of how funny/menacing/tender this world is.

The mark of a great line is density: it does three of the four simultaneously. "We're gonna need a bigger boat" (Jaws) reveals character (deadpan understatement under terror), advances plot (the threat just got bigger), creates conflict (they're outmatched), and sets tone (gallows humor) — all four, in five words. That density is the target.

→ SHORT-FORM: In a teaser the tone-setting function often outranks the others — the line's job is to tell you what kind of ride this is. A single line can be the entire genre signal.

→ AI APPLICATION: Build the four functions into a review pass: for each generated line, label which functions it serves. Cut or merge any line scoring zero. This turns a vague "make it tighter" into a mechanical filter the model can apply.


6. The Line-Level Toolkit

Craft lives in these moves:

  • Economy. Cut every word the line can survive without. "I think that maybe we should probably consider leaving soon" → "We should go." Then ask whether the line itself is needed. The most powerful version of economy is cutting the line entirely and letting an action or a look carry it.
  • The turn within a line. A line can reverse on itself: "I love you — which is why I can't be here." The pivot ("which is why") does the dramatic work inside one breath. Land the charged word last — the end of the line is the position of emphasis. "You, of all people, I trusted" hits harder than "I trusted you, of all people."
  • Interruption & overlap. People talk over each other when stakes rise. Cutting a character off mid-sentence ("—") shows dominance, urgency, or contempt without a stage direction. Overlapping dialogue (Altman, The Social Network's opening) signals speed and intelligence and forces the viewer to lean in.
  • Silence / the unsaid. The most powerful response is often no line. A beat, a pause, a refusal to answer. The Pinter pause (§2) weaponizes silence. On screen, a held reaction shot often beats any retort.
  • The button line. The sharp last line that ends a scene — a punchline, a reversal, a door slammed in words. The button gives the scene a clean exit and (in episodic/short work) a launchpad into the next beat. Write toward the button: know how the scene snaps shut.

→ SHORT-FORM: Short-form is all buttons. A 20-second piece is essentially a hook line and a button line with a visual in between. The last spoken/on-screen line is the one that gets quoted, screenshotted, and shared — engineer it like a punchline (see §8).

→ AI APPLICATION: These tools map to explicit instructions: "End the scene on a sharp button line." "Use an interruption (em-dash cutoff) when the stake rises." "Replace one line of dialogue with a beat of silence and a described action." Specifying the device gets better results than asking for "better dialogue."


7. The Test: Reading Aloud / The Table Read

Dialogue is written for the ear, not the eye. The non-negotiable test is to read it aloud — better, have it read by other voices (the table read). What looks fine silently betrays itself when spoken: tongue-twisters, lines no human breath can sustain, every character sharing the writer's cadence, jokes that don't land, exposition that clunks.

What you're listening for:

  • Breathability — can a person say this on one breath, with natural stress? If you run out of air, so will the actor/VO.
  • Distinct voices — close your eyes; can you tell the speakers apart? (The §3 cover-the-names test, performed live.)
  • Dead air vs. live silence — pauses that feel earned vs. pauses that feel like the scene died.
  • The clunk — any line where you wince. Trust the wince.

→ AI APPLICATION: This is directly automatable now. Pipe the draft through TTS (ElevenLabs — see 18-ai-audio-vo-music-sfx-2026.md) and listen. The synthetic read exposes unspeakable lines and rhythm problems instantly, and doubles as a delivery test for the actual VO. Make "generate → TTS → listen → revise" a loop, not a one-shot.


8. Writing for the Ear: VO, Narration & Hook Copy

Short-form and AI film lean heavily on voice-over and narration — often there's no on-screen dialogue at all. This is its own craft, and it has its own traps.

Voice-over: power vs. crutch

VO is power when it does something the image can't: irony (saying one thing while we see another), interiority (a thought no camera can shoot), compression (collapsing time — "Three years later, I'd lost everything"), or a distinct narrating voice as character (Goodfellas, Fight Club). VO is a crutch when it merely narrates what we already see ("She walked into the room, looking sad" over a shot of a sad woman walking in). The rule: VO should counterpoint or transcend the image, never duplicate it. If the VO and the picture say the same thing, delete one.

The documentary-narrator cadence

The trusted-explainer voice (think nature docs, the trailer voice, the explainer) has a recognizable rhythm: declarative, present or simple past, short clauses, a measured pace with deliberate pauses before the payoff word. "Every year… they return. Not because they choose to. Because they must." Note the fragments — written for the ear, fragments are fine; they create the pauses where breath and gravity live. Spoken rhythm is not written grammar.

Writing for spoken rhythm

  • Short sentences. Fragments allowed. The ear can't parse a 40-word subordinate-clause sentence; the eye can.
  • Front-load or back-load the key word — the ear remembers beginnings and ends (the serial-position effect; cf. peak-end in 05-neuroscience-honest.md).
  • Read it on a breath (§7). Punctuation is breath notation for VO: a period is a stop, a comma a catch, an ellipsis a held beat.
  • Avoid homophones and garden-path lines — "their there" confusion vanishes on the page but garbles in the ear.

The second person

"You" is the most intimate and most coercive pronoun — it implicates the viewer directly ("You think you're safe. You're not."). It's the engine of a huge fraction of short-form hooks and ads because it dissolves the fourth wall and makes the story about the viewer. Use it to accuse, dare, or promise. Overuse flattens into ad-speak; one well-placed "you" is a scalpel.

Hook copywriting — the first line that stops the scroll

The opening spoken/on-screen line in short-form has one job: stop the scroll in under 2 seconds. It is a headline, not a sentence. Workhorse patterns (use, don't worship):

Pattern Shape Example
The contradiction States something that shouldn't be true "I deleted my best-performing video on purpose."
The open loop / curiosity gap Promises a payoff it withholds "Nobody tells you the third thing."
Direct address (2nd person) Implicates the viewer "You're holding your phone wrong."
The stakes/cost Names a loss "This mistake cost me $40,000."
The bold claim Asserts something disputable "Most productivity advice makes you slower."

The curiosity gap (§ from 04-engagement-psychology-hooks.md) is the underlying mechanism: open a question the brain needs closed, then withhold. And every short piece needs a CTA line — the closing instruction (follow, watch, click). The best CTAs are continuous with the story's tone, not a tonal cliff into salesman-voice ("…if you want the other three, they're in part two" reads as story; "LINK IN BIO SMASH THAT FOLLOW" reads as noise).

→ SHORT-FORM: This entire section is the short-form discipline. Most short content is 1–6 lines of VO or zero dialogue with captions. The whole game is: one killer hook line, a visual middle, one button/CTA line. Resist the gravity of feature theory — do not turn a 30s teaser into a talky micro-film with scenes of people exchanging dialogue. The compression isn't "shorter scenes," it's fewer words doing more: a single line of VO over a single charged image beats three lines of on-the-nose chat every time.

→ MUTED VIEWING / CAPTIONS: Assume the sound is off — the majority of feed viewing is muted. Your hook must work as on-screen text too. Captions aren't a transcript; they're typographic performance: break lines on the beat, reveal the key word last (kinetic type), keep 3–6 words per card. The spoken line and the caption can differ — the caption is the version optimized for the eye. (Caption rendering and karaoke timing: the HyperFrames/video-edit pipeline in 13-production-pipeline.md and 18-ai-audio-vo-music-sfx-2026.md.)

→ AI APPLICATION — VO tuned for ElevenLabs: Write the script for the synth voice you'll use. Punctuation drives delivery: ellipses for held beats, em-dashes for sharp cuts, line breaks for breaths, ALL-CAPS sparingly for emphasis, italics for stress. Spell tricky words phonetically if the model mispronounces them. Generate, listen (§7 loop), and tune the punctuation until the rhythm lands — you are conducting the voice with typography. Keep sentences short; synth voices, like human ones, fall apart on long subordinate clauses. And the recurring instruction across this whole chapter: stop the model from over-explaining — its native register narrates and states; you want VO that counterpoints, withholds, and trusts the image.


Quick Reference — the dialogue debugging checklist

Run this over any draft (human or AI):

  1. Verbal action — does each line try to get something, or just inform? (§1)
  2. Subtext — does anyone say what they actually mean? Fix the on-the-nose lines. (§2)
  3. The "as you know" scan — any character telling another a fact they both know? (§4)
  4. Cover the names — can you tell who's speaking? (§3)
  5. Four functions — does each line do ≥1 (ideally 2–3)? Cut the zeros. (§5)
  6. Button — does the scene end on a sharp last line? (§6)
  7. Read it aloud / TTS it — does it breathe? Does it clunk? (§7)
  8. VO ≠ image — does the narration duplicate what we already see? Delete one. (§8)

Sources

Source: SKILL.md on GitHub

No alerts3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    This skill acts as a comprehensive educational and procedural resource for AI filmmaking. It implements a structured 'Director's Bible' consisting of 28 chapters of professional craft knowledge and a gated 'Grilling Workflow' that guides the user through story development, shot planning, and production. No malicious patterns, obfuscation, or data exfiltration attempts were detected.

  • Socket3mo

    No alerts

  • Snyk3mo

    Risk: LOW · No issues

Signed by skilld at 6c79770. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Steadyupdated 4 months ago

README badge

README badge for hoodini/ai-agents-skills/director