All skills
hypit-ai avatar

/hypit

@04405de
by Hypit.AIhypit-ai/hypit18k stars
2,113

Direct and produce videos with Hypit, primarily through generated media and authored components while incorporating briefs, references or supplied material; includes SVML/SVS/SVRun authoring and Runtime or credential setup.

Use this Skill: https://skilld.dev/gh/hypit-ai/hypit/hypit

This session only. Nothing lands on disk.

referencesplaybooksformatstwo-person-podcast.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Two-person podcast

The relationship between two people carries the argument: a question invites a claim, a reaction changes its meaning, and a reply pays it off. Let that relationship motivate framing and performance.

Build the pair from a useful first view

Generate the main host in the actual setting first. Use that image to derive the other host's complementary view: opposite screen position and gaze, microphone entering from the other side, and another visible sector of the same location. The first host view establishes the shared world; the complementary view changes what the other camera can see.

These idle views hold an easy conversational posture: each host turns toward their partner, with their face readable from the chosen camera. Their attitude is already present and can develop through different speaking turns.

The worked restaurant example has this reference structure:

main host in restaurant ──> other host in restaurant
main host + product ──────> main host holding product
other host + product ─────> other host holding product
main host (+ logo) ───────> each lifestyle scene
both host views ──────────> optional split-screen opening image

Most new views are one hop from their useful parent. Derive lifestyle scenes from the main host, not from one another. Product-holding views inherit their own host's camera and add the actual product reference. Preserve gaze, the established setting and microphone placement while changing the intended prop state. File imports are valid inputs even when those images were generated outside Source.

Podcast and interview image examples preserve the concrete English production directions and explain what each reference is responsible for.

Direct conversation inside a Take

podcast-v1 takes two final camera views and two voices, in matching A/B order. Map the Script's actual Role names to those hosts explicitly in action direction. Script remains the sole dialogue text; a single Segment can contain several turns and speaker cuts.

The Kit's Recipe supplies recurring framing, editing, pace, performance, reaction and gesture choices. Action adds this passage's attention and interaction: who is making the claim, who is skeptical, when a reaction earns the other view, and who holds the product. Keep silent listeners physically alive through a posture adjustment, a brief glance down and back up, or an attentive expression while preserving their silence.

A handoff needs legible possession. The main host offers the product toward their partner; the partner receives it in the complementary view; later shots reflect who now has it. A few clear relations are more useful than prescribing every finger. Cuts can carry this exchange without requiring serialized frame-continuation generation.

For example, action accompanying a product-offer passage can say:

GIRL is Host A and GUY is Host B. She treats the recommendation with amused certainty; his reaction
is playful disbelief, not outrage. She offers the product toward the left edge of her view. In his
complementary view it arrives from the right; he receives it and keeps it. When returning to her,
the product is no longer in her hand. Keep the exchange casual, with compact gestures and no
extra dialogue.

This particular direction assumes those view and prop inputs. A phone link sent digitally calls for the corresponding screen or conversational interaction instead of a physical handoff.

An opening can establish both people at once

A top/bottom split screen can immediately reveal the pair before the first question finishes. The speaker can occupy one panel while the listener in the other looks down, adjusts their seat and looks back up. Those small silent actions establish a conversation already in progress.

Keep the supplied top-and-bottom composition. The man in the upper panel delivers the question to
his partner. In the lower panel, the woman briefly looks down, settles into her seat and looks back
toward him as she listens. She remains silent. Let both panels feel like the same conversation
already underway, without large gestures or a frozen listener.

The split reference and its matching video direction form one authored opening Take that can carry both people and a speaker change. The ordinary two-view Podcast Kit suits the later exchange. Use a split-screen opening when that attention device serves the work; use another component or Take shape for a different layout.

Coverage can follow a thought rather than every noun

The example generates three lifestyle scenes together as a short silent montage, then places that one clip over a Selection spanning the host's examples and the start of the partner's response. The viewer understands the broader habit even when each visual cut is not synchronized to its noun. Give the montage enough room to play through once at native speed, so its last scene is not cut off. The Selection may slightly outlast the clip; let the underlying picture return when playback ends. The partner's incoming voice can precede that return: a useful J-cut relationship. Avoid a frozen tail or retiming merely to fill the Selection. Use the B-roll craft's current playback guidance rather than treating an existing project's Recipe as a required choice.

Use separate clips when exact scene-to-cue correspondence matters, especially in reconstruction. Generate each within the selected model's actual range and consume only the needed section. Read B-roll for source duration, stretch, short windows and gap-free boundaries.

Host-colored Caption can clarify the speaker even over full-frame B-roll. Coordinate it with the scene palette and product treatment; choose the colors for this pair. Judge the whole exchange for responsive timing, restrained reactions, prop continuity and a payoff that feels earned.

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The hypit skill provides comprehensive documentation and automation logic for a video production toolchain. It manages environment setup using standard package managers (npm, brew, uv), handles credentials securely through native OS lockers, and executes production builds. Security findings are primarily related to the tool's intended automation capabilities, such as local script execution for web capture and the inherent risk of indirect prompt injection when processing external media links and website content.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 04405de. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 13 hours ago.

Activeupdated 2 weeks ago

README badge

README badge for hypit-ai/hypit