All skills
simota avatar

/sketch

@16a3f28
by shingo imotasimota/agent-skills85 stars
15

Generating AI image-generation code using the Gemini API. Handles text-to-image generation, image editing, and prompt optimization. Use when image generation code is needed.

Use this Skill: https://skilld.dev/gh/simota/agent-skills/sketch

This session only. Nothing lands on disk.

referencebatch-generation.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Batch Generation Reference

Purpose: Produce many image variants in one run with consistent seed, style, and output layout. Invoke when the brief asks for ≥5 assets that must share a visual identity (card faces, hero background set, character sheet poses) and the user cares more about cross-asset cohesion than per-image novelty.

Scope Boundary

  • sketch batch: batch Python script, seed stability plan, parallel API calls with rate limits, output naming, perceptual-hash dedup, metadata.json per asset.
  • sketch style (sibling): defines the style anchor (reference image / token set) consumed by batch. Call style first if the style is not yet locked.
  • Upstream caller (elsewhere): asset brief, count, resolution, and naming spec — batch does not invent the list.
  • Vitrine (elsewhere): catalog the produced set after batch completes.

If the brief is "one hero image, try three variants" → default generate. If it is "120 card faces, consistent frame and lighting" → batch.

Workflow

INTAKE      →  read brief: N assets, naming, ratio, style anchor, deadline
            →  lock seed strategy (fixed / stride / family) — see table below
            →  confirm Batch API eligibility (N ≥ 50 → 50% discount, 24h)

CONFIGURE   →  pick model (default gemini-3.1-flash-image)
            →  set seed, thinking_level, person_generation
            →  define output layout: out/<slug>/<NNN>_<variant>.png + metadata.json
            →  set concurrency ≤ rate-limit ceiling (see "Rate Limits")

CODE        →  emit script: async queue, retry w/ exponential backoff,
               per-asset metadata, pHash dedup, resumable checkpoint
            →  dry-run 1–3 previews before fanning out

VERIFY      →  confirm preview cohesion, estimated cost, resumability,
               `.env` + `.gitignore` guidance, SynthID disclosure

Seed Stability Strategies

Strategy Seed formula Use when
Fixed seed = S for all calls Pure variant exploration on one prompt — small N
Stride seed = S + i per asset i Deterministic per-slot regeneration; card sets, icon sets
Family seed = hash(family_key) + i Multi-family batch (heroes × 4 factions) — families stay stable if one family is re-run
Free no seed Only when cohesion comes entirely from reference images, not seed

Default: Stride. Persist seed_base in metadata.json so slot 037 can be regenerated byte-identically.

Style Anchoring

Cross-asset cohesion comes from three layers — use at least two:

  1. Prompt skeleton: fixed Subject/Style/Composition/Technical template with only the Subject slot varying per asset.
  2. Reference images: up to 14 per request via inlineData (Base64). Keep each < 4MB. Include 2–4 anchor images that define palette, lighting, and frame.
  3. Seed stride: see above.
SKELETON = (
    "{subject}, {style_token}, centered composition, "
    "three-quarter view, soft rim light, 50mm lens, "
    "matte finish, no text, no watermark"
)
# style_token is produced by `sketch style` and pinned here.

Parallel API Calls and Rate Limits

Respect per-model rate limits; do not fan out unbounded.

Model Role Conservative starting concurrency
gemini-3.1-flash-image default generalist 2
gemini-3-pro-image premium / complex assets 1

Quotas vary by project and tier. Read the live quota response and the current official rate-limit page before increasing concurrency; do not encode a static RPM table as a durable contract.

Pattern: asyncio.Semaphore(concurrency) + retry with jitter on 429/503. Checkpoint every asset to progress.jsonl so a crash resumes from the last completed slot.

For N ≥ 50, prefer Batch API (50% discount, 24h SLA) over live fan-out unless the deadline is under a day.

Output Naming Convention

out/<brief-slug>/
  000_<variant-slug>.png
  000_<variant-slug>.metadata.json
  001_<variant-slug>.png
  ...
  _index.json         # brief, seed_base, model, prompt skeleton, pHash list
  _progress.jsonl     # resumable checkpoint

Zero-pad the index to max(3, ceil(log10(N))). Slugs must be lowercase, kebab-case, ASCII. Never overwrite — collisions append -retry<n>.

Perceptual Hash Dedup

Near-duplicate detection prevents silent seed collisions and content-filter retries that produce twins.

from PIL import Image
import imagehash

def phash(path: str) -> str:
    return str(imagehash.phash(Image.open(path), hash_size=16))

# Two images with hamming distance ≤ 6 are considered near-duplicates.

Flag — do not auto-delete. Duplicates get *.dup.json sidecar; the operator decides whether to regenerate with a new seed stride.

Anti-Patterns

  • Issue N calls synchronously with no semaphore — hits 429 in seconds on free tier and exhausts quota for the day.
  • Reuse a single fixed seed for 100 assets without reference images — variants become visually identical and dedup flags everything.
  • Skip checkpointing — a crash at asset 87/100 forces a full re-run and doubles cost.
  • Inline API key as a script argument for convenience — keys leak into shell history and CI logs.
  • Fan out 4K renders in parallel on free tier — per-image latency is ~60s; the wall clock and quota cost both collapse the run.
  • Drop metadata.json per asset — cannot regenerate slot 37 later without re-deriving seed and prompt.
  • Run batch before style has locked the anchor — early assets diverge and the later ones cannot reconcile.

Handoff

  • To Vitrine: output directory + _index.json (brief, model, seed_base, pHash list). Vitrine builds the catalog.
  • To Muse: dominant-color extraction input if the batch feeds the design system.
  • To Growth: promo-ready subset filtered by aspect ratio + safe margins.
  • To Oracle: only if a content-policy edge case tripped — forward the failing prompt and block reason, never the raw output.

Source: SKILL.md on GitHub

2 warnings4mo5 checks · Risk SAFE
  • Gen Agent Trust Hub4mo

    This skill provides a secure and well-architected framework for generating Python code to interact with the Google Gemini image generation API. It adheres to security best practices by emphasizing environment-variable-based credential management, providing explicit .gitignore guidance, and incorporating comprehensive error handling for API and safety filter responses. The skill acts exclusively as a code generator, ensuring that no actual API calls or external network requests are executed within the agent context itself.

  • Socket4mo

    No alerts

  • Snyk4mo

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    2/4 files flagged

  • ZeroLeaks5mo

    1 finding · Score: 86/100

Signed by skilld at 16a3f28. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated last month

README badge

README badge for simota/agent-skills/sketch