All skills
oaustegard avatar

/optimizing-skills

@59f45d6

Disciplined, validation-gated revision of an EXISTING skill so each edit is a measured improvement rather than a guess. Use when editing, revising, or tuning a skill that already exists and there is evidence it underperforms (observed failures, drift, complaints) — invoke by name, or have versioning-skills / creating-skill defer to it before applying edits. Not for authoring a brand-new skill from scratch (use creating-skill) or one-off prose.

Use this Skill: https://skilld.dev/gh/oaustegard/claude-skills/optimizing-skills

This session only. Nothing lands on disk.

referencesskillopt-provenance.md

≈812 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Provenance & the Agent-tool recipe

This skill is distilled from SkillOpt (microsoft/SkillOpt, arXiv:2605.23904, MIT). SkillOpt trains a skill document as the external state of a frozen agent: a separate optimizer LLM turns scored rollouts into bounded add/delete/replace edits, accepted only when they strictly improve a held-out score. We keep the discipline and drop the training harness — no Azure, no rollout infra, no weight-training analogy required.

Mapping: SkillOpt mechanism → this skill

SkillOpt Here
Validation gate (gate.py, strict >, ties rejected) "The gate" — accept only if candidate strictly beats best
Two-tier current/best skill state candidate / best
Textual learning rate = integer cap on edits/step, cosine-annealed "Bounded edits", ~4 default, fewer as it matures
Failure-minibatch + success-minibatch reflection, failures win merge "Reflect: failures first, then successes"
Rank-and-select top-L by impact (clip.py, ranking.md) "Rank when over budget"
Protected slow-update region (<!-- SLOW_UPDATE_START/END -->) "Protect the hard-won core" + longitudinal review
Optimizer-side meta-skill (meta_skill.md) "Carry memory across revisions" via remember()/recall()
Literal str.replace(target, content, 1) edits "Edit mechanics" — exact-match Edit tool ops

Recipe: dispatch scoring/reflection to the Agent tool

The one piece that needs compute is scoring the check set. There is no ANTHROPIC_API_KEY in this environment — route through the Agent tool, not a standalone SDK call.

  1. Score a version. For each check task, spawn subagent_type=general-purpose (model sonnet for cheap, opus for hard checks) with the skill version pasted into the prompt and the task. Collect hard pass/fail. Run best and candidate the same way for a fair comparison. Parallelize independent tasks in one message.
  2. Reflect (optional, for large failure sets). Hand the failing transcripts to an optimizer subagent using the prompt below; it returns bounded edits. You still apply them with the Edit tool and re-score through the gate.

Adapted failure-reflection prompt (optimizer subagent)

You are a failure-analysis agent for an existing skill document.
Given the current skill and MULTIPLE failed task transcripts, identify the
most important COMMON failure pattern across them (not one-off edge cases) and
propose AT MOST <budget> generalizable edits. Do not hardcode task-specific
values. Do not duplicate content already in the skill — patch gaps only.
Return JSON: {"failure_summary": [...], "edits": [{"op": "append|insert_after|
replace|delete", "target": "<exact text, if needed>", "content": "<markdown>"}]}

Adapted ranking prompt (when edits exceed budget)

Rank the proposed edits and select the top <budget>, by: (1) systematic impact
on recurring failures, (2) complementarity / fills a gap, (3) generality as a
principle, (4) actionability. Return {"selected_indices": [...]} in priority order.

Keep the optimizer subagent separate from the target subagent — the model that proposes edits should not be the one being measured by them.

Source: SKILL.md on GitHub

No alerts3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    The skill provides a disciplined framework for optimizing AI agent skills through evidence-based revision and validation. It leverages platform-provided tools for subagent execution and persistent memory. A minor security risk exists in the form of indirect prompt injection, as the skill processes untrusted test tasks and transcripts without explicit boundary markers or sanitization when passing them to subagents for evaluation and reflection.

  • Socket3mo

    No alerts

  • Snyk3mo

    Risk: LOW · No issues

Signed by skilld at 59f45d6. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 weeks ago
metadata
{
  "version": "0.3.0"
}

README badge

README badge for oaustegard/claude-skills/optimizing-skills