All skills
mblode avatar

/agent-skills-creator

@b42cf43
by Matthew Blodemblode/agent-skills136 stars
12

Creates and improves portable Agent Skills with a validator, routing scenarios, and evidence-based keep, cut, merge, or retire decisions. Use when asked to "write a skill", "update all skills", "audit my SKILL.md", "remove redundant instructions", or fix skill triggering. For AGENTS.md or CLAUDE.md use agents-md.

Use this Skill: https://skilld.dev/gh/mblode/agent-skills/agent-skills-creator

This session only. Nothing lands on disk.

referencescapability-delta.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Capability Delta

Use when simplifying a skill, responding to a model upgrade, or auditing a collection.

Retention test

A skill supplies something the task, tools, repository, and host do not already supply. Classify each section before expanding it:

Content Default action Evidence to preserve
Generic competence: reason carefully, write clear prose, inspect code, fix mistakes Delete None unless a specific regression justifies a targeted instruction
Host behavior: tool syntax, permissions, progress updates, memory, delegation Delete duplication; scope necessary adapters to the host Actual tool interface or host documentation
Public domain knowledge Cut tutorial prose; retain a compact rubric when an explicit audit needs repeatable coverage Rule applicability, detection method, false positives
Team taste or product policy Keep in one authoritative location User preference or repository convention
Operational contract or observed failure Keep the minimum reproducible guidance Command, schema, failure consequence, or regression case
Volatile facts: quotas, prices, SDK APIs Resolve from current official sources when used Source and date; no undated snapshot presented as live

A skill in a shared repository is read by several contributors' agents running different models, so it cannot be tuned to one of them. Step lists calibrated to the gaps of the weakest target overconstrain the strongest; a stated outcome and a stated scope of done are actionable at every tier and are what travels.

A model upgrade moves two things at once, in opposite directions, and an audit checks both. Guidance written to rein in an older model (step-by-step recipes, blanket confirmation, a stop for review after the first draft) now overconstrains and makes the newer model stop early or defer decisions it could take. At the same time the newer model is more tentative about how far to take a task, so the scope of done and the permission for known-safe loops need to be stated where they were previously left implicit. Cutting the first without adding the second leaves a skill that stops sooner than the one it replaced.

Strong model performance is a reason to revisit instructions, not proof that removing a particular contract preserves behavior. Do not infer what was in a vendor's post-training from an announcement. Label a static deletion judgement as such; reserve measured claims for actual runs.

Keep, cut, merge, retire

Keep a skill with a distinct trigger only when with-skill evaluations show a worthwhile gain over without-skill on the target job, accounting for tokens and time. Install counts can guide discovery, but aggregate counters do not establish unique users or independent choices and cannot pass this keep test. Without comparative runs, label retention provisional rather than measured.

Cut generic explanations inside the skill. Merge when the remaining payload shares an existing skill's trigger and output contract. Retire when nothing unique remains; record the replacement or native capability and remove routing pointers, README entries, and obsolete fixtures together.

A domain checklist can remain useful even when every rule is familiar: the user requested consistent coverage. Prefer applicability and detection recipes over lectures explaining the concept. Never delete a shipped application's security, accessibility, or data-integrity requirement merely because the model knows its name.

Collection workflow

  1. Pull safely and record the baseline revision, dirty paths, skill inventory, and validator result.
  2. Update the creator's retention criteria first. Audit each skill's entry point, references, scripts, routing neighbors, and evaluation coverage against those criteria.
  3. Audit the descriptions as one listing before editing any body. Concatenate every description in the order the host lists them and read the result as the model does: total length against the host's listing budget, pairs that could claim the same prompt, and any description that over-emphasizes its own applicability. Shorten before adding; a listing that runs over budget is trimmed from every tail at once, so one long description degrades routing for the whole collection.
  4. Keep a collection ledger with one row per skill: unique payload, concrete change or retention reason, and verification status. For large rule sets, record which categories were sampled and expand inspection when a sample fails.
  5. Fill gaps with a concrete contract, tool, or regression scenario. Do not add a new skill simply to cover a topic a frontier agent already handles.
  6. Validate every changed skill and the collection; check stale paths after deletions. Update descriptions and README entries to match final behavior.
  7. Report static checks separately from behavior runs. If target models or a runner are unavailable, ship reviewable edits and scenarios with that limitation explicit. Do not fabricate scores or call authored assertions passing tests.

Behavioral comparison

Use identical task inputs, repository state, tool access, and effort settings in fresh contexts for no-skill, previous-skill, and revised-skill arms. Record model identifier, host, date, loaded files, output artifact, assertion evidence, and failures. Compare task success, preference conformance, tool calls, latency, and context cost. Repeat borderline results before a destructive retirement decision.

User-named target models define the matrix. Unavailable models stay untested; a different model cannot stand in for them. Preserve contract assertions even when both arms pass, since a later edit can regress them. Repeat the keep test when the target model, task, or host changes: a stronger no-skill baseline can erase a previously useful gain.

Source: SKILL.md on GitHub

1 warning1d4 checks · Risk SAFE
  • Gen Agent Trust Hub1d

    The skill is a comprehensive development tool for creating and auditing other agent skills. It includes a shell-based validator script that uses Ruby and Perl to parse metadata and content. Because its primary function involves processing external, potentially untrusted skill files, it possesses an inherent attack surface for indirect prompt injection.

  • Socket1d

    No alerts

  • Snyk1d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    1/4 files flagged

Signed by skilld at b42cf43. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 1 hour ago.

Activeupdated 3 days ago
Other metadata
compatibility
Repository validation requires Bash, Ruby with YAML and JSON, Perl, and standard Unix utilities.

README badge

README badge for mblode/agent-skills/agent-skills-creator