All skills

Build and audit deterministic verification gates — a check that blocks a pipeline and can be shown to go red. Use when writing a calibration gate, CI check, validation script or pre-publication check for a numeric or empirical result; when a plausible-but-wrong value would survive review; when asking whether an existing test, linter rule or check could actually fail; and when a suite passes first try, passes suspiciously often, or was written by whatever produced the thing it checks. Triggers on "can this check fail", "known-bad", "negative control", "calibration gate", "sanity check my results", "is this test actually testing anything".

  • 7 files
  • 47.9 KB
  • Updated last month
  • GitHub

Use this Skill: https://skilld.dev/gh/oaustegard/claude-skills/gating

This session only. Nothing lands on disk.

README.md

≈600 tokens on demand. Your agent reads this file only when SKILL.md points to it.

gating

Build and audit verification gates — deterministic checks that block a pipeline and can be shown to go red.

The characteristic failure of a gate is not a wrong check. A wrong check gets noticed. It is a check that cannot fail, which reports PASS forever and is indistinguishable from a working one from the outside.

So the object under suspicion here is the check itself.

The three obligations

Anchor something your own code did not produce published constant, closed form, invariant, theoretical bound
Known-bad a case the gate demonstrably rejects break the subject plausibly, run the gate, confirm red
Coverage limit what it cannot catch, in writing holes are invisible from inside a green run

scripts/gate.py enforces the last two: a gate that registers no known-bad or no coverage limit reports INCONCLUSIVE and exits 2, not PASS.

Scripts

# mutation pass — which single-token changes does the gate NOT notice?
python3 scripts/mutate.py --target src/codec.py -- python3 calibrate.py

mutate.py is stdlib-only, uses tokenize (so it never corrupts strings or comments), restores the file even on interrupt, and refuses to run when the gate is already red. It works against any gate command; mutmut and cosmic-ray go deeper once it stops finding survivors.

gate.py is a ~150-line harness: anchor(), bracket(), known_bad(), coverage(), note(), and a report() that returns a process exit code.

Relationship to challenging

challenging asks would a careful reader object? — LLM judgement against a persona, verdicts of SHIP/REVISE/RETHINK. gating asks can this check go red? — deterministic, anchored, exit-coded.

Complements with different blind spots: an adversarial reviewer will not recompute your constants, and a gate will not notice that your framing is wrong. Note also that a same-model reviewer is an independent context, not an independent reviewer; where shared priors are the risk, an anchor beats a reviewer because an anchor is not negotiable.

See SKILL.md for the build and audit procedures and the anti-pattern table, references/anchors.md for choosing an oracle, and references/auditing.md for the six-pass "can this fail?" sweep over an existing suite.

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 04bfd5b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last month
metadata
{
  "version": "0.3.0"
}

README badge

README badge for oaustegard/claude-skills/gating