All skills
kunchenguid avatar

/judge-evaluation

@d1655fa

Use when changing or evaluating judge prompts, scoring thresholds, usage floors, or judgment eval corpora.

  • 1 file
  • 747 B
  • Updated 2 weeks ago
  • GitHub

Use this Skill: https://skilld.dev/gh/kunchenguid/compact-adviser/judge-evaluation

This session only. Nothing lands on disk.

SKILL.md

≈31 tokens always: the name and description. ≈137 when used: this file.

  • Judge design rule, measured: two one-sentence atomic questions composed in code beat any single question that folds two judgments together (TypeSafe's own guidance). Hill-climb prompt changes with eval/tools/earn.py (paired bootstrap) and read the usage-floor ladder with eval/tools/schedule.py; a clause stays only if it earns its place.
  • Judgment eval harness: packages/pi-extension/eval/. Session transcripts, labels, worksheets, and results stay in gitignored eval/local/. See that README for the corpus contract and label rubric.

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at d1655fa. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 weeks ago
user-invocable
false
metadata
{
  "internal": true
}

README badge

README badge for kunchenguid/compact-adviser/judge-evaluation