- Judge design rule, measured: two one-sentence atomic questions composed in code beat any single question that folds two judgments together (TypeSafe's own guidance). Hill-climb prompt changes with
eval/tools/earn.py(paired bootstrap) and read the usage-floor ladder witheval/tools/schedule.py; a clause stays only if it earns its place. - Judgment eval harness:
packages/pi-extension/eval/. Session transcripts, labels, worksheets, and results stay in gitignoredeval/local/. See that README for the corpus contract and label rubric.
/judge-evaluation
@d1655faUse when changing or evaluating judge prompts, scoring thresholds, usage floors, or judgment eval corpora.
- 1 file
- 747 B
- Updated 2 weeks ago
- GitHub
Ask your Agent
Use this Skill: https://skilld.dev/gh/kunchenguid/compact-adviser/judge-evaluation
This session only. Nothing lands on disk.
≈31 tokens always: the name and description. ≈137 when used: this file.
Source: SKILL.md on GitHub
Third-party checks
No third-party reports yet.
Provenance
Signed by skilld at d1655fa. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.
Last checked against GitHub yesterday.
Activeupdated 2 weeks ago
Capability
- user-invocable
- false
- metadata
{ "internal": true }