All skills

Build and run clear tests for Claude Code tasks. Use this skill to set pass or fail rules, test task success, check for regressions, measure pass@k and pass^k, or compare prompts and model versions.

  • 1 file
  • 8.6 KB
  • Updated last week
  • GitHub

Use this Skill: https://skilld.dev/gh/agenticluke/agent-eval-harness-plus/skill

This session only. Nothing lands on disk.

SKILL.md

≈51 tokens always: the name and description. ≈2.1k when used: this file.

Eval Harness

Original work by ECC. Credit to ECC for the core idea and design.

Use tests to check if an AI task works. Write the tests before you change the code or prompt.

Core Rules

  1. Define success before you start.
  2. Use clear pass or fail checks.
  3. Test both normal and edge cases.
  4. Run old tests after each change.
  5. Save the setup, result, and reason.
  6. Never change a test only to make a bad result pass.
  7. Do not use web calls, tracking, or data collection.
  8. Do not expose keys, private data, or user files in logs.

Eval Types

Capability Eval

Use this to test a new skill or feature.

# Capability Eval: <name>

Task:
<What Claude must do>

Setup:
<Files, tools, and starting state>

Success Checks:
- [ ] <One clear check>
- [ ] <One clear check>
- [ ] <One clear check>

Edge Cases:
- [ ] <Empty, missing, or bad input>
- [ ] <A limit or rare case>

Expected Output:
<Files, text, or state that must exist>

Fail If:
- <A clear reason to fail>

Each check must be easy to judge. Avoid words such as "good," "clean," or "better" unless the eval explains what they mean.

Regression Eval

Use this to make sure old work still works.

# Regression Eval: <name>

Baseline:
<Git SHA, version, or saved checkpoint>

Tests:
- [ ] <Old behavior that must still work>
- [ ] <Old behavior that must still work>
- [ ] <Old behavior that must still work>

Result:
<X>/<Y> passed

Change From Baseline:
<No change, better, or worse>

Do not claim a regression unless the setup matches the baseline.

Graders

Use the simplest grader that can give a fair result.

1. Code Grader

Use code for facts that have one clear answer.

# Check for expected text
grep -q "export function handleAuth" src/auth.ts \
  && echo "PASS" || echo "FAIL"

# Run one test group
npm test -- --testPathPattern="auth"

# Check the build
npm run build

Check the command exit code. Do not trust printed text alone.

Mark a timeout, crash, or missing tool as ERROR, not FAIL. A fail means the task ran and did not meet the rule.

2. Rule Grader

Use exact rules for text or files.

Rules:
- Output has no more than 100 words.
- Output contains one heading.
- Output does not contain private keys.
- Every file path exists.

State if matching is case-sensitive. Test both valid and invalid samples.

3. Model Grader

Use a model only when code and rules cannot judge the work.

# Model Grader Prompt

Judge the result against the rules below.

Rules:
1. It solves the stated task.
2. It handles each listed edge case.
3. It does not add work outside the task.
4. Its errors tell the user what went wrong.

Score each rule:
- 0: Does not meet the rule
- 1: Partly meets the rule
- 2: Fully meets the rule

Pass Rule:
Pass only if every rule scores 2.

Return:
- Scores
- PASS or FAIL
- One short reason for each score

Hide the expected answer when possible. Do not let the model grade its own hidden thoughts. If the result is close or unclear, ask for human review.

4. Human Grader

Use a person for safety, taste, unclear output, or high-risk work.

# Human Review Required

Change:
<What changed>

Review:
<What the person must check>

Reason:
<Why code or a model cannot judge it safely>

Risk:
<Low, Medium, or High>

Human review is required for security claims, private data rules, and major release choices.

Reliability Metrics

pass@k

pass@k asks: Did at least one of k tries pass?

  • pass@1 measures first-try success.
  • pass@3 measures success within three tries.
  • Use the same task and setup for each try.
  • Reset changed files and state before each try.
  • Do not count a retry that uses hints from an earlier try.

Example:

Task A: FAIL, PASS, FAIL = pass@3 success
Task B: FAIL, FAIL, FAIL = pass@3 failure

pass@3 = 1 successful task / 2 tasks = 50%

Do not report pass@3 from only one task as a broad success rate.

pass^k

pass^k asks: Did all k tries pass?

Use it for work that must be steady.

Task A: PASS, PASS, PASS = pass^3 success
Task B: PASS, FAIL, PASS = pass^3 failure

pass^3 = 1 stable task / 2 tasks = 50%

A common goal for a release path is pass^3 = 100%.

Workflow

1. Define

Before making changes, create:

.claude/evals/<feature>.md

Include:

  • The task
  • The fixed setup
  • Clear success checks
  • Edge cases
  • Regression checks
  • The grader for each check
  • The number of tries
  • The pass rule
  • Time and cost limits, if needed

2. Run a Baseline

Run the eval before the change when possible. Save the result. This shows what the change fixed and what it may have broken.

3. Make the Change

Change only what is needed for the task. Do not edit the eval after seeing a poor result unless the eval is wrong. If you fix the eval, record why.

4. Run the Evals

Run capability tests first. Then run regression tests.

For each run, save:

  • Date and time
  • Git SHA or file version
  • Model and settings, if known
  • Test setup
  • Attempt number
  • PASS, FAIL, or ERROR
  • Short reason
  • Run time, if useful

5. Review Failures

For each failure:

  1. Check that the setup was reset.
  2. Check that the grader is stable.
  3. Find the real cause.
  4. Fix the code or prompt.
  5. Run all needed tests again.

Do not hide failed tries.

6. Report

# Eval Report: <feature>

Baseline:
<Version or checkpoint>

Capability:
- <Test>: PASS on try 1
- <Test>: PASS on try 2
- <Test>: FAIL after 3 tries

Regression:
- <Test>: PASS
- <Test>: PASS

Errors:
- <Test>: ERROR because <reason>

Metrics:
- pass@1: <passed>/<total> = <rate>
- pass@3: <passed>/<total> = <rate>
- pass^3: <passed>/<total> = <rate>

Limits:
- Run time: <value>
- Cost: <value or not tracked>

Open Risks:
- <Risk or "None">

Status:
<Ready, Not Ready, or Needs Human Review>

Use Ready only when every release rule passes.

File Layout

.claude/
  evals/
    <feature>.md
    <feature>.log
    baseline.json

docs/
  releases/
    <version>/
      eval-summary.md

Keep logs small. Do not store secrets, full private prompts, or private user data.

Command Pattern

If the project supports eval commands, use:

/eval define <feature-name>
/eval check <feature-name>
/eval report <feature-name>

If these commands do not exist, create and run the files with local project tools. Do not claim that a command ran when it is only an example.

Edge Cases

Always plan for these when they apply:

  • Empty input
  • Missing files
  • Bad file types
  • Very long input
  • Unicode text
  • Spaces in file paths
  • Tool errors
  • Timeouts
  • Partial output
  • Changed working state
  • Tests that pass only when run in a set order
  • Random or unstable output
  • A grader that gives different scores for the same result
  • A task that passes but breaks old work

If a test is flaky, fix or replace it before using it as a release rule.

Concrete Example

A user asks: "Add email and password sign-up without breaking login."

Create .claude/evals/add-sign-up.md:

# Eval: Add Sign-Up

## Setup

- Start from Git SHA `abc123`.
- Use a new local test database for each try.
- Run each capability test three times.

## Capability Checks

- [ ] A valid email and password create one user.
- [ ] A bad email is rejected.
- [ ] A password under 12 chars is rejected.
- [ ] A second sign-up with the same email is rejected.
- [ ] The saved password is not plain text.
- [ ] An empty request returns a clear error.

## Regression Checks

- [ ] A current user can still log in.
- [ ] A bad password is still rejected.
- [ ] Log out still clears the session.
- [ ] Public pages still load.

## Graders

- Use code tests for all checks.
- Require human review of password storage.
- Treat a test crash as ERROR.

## Pass Rules

- Capability: pass@3 at least 90%.
- Regression: pass^3 equals 100%.
- No open security issue.

Run the baseline, make the change, and run all checks. A valid report may end with:

Metrics:
- pass@1: 9/10 = 90%
- pass@3: 10/10 = 100%
- pass^3: 4/4 = 100%

Open Risks:
- Password storage needs human review.

Status:
Needs Human Review

Release Rules

Before marking work ready:

  • All required checks pass.
  • Regression checks meet their set limit.
  • Errors are fixed or explained.
  • Flaky graders are not release gates.
  • Security work has human review.
  • The report names any open risk.
  • The eval and code are saved in the same version.

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at ed002d2. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated last week
origin
ECC
tools
Read, Write, Edit, Bash, Grep, Glob

README badge

README badge for agenticluke/agent-eval-harness-plus