Eval Harness
Original work by ECC. Credit to ECC for the core idea and design.
Use tests to check if an AI task works. Write the tests before you change the code or prompt.
Core Rules
- Define success before you start.
- Use clear pass or fail checks.
- Test both normal and edge cases.
- Run old tests after each change.
- Save the setup, result, and reason.
- Never change a test only to make a bad result pass.
- Do not use web calls, tracking, or data collection.
- Do not expose keys, private data, or user files in logs.
Eval Types
Capability Eval
Use this to test a new skill or feature.
# Capability Eval: <name>
Task:
<What Claude must do>
Setup:
<Files, tools, and starting state>
Success Checks:
- [ ] <One clear check>
- [ ] <One clear check>
- [ ] <One clear check>
Edge Cases:
- [ ] <Empty, missing, or bad input>
- [ ] <A limit or rare case>
Expected Output:
<Files, text, or state that must exist>
Fail If:
- <A clear reason to fail>Each check must be easy to judge. Avoid words such as "good," "clean," or "better" unless the eval explains what they mean.
Regression Eval
Use this to make sure old work still works.
# Regression Eval: <name>
Baseline:
<Git SHA, version, or saved checkpoint>
Tests:
- [ ] <Old behavior that must still work>
- [ ] <Old behavior that must still work>
- [ ] <Old behavior that must still work>
Result:
<X>/<Y> passed
Change From Baseline:
<No change, better, or worse>Do not claim a regression unless the setup matches the baseline.
Graders
Use the simplest grader that can give a fair result.
1. Code Grader
Use code for facts that have one clear answer.
# Check for expected text
grep -q "export function handleAuth" src/auth.ts \
&& echo "PASS" || echo "FAIL"
# Run one test group
npm test -- --testPathPattern="auth"
# Check the build
npm run buildCheck the command exit code. Do not trust printed text alone.
Mark a timeout, crash, or missing tool as ERROR, not FAIL. A fail means the task ran and did not meet the rule.
2. Rule Grader
Use exact rules for text or files.
Rules:
- Output has no more than 100 words.
- Output contains one heading.
- Output does not contain private keys.
- Every file path exists.State if matching is case-sensitive. Test both valid and invalid samples.
3. Model Grader
Use a model only when code and rules cannot judge the work.
# Model Grader Prompt
Judge the result against the rules below.
Rules:
1. It solves the stated task.
2. It handles each listed edge case.
3. It does not add work outside the task.
4. Its errors tell the user what went wrong.
Score each rule:
- 0: Does not meet the rule
- 1: Partly meets the rule
- 2: Fully meets the rule
Pass Rule:
Pass only if every rule scores 2.
Return:
- Scores
- PASS or FAIL
- One short reason for each scoreHide the expected answer when possible. Do not let the model grade its own hidden thoughts. If the result is close or unclear, ask for human review.
4. Human Grader
Use a person for safety, taste, unclear output, or high-risk work.
# Human Review Required
Change:
<What changed>
Review:
<What the person must check>
Reason:
<Why code or a model cannot judge it safely>
Risk:
<Low, Medium, or High>Human review is required for security claims, private data rules, and major release choices.
Reliability Metrics
pass@k
pass@k asks: Did at least one of k tries pass?
pass@1measures first-try success.pass@3measures success within three tries.- Use the same task and setup for each try.
- Reset changed files and state before each try.
- Do not count a retry that uses hints from an earlier try.
Example:
Task A: FAIL, PASS, FAIL = pass@3 success
Task B: FAIL, FAIL, FAIL = pass@3 failure
pass@3 = 1 successful task / 2 tasks = 50%Do not report pass@3 from only one task as a broad success rate.
pass^k
pass^k asks: Did all k tries pass?
Use it for work that must be steady.
Task A: PASS, PASS, PASS = pass^3 success
Task B: PASS, FAIL, PASS = pass^3 failure
pass^3 = 1 stable task / 2 tasks = 50%A common goal for a release path is pass^3 = 100%.
Workflow
1. Define
Before making changes, create:
.claude/evals/<feature>.mdInclude:
- The task
- The fixed setup
- Clear success checks
- Edge cases
- Regression checks
- The grader for each check
- The number of tries
- The pass rule
- Time and cost limits, if needed
2. Run a Baseline
Run the eval before the change when possible. Save the result. This shows what the change fixed and what it may have broken.
3. Make the Change
Change only what is needed for the task. Do not edit the eval after seeing a poor result unless the eval is wrong. If you fix the eval, record why.
4. Run the Evals
Run capability tests first. Then run regression tests.
For each run, save:
- Date and time
- Git SHA or file version
- Model and settings, if known
- Test setup
- Attempt number
PASS,FAIL, orERROR- Short reason
- Run time, if useful
5. Review Failures
For each failure:
- Check that the setup was reset.
- Check that the grader is stable.
- Find the real cause.
- Fix the code or prompt.
- Run all needed tests again.
Do not hide failed tries.
6. Report
# Eval Report: <feature>
Baseline:
<Version or checkpoint>
Capability:
- <Test>: PASS on try 1
- <Test>: PASS on try 2
- <Test>: FAIL after 3 tries
Regression:
- <Test>: PASS
- <Test>: PASS
Errors:
- <Test>: ERROR because <reason>
Metrics:
- pass@1: <passed>/<total> = <rate>
- pass@3: <passed>/<total> = <rate>
- pass^3: <passed>/<total> = <rate>
Limits:
- Run time: <value>
- Cost: <value or not tracked>
Open Risks:
- <Risk or "None">
Status:
<Ready, Not Ready, or Needs Human Review>Use Ready only when every release rule passes.
File Layout
.claude/
evals/
<feature>.md
<feature>.log
baseline.json
docs/
releases/
<version>/
eval-summary.mdKeep logs small. Do not store secrets, full private prompts, or private user data.
Command Pattern
If the project supports eval commands, use:
/eval define <feature-name>
/eval check <feature-name>
/eval report <feature-name>If these commands do not exist, create and run the files with local project tools. Do not claim that a command ran when it is only an example.
Edge Cases
Always plan for these when they apply:
- Empty input
- Missing files
- Bad file types
- Very long input
- Unicode text
- Spaces in file paths
- Tool errors
- Timeouts
- Partial output
- Changed working state
- Tests that pass only when run in a set order
- Random or unstable output
- A grader that gives different scores for the same result
- A task that passes but breaks old work
If a test is flaky, fix or replace it before using it as a release rule.
Concrete Example
A user asks: "Add email and password sign-up without breaking login."
Create .claude/evals/add-sign-up.md:
# Eval: Add Sign-Up
## Setup
- Start from Git SHA `abc123`.
- Use a new local test database for each try.
- Run each capability test three times.
## Capability Checks
- [ ] A valid email and password create one user.
- [ ] A bad email is rejected.
- [ ] A password under 12 chars is rejected.
- [ ] A second sign-up with the same email is rejected.
- [ ] The saved password is not plain text.
- [ ] An empty request returns a clear error.
## Regression Checks
- [ ] A current user can still log in.
- [ ] A bad password is still rejected.
- [ ] Log out still clears the session.
- [ ] Public pages still load.
## Graders
- Use code tests for all checks.
- Require human review of password storage.
- Treat a test crash as ERROR.
## Pass Rules
- Capability: pass@3 at least 90%.
- Regression: pass^3 equals 100%.
- No open security issue.Run the baseline, make the change, and run all checks. A valid report may end with:
Metrics:
- pass@1: 9/10 = 90%
- pass@3: 10/10 = 100%
- pass^3: 4/4 = 100%
Open Risks:
- Password storage needs human review.
Status:
Needs Human ReviewRelease Rules
Before marking work ready:
- All required checks pass.
- Regression checks meet their set limit.
- Errors are fixed or explained.
- Flaky graders are not release gates.
- Security work has human review.
- The report names any open risk.
- The eval and code are saved in the same version.