All skills

Use when measuring how well an AI agent handles a recurring kind of task. Builds a small eval set, runs it, and raises the difficulty once the agent passes.

  • 2 files
  • 2 KB
  • Updated 10 hours ago
  • GitHub

Use this Skill: https://skilld.dev/gh/erkamyaman/dhh-p-bloom-agent-skills/agent-evals

This session only. Nothing lands on disk.

referenceseval-template.md

≈135 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Eval template

## Task <id>
Prompt: <exactly what the agent is given>
Starting state: <repo commit or fixture>
Pass check: <command that exits 0 on success>
Difficulty: easy | medium | hard
Notes: <known pitfalls>

Results table

Task Run 1 Run 2 Notes
01 pass pass
02 fail pass flaky check?

Raising the bar

  • Combine two tasks into one.
  • Remove hints from the prompt.
  • Use a larger or messier codebase.
  • Add a performance or size requirement to the pass check.

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 57c6daa. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 4 hours ago.

Activeupdated 10 hours ago

README badge

README badge for erkamyaman/dhh-p-bloom-agent-skills/agent-evals