---
title: "skill by agenticluke · skilld"
canonical_url: "https://skilld.dev/gh/agenticluke/eval-first-coding-plus"
meta:
  description: "Use an eval-first workflow, small tasks, risk checks, and cost-aware model choice for AI-led software work. From agenticluke/eval-first-coding-plus."
  "og:description": "Use an eval-first workflow, small tasks, risk checks, and cost-aware model choice for AI-led software work. From agenticluke/eval-first-coding-plus."
  "og:title": "skill by agenticluke"
  "twitter:description": "Use an eval-first workflow, small tasks, risk checks, and cost-aware model choice for AI-led software work. From agenticluke/eval-first-coding-plus."
  "twitter:title": "skill by agenticluke"
---

`

[All skills](https://skilld.dev/skills)

[![agenticluke avatar](https://skilld.dev/_img/avatar?url=https%3A%2F%2Fgithub.com%2Fagenticluke.png%3Fsize%3D96)](https://skilld.dev/gh/agenticluke)

# **/skill**

[@37d231e](https://github.com/agenticluke/eval-first-coding-plus/commit/37d231e8386e962eea72d6ff6cbb68d0a8d22391 "Your agent reads SKILL.md at commit 37d231e")

by [agenticluke](https://skilld.dev/gh/agenticluke)· [agenticluke](https://skilld.dev/gh/agenticluke)/ [eval-first-coding-plus](https://skilld.dev/gh/agenticluke/eval-first-coding-plus)

Use an eval-first workflow, small tasks, risk checks, and cost-aware model choice for AI-led software work.

- 1 file
- 4.2 KB
- Updated last month
- [GitHub](https://github.com/agenticluke/eval-first-coding-plus/blob/main/skill/SKILL.md "View SKILL.md on GitHub")

## SKILL.md

4.2 KB

**≈29** tokens always: the name and description. **≈1k** when used: this file.

## Agentic Engineering

> Original skill by ECC. This version keeps ECC's core method and gives full credit to ECC.

Use this skill when an AI agent will do most of the coding. A person must still control quality, safety, and risk.

### Core Rules

1. Write clear done rules before coding.
2. Split the work into small tasks.
3. Pick a model that fits each task.
4. Test before and after each change.
5. Stop when the result is unsafe or unclear.

### Eval-First Loop

For each change:

1. Write tests for the new skill or feature.
2. List old tests that must still pass.
3. Run the tests before editing.
4. Save the failed test names and key error text.
5. Make one small change.
6. Run the same tests again.
7. Compare the results.
8. Mark the task done only when its done rules pass.

If no test system exists, use a small check that can be run again. This may be a command, sample input, or short review list.

Do not change a test just to make bad code pass. Change a test only when the planned behavior has changed.

If the first run already passes, check that the test can fail for the right reason. A test that can never fail gives no proof.

### Split the Work

Aim for tasks that take about 15 minutes.

Each task must:

- Have one clear goal.
- Have one main risk.
- Be safe to test on its own.
- Have a clear done rule.
- List the files it may change.

Split a task again when it touches many parts, has more than one main risk, or cannot be checked alone.

Keep linked changes together when splitting them would leave broken code.

### Pick a Model

Use the lowest model level that can do the task well:

- Haiku: Sort items, fill simple code forms, or make one small edit.
- Sonnet: Build features, fix common bugs, or clean up code.
- Opus: Plan system design, find hard root causes, or guard rules across many files.

Move up one level only when the lower level fails due to weak thought or missed links.

Do not move up for a bad prompt, missing files, broken tools, or unclear done rules. Fix those first.

Move down when the task becomes small and clear.

### Session Rules

- Keep one session for tasks that share the same files and facts.
- Start a new session after a major phase, such as plan, build, or review.
- Save a short handoff before starting a new session.
- Shorten context only after a task or milestone is done.
- Do not shorten context during an active bug hunt.

A handoff should list:

- The goal.
- What changed.
- What passed.
- What failed.
- Known risks.
- The next task.

### Review AI-Written Code

Check these first:

- Rules that must always stay true.
- Empty, missing, large, and bad inputs.
- Failures, timeouts, and partial work.
- Login, access, and secret handling.
- Links between files or services.
- Safe release and rollback steps.
- Tests for each fixed bug.

Do not spend review time on style that a formatter or linter will fix.

Do not accept code only because tests pass. Read risky code and check that the tests cover the real danger.

Stop and ask a person when:

- The change may delete or expose data.
- Access rules are unclear.
- A release cannot be rolled back.
- The done rules conflict.
- A key choice needs product or legal input.

### Cost Rules

For each task, keep a short local note with:

- Model used.
- Rough token count.
- Retry count.
- Time used.
- Pass or fail.
- Why a higher model was used, if any.

Do not add tracking code, analytics, telemetry, or outside calls.

Set a retry limit before work starts. After two failed tries with the same plan, stop and find the cause. Change the plan before trying again.

### Example

Task: Add a rule that blocks an empty user name.

Done rules:

- An empty name returns a clear error.
- A valid name still works.
- All old user tests pass.
- Only the user check and its tests may change.

Steps:

1. Add a test with an empty name.
2. Run it and save the failure.
3. Use Sonnet because this is a small code change.
4. Add the name check.
5. Run the new test and all old user tests.
6. Review empty, blank, and very long names.
7. Record the result and mark the task done only if all done rules pass.

Source: [SKILL.md on GitHub](https://github.com/agenticluke/eval-first-coding-plus/blob/main/skill/SKILL.md)

## Third-party checks

No third-party reports yet.

## Provenance

[Signed by skilld at 37d231e.](https://github.com/agenticluke/eval-first-coding-plus/commit/37d231e8386e962eea72d6ff6cbb68d0a8d22391 "37d231e8386e962eea72d6ff6cbb68d0a8d22391") This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Activeupdated last month

## Capability

<dl>

<dt>origin</dt>
<dd>ECC</dd>

</dl>

## README badge

![README badge for agenticluke/eval-first-coding-plus](https://skilld.dev/b/agenticluke/eval-first-coding-plus?theme=light&label=0)

## Related skills

-
-
-
-
-
-