---
title: "skill by agenticluke · skilld"
canonical_url: "https://skilld.dev/gh/agenticluke/build-review-loop-plus"
meta:
  description: "A GAN-style work loop that keeps building and review separate to improve app quality. From agenticluke/build-review-loop-plus."
  "og:description": "A GAN-style work loop that keeps building and review separate to improve app quality. From agenticluke/build-review-loop-plus."
  "og:title": "skill by agenticluke"
  "twitter:description": "A GAN-style work loop that keeps building and review separate to improve app quality. From agenticluke/build-review-loop-plus."
  "twitter:title": "skill by agenticluke"
---

`

[All skills](https://skilld.dev/skills)

[![agenticluke avatar](https://skilld.dev/_img/avatar?url=https%3A%2F%2Fgithub.com%2Fagenticluke.png%3Fsize%3D96)](https://skilld.dev/gh/agenticluke)

# **/skill**

[@135a57f](https://github.com/agenticluke/build-review-loop-plus/commit/135a57f19754f363ede6240d5f97efbdf1381352 "Your agent reads SKILL.md at commit 135a57f")

by [agenticluke](https://skilld.dev/gh/agenticluke)· [agenticluke](https://skilld.dev/gh/agenticluke)/ [build-review-loop-plus](https://skilld.dev/gh/agenticluke/build-review-loop-plus)

A GAN-style work loop that keeps building and review separate to improve app quality.

- 1 file
- 9.6 KB
- Updated 3 weeks ago
- [GitHub](https://github.com/agenticluke/build-review-loop-plus/blob/135a57f19754f363ede6240d5f97efbdf1381352/skill/SKILL.md "View SKILL.md on GitHub")

## SKILL.md

9.6 KB

**≈23** tokens always: the name and description. **≈2.4k** when used: this file.

## GAN-Style Harness Skill

> **Credit: This skill comes from the ECC Community.**
>
> It was inspired by [Anthropic's Harness Design for Long-Running Application Development](https://www.anthropic.com/engineering/harness-design-long-running-apps), published March 24, 2026.

Use separate agents to plan, build, and review an app. The builder must not review its own work.

This works like a GAN. One agent makes the app. Another agent finds flaws. The builder then fixes those flaws.

### When to Use

Use this skill when:

- A short prompt must become a full app.
- The app must look polished.
- Many parts must work together.
- You can afford several build and review rounds.
- A strict review will improve the result.

Do not use this skill when:

- The task is one small file change.
- The budget is under $10.
- The task is a simple code cleanup.
- Clear tests already define the whole task.
- You cannot run or test the app.

### Main Rule

Keep the Generator and Evaluator separate.

The Generator builds the app. The Evaluator tests it. The Evaluator must not edit the app. The Generator must not change review scores.

### Work Flow

```
Short prompt
    |
    v
Planner writes the spec
    |
    v
Generator builds one part
    |
    v
Evaluator tests the live app
    |
    v
Generator fixes clear issues
    |
    +---- repeat until the app passes or the limit is reached
```

Use 5 to 15 rounds. Stop early when the app passes all required checks.

### The Three Agents

#### 1. Planner

The Planner turns a short prompt into a clear product spec.

It must:

- List the users and their main goals.
- List the features in order of need.
- Split large work into small rounds.
- State the look and feel.
- State how each feature will be tested.
- Mark must-have and nice-to-have work.
- Note limits, risks, and facts that are not known.

Do not force every project to have 16 features. Pick a size that fits the prompt, time, and budget.

Use Sonnet by default. Set `GAN_PLANNER_MODEL=opus` when the plan needs deeper thought.

#### 2. Generator

The Generator builds from the spec.

Before each round, it must write a short round plan with:

- What it will build.
- What files it may change.
- How it will test the work.
- What it will not build yet.

The Generator must:

- Read the spec and the latest review.
- Fix high-risk bugs first.
- Keep old working parts safe.
- Run useful tests after each change.
- Save work in Git when Git is present.
- Never delete user work without clear approval.
- Record any issue it cannot fix.

Use Sonnet by default. Set `GAN_GENERATOR_MODEL=opus` for hard coding work.

#### 3. Evaluator

The Evaluator tests the running app, not only the code.

It must:

- Open the live app.
- Try each required user task.
- Click links and buttons.
- Fill and send forms.
- Test bad input and empty input.
- Test loading, empty, error, and success states.
- Check small and large screen sizes.
- Check keyboard use and clear focus marks.
- Check that text is easy to read.
- Test API routes when the app has an API.
- Save clear steps for every bug.
- Give proof, such as a page, action, error, or screen image.

Use Playwright when it is set up. If it is not set up, use the best local test tools. State what could not be checked.

The Evaluator must not praise weak work. It must also not invent flaws. Each issue needs proof.

Use Sonnet by default. Set `GAN_EVALUATOR_MODEL=opus` when review work is hard.

### Review Score

Score each item from 1 to 10.

#### Design Quality, weight 0.3

- 1 to 3: Looks broken or copied from a plain template.
- 4 to 6: Clear, but plain or uneven.
- 7 to 8: Has a strong and steady style.
- 9 to 10: Looks ready for real users.

#### Originality, weight 0.2

- 1 to 3: Uses only common colors and layouts.
- 4 to 6: Has a few custom choices.
- 7 to 8: Has a clear and fresh idea.
- 9 to 10: Feels new, useful, and well judged.

#### Craft, weight 0.3

- 1 to 3: Has broken layout or missing states.
- 4 to 6: Works, but feels rough.
- 7 to 8: Feels smooth and works on many screen sizes.
- 9 to 10: Every small part feels careful and complete.

#### Function, weight 0.2

- 1 to 3: Main tasks fail or are missing.
- 4 to 6: Main path works, but bad input or rare cases fail.
- 7 to 8: All required tasks work with clear errors.
- 9 to 10: Works well across all tested cases.

Use this formula:

```
score =
  design × 0.3 +
  originality × 0.2 +
  craft × 0.3 +
  function × 0.2
```

The default pass score is `7.0`.

A pass also requires:

- No open high-risk bug.
- Every must-have feature works.
- Required tests pass.
- No score is below `5`.
- The Evaluator tested the latest build.

A high total score cannot hide a broken main feature.

### Review Report Format

The Evaluator must return:

```
# Review 003

Build tested: <Git commit or build name>
URL tested: <local URL>
Total score: 7.2
Result: FAIL

## Scores

- Design Quality: 8/10
- Originality: 7/10
- Craft: 7/10
- Function: 6/10

## Must Fix

1. Task cards vanish after a page reload.
   - Steps: Add a card, reload the page.
   - Expected: The card stays.
   - Found: The card is gone.

## Should Fix

1. The save button has no focus mark.

## Checks That Passed

- A user can create a board.
- The layout works at 375 px and 1440 px.

## Not Tested

- Email login. No test email service was set up.
```

Number review files in order, such as `feedback-001.md` and `feedback-002.md`.

### Stop Rules

Stop when one of these is true:

- The app meets every pass rule.
- The round limit is reached.
- The budget limit is reached.
- The same blocked issue appears in three reviews.
- A needed tool, key, service, or user choice is missing.
- A fix would cause data loss or a large change outside the spec.

When stopped without a pass, report:

- What works.
- What is still broken.
- Why work stopped.
- The safest next step.

Do not weaken the score just to finish.

### Edge Cases

- If the app will not start, give Function a score from 1 to 3. Fix startup first.
- If the app has no user screen, skip visual scores only when the spec says it is API-only. Reweight the remaining items so their weights total `1.0`.
- If a feature needs a paid service, use a safe mock only when the spec allows it.
- If test data may harm real data, use a local or test account.
- If two reviews give opposite advice, follow the spec and ask the Planner to settle the conflict.
- If a review finds a security or data-loss risk, fix it before style issues.
- If the working tree has user changes, keep them. Do not reset or replace them.
- If the app uses random output, test it more than once.
- If a fix lowers another score, record the tradeoff and test both parts again.

### Commands

#### Full Harness

```
/project:gan-build "Build a project board with columns, task cards, team notes, and dark mode"
```

#### Custom Limits

```
/project:gan-build "Build a recipe sharing site" \
  --max-iterations 10 \
  --pass-threshold 7.5
```

#### Design Only

Use the Generator and Evaluator when a full plan is not needed:

```
/project:gan-design "Create a landing page for a money tracker"
```

#### Shell Script

```
GAN_MAX_ITERATIONS=10 \
GAN_PASS_THRESHOLD=7.5 \
GAN_EVAL_CRITERIA="functionality,performance,security" \
./scripts/gan-harness.sh "Build a task API"
```

### Concrete Example

Prompt:

```
Build a small book tracker. A user can add a book, mark it read, search by title, and use dark mode.
```

The Planner writes:

```
Round 1: Add, list, and save books.
Round 2: Mark books as read and search by title.
Round 3: Add dark mode, small-screen layout, and error states.
Pass check: All four user tasks work after a page reload.
```

The Generator builds Round 1 and starts the app on port `3000`.

The Evaluator tests:

```
1. Add a book.
2. Reload the page.
3. Check that the book still exists.
4. Add a blank title.
5. Check that a clear error appears.
```

The Evaluator finds that books vanish after reload. It writes `feedback-001.md` with the exact steps.

The Generator reads that file, adds saved storage, and runs the tests again.

The Evaluator tests the new build. The loop ends only when all must-have tasks work and the score is at least `7.0`.

### Manual Claude Code Flow

```
# 1. Plan
claude -p --model sonnet \
  "Act as the Planner. Turn this brief into a clear spec: Build a book tracker. Write spec.md."

# 2. Build
claude -p --model sonnet \
  "Act as the Generator. Read spec.md. Build Round 1. Test it. Start the app on port 3000."

# 3. Review
claude -p --model sonnet \
  --allowedTools "Read,Bash,mcp__playwright__*" \
  "Act as the Evaluator. Test http://localhost:3000. Use the score rules. Write feedback-001.md."

# 4. Fix
claude -p --model sonnet \
  "Act as the Generator. Read spec.md and feedback-001.md. Fix every Must Fix item. Run the tests."

# Repeat review and fix steps until the app passes or a stop rule applies.
```

### Model Changes

Use less structure when models can handle longer work well.

#### Stage 1: Less Capable Models

- Use small rounds.
- Start a fresh context for each round.
- Give each agent short and exact tasks.
- Save the plan, work state, and review in files.

#### Stage 2: More Capable Models

- Use the full Planner, Generator, and Evaluator flow.
- Agree on each round before coding.
- Use fewer, larger rounds when the risk is low.
- Keep separate review files.

#### Stage 3: Future Models

- Merge planning into the Generator only if the plan stays clear.
- Keep the Evaluator separate.
- Keep proof for each bug.
- Keep pass rules and stop rules.
- Remove steps only after tests show they are no longer needed.

Source: [SKILL.md on GitHub](https://github.com/agenticluke/build-review-loop-plus/blob/135a57f19754f363ede6240d5f97efbdf1381352/skill/SKILL.md)

## Third-party checks

No third-party reports yet.

## Provenance

[Signed by skilld at 135a57f.](https://github.com/agenticluke/build-review-loop-plus/commit/135a57f19754f363ede6240d5f97efbdf1381352 "135a57f19754f363ede6240d5f97efbdf1381352") This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 weeks ago.

Activeupdated 3 weeks ago

## Capability

<dl>

<dt>origin</dt>
<dd>ECC-community</dd>

<dt>tools</dt>
<dd>Read, Write, Edit, Bash, Grep, Glob, Task</dd>

</dl>

## README badge

![README badge for agenticluke/build-review-loop-plus](https://skilld.dev/b/agenticluke/build-review-loop-plus?theme=light&label=0)

## Related skills

-
-
-
-
-
-