---
title: "skill by agenticluke · skilld"
canonical_url: "https://skilld.dev/gh/agenticluke/dual-review-gate-plus"
meta:
  description: "Adversarial review with a fix loop. Two independent reviewers must both pass the work before it can ship. From agenticluke/dual-review-gate-plus."
  "og:description": "Adversarial review with a fix loop. Two independent reviewers must both pass the work before it can ship. From agenticluke/dual-review-gate-plus."
  "og:title": "skill by agenticluke"
  "twitter:description": "Adversarial review with a fix loop. Two independent reviewers must both pass the work before it can ship. From agenticluke/dual-review-gate-plus."
  "twitter:title": "skill by agenticluke"
---

`

[All skills](https://skilld.dev/skills)

[![agenticluke avatar](https://skilld.dev/_img/avatar?url=https%3A%2F%2Fgithub.com%2Fagenticluke.png%3Fsize%3D96)](https://skilld.dev/gh/agenticluke)

# **/skill**

[@ca8f72d](https://github.com/agenticluke/dual-review-gate-plus/commit/ca8f72d4c2e084fe3be8ab65d6c8e81eac94b6c8 "Your agent reads SKILL.md at commit ca8f72d")

by [agenticluke](https://skilld.dev/gh/agenticluke)· [agenticluke](https://skilld.dev/gh/agenticluke)/ [dual-review-gate-plus](https://skilld.dev/gh/agenticluke/dual-review-gate-plus)

Adversarial review with a fix loop. Two independent reviewers must both pass the work before it can ship.

- 1 file
- 9.3 KB
- Updated 2 weeks ago
- [GitHub](https://github.com/agenticluke/dual-review-gate-plus/blob/main/skill/SKILL.md "View SKILL.md on GitHub")

## SKILL.md

9.3 KB

**≈28** tokens always: the name and description. **≈2.3k** when used: this file.

## Santa Method

The Santa Method checks important work before it ships.

One agent makes the work. Two new agents review it alone. Both reviewers use the same rules. The work ships only if both reviewers pass it.

If either reviewer finds a real problem, fix the work. Then ask two new reviewers to check the full work again.

### When to Use It

Use this method when:

- The work will be public.
- The work will go to users.
- The work must follow legal, safety, brand, or policy rules.
- Code may ship without a person checking it.
- Facts, quotes, links, or API details must be right.
- A large batch may repeat the same error.
- A false claim could cause harm.

Do not use it for:

- Early drafts.
- Brainstorming.
- Low-risk notes.
- Checks that tools can prove, such as tests, builds, type checks, or lint.

Run tool-based checks first. Use Santa after those checks pass.

### Core Rule

Both reviewers must return `PASS`.

If one reviewer returns `FAIL`, do not ship.

Do not average the two results. Do not let one pass cancel one fail.

### Roles

#### Generator

Creates the first output from the task rules.

#### Reviewer B

Checks the output alone.

#### Reviewer C

Checks the same output alone.

#### Fixer

Fixes only the issues found by the reviewers.

The generator and fixer may be the same agent. The reviewers must not be the generator or fixer.

### Review Rules

Each reviewer must receive:

- The full task.
- The full output.
- The same review list.
- Any source files needed to check facts.

The reviewers must not receive:

- The other review.
- Notes from an earlier review round.
- The generator's hidden reasoning.
- A hint that the output should pass.

Run both reviews at the same time when possible.

Use new reviewer agents for every round. This helps stop old views from shaping the next review.

### Review List

Write clear pass rules before review starts. Each rule must be easy to test.

Common rules:

| Check | Pass rule | Fail examples |
| --- | --- | --- |
| Facts | Each claim matches a trusted source or given text. | Made-up facts, dates, quotes, or links. |
| Full scope | Every task rule is met. | Missing part or skipped edge case. |
| Internal match | No part conflicts with another part. | One part says yes and another says no. |
| Policy | All required rules are followed. | Banned words, missing notice, wrong tone. |
| Code | Code is sound and tool checks pass. | Bug, bad input check, test failure. |
| Safety | No secret or unsafe action is exposed. | API key in code or unsafe command. |

Add rules that fit the task.

For code, also check:

- Errors are handled.
- Empty and bad input are handled.
- Secrets are not in the output.
- User input is checked.
- New paths have tests when needed.
- The change does not break old behavior.

For public text, also check:

- Names and links are right.
- Claims have support.
- The tone fits the brand.
- Required notices are present.
- Calls to action point to the right place.

For legal, money, or health text, also check:

- No result is promised without proof.
- Required warnings are present.
- Terms fit the right place and law.
- A human review is used when rules require it.

### Reviewer Prompt

Use this prompt for both reviewers:

```
You are an independent quality reviewer.

You have not seen any other review. Do not guess what another reviewer may say.

TASK
{task_spec}

OUTPUT TO CHECK
{output}

SOURCES
{sources}

PASS RULES
{rubric}

Check every pass rule.

Return valid JSON only:

{
  "verdict": "PASS",
  "checks": [
    {
      "criterion": "Name of the rule",
      "result": "PASS",
      "detail": "Short reason with exact proof"
    }
  ],
  "critical_issues": [],
  "suggestions": []
}

Use FAIL for the main verdict if any rule fails.

For each failed rule:
- Name the exact problem.
- Point to the exact part of the output.
- Say what must change.
- Do not fail work for taste or style unless the pass rules cover it.

Put only ship-blocking problems in critical_issues.
Put optional ideas in suggestions.
Do not invent facts, rules, or missing needs.
```

If a reviewer returns broken JSON, an unknown verdict, or misses a rule, treat that review as invalid. Run that review again with a new agent. Do not count an invalid review as a pass.

### Verdict Gate

```
def santa_verdict(review_b, review_c):
    if review_b["verdict"] == "PASS" and review_c["verdict"] == "PASS":
        return {
            "verdict": "NICE",
            "issues": [],
        }

    issues = dedupe(
        review_b.get("critical_issues", [])
        + review_c.get("critical_issues", [])
    )

    return {
        "verdict": "NAUGHTY",
        "issues": issues,
    }
```

A reviewer may fail a rule but forget to add it to `critical_issues`. The gate must also collect failed checks.

If reviewers disagree, keep the failure unless a trusted tool or source proves it wrong. Record why it was rejected. Do not ask the reviewers to debate each other.

### Fix Loop

Use a small limit. Three fix rounds is a good default.

```
MAX_FIX_ROUNDS = 3

output = generate(task_spec)

for round_number in range(MAX_FIX_ROUNDS + 1):
    review_b, review_c = run_two_new_reviews(
        task_spec=task_spec,
        output=output,
        rubric=rubric,
        sources=sources,
    )

    result = santa_verdict(review_b, review_c)

    if result["verdict"] == "NICE":
        return ship(output)

    if round_number == MAX_FIX_ROUNDS:
        return ask_for_human_review(
            output=output,
            issues=result["issues"],
        )

    output = fix_output(
        output=output,
        issues=result["issues"],
        instruction=(
            "Fix every listed issue. Keep correct parts unchanged. "
            "Do not add work that the task did not ask for."
        ),
    )
```

After each fix:

1. Run tool checks again if the change can affect them.
2. Give the full new output to two new reviewers.
3. Check every rule again, not only the old failures.
4. Ship only after both reviewers pass.

Stop and ask a person for help when:

- The limit is reached.
- The reviewers need a source they do not have.
- Two rules conflict.
- A fix needs a choice only the user can make.
- A legal, safety, or policy rule calls for human approval.

### Large Batches

For high-risk work, review every item.

For lower-risk batches of 100 or more items:

1. Split items by type, source, or template.
2. Pick at least five items from each group.
3. Include edge cases, not just random items.
4. Run the full Santa Method on each picked item.
5. If one item fails, find every item that may share the same fault.
6. Fix those items.
7. Pick a new sample and review again.
8. Do not ship the batch if the risk is still unclear.

Sampling can miss rare errors. Do not use it when one bad item could cause serious harm.

### Fallback Without Subagents

If separate agents are not available:

1. Save the task, output, sources, and review list.
2. Start a clean context.
3. Run Reviewer B.
4. Save its JSON.
5. Clear the context.
6. Start another clean context.
7. Run Reviewer C.
8. Apply the same verdict gate.

This is weaker than two separate agents because old context may leak into later work. Say when this fallback was used.

### Common Problems

| Problem | Fix |
| --- | --- |
| Reviews never end | Stop after the set limit and ask a person. |
| Reviewers pass all work with no proof | Require one result and reason for every rule. |
| Reviewers flag personal taste | Use clear pass rules. Reject notes outside those rules. |
| A fix breaks another part | Review the full output again with new agents. |
| Both reviewers miss the same issue | Use better sources, add a third reviewer, or ask a person for high-risk work. |
| Reviewers copy each other | Keep their prompts and results apart. |
| A reviewer invents a rule | Reject that finding unless the task or review list supports it. |
| A source cannot be checked | Mark the claim as unproved. Remove it, add a source, or ask a person. |
| The output is too large | Split it into clear parts, then review the final joined output too. |

### Concrete Example

Task:

```
Write a setup guide for a command-line tool.
It must use version 4.2.
It must cover Linux and macOS.
It must not claim Windows support.
```

First output:

```
Install Tool 4.1 with the command below.
The same steps work on Linux, macOS, and Windows.
```

Reviewer B:

```
{
  "verdict": "FAIL",
  "checks": [
    {
      "criterion": "Use version 4.2",
      "result": "FAIL",
      "detail": "The output says version 4.1."
    },
    {
      "criterion": "Do not claim Windows support",
      "result": "FAIL",
      "detail": "The output says the steps work on Windows."
    }
  ],
  "critical_issues": [
    "Change version 4.1 to 4.2.",
    "Remove the Windows support claim."
  ],
  "suggestions": []
}
```

Reviewer C may find the same faults or a different fault. Since at least one review failed, the work does not ship.

The fixer changes the guide. Then two new reviewers check the full guide. It ships only if both return `PASS`.

### Final Ship Check

Before shipping, confirm:

- Both final reviews are valid.
- Both final verdicts are `PASS`.
- Both reviews used the same task and rules.
- Both reviewers saw the latest full output.
- Tool checks pass.
- No required human approval is missing.

Source: [SKILL.md on GitHub](https://github.com/agenticluke/dual-review-gate-plus/blob/main/skill/SKILL.md)

## Third-party checks

No third-party reports yet.

## Provenance

[Signed by skilld at ca8f72d.](https://github.com/agenticluke/dual-review-gate-plus/commit/ca8f72d4c2e084fe3be8ab65d6c8e81eac94b6c8 "ca8f72d4c2e084fe3be8ab65d6c8e81eac94b6c8") This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 weeks ago.

Activeupdated 2 weeks ago

## Capability

<dl>

<dt>origin</dt>
<dd>Ronald Skelton - Founder, RapportScore.ai</dd>

</dl>

## README badge

![README badge for agenticluke/dual-review-gate-plus](https://skilld.dev/b/agenticluke/dual-review-gate-plus?theme=light&label=0)

## Related skills

-
-
-
-
-
-