Rubber-Duck Review
Use this when the user has an intuition but the factory frame is not yet sharp.
Question Ladder
Ask only the questions that change the next action.
- Who exactly has the pain?
- What do they do today instead?
- What repeated complaint, workaround, bad review, or spending proves the pain?
- What is the smallest artifact that can test the riskiest assumption?
- What would make this idea not worth doing?
- What platform, legal, privacy, or cost boundary can block it?
- What metric would make the next cycle obvious?
Review Frames
| Frame | What to catch |
|---|---|
| Audience | Too broad, unreachable, or imaginary user |
| Pain | Nice-to-have disguised as urgent need |
| Substitute | Existing tool already solves it well enough |
| Distribution | No believable way to reach first users |
| Scope | MVP still too large for one cycle |
| Risk | Store policy, TOS, privacy, copyright, safety, or payment issue |
| Metric | Outcome cannot be measured without guessing |
Response Shape
Return the review in this shape:
## Sharpest Version
<one sentence opportunity statement>
## Assumptions to Test
- <assumption> -> <how to test>
## Kill Risks
- <risk> -> <early signal>
## First Queue Slice
- discover: <task>
- research: <task>
- build/design: <task>
- review/track: <task>Opportunity Statement Pattern
For <specific audience> who struggle with <repeated pain>, create <small artifact> that helps them <measurable outcome>, validated by <metric>.Bias Checks
- If the idea starts from technology, force one pass from user pain.
- If the idea starts from money, force one pass from distribution.
- If the idea starts from a cool mechanic, force one pass from retention or replay value.
- If the idea starts from automation, force one pass from policy and trust risk.
Role-Level Critic Checkpoints (Layer 1: Light Self-Critic)
上のセクションは opportunity ideation review (workspace 全体の meta critic)。ここから下は 4 role の各 final-decision 前 に走らせる短い self-critic checkpoint。1 role あたり最大 5 質問、通常 1-2 分の pause だけ。全 role で共通の判定枠。
Layer 位置付け
| Layer | 実行者 | 起動タイミング | 出力 | 詳細 |
|---|---|---|---|---|
| Layer 1 Light Self-Critic | 各 role 自身 (context 内で観点切替) | 各 role の final decision 前、常時 | 内部 log (dashboard-state.criticLog に summary) | 本ページ |
| Layer 2 Advisory Critic | Critic role invocation (rubber-duck 深掘り) | Fallback lane #3、Persistence Exhaustive の task | critic-report.md (advisory) |
Fallback-lane.md #3 参照 |
| Layer 3 Blocking Critic | Critic role invocation (blocking) | 重要 gate: portfolio promote / focus theme apply / prompt-self-improvement commit / external publish | critic-report.md (blocking, verdict=pass 必須) |
本ページ末尾 + prompt-self-improvement.md 参照 |
- Advisory は verdict = reject でも proceed 可、log 化必須
- Blocking は verdict = pass 以外は proceed 不可、user 明示 override は security-approve 経由
Role 別 Checkpoint (Layer 1)
Commander (task selection bias check)
Final task dispatch 前に自問:
- この task は north-star / focus theme に照らして今 top priority か? 他の候補と比較したか?
- Portfolio Top-N に対して機会損失を起こしていないか?
- Persistence profile は task class に妥当か?
- Anti-pattern registry の同一 fingerprint を warn していないか?
- Fallback lane #3 (advisory critic) を過去 5 サイクル呼んだか? (呼んでなければ dispatch 対象に加える)
Worker (approach sanity check)
Final artifact commit / handoff 前に自問:
- 選んだ approach は anti-pattern registry の count=3 に触れないか? (触れるなら別 approach)
- Reversible な操作か? Backup-First 手順は取ったか?
- Real-surface で verify 可能か? mock だけで完了扱いにしていないか?
- Approval bucket は auto か security か? security ならなぜ user 承認を skip して commit しようとしているか?
- Attempt-log に verdict と evidenceRef を記録したか?
Reporter-Learner (confirm bias check)
Learning append 前に自問:
- 成功事例だけを拾っていないか? 直近 failure fingerprint を registry に反映したか?
- Evidence は "実測 / 推定 / 仮説" に分けたか?
- Next queue seed は learning から演繹しているか? それとも既存 candidate の焼き直しか?
- Diminishing returns / persistence 閾値超過を見落としていないか?
- Critic verdict (Layer 2/3) を learning materials に統合したか?
Workflow-Review (meta critic)
Cadence / prompt / gate 変更提案前に自問:
- 提案は直近 3 サイクル以上の evidence を持つか? (単発の思いつきではないか)
- 外部 evidence (別 workspace / documented pattern / benchmark) を 1 件以上参照したか?
- 変更は reversible か? Rollback 手順を pair で用意したか?
- CriticLog の verdict 分布を分析したか? (Rubric 調整の signal がないか)
- 変更 apply 後の smoke test / validate script pass を必須にしたか?
Layer 1 Output Contract
各 role は checkpoint 完了で dashboard-state.criticLog に append:
{
"ts": "ISO-8601",
"layer": 1,
"role": "commander|worker|reporter-learner|workflow-review",
"task_id": "...",
"questions_passed": 5,
"questions_failed": [],
"verdict": "proceed|revise|escalate-to-layer2",
"note": "short"
}- Verdict =
proceed: そのまま次アクション - Verdict =
revise: role 内で 1 回だけ再検討、再 checkpoint - Verdict =
escalate-to-layer2: fallback lane #3 (Critic Role invocation) を明示 request
Layer 3 Blocking Critic (重要 gate) — SSOT
このリストが Layer 3 blocking gate の SSOT。他 references (tunable-defaults.md, prompt-self-improvement.md, workspace-setup.md, SKILL.md) はこの節を参照する。gate 追加/変更は本節でのみ行う。
以下の task は Layer 3 を 必ず経由 させる (Layer 1 だけでは不十分):
- Portfolio candidate の promote (Top-N 昇格)
- Focus theme apply (3 ヶ月更新の確定 — 期間は tunable、apply gate は hard rule)
- Prompt self-improvement の commit (壊れた prompt 抑止)
- External publish 承認直前 (Marketplace / npm / SNS 等)
- Small-Bet Pilot → 本実装 昇格
上記 5 種は hard rule (変更禁止)。詳細: references/tunable-defaults.md の Hard Rule 一覧。
Repair re-review is not a sixth Layer 3 gate. It is a separate quality-continuation subgate that starts only after a blocking finding or required fix, and its different-family requirement is governed by the Repair -> Re-review Contract below.
Layer 3 契約
- Critic Role を新規 context (chat thread 分離 or subagent) で起動
- Input: 対象 artifact + 判定 rubric (下記) + 関連 anti-pattern entries
- Output:
critic-report.md(verdict: pass | reject | conditional) - Verdict =
pass: proceed - Verdict =
reject: task 停止、security-approve経由で user override 可 - Verdict =
conditional: 修正条件明示、worker が対応してから re-critic
Repair -> Re-review Contract
This contract starts only for (a) a Layer 3 conditional or reject, or (b) a queue kind review whose ## required fixes contains at least one stable finding ID. Layer 2 remains advisory: it may suggest a repair but does not make the task blocking.
Two Separate Counters
- A critic consultation round is the short producer-critic exchange. Its existing 3-round anti-oscillation cap remains local to that consultation.
- A workflow repair round is durable work that may span scheduled runs: repair one finding subset, validate it, then re-review the changed artifact. The default cap is 10 workflow rounds.
- When the commander creates and claims a repair atomically, it reserves one parent attempt and writes
nextState: repair-started. Deterministic validation, whether pass or fail, completes that same reserved attempt. Its effective cap ismin(reviewRepairRounds, profile.maxIteration - iterationsUsed). - If the repair claim expires or the worker fails before validation, the health reconciler finalizes the same round as
repair-start-failed, preserves partial artifacts, and consumes the reserved parent attempt. It may requeue only while the effective cap remains; it never silently returns the repair to pending. - Ten is a fail-safe, not the primary control. If the same blocking finding remains unresolved for two eligible workflow rounds, or the repair has no substantive artifact and evidence change, use the existing commander replan path before another repair task.
- If
profile.maxIteration - iterationsUsed <= 0, do not create a repair. Transition directly todeferred-exhaustedwithreason: persistence-exhausted, distinct fromreview-exhausted.
Durable Round Record
Append one record to dashboard-state.criticLog for every review, repair, re-review, and recovery decision. Retain the existing critic fields and add:
{
"parentTaskId": "task that owns persistence",
"workflowRound": 1,
"inputHash": "sha256",
"outputHash": "sha256 or null",
"findingIds": ["R1", "R2"],
"findingResolution": [
{
"id": "R1",
"state": "resolved|withdrawn|not-reproducible",
"evidenceRef": "path"
}
],
"validationResults": [
{
"id": "AC-1",
"expected": "...",
"actual": "...",
"result": "pass|fail",
"evidenceRef": "path"
}
],
"repairTaskId": "optional task id",
"receiptSource": "adapter-execution-record|scheduler-history|subagent-receipt",
"receiptRef": "immutable adapter or harness receipt",
"receiptHash": "sha256",
"nextState": "repair-started|repair-start-failed|repair-planned|validation-failed|replan|blocked-independence|parked-independence|overridden-independence|deferred-exhausted|complete|rejected",
"reason": "short",
"evidenceRef": "path or hash"
}Finding IDs remain stable across workflow rounds. A changed hash is necessary but insufficient: each blocking finding needs resolved, withdrawn, or not-reproducible with evidence, and re-review must confirm that resolution before pass.
repair is a child task, never a separate persistence target. It carries parentTaskId but has no independent profile or attempt budget. The parent attempt is reserved at repair claim, then finalized as validation pass/fail/start-failed; blocked-independence consumes no additional attempt. Every acceptance check must report a machine-comparable expected, actual, result, and evidenceRef; prose alone cannot close a repair round.
nextState is one of repair-started|repair-start-failed|repair-planned|validation-failed|replan|blocked-independence|parked-independence|overridden-independence|deferred-exhausted|complete|rejected. Terminal states are repair-start-failed|complete|deferred-exhausted|rejected; every other state remains open and is excluded from criticLog rotation.
State Transitions and Recovery
conditionalorrequired fixes: create onerepairtask for a fixed finding subset with acceptance checks. The worker changes the artifact, runs deterministic validation, records its output hash, validation results, and finding resolutions, then dispatches the required independent re-review. A failed or missing acceptance check recordsnextState: validation-failed, consumes the parent attempt, and returns to repair; it cannot be explained away in prose.reject: do not retry the same approach as a repair. The commander may use the existingreplanpath only with a new hypothesis and new evidence; otherwise preserve the rejection.- If re-review confirms every blocking finding ID resolved, mark the task complete. A repair worker never self-certifies its own output.
- Repair re-review is a blocking quality gate: it requires a new context and
independenceVerdict: different-family. The commander derives that verdict from an adapter or harness receipt identified byreceiptSource, immutablereceiptRef, andreceiptHash: producer and critic models must differ, both families must be known and differ,familyResolvermust identify the resolver, and router / auto models are ineligible. A worker may carry candidate fields but cannot author the receipt. Missing, null, same-family, unresolved, degraded, or self-declared-only evidence isblocked-independence. - A
blocked-independencerecord does not consume a workflow repair round, but it is not free: count it byparentTaskIduntil a valid independent re-review occurs. The configuredindependenceBlockLimit(default 3, allowed 1-5) transitions the parent toparked-independence, appends asecurity-approvedecision topendingApprovals, and stops further scheduled re-review dispatch. A user may approveoverridden-independence, but its quality verdict never becomespass; it queues exactly one follow-upreviewtask with the same parent and unresolved finding IDs when an eligible independent critic is available. Override does not reset the independence block counter. - At the effective cap, transition to
deferred-exhaustedwithreason: review-exhausted, preserve unresolved finding IDs, and continue with a safe fallback task. A human override may advance an approved boundary but never rewritesconditionalorrejectintopass. - After an interrupted run, reconcile
outputHashwithcriticLog. If repair output exists without re-review, resume the same workflow round without incrementing it. If the repair never reached validation, finalizerepair-start-failedfor the already-reserved attempt before deciding whether a bounded requeue is allowed. If upstream evidence or the reviewed input hash changed, discard the pending finding set and create a new review record before repairing.
re-review-after-substantive-repair, durable workflow-round records, and no self-certification are hard invariants. Only the numeric cap is tunable.
Rubric (severity 5 段階)
段階名 (info/minor/major/critical/blocking) は hard rule (verdict 比較保護のため変更禁止)。各段階の判定基準は tunable、workflow-review が調整可。
| Severity | 基準 | Verdict 影響 |
|---|---|---|
| info | 参考情報、改善余地 | pass 可 |
| minor | 軽微、非阻害 | pass 可 |
| major | 修正推奨、非致命 | conditional |
| critical | 修正必須、実害あり | reject |
| blocking | 反対理由あり、進めるべきでない | reject |
- Critic は severity 分布を artifact 末尾に必ず出力
- Workflow-review が criticLog の severity 分布を追跡し、rubric 調整の signal を検出
Independence 契約 (Layer 2/3)
同じ LLM が producer と critic を演じても、context 分離が不十分だと bias が残る。
- Critic dispatch 時は 完全に新規 context (chat thread 切替 or subagent 呼び出し) で起動
- Persistence = Exhaustive の task では duck-critic skill 経由の別モデル critic を default
- Layer 3 は別モデルファミリの critic を必須とする (context 分離だけでは不十分)。同一 model が producer と critic を兼ねると、自分の成果物は価値あるように見えて gate が室温化する。同じ family 内の別 tier (例: Opus と Sonnet) も同一扱いとする
- Producer / critic の model 名と family 判定結果を
critic-report.mdとcriticLogに記録する (producerModel/criticModel/producerFamily/criticFamily/familyResolver/independenceVerdict)。自己申告の文字列一致ではなく family 判定で見る。family 不定の model (auto / router 系) は critic に使えない - Layer 3 は fail-closed。
independenceVerdictがdifferent-family以外 (same-family/unresolved/degraded) のとき、Layer 3 は verdict = pass を出せない (= proceed 不可)。独立性は rubric severity とは別の前提条件で、verdict 語彙を増やさない。degradedは Layer 2 専用で、Layer 3 では user 明示 override (security-approve 経由) がない限り進めない - Blocking gate を消費する側 (portfolio promote / focus theme apply / prompt-self-improvement commit / external publish) は、
verdict = passだけでなくindependenceVerdict = different-familyも確認してから進む - Critic verdict は producer の context には戻さない、
critic-report.mdを artifact として handoff
Iteration 抑制
Layer 1 は毎 role 1 回、Layer 2 は fallback lane #3 経由、Layer 3 は 5 gate のみ。iteration 倍増を避けるため:
- Layer 1 は 1-5 質問だけ、深掘りしない
- Layer 2 は 1 artifact に 1 回だけ (再 critic は artifact hash 変更時のみ)
- Layer 3 は上記 5 gate に限定、他 task は Layer 1 / 2 で完結
- Blocker Test (fallback-lane.md 参照) の 4 問 gate は Layer 1 と別、blocker 認定専用
Override (Layer 3 blocking)
Blocking verdict = reject でも user は override 可:
- User が
override --reason "..."を dashboard 経由で送る - approvalLog と criticLog 両方に reject verdict + override reason を残す
- Override 後の task は特別 tag
overridden-critic-rejectで追跡 - Workflow-review が reject override と
overridden-independenceを合わせた override 頻度を監視 (5% 超過で rubric 見直し提案)
Log Retention
- Layer 1 criticLog: 30 日で rotation (直近だけ意味あり)
- Layer 2/3 critic-report: 90 日、archive に押し出す
- Open repair criticLog records(terminal state 以外)は terminal state になるまで rotation しない
- Rubric 調整 event: 恒久 (workflow-review 学習資産)
See Also
references/fallback-lane.md: Layer 2 (fallback lane #3 Advisory Critic)references/persistence-profile.md: Exhaustive task の critic 強制references/prompt-self-improvement.md: Layer 3 commit gatereferences/approval-policy.md: security-approve 経由 overrideassets/prompts/*.md: 各 role prompt に Layer 1 checkpoint を埋め込む