Inline Recipes — kaizen / essential / killer / trim
Purpose: Full phase contracts for the four Recipes that have no separate reference file. SKILL.md ## Subcommand Dispatch keeps only a summary; this file holds the complete contracts.
Read when: Executing one of kaizen, essential, killer, or trim Recipes.
kaizen — Multi-axis iterative improvement of an existing feature
Kaizen is a PDCA loop, not a single pass. The contract below runs Phase 3↔4 as a bounded improvement cycle against a quantified target, and stops on target-met OR diminishing-returns — never "one improvement and done". Default Mode:
AUTORUN(each cycle is scope-bounded and reversible); escalate toGUIDEDwhen an axis touches 10+ files or structural boundaries (then Magi confirms the axis plan before Phase 3). Codex owns code-gen (Bolt/Tuner/Zen/Sweep/Artisan/Builder); Claude owns judgment/measurement (Lens/Pulse/Echo/Magi/Radar-gate/Guardian) — engine-routing principle shared by all four recipes below (§ Contract elements).
- Phase 1 DIAGNOSE + BASELINE (parallel) — Lens[claude map-current-implementation] unconditionally; conditionally add Pulse[claude KPI-measure] if metrics instrumentation exists, Echo[claude UX-walkthrough] if UI surface, Voice[claude sentiment]/trace[claude session-replay] if user-feedback or session data is available. Capture a quantified baseline per candidate axis (perf: latency/render numbers; code-quality: complexity/coverage; UX: friction/task-success; feature-extension: usage gap). Non-instrumented surfaces → record a qualitative baseline (Echo/Voice rubric score). Goal: multi-signal picture with a measurable starting point to compare against later.
- Phase 2 PROPOSE + TARGET GATE (sequential) — Spark[claude improvement-spec, constrained to enhancing existing data/logic — not new feature ideation] → Magi[claude axis-prioritize] selects one or two axes from
{perf, UX, code-quality, feature-extension}AND, for each, fixes a stop criterion: a target value (e.g. "p95 < 200ms", "coverage ≥ 80%") and a loop cap (loop ≤ 3 cycles (default 3), per recipe-contract §2). Rejects "improve everything" plans; rejects axes with no measurable target ("improve without measuring" cannot enter Phase 3). - Phase 3 IMPROVE (axis-bounded, parallel within axis) — perf → Bolt[codex frontend]/tuner[codex explain]; UX → Palette[codex usability]/prose[codex microcopy]/flow[codex motion]; code-quality → Zen[codex refactor]/sweep[codex dead-code]; feature-extension → Artisan[codex component]/builder[codex api]. Independent sub-axes parallel; dependent ones serialize.
- Phase 4 VERIFY + CROSS-AXIS GUARD (the loop gate) — Radar[claude-gated regression] blocks any regression of existing behavior. Re-measure the Phase 1 baseline metric for the worked axis (Pulse/Echo Before/After). Cross-axis guard: confirm the improvement did not regress a non-worked axis (e.g. a perf change that hurt UX, or a refactor that dropped coverage) — if it did, the cycle fails the gate. Then branch:
- Target met (metric reached the Phase 2 stop value, no regression) → exit loop → Phase 5.
- Target not met AND iterations remain AND last cycle's marginal gain ≥ threshold → loop back to Phase 3 for the next cycle (carry forward the delta, not a fresh diagnose).
- Diminishing returns (marginal gain below threshold,
Δ < ε) OR cap-reached (per the Phase 2 stop criterion) → Void[claude] confirms the stop, records "remaining gap vs target", → Phase 5 with partial-improvement note. Never burn cycles past the point of marginal value.
- Phase 5 SHIP — Guardian[claude PR-prep] produces PR with embedded Before/After report: baseline → final per axis, cycles run, stop reason (
target-met|diminishing-returns (Δ < ε)|cap-reached), and any residual gap. - Boundaries: vs
refactor(internal-only, no external delta) — kaizen explicitly improves externally-observable quality alongside internal hygiene. vsoptimize(perf-only) — kaizen treats perf as one axis. vsfeature(new capability) — kaizen polishes a shipped feature. vsapex/loop recipes — kaizen's loop is scope-bounded to chosen axes against a fixed target, not open-ended discovery. - Anti-patterns prevented (additional to § Failure Modes Prevented below, which owns the measurable-target and diminishing-returns mitigations): (1) "rewrite the whole module under improvement banner" (Magi axis-cap), (2) "improvement that regresses something else" (Phase 4 Radar + cross-axis guard), (3) "single-pass masquerading as kaizen" (Phase 3↔4 loop with quantified stop criterion).
- Add-ons: +Scout for deeper root-cause when Lens insufficient, +Atlas for structural change, +Ripple for cross-module impact before committing to an axis, +Experiment to A/B-validate a user-facing improvement against the baseline when traffic allows.
essential — Single must-have verdict + conditional implementation
- Phase 1-4 Verdict (sequential refinement funnel) — Echo[demand: claude pain-extraction] → Spark[claude spec] → Magi[claude necessity-arbitration] → Rank[claude MoSCoW-must]. Each step narrows the previous output; parallelization would force redundant re-synthesis. Subtraction-oriented — Magi's Sophia filters "Should-have" posing as "Must-have".
- Convergence rule: Rank's MoSCoW output filtered to the single top Must-have (highest necessity score). Tie-break: when candidates tie on necessity, Magi prefers the one that is necessary AND distinctive over the safest generic option (essential ≠ timid); still tied → escalate to user via AskUserQuestion with tied candidates.
- Ambition check (essential ≠ timid): "must-have" = the feature without which the product fails — that is sometimes a bold, differentiated bet, not the blandest table-stakes item. Subtraction targets scope, not ambition: Magi's necessity-arbitration must reject the default failure mode of converging on an obvious commodity feature when a bolder must-have would better serve the core job. If the verdict is a safe table-stakes feature, the card's "Why" must justify why a bolder must-have lost, not merely assert minimality. Optional +Flux in Phase 2-3 to reframe the must-have candidate boldly before Magi arbitrates.
- Phase 5 DELIVER verdict via AskUserQuestion — verdict card + Yes/No/Modify branches per
reference/verdict-gate.md(the card is contract-level, stops even under AUTORUN). Recipe-specific card fields:Recommended must-have: <single feature>,Source of conviction: Echo[demand]→Spark→Magi→Rank summary(no flag/KPI — essential ships table-stakes, not a bet). - Phase 6 Conditional Implementation (the Yes branch's build): Sherpa[claude atomic-decomposition] → Builder[codex] → Radar[codex] → Guardian[claude] → DELIVER working feature + tests + PR, inheriting the single-feature scope constraint (verdict-gate.md §2) — Codex owns code-gen, Claude owns judgment (§ Contract elements).
- Failure mode prevented (additional to § Failure Modes Prevented below, which owns the unbounded-build mitigation shared with
killer): over-engineering (Phase 1-5) + timid-convergence (the Ambition check stops the funnel from defaulting to a commodity must-have when a bolder one is genuinely more essential). - Add-ons: +Void for aggressive scope cut, +Scribe[unified] for atomic-unit specs in Phase 1-4, +Flux to reframe the must-have boldly before Magi's necessity-arbitration.
killer — Single differentiator verdict + conditional flagged implementation
- Phase 1 (parallel hub-spoke, cross-engine triangulation, dual-engine baseline) — Default baseline distributes Phase 1 branches across both Claude and Codex to preserve perspective independence at the model-priors level (not just prompt-frame level): Compete[claude + WebSearch tool for current market gap-analysis, framed as "industry analyst"] ‖ Flux[codex sandbox-execution priors, framed as "what would the market gap look like in code / infrastructure / developer-experience terms" — Codex's GitHub-heavy training surfaces gaps a market-focused model misses] ‖ Echo[demand: claude empathy/latent-needs, framed as "user advocate"]. Optional agy lift (when AVAILABLE at PREFLIGHT): Compete adds a second branch on agy (Search-grounded for fresher market data than Claude's training cutoff), Flux adds a second branch on agy (the verified authorized model at High effort tier for cross-domain analogy generation — the verified authorized model mandated,
_common/CLI_COMPATIBILITY.md §4 ‡, no Deep Think) — agy's training-data priors give an additional independence axis. Failure mode prevented: model-monoculture in the triangulation step. Engine-attribution tags:[claude-compete],[codex-flux],[claude-echo[demand]]for the dual-engine baseline; add[agy-search],[agy-flash-high]when the optional lift is active. - Phase 2 (sequential synthesis) — Spark[claude] aggregates the independent perspectives into the single most decisive killer feature (one feature, not a ranked list). Pick the boldest viable differentiator, not the safest one — a "killer" feature that any competitor would also obviously build is table-stakes, not a killer. If synthesis regresses to an incremental, easily-copied feature, it has failed the recipe's purpose; the cross-domain (Flux) and latent-need (Echo[demand]) inputs exist precisely to push past the obvious. Carry the boldest candidate into Phase 3's gate rather than pre-filtering it for "riskiness" — the gate, not Spark, decides survival.
- Phase 3 DEFEND + VERIFY (the killer-grade gate — a decisive feature that can't survive this is just a nice-to-have) —
- Moat check — Compete[claude] classifies defensibility: time-to-copy (can a competitor replicate within one release cycle?) and moat class (data / network-effect / integration-depth / switching-cost / brand / buildable-emergent / none). Distinguish "no moat" from "moat forms after launch." A novel feature with no current moat but a credible buildable/emergent moat (first-mover data accrual, network effects that compound with adoption, switching-cost that grows post-launch) is NOT auto-downgraded — it passes as GO-with-flag, provided the Phase 5 rollout plan carries a moat-building milestone. Only
moat=none AND no path to a moat AND time-to-copy < one cycledowngrades from "killer" to "nice-to-have." Magi weighs moat class and trajectory (not just present-tense moat) as a first-class GO criterion. Rationale: the most ambitious killer features are novel precisely because no one has built the moat yet — penalizing them for lacking a moat that can only form post-launch is the conservatism trap this recipe must avoid. - Adversarial refutation (refute-polarity, high-stakes) — run the skeptic panel per
_common/ADVERSARIAL_REFUTATION.md(2-3 cross-engine skeptics, evidence-vs-novelty discipline, default-to-refuted-on-evidence-claims-only, majority-on-evidence aggregation, GO-with-flag for merely-unproven-because-new). Killer-specific attack angles: "this gap is already served by<X>" (market), "users don't actually want this enough to switch" (demand), "infeasible at our scale/timeline" (delivery). Survival feeds the moat verdict below — a claim that survives evidence-based refutation but is unproven-because-new is exactly the killer bet the moat/trajectory gate protects. - Then Magi[claude] binary Go/No-Go incorporating moat class + refutation survival.
- Moat check — Compete[claude] classifies defensibility: time-to-copy (can a competitor replicate within one release cycle?) and moat class (data / network-effect / integration-depth / switching-cost / brand / buildable-emergent / none). Distinguish "no moat" from "moat forms after launch." A novel feature with no current moat but a credible buildable/emergent moat (first-mover data accrual, network effects that compound with adoption, switching-cost that grows post-launch) is NOT auto-downgraded — it passes as GO-with-flag, provided the Phase 5 rollout plan carries a moat-building milestone. Only
- Convergence rule: Spark MUST synthesize one feature (not a ranked list), choosing the boldest viable differentiator; Magi delivers binary Go/No-Go gated on moat-class/trajectory + refutation-survival (survival defined per
_common/ADVERSARIAL_REFUTATION.md§2 — withstood evidence-based refutation; merely-unproven-because-new is not a refutation). Tie-break: Magi forces selection via strategic-impact criterion (market timing × differentiation depth × moat durability or buildability × feasibility); NO-GO still surfaces runner-up with "weakest-link" annotation. - Phase 4 DELIVER verdict via AskUserQuestion — verdict card + Yes/No/Modify branches per
reference/verdict-gate.md(always-flagged: killer ships a bet, so the flag clause of §3 always applies). Killer-specific card fields: perspective-attributed evidence (mark which frame produced which insight; note which branches were agy-backed vs Claude-frame-only); moat class + trajectory (current vs buildable-emergent) + time-to-copy + which refutations the feature survived, and whether any open risk is "refuted-on-evidence" vs "unproven-because-new"; Magi verdict (GO confidence H/M/L | GO-with-flag for buildable-moat/unproven-but-bold | NO-GO reason). When the verdict is GO-with-flag on an unproven-but-bold bet, the card states plainly that this is a deliberate aggressive bet whose kill-criterion is the real test. - Phase 5 Conditional Implementation (the Yes branch's build): Sherpa[claude decomposition] → if
ui_dimension != none: Forge[codex prototype-validation] → Artisan[codex frontend-production] → Builder[codex backend/logic] → Radar[codex edge cases for differentiator] → judge[claude multi-engine review — killer features are high-stakes] → Guardian[claude] with feature-flag recommendation for controlled rollout (differentiation risk) → DELIVER working feature + tests + PR + flag config + rollout plan. The flag carries the flag+KPI+kill structure ofreference/verdict-gate.md§3, where killer's differentiation KPI is the measurable form of the Phase 3 killer hypothesis (the adoption / retention / switching signal that proves the edge is real). Hand off togrowth-acceptancewhen the +14/+30/+90d measurement loop is warranted. - If No: per verdict-gate.md (auditable "decided-not-to-ship" record) — killer's record includes the moat verdict + surviving/failed refutations.
- If Modify: per verdict-gate.md (bounded loop-back, 2-Modify cap then escalate, carry refuted items forward as exclusions) — killer loops back to Phase 1 with the modification as an added constraint (e.g. "reframe around X constraint" → Flux re-runs with updated directive).
- Add-ons: +Flux for iterative deep-dive on Spark output in Phase 2, +Field for additional market trend grounding in Phase 1, +Omen for a pre-mortem on the differentiation bet before Phase 5 commits.
trim — Dead-weight feature removal verdict + conditional excision
trimis the inverse ofessential/killer. Where those decide what THE ONE feature to build is,trimdecides which existing features to remove. It applies the essential axis (is this a must-have for the core job?) and the killer axis (is this a defensible differentiator?) as a 2×2 filter against the live feature set: a feature survives if it is essential OR killer; only a feature that is neither — and carries real cost — becomes a removal candidate. Core engine isvoid(YAGNI / Feature Sunset / CoK / blast radius);trimadds the dual-axis judgment plus multi-agent execution that void's own propose-only recipes lack. Default Mode:AUTORUNwith the Phase 4 verdict gate (removal is semi-destructive); escalate toGUIDEDwhen any target's blast radius isPUBLIC_API/DATAor the slate touches 10+ files. Codex owns removal code-gen (Sweep/Builder/Radar); Claude owns judgment/evidence (PDM/Lens/Void/Magi/Compete/Guardian) — engine-routing principle shared by all four (§ Contract elements).
- Target resolution —
trim <target>scopes the audit to the named feature / module / area.trimwith no target → whole-project auto-scan (the proactive form of trim, analogous to/Nexusno-args proactive mode but scoped to removal): PDM[claude] builds the full feature inventory from specs / routes / feature-flags / nav / code, and Void[claude] ranks the entire surface by carrying cost, so Phase 1 sweeps every shipped feature for dead-weight candidates instead of waiting for a named target. No-target runs default to GUIDED (a whole-project removal slate is higher-stakes than a scoped one) and cap the Phase 4 slate to the top-N by CoK (default 10) with the dropped long-tail noted, never a silent truncation. - Phase 1 INVENTORY + EVIDENCE (parallel) — Build the candidate feature set (per Target resolution: the named scope, or the whole project when no target was given) and ground each one in real carrying-cost evidence. PDM[claude feature-inventory] or Lens[claude map-current-implementation] enumerates the existing features in scope; Void[claude] gathers per-feature usage telemetry / git change-frequency / bug-density / CoK (0-10). Conditionally add Trace[claude session-replay] / Voice[claude sentiment] when usage or feedback data exists. Cohort-segment the usage thresholds (void's rule): admin-only / operator / compliance features are judged against their own cohort, not total users — a
<1%-of-all-users feature can be healthy when the denominator is the wrong cohort. Goal: a feature list where each entry has a measurable carrying cost and a usage signal — never a hunch. (Evidence is the entry condition for Phase 2; "remove without measuring" cannot proceed.) - Phase 2 DUAL-AXIS SCORING (the essential×killer gate) — Score each feature on two independent axes:
- Essential axis — Magi[claude necessity-arbitration]: would removing this break the core job-to-be-done? (the
essentialtest, inverted). Table-stakes-but-necessary counts as essential even when unexciting — Sophia filters "load-bearing must-have" from "feature that merely feels important". - Killer axis — Compete[claude differentiation/moat]: does this feature carry a defensible differentiation / moat (the
killertest)? Classify time-to-copy + moat class (data / network-effect / integration-depth / switching-cost / brand / buildable-emergent / none). A quiet, low-usage feature that is a genuine differentiator passes the killer axis and is protected from pure-usage YAGNI. - Combine into a 2×2 verdict:
KEEPif essential OR killer;SIMPLIFYif essential-but-overbuilt (necessary core, bloated implementation);REMOVE-CANDIDATEonly if neither essential nor killer AND CoK ≥ 7 (void's strong-removal threshold). Never let a single axis alone trigger removal.
- Essential axis — Magi[claude necessity-arbitration]: would removing this break the core job-to-be-done? (the
- Phase 3 SAFETY + REFUTATION GATE (a removal that can't survive this stays) — the removal slate passes a Sentinel security guard, adversarial refutation ×2-3 (must-stay polarity), and a blast-radius check before any removal verdict: run the skeptic panel in defend-polarity per
_common/ADVERSARIAL_REFUTATION.md§3 (skeptics argue the feature must STAY; default-to-keep on evidence claims; majority must-stay-on-evidence → downgrade toKEEP/DEFER) and apply its §5 hard exclusions (safety-critical auth/encryption/input-validation excluded from removal without a Sentinel security review; confidence<60%→ do not propose;PUBLIC_API/DATAblast radius → Ask First). Trim-specific stay-angles: "a critical cohort silently depends on this" (load-bearing), "this is a compliance / contractual obligation" (mandate), "an integration or downstream feature breaks without it" (entanglement). Distinguish genuinely-dead-weight from low-usage-but-load-bearing — judged against the Phase 1 cohort-segmented usage, not total users. Classify blast radius per target (internal / team / PUBLIC_API / DATA). +Ripple here when a target's entanglement is unclear. - Phase 4 DELIVER verdict via AskUserQuestion — verdict gate + Yes/No/Modify branches per
reference/verdict-gate.md(incl. §4 destructive-verdict rules:PUBLIC_API/DATA/irreversible rows flagged in-card and confirmed before excision even under AUTORUN;SIMPLIFYroutes to kaizen/Zen). Trim-specific card shape — a Removal slate (not a single-item card), one row per candidate:feature / essential? (Y/N + why) / killer? (Y/N + moat class) / CoK / blast radius / evidence (usage %, last-meaningful-change, bug density) / verdict (REMOVE | SIMPLIFY | KEEP-WITH-WARNING) / reversibility (flag-off vs hard-delete). Header summarizes count + estimated carrying-cost reclaimed (hours/sprint, lines, deps). - Phase 5 Conditional Excision (only if Yes) — For confirmed
REMOVEtargets: Sherpa[claude phased-decomposition — one feature per atomic step, flag-off before hard-delete when reversibility matters] → Sweep[codex dead-code/file deletion] ‖ Builder[codex remove feature code / routes / flags / config] → Radar[codex verify-green-after-removal — no regression in kept behavior] → Guardian[claude PR with removal report: what was cut, carrying cost reclaimed, reversibility/rollback note, and the essential×killer verdict per item]. Phased, small-scope removals (void's ≈60% fewer-regression-bugs rule) — never a big-bang multi-feature delete in one commit.SIMPLIFYverdicts route tokaizen/zen, not deletion. - If No: per verdict-gate.md (auditable "decided-to-keep" record) — trim embeds the dead-weight audit (essential×killer verdict + CoK + evidence per item) so the decision is re-evaluable later.
- If Modify: per verdict-gate.md (bounded loop-back, carry rejected items forward as exclusions) — trim captures which features to spare or add and loops back to Phase 2 with the adjusted set as the new scope (carry forward Phase 1 evidence; re-score only changed items).
- Boundaries: vs
essential(picks THE ONE must-have to build) /killer(picks THE ONE differentiator to build) — trim is their inverse, finding existing features to remove using the same two axes as the filter. vskaizen(improves a kept feature against a target) — trim removes rather than polishes; aSIMPLIFYverdict hands off to kaizen. vs void'sprune/cutrecipes (single-agent proposal, no execution) — trim adds the essential×killer dual-axis judgment and multi-agent execution (Sweep/Builder/Radar/Guardian). vsmigrateDECOMMISSION (removes replaced old code after a parity-proven migration) — trim removes unjustified features regardless of any replacement. - Anti-patterns prevented (additional to § Failure Modes Prevented below, which owns the load-bearing-removal mitigation): (1) removing safety-critical code (Phase 3 safety exclusion), (2) hunch-based removal (Phase 1 evidence is a mandatory entry condition), (3) single-axis over-deletion (REMOVE requires failing both axes, not just low usage — the killer axis protects quiet-but-defensible features), (4) sweeping big-bang deletion (Phase 5 phased small-scope + flag-off-before-delete).
- Add-ons: +Ripple for blast-radius/impact before committing a target, +Omen for a pre-mortem on the full removal slate before Phase 5, +Trail to confirm a low-usage feature isn't a recently-added one still ramping (don't trim before it's had a fair chance), +Flux to reframe "is this dead weight, or just undermarketed / undiscovered?" before condemning it.
Contract elements shared by all four
Per reference/recipe-contract.md §1 — the elements these four inline recipes hold in common, so each section above states only its own specialization:
| # | Element | Contract |
|---|---|---|
| 2 | Engine routing | All four follow summit principles: Codex owns code-gen, Claude owns judgment/measurement (reference/summit-recipe.md); per-recipe agent-role mapping is stated in each recipe's own section above. |
| 3 | Resume | N/A for essential / killer / trim — each is a short verdict pass (2-5 agents, one forward run); re-invoking is cheaper than checkpointing. kaizen uses checkpoint-resume: the baseline measurement and each cycle's rubric scores persist at cycle boundaries, so an interrupted run resumes mid-loop with its trajectory intact. |
| 4 | Output report | kaizen → named Before/After Report (baseline vs final per axis, cycles run, exit reason, residual gap). The three verdict recipes → the named Verdict Card defined in reference/verdict-gate.md (verdict · evidence · Yes/No/Modify branch · flag + KPI + kill criteria where applicable). |
| 5 | Failure Modes Prevented | See table below. |
| 7 | Scale | kaizen 4-10 agents × ≤3 cycles, mid cost. essential / killer / trim 2-5 agents, single pass, low cost — the verdict is cheap; only the conditional implementation that follows a Yes carries real cost, and it is scoped by the branch taken. |
Failure Modes Prevented
| Failure | Mitigation | Applies to |
|---|---|---|
| "Improve it" with no measurable target → taste-driven churn | Baseline measurement is the entry condition; the rubric target is fixed before any change | kaizen |
| Endless polish past the point of value | loop ≤ N cycles (default N=3) + diminishing-returns (Δ < ε) exit; non-ACCEPT exits report best-so-far |
kaizen |
| The producer grading its own improvement | Generator-Evaluator separation per reference/evaluator-loop-protocol.md |
kaizen |
| A verdict asserted from intuition rather than evidence | Verdict Card requires cited evidence per reference/verdict-gate.md; an unsupported verdict cannot be issued |
verdict trio |
| Confirmation bias toward the answer the asker wants | Adversarial refutation panel per _common/ADVERSARIAL_REFUTATION.md (polarity set against the proposed verdict) |
killer, trim |
| A "Yes" silently becoming an unbounded build | The Yes branch is conditional implementation, scoped by the card — killer additionally ships behind a feature flag with a stated KPI and kill criterion |
essential, killer |
| Removing something still load-bearing | trim requires a dependency/usage check before excision; an unresolved dependency turns the verdict into Modify, never Yes |
trim |
| A verdict that neither ships nor closes | Every card ends in exactly one of Yes / No / Modify — an undecidable case exits BLOCK with the missing evidence named |
verdict trio |