Nexus Task Battery — Routing Verification Suite
Purpose: Standing regression battery for the internal CLASSIFY phase (RESOLVE → GATE → MULTI? → REDIRECT? → SELECT → LADDER? → CHAIN_SELECT). Re-run this battery whenever routing-matrix.md, signal-keywords.md, or the Recipe registry in nexus/SKILL.md changes, to prove routing behavior didn't regress.
Read when: Verifying a routing-machinery change (LADDER wiring, Recipe additions, Signal Keyword edits) before merge, or re-running the Wish #1 harness regression check.
Verification depth: All items verified at routing/chain-selection level (does the item resolve to the expected chain — no full execution required). Items marked [E2E] are candidates for real end-to-end execution when validating the harness end-to-end, not just its routing decisions. Executed status: Cycle-1 ran 2 live end-to-end probes matching two [E2E] items' shapes — Probe-Redirect (item 7's shape: "improve the design" REDIRECT walk) and Probe-Ladder (item 29's shape: full LADDER walk on the USPTO patent-filing out-of-coverage case). Both are marked [E2E executed] below; the remaining [E2E] items (5, 12) are [E2E pending] — routing-level only so far, real end-to-end execution not yet run against them.
Battery (64 items — IDs run 1-65 with no #14)
What each block exists to prove: 1-30 Recipe-Family routing + ambiguous/edge inputs (original battery); 31-35 2 ambiguous-anchor stress tests + 3 extra out-of-coverage cases, broadening the dim-3 Coverage-closure proof past two data points; 36-37 FIGURE_CHANNELING named-figure routing; 38-39 the eureka invention/build split (ship=true fires only on explicit build intent, never inferred); 40-43 the Quality-Max design wing (crucible/silhouette/lattice/chorus) — each proves it is not absorbed by its nearest neighbor and that its precondition fires; 44 assay — the prove-vs-fix boundary against anneal, the untested-vs-unnecessary distinction, and that apply=true is never inferred; 45-46 quell — the external-reviewer-to-zero oracle and its behavior-preserving profile; 47-48 burnish — the same oracle over a rendered surface, proving the split oracle fires and that profile=faithful blocks rather than degrades without a reference of record; 49-51 the rest of the _common/FINDING_LEDGER.md family — whet (a mutation engine as the evaluator, and the three dishonest closures its integrity rules block) and security mode=to-zero (a scanner sweep, and that RISK-ACCEPTED is not a free exit); 52 optimize mode=to-zero; 53-56 verity; 57-59 abide; 60-61 PROMPT_SPEC; 62-64 gated SPECIFY; 65 deliver — product-from-scratch routing with over-capture guards against feature and apex. The per-item rows below carry the specific neighbor and precondition for each.
Numbering note: items 45-46 previously carried the IDs 40-41, colliding with the Quality-Max design-wing rows that the block description above assigns to 40-44. They were renumbered when 47-48 were added; _common/scripts/task-battery-check.py tracks the current IDs.
Items 40-44 — routing-layer E2E executed (Probe-DesignWing). 10 blind probes, each a fresh subagent with no authoring context, given only the user phrasing and told to route via SKILL.md: 5 positive (one per new recipe) + 5 negative over-capture probes asserting the new recipes do not absorb their neighbors' work — "the design drifted, brush it up" → anneal (not assay) · "polish the UI, make it look better" → restyle (not silhouette/crucible) · "apply the finalized brand book everywhere" → rebrand (not lattice) · "add aria-labels and a focus ring" → palette direct (not crucible) · "plan the web→iOS/Android port with a parity matrix" → port (not chorus). Result: 10/10 correct, zero over-capture; ~622k subagent tokens, 10 agents. Scope of the claim: this verifies the routing layer only — the instrument layers (crucible condition cells, silhouette attribution panels, lattice residue scan, chorus two-sided gate, assay four instruments) remain UNVERIFIED, because none of the five preconditions can be satisfied by this repo (no running UI, no product surface/brand, no token system, no second platform, no test suite). Instrument-layer E2E requires a target repository. Author-bias residual: the probe phrasings were written by the same author as the recipes; an independent phrasing set would strengthen the result. C's ecosystem-l10n item was explicitly discarded (S9: "partial overlap" with polyglot, not a clean out-of-coverage case).
| # | Input | Family | Expected routing |
|---|---|---|---|
| 1 | /nexus bug — "login returns 500 after deploy" |
Fix | bug Recipe directly (subcommand match) |
| 2 | "there's a memory leak in the worker pool" | Fix | CLASSIFY phase → Signal Keywords → bug (memory leak carve-out, not optimize) |
| 3 | "CVE-2026-xxxx in our lodash dep" | Fix | security |
| 4 | "clean up this module, no behavior change" | Improve | refactor |
| 5 | "make the dashboard load faster" [E2E pending] | Improve | optimize (measure-first) |
| 6 | "polish the checkout flow" | Improve | kaizen (overloaded polish/improve → REDIRECT per signal-keywords.md, resolves to kaizen given "existing feature, multi-axis") |
| 7 | "improve the design of this component" (ambiguous) [E2E executed — Probe-Redirect] | Improve | REDIRECT one-question: code-design → anneal, UI/look-and-feel → restyle |
| 8 | "audit this codebase's architecture for weaknesses" | Improve | anneal |
| 9 | "redesign the settings screen, make it modern" | Improve | restyle |
| 10 | "run this until it's done, don't stop" | Loop | shape-resolve per _common/LOOP_PRECONDITIONS.md → goal/converge/project-local orbit/apex; unavailable orbit falls back to goal or apex, and the five-point gate runs at the resolved owner |
| 11 | "set up a goal to keep the test suite green nightly" | Loop | goal (checks for machine-checkable oracle + hard-stop bound) |
| 12 | "add OAuth login" [E2E pending] | Build | feature |
| 13 | "build this whole idea end to end, 8-25 agent budget ok" | Build | apex |
| 15 | "spec out a notifications feature before building" | Discover→build | spec (stops at spec; pairs with feature) |
| 16 | "give our repo a self-driving team+work plan" | Discover→build | charter (stops at doc; pairs with enact) |
| 17 | "think through whether microservices are worth it here, no code" | Reason | gedanken |
| 18 | "evolve the checkout feature — no code, just direction" | Reason | delve (overloaded "evolve a feature" → REDIRECT resolves here since no-code stated explicitly) |
| 19 | "map how this repo works across the 3 services" | Comprehend | cartograph |
| 20 | "how did this codebase get the way it is, from git history" | Comprehend | chronicle |
| 21 | "what's the one must-have feature here" | Verdict | essential |
| 22 | "what feature is dead weight, safe to remove?" | Verdict | trim |
| 23 | "clone this competitor's onboarding flow" | Reproduce/Synth | clone |
| 24 | "this is a once-in-a-lifetime ask, spare nothing" | Quality-Max | wish (always-confirm gate fires) |
| 25 | "give me the best possible design for our flagship screen" | Quality-Max | runway |
| 26 | "generate a full research + legal + saas doc package" | Document package | package |
| 27 | /nexus with no arguments |
Meta | proactive mode (non-Recipe) |
| 28 | "switch to the security-focused skill profile for this sprint" | Meta | pack |
| 36 | "how would Feynman approach debugging this?" | Thinking (named figure) | FIGURE_CHANNELING (Magi[advisor]) — a real named individual is the trigger; use the advisory Recipe because no verdict was requested |
| 37 | "get me a panel of Buffett and Munger on this pricing decision, then tell me what to do" | Thinking → Decision | FIGURE_CHANNELING chain decide: Magi[advisor] → Magi[decide] → Builder — the first Recipe contrasts the named experts; the second issues the requested verdict |
| 29 | [out-of-coverage, E2E executed — Probe-Ladder] "negotiate and file the actual patent application for this algorithm with the USPTO" | none of 90 global | CLASSIFY phase → REDIRECT fails, SELECT fails (no matrix row: patent prosecution is not code/design/research-artifact work any currently registered skill claims) → LADDER: compass(recommend) returns Gap mode → architect gap-fill proposal artifact presented to user (e.g., "no patent/IP-filing skill exists; nearest partial-fit is canon[legal] for ToS/Privacy legal review, not patent prosecution — propose new skill or route to human patent counsel"); fallback_taken: architect-invoked |
| 30 | [out-of-coverage] "book flights, hotels, and a group dinner reservation for the 12-person team offsite" | none of 90 global | Same LADDER path — compass Gap mode (no travel/logistics-booking skill in the current catalog) → architect proposal or explicit "route to a human travel coordinator; no in-repo skill covers real-world booking transactions" — never silently answered as generic travel advice; fallback_taken: architect-invoked |
| 38 | "invent a genuinely new way to do X — nobody has solved this" | Reproduce/Synthesize & Invent | eureka default (ship=false) — dossier only. The invention ask alone must NOT be read as a build ask; ship=true is never inferred (eureka-recipe.md §1a) |
| 39 | "invent something new here and build it end-to-end" | Reproduce/Synthesize & Invent | eureka ship=true — the explicit build intent is the trigger. Two confirms (Phase 0 framing gate + Phase 8.5 Ship Gate with a separate build envelope); a non-ACCEPT invention exit ends at the dossier with the failed precondition named, and never silently degrades into a plain apex run |
| 45 | "run codex review and keep fixing until there are zero findings" / "loop the fix until the review comes back clean" | Loop | quell — the completion oracle is an external reviewer's finding count, not a rubric score (converge) and not a merge decision (acceptance). "Never stops" resolves to announce-and-proceed, never to an unbounded loop: _common/LOOP_PRECONDITIONS.md #2 still binds (quell-recipe.md §2) |
| 46 | "refactor this module and keep going until codex review is clean" | Loop / Improve | quell profile=refactor — the review loop, not plain refactor (which ships once) and not general profile: behavior preservation is enforced as an Equivalence Gate + frozen test files + DEFERRED (behavior-changing), so the loop cannot reach zero by editing tests or by "fixing it properly" (quell-recipe.md §5a) |
| 47 | "keep fixing the checkout screen until the design review comes back with nothing" | Loop | burnish — an external multimodal reviewer over the rendered surface, not quell (whose object is a code diff) and not restyle (internal evaluators against its own Brief). The split oracle must fire: hard findings to zero, soft axes to ≥ 2 — a run that promises literal zero on taste axes has mis-read §2 |
| 48 | "the direction is settled, now make the settings screen match the Figma frame exactly" | Loop / Improve | burnish profile=faithful — REFERENCE-DRIFT blocks at any severity and a "better-looking" divergence is DEFERRED, never applied. With no declared reference of record the run must exit BLOCK, never fall back to profile=general. Direction not yet settled → restyle first (over-capture guard) |
| 49 | "the tests all pass but I don't trust them — prove they'd actually catch a regression" | Loop | whet — the oracle is a mutation engine's surviving-mutant set. Must not resolve to radar (coverage counts execution, not detection) nor to siege direct (which measures once and recommends). The floor must be a per-partition threshold contract, never a chase to 100% (MA-02) |
| 50 | "our mutation score is 62%, get it up" | Loop | whet — and the integrity rules must fire: EQUIVALENT shrinks the score's denominator, so it needs a failed distinguishing-test attempt and may never be ratified by the test author; a test that mirrors the mutated expression is TAUTOLOGICAL-KILL; deleting the host code closes as CLOSED-BY-REMOVAL and never counts as a kill. A flaky suite must BLOCK at BASELINE, not run anyway |
| 51 | "clear the security backlog — fix every finding the scanner reports" | Fix / Loop | security mode=to-zero — the sweep, not the single-pass phases (one named vuln) and not quell (whose evaluator is a code reviewer on a diff). RISK-ACCEPTED without a named owner and expiry is invalid, CRITICAL/HIGH acceptance is Ask First, and a lapsed acceptance re-opens rather than persisting |
| 52 | "nine of our fourteen routes are over their Core Web Vitals budget — get them all under" | Improve / Loop | optimize mode=to-zero — a set of budget violations, not one target to one number (the single-pass phases) and not kaizen (one feature vs a target). Budgets are frozen at BASELINE, so BUDGET-RAISED is a blocking finding only the adjudicator ratifies; a close needs the declared sample count, never one favorable run (FLAKE-CLOSED); work deferred past the measured window is METRIC-GAMED. "make the dashboard load faster" (item 5) must still resolve to plain optimize — one target, one number |
| 31 | [ambiguous stress test] bare "optimize" (no object, no target, no metric named) | Improve (ambiguous) | CLASSIFY phase → GATE fires (context_confidence < 0.60 or 2+ valid interpretations: perf tuning? DB query? cost? team process?) → ONE focused clarifying question before any chain is selected — never silently guesses optimize |
| 32 | [ambiguous stress test] bare "landing page" (no verb, no scope) | Build/Improve (overloaded) | CLASSIFY phase → REDIRECT fires on the overloaded landing page anchor (routine LP → funnel[premium]/funnel; wish-grade one-shot LP → marquee — per signal-keywords.md marquee row) → one-question REDIRECT disambiguation, not a silent default to either |
| 33 | [out-of-coverage] "prove this sorting algorithm terminates and is correct using formal methods / theorem proving" | none of 90 global | CLASSIFY phase → REDIRECT fails, SELECT fails (no matrix row: formal verification / theorem-proving toolchains, e.g. Coq/Lean/TLA+, are not claimed by any currently registered skill — radar/attest cover empirical test/spec conformance, not machine-checked proof) → LADDER: compass Gap mode → architect proposal (e.g. "no formal-verify skill exists; nearest partial-fit is attest for spec-conformance testing, not machine-checked proof — propose new skill or route to a formal-methods specialist"); fallback_taken: architect-invoked |
| 34 | [out-of-coverage] "negotiate and draft the commercial lease terms for our new office space" | none of 90 global | CLASSIFY phase → REDIRECT fails, SELECT fails (no matrix row: real-estate lease negotiation is not code/design/research-artifact work; canon[legal] covers ToS/Privacy/Tokushoho only, explicitly excludes other contract types) → LADDER: compass Gap mode → architect proposal or explicit "route to a commercial real-estate attorney; no in-repo skill covers lease negotiation"; fallback_taken: architect-invoked |
| 40 | "the checkout works great on my machine — prove it still works on a bad network and with a screen reader" | Quality-Max (design floor) | crucible — the floor oracle (binary task completion per condition cell), not runway (ceiling) and not canon (standards conformance is one axis here). A variant that fails under ideal conditions must exit to bug at Phase 2, never be burned as a condition finding |
| 41 | "our product looks like every other SaaS — would anyone know it's ours?" | Quality-Max (design distinction) | silhouette — blind attribution vs K ≥ 3 competitors against a pre-committed Sameness Ledger. Must not resolve to runway (a craft-ceiling winner can be generic) nor to hallmark unless no settled brand exists, in which case hallmark runs first |
| 42 | "we have a design system but half the screens use hardcoded colors" | Quality-Max (system conformance) | lattice — steady-state conformance with RESIDUE-GATE to dry; not rebrand (no identity change) and not muse (the system already exists). With no system of record the run must stop and route to muse/vitrine, never invent the system mid-sweep |
| 43 | "the iOS and Android apps have drifted apart — make them feel native but like the same product" | Quality-Max (cross-platform) | chorus — two-sided gate (Idiom × Kinship) over a ratified Invariant/Variant Contract; not port (that plans a port; port→chorus is a plan→prove pair) and not lattice (within one platform). A single shipped platform must fail the precondition and route to native/restyle |
| 44 | "everyone says our architecture is good — prove it. is the repository abstraction actually necessary?" | Quality-Max (code design) | assay — a claim is in question, not an undiagnosed weakness, so it must not resolve to anneal (which owns "the design drifted, clean it up") nor to void (propose-only). The SUBTRACT instrument must report untested, not unnecessary, when the removal's zero diff sits on an element the suite does not exercise; a scope with no adequate element at all must exit BLOCK (oracle-inadequate) with the blind-spot map rather than run instruments anyway. apply=true must never be inferred from the phrasing |
| 53 | "go through the codebase and the docs and list everywhere they contradict each other" | Comprehend | verity — the coherence axis of the family; not cartograph/chronicle (those document what is / how it got here) and not judge (a diff review). The run must stay report-only: an attempt to edit a doc mid-run is out of contract |
| 54 | "nobody knows why this timeout is 4300ms — find everything like that before the maintainer leaves" | Comprehend | verity classes=unexplained — the UNEXPLAINED class, CORRELATE-REVERSE only. Every entry must carry a complete Provenance Search Record; an entry without one must be withdrawn at GATE, never shipped as "nobody knows" |
| 55 | [over-capture guard] "the spec and the code have drifted — align them and clean it up" | Improve | anneal — not verity. Bare spec-code drift splits on register-vs-fix: "list every divergence" → verity, "align and clean up" → anneal (verity→anneal is the audit→fix pair) |
| 56 | [over-capture guard] "what's actually built vs what the roadmap says?" | — | PROJECT_STATUS (PDM) direct — not verity. docs-vs-code drift splits on the question: planned-vs-shipped scope → pdm; a register of contradictions/stale/unexplained across the record → verity |
| 57 | "audit the recent changes and check whether they've drifted from our ADRs" | Comprehend | abide — change-anchored, so not verity (no change anchor) and not judge (diff bugs/quality). The governance corpus must be frozen independently of the change set: an ADR the change never touched still binds, which is precisely what verity since=<tag> cannot see. Report-only: a SUPERSEDES-SILENTLY finding routes to atlas for a new superseding record and must never edit the original |
| 58 | "we shipped this but nobody wrote an ADR for it" | Comprehend | abide classes=ungoverned — the UNGOVERNED class, the hunks→rules join direction. Every entry must carry a complete Governance Search Record and a stated architectural-significance trigger; an entry with neither must be withdrawn at GATE, never shipped as "no ADR exists". The VIOLATES/SUPERSEDES-SILENTLY fork must resolve to UNDECIDED rather than default to "deliberate" when no intent evidence settles it |
| 59 | [over-capture guard] "go through the repo and list everywhere the docs and the code disagree" | Comprehend | verity — not abide. The split is the anchor, not the subject: no change set named → verity (artifact inventory at a pinned HEAD); "the recent changes / this PR / this release vs the docs" → abide (change set vs the standing record) |
| 60 | "this prompt is too vague — 'you are a world-class UX designer, design a modern high-quality admin screen, make it look good'. turn it into something the model can actually execute" | Specialist | PROMPT_SPEC (Chisel) direct — not oracle (no prompt-system question: no few-shot policy, schema, versioning, eval gate, or cost decision is in play) and not scribe (the artifact stays a prompt, not a spec document). The persona line must be decomposed into evaluation axes, not retitled, and the axes it implied must survive its deletion; every derived rule must be a bound, an observable behavior, or a third-party-scorable criterion — a vague-for-vague swap is the failure this item catches |
| 61 | [over-capture guard] "make this better, I want something high quality" (no prompt text supplied) | Improve | Must terminate in exactly one question, and must not be Chisel. Two blockers compound and must be batched into a single ask (Three Laws #2): (a) the referent is unresolved — "this" names no target and no git/PROJECT.md/history context fixes it (intent-clarification.md § Uncertainty Typing, Referent dimension); (b) make it better is a listed overloaded anchor (intent-clarification.md:106 § Overloaded-Anchor REDIRECT → delve · kaizen · optimize · refactor · restyle · converge). Reaching the ask via that REDIRECT or via a sub-floor CLASSIFY GATE are both passes — asking once is the invariant; asking twice is also a failure. Note "high quality" does not resolve (b): it names no axis. Failing is any silent resolution, and specifically any route to Chisel: the words "high quality" and "make this better" are exactly Chisel's lexicon, but Chisel requires supplied prompt text as the object. A regression here means Chisel started absorbing intent clarification, putting every vague user request through a rewrite it never asked for |
| 62 | [gate guard] "fix the typo in the footer copyright year" | Bug | bug with SPECIFY skipped — a single trivial spawn with no load-bearing ambiguity. The gate firing here is the Core Rule #1 regression this item exists to catch: an unconditional SPECIFY adds a spawn to every chain, including ones with nothing to specify. The skip must be logged (Specify: skipped (reason: single trivial spawn)), never silent |
| 63 | [gate guard] "make the onboarding flow high-quality and thorough — spec it, build it, test it, and ship a PR" | Feature | feature with SPECIFY applied — load-bearing ambiguity ("high-quality", "thorough") sitting in a goal position, on a ≥3-spawn chain. The Specified Brief must reach Sherpa, Builder, Radar, and Guardian verbatim, and its delegated list must be non-empty — an empty one means the hub decided everything and the specialists became clerks (AS-09). SPECIFY must not fire before GATE nor replace it |
| 64 | [ordering guard] "our SKILL.md files have drifted — wording is inconsistent across them. sweep the repo and tidy them up" | Specialist | Gauge[audit] → Builder → Guardian? — not Chisel (a bare SKILL.md is excluded from PROMPT_SPEC; the object must be supplied prompt text). The load-bearing assertion is ordering: the Ask First gate (10+ files) resolves before SPECIFY runs, never after. The confirmation's answer sets the scope — all files or a subset, auto-apply or report-first — so a brief written ahead of it is written against the wrong target. Detected independently by two probes as the natural-but-wrong reading, which is why it is pinned here |
| 65 | "build and ship this MVP from scratch; size the process to the repo" | Build | deliver — scope-adaptive product/MVP delivery. One bounded capability must remain feature; an explicit high-investment 8-25-agent discovery-to-ship request must remain apex. The selected tier and evidence must appear in the Delivery Report. |
| 35 | [out-of-coverage] "configure the BACnet/Modbus points list for the building's HVAC controllers" | none of 90 global | CLASSIFY phase → REDIRECT fails, SELECT fails (no matrix row: OT/building-automation protocol configuration — BACnet/Modbus point mapping — is outside every skill's declared scope, including builder/gear/scaffold which are dev-environment/CI infra, not physical OT/ICS) → LADDER: compass Gap mode → architect proposal or explicit "route to a building-automation/OT engineer; no in-repo skill covers BACnet/Modbus point configuration"; fallback_taken: architect-invoked |
Items 29-30 and 33-35 are the dim-3 proof points (5 total, exceeding the rubric's ≥2 minimum): each must terminate in a named compass Gap-mode output and a named architect proposal artifact (or an explicit, logged decline recorded as fallback_taken: neither — reason: <reason>), never a quietly-generic answer.
Regression discipline
- Re-run items 1-28 whenever
routing-matrix.md,signal-keywords.md, ornexus/SKILL.md's Recipe Registry orreference/recipes-index.mdchanges. Items 1-28 must resolve to byte-identical chain selections across the change — they prove REDIRECT/SELECT priority is untouched. - Items 29-30 and 33-35 (out-of-coverage) are the only items expected to exercise the LADDER path; re-run them whenever
routing-matrix.md§ LADDER,compass/SKILL.mdOutput Routing "No matching skill" row, orarchitect'sARCHITECT_TO_NEXUS_HANDOFFschema changes. - Items 31-32 (ambiguous stress tests) are the only items expected to hard-stop at
GATE/REDIRECTwith a clarifying question — a regression here means an ambiguous anchor started silently resolving without a question, which is the DC-1-adjacent failure this battery exists to catch.