All skills
oaustegard avatar

/agent-routing

@576f2c8

Decide which model, effort level, and cascade shape each subagent gets, and how to keep improvement loops safe (evaluator-as-selector, stop on regression). Routes on measured cost-per-completed-task rather than per-token price, because a tier's token count varies more by task shape than price varies across tiers. Covers per-model effort semantics, the concision lever, cascade preconditions, context handoff, and watching a subagent fan-out live. Use when spawning subagents via the Agent or Workflow tools, when choosing how to escalate a failed attempt, when fanning out more than a handful of agents, or when asked which model or effort a task should get. Grounded in measured calibration (references/calibration-2026-07-15.md), a 2026-08 coding-cost study, and a 2026-09 agentic-repair battery that measured the cascade rungs directly; Managed Agents API specifics are operational, not calibrated.

Use this Skill: https://skilld.dev/gh/oaustegard/claude-skills/agent-routing

This session only. Nothing lands on disk.

referencescalibration-2026-07-15.md

≈1.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Calibration evidence — 2026-07-15

Data behind the routing heuristics in SKILL.md. Collected in one CCotw session (Fable 5 orchestrator, Workflow tool, session bf6d02ee), ~12.3M subagent tokens total across four workflow runs, zero agent errors. All numeric grading was deterministic (Python scorers against generated ground truth); the only judge tokens were Sonnet scoring open-ended text in Experiment 1.

Experiment 1 — Looped Haiku subagents (Looped-Mamba analog, arXiv 2607.10110)

Design: answer_r = Haiku(task, answer_{r-1}), R=5, out-of-band eval per loop.

Task family Instances × loops Result
Nested mod-23 arithmetic, depth-3 trees 5 × 5 100% at r=1, flat, 0 churn
p-hop function chains (n=7, 8–12 hops) 5 × 5 100% at r=1, flat, 0 churn
Harder: depth-4 trees (16 leaves) 8 × 5 100% at r=1, flat, 0 churn
Harder: n=12 chains, 20–30 hops 8 × 5 100% at r=1, flat, 0 churn
Answer-only control (CoT suppressed, effort=low), depth-4 trees 8 × 5 still 100% at r=1, 0 churn
Open-ended (Sonnet judge /10): prime one-liner 1 × 5 [10,10,10,10,10]
Open-ended: sky-blue explanation 1 × 5 [9,9,9,9,9]
Open-ended: 5-7-5 haiku 1 × 5 [9,3,3,3,3] — loop 2 broke the middle line to 8 syllables, loops 3–5 froze on identical broken text

Interpretation: a single LLM call spends variable internal compute (CoT), so answer_1 is already a fixed point on checkable tasks — the paper's fixed-compute-per-step premise doesn't hold for LLM subagents. The exit-gate idea does transfer: the evaluator-as-selector recovers the loop-1 haiku and skips wasted loops elsewhere. Answer churn across all deterministic loop-pairs: 0/96 flips.

Experiment 2 — Single-shot routing calibration (60 agents)

20 tasks × 3 configs, single shot, deterministic local scoring:

Tasks: 9 constraint-stack sentences (K=3/4/5 simultaneous constraints: exact word count, begin-word, include-word, end-word, no letter 'e'), 4 trap-laden word-math problems (distractor numbers), 3 state-tracking problems (3 boxes, 25 operations), 4 constraint-preserving revisions (exact N words, fixed first/last word, ≥2 changes).

Config stack trap state revision total
haiku, effort=low 9/9 4/4 3/3 4/4 20/20
haiku, effort=high 9/9 4/4 3/3 4/4 20/20
sonnet, effort=low 6/9 4/4 3/3 4/4 17/20

Sonnet-low failures (hand-verified, real):

  • stack-K3-1: 15 words where exactly 14 required
  • stack-K4-0: 13 words where exactly 12 required
  • stack-K5-0: used "quiet" in a no-letter-'e' sentence

Notes:

  • n=20/config: 17 vs 20 is not statistically significant. The defensible claim is "no evidence of up-tier benefit on mechanical tasks," which is enough to invert the default (burden of proof on routing up).
  • Haiku's K5 lipogram outputs were flawless, e.g. "Our distant stars hang bright and high throughout dark night" (10 words, begins "our", includes "stars", ends "night", zero e's).
  • Effort had no measurable effect on Haiku here (both 100%) — consistent with Experiment 1's answer-only/effort-low control also scoring 100%.

Pricing basis (2026-07, per MTok in/out)

Haiku 4.5 $1/$5 · Sonnet 5 $2/$10 · Opus 5 $5/$25.

Superseded 2026-08-17: this section previously read Sonnet 5 at $3/$15 with $2/$10 as an intro rate expiring 2026-08-31. $2/$10 is now the standing price (user-confirmed). Any analysis computed at $3/$15 understates Sonnet by ~1/3.

The p_fail-only break-even formerly stated here (Haiku-first beats Sonnet-direct while p_fail(Haiku) < 1 − c_H/c_S ≈ 2/3) is wrong in the general case. It assumes c_H < c_S per task, which holds only when outputs are short. On long-output work Haiku's verbosity inverts it: measured 2026-08-17, Haiku cost $0.067/task against Sonnet's $0.031, so Haiku-first loses at every p_fail, including zero. Check cost-per-task first; see the cascade precondition in SKILL.md. The verifier-is-near-free assumption still holds — all Experiment 2 scoring was local Python.

What has NOT been measured

  • A deterministic task family where Haiku actually fails (the cliff).
  • Multi-turn agentic tool-use quality per tier.
  • Same-model self-judging reliability.
  • Haiku at higher K (>5 simultaneous constraints) or longer state chains.
  • Multi-source consistency/reconciliation (do two extracts of the same period agree?). One anecdote favors up-tiering; untested.

Re-run: generators + scorer live in the session scratchpad pattern (gen.py/gen2.py/gen3.py, score.py); regenerate with new seeds and a current model rev before trusting the table across model versions.

Source: SKILL.md on GitHub

No alerts12d3 checks · Risk SAFE
  • Gen Agent Trust Hub12d

    The skill provides technical guidelines for routing tasks between different AI models and managing subagent orchestration. It is generally safe; however, it encourages a workflow where untrusted data, such as file scan outputs, is passed directly into subagent prompts, which creates a surface for indirect prompt injection.

  • Socket12d

    No alerts

  • Snyk12d

    Risk: LOW · No issues

Signed by skilld at 576f2c8. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 3 days ago
metadata
{
  "author": "Oskar Austegard and Claude",
  "version": "2.3.0"
}
Other metadata
compatibility
Designed for Claude Code / Claude Code on the Web — assumes an orchestrator with Agent/Workflow subagent tools. Only the Workflow tool sets a subagent's effort; the Agent tool sets its model. Not applicable to claude.ai chat use.

README badge

README badge for oaustegard/claude-skills/agent-routing