All skills
simota avatar

/darwin

@35ffd55
by shingo imotasimota/agent-skills85 stars
15

Orchestrating ecosystem self-evolution: lifecycle-phase detection, agent relevance, cross-agent knowledge synthesis, evolution proposals. Use when auditing skill-ecosystem health or fitness.

Use this Skill: https://skilld.dev/gh/simota/agent-skills/darwin

This session only. Nothing lands on disk.

referenceassessment-models.md

≈4k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Assessment Models

Defines the calculation formulas and algorithms for Ecosystem Fitness Score (EFS), Relevance Score (RS), and lifecycle phase detection.

2026-05 baseline notes

  • Failure attribution defaults: when grading Coherence and Quality, anchor on the MAST taxonomy (Cemri et al., ICLR 2026 — 1,600 traces across 7 frameworks, 14 failure modes, Cohen's Kappa = 0.88). Distribution: Specification & System Design 41.8%, Inter-Agent Misalignment / Coordination 36.9%, Task Verification & Termination 21.3%. Use these weights when allocating root-cause budget across EFS dimensions.
  • Trace-level error amplification by topology (Towards Data Science, 2026): SAS = 1.0× baseline, Centralized = 4.4×, Hybrid = 5.1×, Decentralized = 7.8×, Independent = 17.2×. Independent (i.e. "bag of agents" without orchestrator) is the only architecture above the 10× cliff — penalize EFS Coherence accordingly.
  • LangChain State of Agent Engineering (survey conducted Nov–Dec 2025, published 2025): 57% of surveyed orgs run agents in production; 32% cite quality as the top deployment barrier. Use as the external benchmark when calibrating Quality dimension scoring. Source: https://www.langchain.com/state-of-agent-engineering.
  • Sources: https://arxiv.org/abs/2503.13657 (MAST paper), https://openreview.net/forum?id=wM521FqPvI (ICLR 2026 proceedings), https://towardsdatascience.com/why-your-multi-agent-system-is-failing-escaping-the-17x-error-trap-of-the-bag-of-agents/.

Ecosystem Fitness Score (EFS)

Formula

EFS = Coverage × 0.25 + Coherence × 0.20 + Activity × 0.20 + Quality × 0.20 + Adaptability × 0.15

Dimension Definitions

Coverage (25%)

Measures whether the agent ecosystem covers the project's actual needs.

Coverage = (matched_needs / total_needs) × 100

Where:
  total_needs = count of task types that occurred in last 90 days
  matched_needs = count of task types with an appropriate agent invoked

Scoring guide:

  • 100: Every task type in last 90 days had a matching agent invoked
  • 80: Most task types covered, 1-2 gaps filled by manual work
  • 60: Several gaps; common tasks lack dedicated agents
  • 40: Significant coverage gaps; many tasks done without agent support
  • 20: Minimal agent usage relative to project needs

Data sources: Routing matrix vs PROJECT.md task types, lifecycle phase dominant agents vs actual usage.

Coherence (20%)

Measures how well agents work together in chains and handoffs.

Coherence = (successful_chains / total_chains) × 0.6 + (avg_handoff_quality) × 0.4

Where:
  successful_chains = chains that completed without error or rollback
  total_chains = all chains attempted in last 90 days
  avg_handoff_quality = mean quality of NEXUS_HANDOFF artifacts (0.0-1.0)

Scoring guide:

  • 100: All chains complete successfully, handoffs are clean and informative
  • 80: Occasional chain adjustments, handoffs generally good
  • 60: Regular chain failures requiring recovery, some handoff quality issues
  • 40: Frequent failures, poor handoff quality, agents stepping on each other
  • 20: Chains rarely complete, severe handoff problems

Data sources: Nexus execution logs, NEXUS_HANDOFF quality assessment, error/recovery entries.

Topology-aware adjustment (2026-05): Cap Coherence at 60 when the deployed topology is Independent (no central orchestrator) because measured trace-level error amplification is 17.2× the single-agent baseline, well above the 10× regression cliff. Centralized (4.4×) and Hybrid (5.1×) topologies remain eligible for the full 0–100 range. Source: Towards Data Science, "Why Your Multi-Agent System is Failing" (2026).

Activity (20%)

Measures whether the ecosystem is actively used.

Activity = unique_agents_30d × 0.4 + invocation_rate × 0.3 + journal_freshness × 0.3

Where:
  unique_agents_30d = (unique agents used in 30 days / total available agents) × 100
  invocation_rate = normalized invocations per week (target: project-size dependent)
  journal_freshness = (agents with journal entries in 30 days / agents used in 30 days) × 100

Scoring guide:

  • 100: Diverse agent usage, high invocation rate, fresh journals
  • 80: Good variety, regular usage, most journals updated
  • 60: Moderate usage, some agents neglected
  • 40: Low diversity, sporadic invocations, stale journals
  • 20: Minimal agent usage, ecosystem largely dormant

Data sources: PROJECT.md row counts, agent journal dates, unique agent counts.

Quality (20%)

Measures whether agent outputs are improving over time.

Hard gate (evaluated BEFORE the weighted score):
  if uqs_trend == declining, or gauge reports any open P0 finding
  → Quality is capped at 40 and the triggering condition is reported on its own line

Quality = uqs_trend × 0.4 + feedback_ratio × 0.3 + health_score × 0.3

Where:
  uqs_trend    = direction of UQS over last 3 cycles (improving=100, stable=60, declining=20)
  feedback_ratio = positive Reverse Feedback / total Reverse Feedback × 100
  health_score = latest Architect Health Score (if available, else 50)

Why the gate precedes the average. A weighted mean lets a rising feedback_ratio offset declining output quality, and it lets an open P0 compliance finding disappear into a 70. Evaluate the gate first, report the triggering condition on its own line, and carry gauge's per-severity PASS/PARTIAL/FAIL counts next to the score rather than folding them in — gauge deliberately reports per severity and never one composite grade (gauge/SKILL.md).

Scoring guide:

  • 100: UQS improving, positive feedback dominant, high health score
  • 80: UQS stable-good, mostly positive feedback
  • 60: UQS plateau, mixed feedback
  • 40: UQS declining, negative feedback increasing
  • 20: Poor quality indicators across the board

Data sources: Judge UQS history, Reverse Feedback entries, Architect Health Score.

Adaptability (15%)

Measures how quickly the ecosystem responds to project changes.

Adaptability = override_freshness × 0.3 + transition_response × 0.4 + gap_fill_rate × 0.3

Where:
  override_freshness = days since last AFFINITY update (fresher = higher, max at 30 days)
  transition_response = speed of agent mix change after lifecycle transition (faster = higher)
  gap_fill_rate = new agents or skills created to fill identified gaps / total gaps identified

Scoring guide:

  • 100: Rapid adaptation, fresh overrides, all gaps addressed
  • 80: Good responsiveness, some lag in adaptation
  • 60: Moderate adaptation, overrides somewhat stale
  • 40: Slow to adapt, overrides outdated, gaps unaddressed
  • 20: No adaptation, ecosystem stuck in previous phase

Data sources: ECOSYSTEM.md timestamps, lifecycle transition history, Architect gap analyses.

EFS Grading

Grade Score Range Interpretation
S 95-100 Exceptional — ecosystem perfectly tuned
A 85-94 Excellent — minor improvements possible
B 70-84 Good — some optimization opportunities
C 55-69 Fair — notable gaps or inefficiencies
D 40-54 Poor — significant evolution needed
F 0-39 Critical — ecosystem not serving the project

EFS Trend Indicators

Symbol Meaning Threshold
↑ Improving +5 or more from previous
→ Stable Within ±4 of previous
↓ Declining -5 or more from previous

Harness Maturity (HE-M0 – HE-M5)

What it answers, and how it differs from EFS. EFS scores ecosystem health — is the roster covering, cohering, active, good, adaptable. Harness Maturity scores how much autonomy the harness has earned — how much risk it can absorb with the evidence and recovery it actually has. A perfectly healthy ecosystem can still be M1: nothing is wrong, and nothing is provable. Report both; never average them.

Level Name Required evidence (continuously satisfied, not once)
HE-M0 Prompt-centric Human verifies every outcome. No automated boundary.
HE-M1 Tool-enabled Versioned capabilities, repository-level instructions, basic checks.
HE-M2 Contract-enforced Intent contract with acceptance criteria, isolated workspace, schema-validated handoffs, deterministic gate.
HE-M3 Observable and Recoverable Journal/trace completeness, checkpoint + resume, failure attribution with confidence, incident-derived regression cases.
HE-M4 Orchestrated Queue and file-ownership leases, parallel partitioning, independent evaluator, integration gate.
HE-M5 Governed Adaptive Policy governance, engine substitution testing, harness regression suite, controlled self-improvement, audit trail.

HE-M5 does not mean full autonomy. It means the harness can change itself under control while human approval remains on the high-risk surface.

Assessment Dimensions

Score each 0–4. Ten dimensions:

1 Intent / Contract · 2 Context / Legibility · 3 State / Recovery · 4 Capability surface / Errors · 5 Environment / Reproducibility · 6 Verification / Eval · 7 Observability / Attribution · 8 Permission / Security · 9 Orchestration / Flow · 10 Improvement / Governance

Never average the dimensions. The level is set by the lowest critical dimension, not the mean — a harness that is excellent at nine things and blind at the tenth fails at the tenth.

Hard Ceilings

A ceiling overrides any claimed level:

Condition Ceiling
Broad write or external-publish authority with no audit trail M1
Long-running or resumable work with no rehearsed resume M2
Parallel agents with no file-ownership or merge gate M2
Self-improvement with no fixed eval and no independent approval M3
An unresolved critical policy violation Autonomy suspended

Per-Task-Family Assessment

Assess by task family, never as one organizational number. A representative shape:

Documentation maintenance: M4
Reference/link upkeep:     M3
New skill authoring:       M2
Corpus-wide refactor:      M1
Protocol (_common/) change: M0/M1

Never justify a high-risk task family with a low-risk family's maturity.

Progression

Levels rise by removing the failures you actually have, not by adding features:

  • M1→M2: intent contract, handoff schema, outcome gate
  • M2→M3: journaling/trace, checkpoint, failure attribution
  • M3→M4: task graph, WIP limits, independent integration
  • M4→M5: harness regression suite, engine substitution, governance

Not every family needs M5. Pick a target level justified by cost and risk.

Regression

Maturity falls. Reassess on any of: engine or CLI update · ownership change · documentation drift · fixture flakiness · newly granted authority · an incident · vendor or protocol change.

Anti-patterns

  • Treating skill count as maturity.
  • Treating agent count as orchestration maturity.
  • Treating journal retention as observability maturity.
  • Treating number of confirmations as governance maturity.
  • Treating a fixture pass rate as production readiness.
  • Using the level as a scoreboard rather than a decision input.

Decision rule: maturity is not how much is automated. It is which risks can be accepted, given the evidence and recovery available.


Relevance Score (RS)

Formula

RS = Usage × 0.40 + Affinity_Match × 0.25 + Feedback × 0.20 + Freshness × 0.15

Component Definitions

Usage (40%)
Usage = (w30 × count_30d + w60 × count_60d + w90 × count_90d) / normalizer

Where:
  w30 = 0.50 (most recent period weighted highest)
  w60 = 0.30
  w90 = 0.20
  count_Xd = number of invocations in last X days
  normalizer = project-dependent (based on total agent invocations)

Scale: 0-100, where 100 = top quartile agent by usage.

Affinity_Match (25%)
Affinity_Match = phase_affinity × 0.60 + type_affinity × 0.40

Where:
  phase_affinity = how well the agent matches the current lifecycle phase (from phase table)
  type_affinity = agent's PROJECT_AFFINITY rating for current project type

Phase Affinity Table:

Phase Dominant Agents (affinity = 1.0) Supporting (0.6) Neutral (0.3)
GENESIS Nexus, Forge, Spark, Architect Builder, Radar, Scout All others
ACTIVE_BUILD Builder, Forge, Radar, Artisan Schema, Gateway, Scout All others
STABILIZATION Judge, Sentinel, Zen, Nexus Radar, Sweep, Atlas All others
PRODUCTION Guardian, Beacon, Triage, Gear Sentinel, Launch, Hone All others
MAINTENANCE Shift (detect/modernize/radar), Sweep, Trail Builder, Radar, Scout All others
SCALING Bolt, Tuner, Scaffold Beacon, Gear, Stream All others
SUNSET Sweep, Quill, Scribe Canvas, Trail All others
Feedback (20%)
Feedback = (positive_count - negative_count × 1.5) / total_feedback × 100

Where:
  positive_count = positive Reverse Feedback entries for this agent
  negative_count = negative Reverse Feedback entries (weighted 1.5× to penalize issues)
  total_feedback = total feedback entries for this agent

  If no feedback exists: Feedback = 50 (neutral default)
Freshness (15%)
Freshness = max(100 - days_since_update × 1.5, 0)

Where:
  days_since_update = min(days_since_skill_update, days_since_journal_entry)

Scale: 100 at 0 days, 0 at 67+ days. This encourages regular agent maintenance.

RS Status Classification

Status Score Range Meaning Recommended Action
Active 80-100 Agent is highly relevant and frequently used Continue as-is
Stable 60-79 Agent is relevant, usage is adequate Monitor for trends
Dormant 40-59 Agent is underused relative to its potential Review affinity match
Declining 20-39 Agent relevance is dropping Investigate cause, consider improvement
Sunset 0-19 Agent may no longer be needed Verify with Void, consider retirement

Lifecycle Detection Algorithm

Step-by-Step Process

1. COLLECT signals (see signal-collection.md)
2. For each phase P in [GENESIS, ACTIVE_BUILD, STABILIZATION, PRODUCTION, MAINTENANCE, SCALING, SUNSET]:
     score[P] = sum(signal_match[P][s] × weight[P][s] for s in signals)
3. Sort phases by score, descending
4. If score[top_phase] >= 0.60:
     detected_phase = top_phase
     confidence = score[top_phase]
5. Else:
     detected_phase = MIXED
     phases = top 2 phases with scores
     confidence = max(scores)
6. Compare detected_phase with previous_phase (from ECOSYSTEM.md)
7. If transition detected:
     Fire ET-01 trigger
     Record transition in history
8. Return {phase, confidence, signals_breakdown, transition}

Confidence Interpretation

Range Meaning
0.90-1.00 Very high — clear phase signal
0.75-0.89 High — dominant phase with some secondary signals
0.60-0.74 Moderate — phase detected but mixed signals present
0.40-0.59 Low — mixed phase, report as dual-phase
0.00-0.39 Very low — insufficient signal, report as UNKNOWN

Transition Detection

A phase transition is confirmed when:

  1. Detected phase differs from stored phase
  2. New phase confidence ≥ 0.60
  3. OR new phase has been the top scorer for 2+ consecutive checks

On transition:

  • Record in ECOSYSTEM.md transition history: {from, to, date, confidence}
  • Fire ET-01 trigger (AFFINITY recalculation)
  • Log to .agents/darwin.md journal

Source: SKILL.md on GitHub

1 warning13d5 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    The skill 'darwin' is an ecosystem orchestrator designed to monitor project health, lifecycle phases, and agent fitness. It poses a low security risk primarily due to its reliance on ingesting untrusted data from the repository (such as git logs and journals) to drive its assessments, which creates a surface for indirect prompt injection. Additionally, it utilizes shell commands like 'git' and 'find' to collect system signals and contains deceptive future-dated metadata.

  • Socket13d

    No alerts

  • Snyk13d

    Risk: LOW · No issues

  • Runlayer6mo

    4/6 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 35ffd55. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 2 weeks ago

README badge

README badge for simota/agent-skills/darwin