All skills
boshu2 avatar

/craft-goal

@833aefb
by Boboshu2/agentops446 stars
42

Draft or lint a bounded persistent goal above a bead graph of RPI experiments. Use when: this goal workflow is explicitly selected; shaping a single change belongs to Plan.

Use this Skill: https://skilld.dev/gh/boshu2/agentops/craft-goal

This session only. Nothing lands on disk.

SKILL.md

≈46 tokens always: the name and description. ≈2.6k when used: this file. ≈998 more on demand in 4 files.

Craft Goal

Craft the autonomy contract above AgentOps RPI. A goal is a persistent Mayor over a bead-shaped experiment graph. Each RPI is one scientific trial; the goal selects the next useful trial, preserves what was learned, and ratchets toward a larger outcome.

Goal / Mayor: observe graph → choose bounded wave → consume verdicts → ratchet
  └─ Bead: durable experiment intent, context, scratch, evidence, and links
       └─ RPI: plan → implement → fresh validate → bounded repair → verdict → report
            └─ Implementation: one RED → GREEN → refactor experiment

The number of RPIs need not be known in advance. The goal is safe when success is decidable, every experiment is bounded, evidence retains its provenance, and the authorization envelope cannot silently renew itself. Beliefs are revisable: new evidence may retract an earlier claim. More stored knowledge is neither progress nor proof that knowledge is correct.

Insight: bounded waves shorten the feedback loop; one hard, non-renewing campaign envelope prevents those waves from becoming infinite continuation.

Authority boundary. The emitted goal prompt and safety report are inert caller-owned text. Crafting one creates no goal, starts no runtime, and mutates no bead; it confers no standing authorization. The prompt drives RPI dispatch only when a caller pastes it into their own goal runtime under their own authority, and only within the non-renewing envelope the caller then sets.

Named failure mode — completion treadmill: discoveries recursively become requirements and activity continues without new information. Its opposite is first-red abandonment: one falsified hypothesis ends a viable campaign. Anti-pattern: choose endless retries or stop on the first red. Corrective: continue while experiments produce a defined ratchet and remain inside the envelope; invoke an andon on churn, judgment, or exhaustion. Stop when the goal reports ACHIEVED, NOT_ACHIEVED, or NEEDS_OPERATOR.

Modes

Caller wording Mode Result
"craft a goal", "turn this into a goal" craft Compile a Mayor-style goal prompt and settings.
"lint/review this goal", "is this safe" lint Return findings and a rewrite when supplied facts permit one.

Stop after 1 compilation pass. Never create a goal or mutate beads.

Admission and sizing

Fuzzy route is acceptable; fuzzy success is not. Before goal creation, the caller must know the outcome, what evidence would prove it, non-goals, and authority. The exact experiment graph may still be unknown. Write each terminal criterion as a Given/When/Then example with an observable result, and name each domain term once; the caller can settle these with Interview first.

  • Return USE_RPI for one shaped experiment with no verdict-driven follow-on.
  • Use a goal for a terminal outcome that may need several related experiments.
  • A shaped goal with no beads may begin with 1 bounded discovery wave that creates the root and initial experiment beads.
  • Return UNSAFE_GOAL when no falsifiable first question or terminal evidence can be named. Route that intent to idea/plan work.
  • Return UNSAFE_GOAL for indefinite monitoring or event reaction; that is an automation, not a terminal goal.

Goals may be different sizes. Size the wave and hard campaign envelopes to the outcome; do not invent one universal budget.

Critical constraints

  • Closed outcome, adaptive route: Freeze terminal acceptance. New facts may change hypotheses and dependencies, never silently enlarge success. Why: discovery should steer the route, not redefine the finish line.
  • Bead knowledge graph: Use the tracker as durable memory, not a parallel goal ledger. Root epic = outer intent; child bead = one experiment/RPI. Why: compaction must not erase the scientific record.
  • RPI membrane: One candidate gets one bounded RPI and an author-distinct fresh validation result. The goal may request durable verdict evidence but never rewrites it. Why: orchestration cannot author its own proof.
  • Brownian ratchet: Continue only when a result adds non-duplicative, decision-relevant knowledge or advances acceptance. Why: activity without information is churn.
  • Two-level bounds: Every RPI is bounded; every dispatch wave is bounded; the full goal also has monotonic hard ceilings. Why: a new wave must not mint a new campaign.
  • Earned andon: Ordinary informative red may change the route within frozen acceptance. Repeated no-information failure, regression, recurrence, oscillation, or scope pressure enters HOLD. Exactly 1 bounded fresh helper per incident may return UNSTUCK or ESCALATE inside the remaining allowance; cancellation, an explicit refusal/judgment lane, or a spent hard budget skips the helper and stops work.
  • Operator legibility: At each wave boundary, report the acceptance matrix, graph frontier, verdicts, ratchets, churn, remaining budget, and next thesis.
  • Exterior self-repair: Repair an unstable factory from an ordinary shell/worktree and use the factory only for a declared bounded canary.

Stop when the goal reports ACHIEVED, NOT_ACHIEVED, or NEEDS_OPERATOR.

Graph walk

Navigate owns the runtime walk: the bead graph contract, edge semantics, what counts as a ratchet, discovery classes and the wave checkpoint. The frozen prompt tells the goal to apply it each wave. Lint that the prompt names a root epic or its bounded bootstrap rule, ties each experiment to an unmet criterion or named blocking uncertainty, keeps all three discovery classes and never counts activity as progress. Craft Goal reads tracker state when present but starts nothing and needs no tracker installed to compile a prompt.

Convergence and andons

Specify both:

  • Wave envelope: RPIs, concurrency, wall time/tokens, live attempts, and a checkpoint at its end.
  • Goal envelope: total RPIs, wall time/tokens, live attempts, compactions, and any patch/surface limit for the whole campaign.

Dispatch budget: every wave declares numeric RPI, token, time, and concurrency limits before any work is selected, including helper and validation costs. Name the native control that enforces each claimed hard limit and how remaining allowance is observed. Objective text is an instruction, not enforcement; do not represent an unmeasured aggregate as a remaining balance. No helper, retry, new subject, compaction, or wave renews the goal allowance.

Continue automatically across waves only while a ratchet exists and the next experiment fits frozen acceptance, authority, and remaining envelope.

Enter HOLD on any declared trigger: repeated blocker, no ratchet for the configured number of RPIs, oscillation between prior approaches, introduced regression, unknown new-defect cause, recurrence, requested acceptance change, or operator-reserved decision. HOLD stops implementation for causal examination. While the caller's remaining allowance admits it, consult exactly 1 bounded fresh-context helper per HOLD incident; rewording the blocker or receiving an automatic continuation does not create a new incident. Supply acceptance, observations, failed approaches, exact evidence, and remaining allowance.

  • UNSTUCK must name a materially different experiment, its discriminating check, and why it fits unchanged acceptance, authority, and remaining bounds; only the selected outer goal may resume. It never revives a spent RPI bound.
  • ESCALATE, an unhelpful helper, or no admissible experiment emits NEEDS_OPERATOR and performs no more implementation or helper dispatch.
  • Cancellation stops immediately. An explicit refusal/judgment lane or a genuinely spent hard time, cost, or quota ceiling skips the helper; report the refusal or NOT_ACHIEVED with the exact gaps. A retry threshold alone is not proof of a spent hard budget.

Native goal objective text and a terminal report do not enforce continuation. No agent-callable native pause or aggregate allowance operation is demonstrated by this contract. Report which native controls were actually observed, any unmeasured allowance, and whether implementation stopped; never claim the goal is paused from prose alone. When operator action is required, report that need truthfully and keep further work stopped. The persistent controller's threshold for recording blocked is separate status bookkeeping, never permission for extra experiments or helpers. Stop when any terminal report is emitted.

Frozen prompt

Read and fill the copy-paste-only goal prompt. Preserve its headings and terminal semantics; replace every angle-bracket field.

Quality

Lead with SAFE_TO_CREATE, USE_RPI, or UNSAFE_GOAL. SAFE_TO_CREATE judges prompt content; it does not certify native enforcement or create a goal. Return the copy-paste prompt, separate goal-tool token budget, assumptions, and one lint line for: outcome, evidence, admission, bead graph, RPI boundary, ratchet, discovery, wave budget, hard budget, breaker, operator andon, scope, self-hosting, and terminal reports.

Output validator — a captured decision must lead with exactly one terminal token:

printf '%s\n' "$decision" | head -n1 | grep -Eq '^(SAFE_TO_CREATE|USE_RPI|UNSAFE_GOAL)\b'

This pins the machine-checkable shape of the output contract. scripts/validate.sh still enforces structural hygiene; the fourteen lint dimensions above stay a human rubric because they judge prompt content that has no persisted artifact at gate time.

Done when:

  • success is finite but the route may adapt;
  • the tracker can reconstruct intent, experiments, evidence, and provenance;
  • informative red can continue but repeated non-information cannot;
  • recursion cannot expand acceptance or reset monotonic ceilings;
  • both successful and non-success terminal reports exist.

Stop after 1 lint pass and zero goal executions. Paired evidence: docs/learnings/2026-07-12-go-cli-goal-stall-tracker-layer-confusion.md and skills/rpi/SKILL.md.

Failure behavior

Return UNSAFE_GOAL with missing decisions. Do not invent acceptance, authority, graph semantics, or campaign size. The caller owns revision and goal creation and can settle the missing decisions first with Interview.

Source: SKILL.md on GitHub

No alerts3d3 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    The skill drafts and reviews autonomy contracts for agent experiments. It is susceptible to indirect prompt injection because it incorporates untrusted user data into a generated prompt that governs downstream agent behavior. It also executes local validation scripts using computed paths.

  • Socket3d

    No alerts

  • Snyk3d

    Risk: LOW · No issues

Signed by skilld at 833aefb. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 19 hours ago.

Activeupdated last week

README badge

README badge for boshu2/agentops/craft-goal