Loop hygiene:
- Run
bun cli.ts run{{rootFlag}} <workspace>in the foreground. The Shell default backgrounds the loop and breaks the heartbeat when your turn ends. - Exit code 100 is a planned checkpoint restart, not an error. Rerun the same command immediately; it resumes from committed
state.json. - Exit code 1 on a non-empty error set is your turn. The loop exited because a task crashed; the script already wrote a synthetic
handoffs/<task>-failure.mdfor each dead worker and anyhandoffs/<task>-finished-no-handoff.mdfor workers that ended without a structured handoff. In-flight workers keep running; the nextrunreattaches viarecoverRunning. - After
runreturns, calltree. If any task is stillpendingorrunning, loop again. - Don't end your turn while this workspace has non-terminal tasks.
Reacting to failure handoffs:
For each task with status: "error" and a matching handoffs/<task>-failure.md, read the Failure mode line and decide:
cap-hitoroom: retry with smaller scope (split into narrower tasks, tighterpathsAllowed, leanerscopedGoal).network-drop: retry as-is; treat as transient.tool-error: retry with a differentmodel.unknown: read theLast activityandSDK errorlines; if no signal, treat as transient and retry as-is; abandon if it fails again. For<task>-finished-no-handoff.md, read the raw snippet athandoffs/<task>.mdand decide whether the worker's intent was recoverable; retry or abandon. Each retry costs another cloud-agent run; budget your decisions. After 2 retries on the same task, prefer abandon (drop the task fromplan.json, replan around it) over a 3rd attempt unless you have specific evidence the next retry will succeed. Updateplan.json, then re-runbun cli.ts run{{rootFlag}} <workspace>to continue.