Enterprise Agent Ops
Original work by ECC. Credit to ECC for this skill and its core ideas.
Use this skill for agents that run in the cloud or stay on for a long time. Do not use it for a short, one-time command.
Main Goals
- Start, pause, stop, and restart the agent.
- Record logs, counts, timing, and task paths.
- Limit access and add a kill switch.
- Ship updates safely and roll them back.
- Keep a record of risky actions.
Required Controls
Before launch:
- Build one fixed release file or image.
- Give the agent only the access it needs.
- Load secrets from the run setting. Do not put secrets in code.
- Hide secrets and private data in logs.
- Set a hard time limit for each task.
- Set a limit on retries and total work.
- Use wait times between retries.
- Stop retries for errors that will not get better.
- Add a kill switch that stops new work.
- Keep an audit log for risky actions.
- Test restart and rollback steps.
- Set limits for cost, task count, and run time.
Track These Measures
- Tasks that end well
- Tasks that fail
- Retries per task
- Time used per task
- Time needed to recover
- Cost per task that ends well
- Failures by type
- Tasks waiting in line
- Work stopped by safety rules
Do not save secrets, full user text, or private data in these records.
Safe Task Rules
For each task:
- Give it a clear ID.
- Set its access scope.
- Set time, retry, and cost limits.
- Record each risky action.
- Mark the task as done, failed, timed out, or stopped.
- Clean up locks and short-term files.
A retry must not run the same risky action twice. Use a task ID or save a safe check first.
Start and Stop Rules
- Start only after health checks pass.
- Pause by stopping new tasks. Let safe work finish.
- Stop new work before a planned shutdown.
- Set a short wait limit for tasks still running.
- Save needed state before restart.
- On restart, check old tasks before running them again.
- Mark tasks with lost or bad state for review.
If the state store is down, stop new work unless the task is safe without saved state.
Update Rules
- Test the new release.
- Run safety and access checks.
- Send a small share of work to it.
- Watch errors, cost, and task time.
- Raise the share in small steps.
- Stop the update if a set limit is crossed.
- Roll back to the last good release.
Do not change code, settings, and data shape in one step when they can be split.
When Failures Rise
- Stop the new update.
- Stop new risky work if harm may grow.
- Save a few useful logs and task paths.
- Hide secrets and private data.
- Find the bad route, tool, model, or release.
- Make the smallest safe fix.
- Run old tests and safety checks.
- Restart with a small share of work.
- Watch the service before full use.
- Write down the cause, fix, and follow-up work.
Use the kill switch at once if the agent may leak data, spend without a limit, harm users, or repeat risky acts.
Edge Cases
- Tool is down: Wait, retry within the limit, or fail in a clear way.
- Rate limit: Slow down and honor the wait time from the tool.
- Bad input: Reject it. Do not retry.
- Task hangs: Stop it at the hard time limit.
- Restart during work: Check saved state before retrying.
- Same task arrives twice: Use the task ID to avoid doing it twice.
- Secret appears in a log: Hide it, block more leaks, and follow the key change plan.
- Cost rises fast: Pause new work and check loops or large inputs.
- Bad update: Roll back. Do not patch the live release by hand.
- Audit log fails: Block high-risk actions until logs work again.
- Kill switch fails: Stop the service at the host or container level.
Concrete Example
A support agent runs all day as a systemd service.
Release: support-agent:1.4.2
Task time limit: 5 minutes
Retry limit: 2
Cost limit: $0.25 per task
Kill switch: AGENT_ACCEPT_NEW_TASKS=false
Update plan: 5%, 25%, 50%, then 100%
Rollback point: support-agent:1.4.1For each task, the service records:
task_id=help-4821
release=1.4.2
result=success
retries=1
time_seconds=42
cost_usd=0.07If failures pass the set limit, stop the update. Keep old tasks safe. Send new tasks to version 1.4.1. Save a few clean logs, find the cause, test the fix, and try again with 5 percent of the work.
Run Tools
This skill can work with:
- PM2
systemd- Container tools
- Build and release checks
The same rules apply to each tool. Use fixed releases, small updates, clear limits, health checks, a kill switch, and a tested rollback.