- agent-evals
Use when measuring how well an AI agent handles a recurring kind of task. Builds a small eval set, runs it, and raises the difficulty once the agent passes.
Updated
- agent-security-basics
Use when letting AI agents run code, install packages or touch real systems. Covers sandboxing, least privilege and quarantine basics.
Updated
- architecture-for-agents
Use when structuring a codebase that AI agents will edit heavily. Weighs duplication against abstraction now that repeating code and keeping copies in sync is cheap.
Updated
- async-delegation
Use when you have several independent jobs for AI agents. Splits work into parallel tasks, hands them off like to coworkers, and reviews results in batches.
Updated
- black-box-review
Use when reviewing code an AI agent wrote, especially in a language or area you don't know well. Judges the result through tests and behaviour instead of reading every line.
Updated
- cli-for-your-app
Use when making an app or library usable by a person's own AI agent. Designs a command-line interface so agents can drive the app without a built-in chatbot.
Updated
Curated setups
Skills grouped for one workflow and installable with one command.