All skills
affaan-m avatar

/agent-harness-construction

@c70874f

Design and optimize AI agent action spaces, tool definitions, and observation formatting for higher completion rates. Use when defining or revising an agent's tool set, action space, or observation format.

  • 1 file
  • 2.1 KB
  • Updated 2 months ago
  • GitHub

Use this Skill: https://skilld.dev/gh/affaan-m/everything-claude-code/agent-harness-construction

This session only. Nothing lands on disk.

SKILL.md

≈58 tokens always: the name and description. ≈454 when used: this file.

Agent Harness Construction

Use this skill when you are improving how an agent plans, calls tools, recovers from errors, and converges on completion.

Core Model

Agent output quality is constrained by:

  1. Action space quality
  2. Observation quality
  3. Recovery quality
  4. Context budget quality

Action Space Design

  1. Use stable, explicit tool names.
  2. Keep inputs schema-first and narrow.
  3. Return deterministic output shapes.
  4. Avoid catch-all tools unless isolation is impossible.

Granularity Rules

  • Use micro-tools for high-risk operations (deploy, migration, permissions).
  • Use medium tools for common edit/read/search loops.
  • Use macro-tools only when round-trip overhead is the dominant cost.

Observation Design

Every tool response should include:

  • status: success|warning|error
  • summary: one-line result
  • next_actions: actionable follow-ups
  • artifacts: file paths / IDs

Error Recovery Contract

For every error path, include:

  • root cause hint
  • safe retry instruction
  • explicit stop condition

Context Budgeting

  1. Keep system prompt minimal and invariant.
  2. Move large guidance into skills loaded on demand.
  3. Prefer references to files over inlining long documents.
  4. Compact at phase boundaries, not arbitrary token thresholds.

Architecture Pattern Guidance

  • ReAct: best for exploratory tasks with uncertain path.
  • Function-calling: best for structured deterministic flows.
  • Hybrid (recommended): ReAct planning + typed tool execution.

Benchmarking

Track:

  • completion rate
  • retries per task
  • pass@1 and pass@3
  • cost per successful task

Anti-Patterns

  • Too many tools with overlapping semantics.
  • Opaque tool output with no recovery hints.
  • Error-only output without next steps.
  • Context overloading with irrelevant references.

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at c70874f. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 10 hours ago.

Activeupdated 2 months ago
metadata
{
  "origin": "ECC"
}
  • agent-design
  • tool-definitions
  • action-space
  • observation-formatting
  • error-recovery
  • context-budgeting
  • ai-agents
  • prompt-engineering
  • react-pattern

README badge

README badge for affaan-m/everything-claude-code/agent-harness-construction

Teaches how to design tool definitions, action spaces, and error recovery paths for AI agents to improve task completion rates. Covers tool granularity (micro, medium, macro), observation formatting with status and next-steps, context budgeting, and ReAct versus function-calling trade-offs. Applies to any agent framework where you control the tool interface and error handling contract.

Generated from the current SKILL.md.

Does this skill apply to specific AI models or frameworks?
No. The skill describes general principles for designing agent action spaces, tool definitions, and observation formatting that work across different models and agent frameworks.
What's the difference between micro-tools, medium tools, and macro-tools?
Micro-tools handle high-risk operations (deploy, migration, permissions) in isolation. Medium tools handle common edit/read/search loops. Macro-tools are used only when round-trip overhead dominates cost.
What should every tool response include?
Each response should include status (success/warning/error), a one-line summary, actionable next_actions, and artifacts (file paths or IDs).
When should I use ReAct versus function-calling?
Use ReAct for exploratory tasks with uncertain paths, function-calling for structured deterministic flows, or hybrid (ReAct planning plus typed tool execution) for most cases.
What metrics should I track to benchmark agent performance?
Track completion rate, retries per task, pass@1 and pass@3, and cost per successful task.

Generated from the current SKILL.md. These answers refresh after source changes.