All skills
mblode avatar

/ax-audit

@57eb304
by Matthew Blodemblode/agent-skills136 stars
12

Audits agentic products for tool parity, authority, approval payloads, recovery, and trust using 27 rules and a ship verdict. Use when asked for an "AX audit", to review an agent approval flow, or whether an agent can operate the product. For human-facing API ergonomics use dx-audit; for ordinary UI use ui-design.

Use this Skill: https://skilld.dev/gh/mblode/agent-skills/ax-audit

This session only. Nothing lands on disk.

rules-archgranularity-workflow-shaped-tool.md

≈915 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Tool bundles decision logic instead of being atomic

A tool like analyze_and_organize(folder) bundles judgment into code. Changing what "organize" means needs a code refactor, not a prompt edit, and the agent can't apply judgment to intermediate steps.

What goes wrong

analyze_and_organize_inbox scans emails, decides importance, files them. User asks "Why did you archive that?" The agent can't change the logic: it's hardcoded. With atomic primitives, the agent decides itself.

Detection

Surfaces: agent-config

Static signals:

  1. Grep tool definitions for compound names (_and_, _then_, _with_).
  2. Check implementations for branching logic making domain decisions.
  3. Count distinct API calls per tool; >1 suggests bundling.

Concrete commands:

rg 'name:\s*["\x27]\w+_(and|then|with)_\w+' --type=ts src/tools/
rg 'name:\s*["\x27](process|handle|manage|analyze|organize|auto)_' --type=ts src/tools/
rg -l 'tool\(|defineTool' --type=ts src/tools/ | while read f; do
  rg -c --with-filename 'if\s*\(|switch\s*\(' "$f"
done | awk -F: '$2>3'

Judgment signals:

  • Consolidation that removes mechanical chaining passes. A schedule_event that finds availability and books is one user action with no step the user would want to veto, and fewer tools of that kind is the guidance Anthropic gives for tool sets. The fail is a bundled decision (which role to assign, what counts as stale): that is the step the user will disagree with and the agent cannot change. Test: is there an intermediate step a user would want to see or override?

False-positive guards:

  • Skip atomic transactions (e.g., transfer_funds) and // ax-audit-ignore:granularity-workflow-shaped-tool.

Fix

Split into atomic primitives. Let the agent decide what to move and where.

// before: analyze_and_organize_inbox: after: atomic primitives
export const listEmails = tool({ name: "list_emails", /* ... */ });
export const readEmail = tool({ name: "read_email", /* ... */ });
export const moveEmail = tool({ name: "move_email", /* ... */ });

Default tier and overrides

Defaults to: fix-this-sprint: works until the user disagrees with a bundled decision.

Surface Tier
Agent tool execution fix-this-sprint
Agent config fix-this-sprint

No tool-execution bump: a bundled tool still does what it says, it just does too much of it. The blocker in that neighbourhood is an ungated action (comm-no-approval-gate), not a coarse one.

Examples

Anti-pattern (fails):

export const processNewUser = tool({
  name: "process_and_configure_new_user",
  execute: async ({ email, name }) => {
    const user = await api.post("/users", { email, name });
    await api.post(`/users/${user.id}/roles`, { role: "member" });  // can't choose role
    await api.post("/emails/send", { to: email, template: "welcome" }); // can't skip
  },
});

Applied (passes):

export const createUser = tool({ name: "create_user", /* ... */ });
export const assignRole = tool({ name: "assign_role", /* ... */ });
export const sendEmail = tool({ name: "send_email", /* ... */ });
// Agent decides: skip welcome email, assign admin role

Suppression

// ax-audit-ignore:granularity-workflow-shaped-tool, atomic transaction
export const transferFunds = tool({ name: "transfer_funds", /* ... */ });

Source: SKILL.md on GitHub

No alerts13d3 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    The skill is a specialized auditing framework for AI agent products, focusing on architectural integrity and user trust. It uses standard shell tools for static analysis of codebases. The analysis found no malicious behavior, obfuscation, or data exfiltration risks.

  • Socket13d

    No alerts

  • Snyk13d

    Risk: LOW · No issues

Signed by skilld at 57eb304. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 32 minutes ago.

Activeupdated 2 weeks ago

README badge

README badge for mblode/agent-skills/ax-audit