All skills
mblode avatar

/ax-audit

@57eb304
by Matthew Blodemblode/agent-skills134 stars
12

Audits agentic products for tool parity, authority, approval payloads, recovery, and trust using 27 rules and a ship verdict. Use when asked for an "AX audit", to review an agent approval flow, or whether an agent can operate the product. For human-facing API ergonomics use dx-audit; for ordinary UI use ui-design.

Use this Skill: https://skilld.dev/gh/mblode/agent-skills/ax-audit

This session only. Nothing lands on disk.

rules-archcontext-no-checkpoint-resume.md

≈780 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Long-running agent with no checkpoint/resume

Agent runs a multi-step task with no durability. Browser closes, network drops, session times out, all progress lost. User starts from scratch. Violates Improvement Over Time: completed work should survive interruption.

What goes wrong

User asks the agent to refactor 15 files. Agent completes 12 over 4 minutes. Laptop sleeps. On reconnect the session is gone, with no record of what was done. Agent redoes all 15, possibly making different choices.

Detection

Surfaces: agent-dashboard

Static signals:

  1. Find execution loops or multi-step task handlers.
  2. Check whether they persist state between iterations.
  3. Check whether resume/recovery logic exists.
  4. Flag multi-step agents with no checkpoint writes.

Runtime signals: Agent runs >60s with no state persistence. Reconnect restarts from scratch.

Judgment signals:

  • A client resumeStream (AI SDK 7) resumes only if the server kept the run; a Claude Agent SDK PreToolUse hook returning defer persists the session across an approval wait. Either is a pass only when the server side is there; a reconnect helper over a run that died with the request is not.

Concrete commands:

rg '(for\s*\(|while\s*\(|for await)' --type=ts -A 5 src/ | rg -B 1 '(toolCall|executeStep|runTool)'
rg '(checkpoint|saveState|persistSession|saveProgress)' --type=ts src/
rg '(resume|recover|restoreSession|loadCheckpoint|resumeStream|defer)' --type=ts src/

False-positive guards:

  • Skip files with // ax-audit-ignore:context-no-checkpoint-resume.
  • Skip single-step agents (no loop, single tool call).
  • Skip agents reliably under 10 seconds.
  • Skip test files and fixtures.

Fix

// before
async function refactorFiles(files: string[]) {
  for (const file of files) await agent.refactor(file);
}

// after: checkpoint after each step, resume on reconnect
async function refactorFiles(sessionId: string, files: string[]) {
  const cp = await loadCheckpoint(sessionId);
  const done = new Set(cp?.completed ?? []);
  for (const file of files) {
    if (done.has(file)) continue;
    await agent.refactor(file);
    done.add(file);
    await saveCheckpoint(sessionId, { completed: [...done], updatedAt: Date.now() });
  }
}

Default tier and overrides

Defaults to: backlog

Surface Tier
Agent tool execution fix-this-sprint
Agent chat backlog
Agent config backlog
Agent dashboard backlog

Examples

Anti-pattern (fails): for (const t of tasks) await agent.execute(t): tab closes at task 8, all lost.

Applied (passes): Loop resumes from loadCheckpoint(id) index, calls saveCheckpoint after each step.

Suppression

// ax-audit-ignore:context-no-checkpoint-resume, sub-second operation
await agent.formatSingleFile(filePath);

Source: SKILL.md on GitHub

No alerts13d3 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    The skill is a specialized auditing framework for AI agent products, focusing on architectural integrity and user trust. It uses standard shell tools for static analysis of codebases. The analysis found no malicious behavior, obfuscation, or data exfiltration risks.

  • Socket13d

    No alerts

  • Snyk13d

    Risk: LOW · No issues

Signed by skilld at 57eb304. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 2 weeks ago

README badge

README badge for mblode/agent-skills/ax-audit