All skills
mblode avatar

/ax-audit

@57eb304
by Matthew Blodemblode/agent-skills134 stars
12

Audits agentic products for tool parity, authority, approval payloads, recovery, and trust using 27 rules and a ship verdict. Use when asked for an "AX audit", to review an agent approval flow, or whether an agent can operate the product. For human-facing API ergonomics use dx-audit; for ordinary UI use ui-design.

Use this Skill: https://skilld.dev/gh/mblode/agent-skills/ax-audit

This session only. Nothing lands on disk.

rules-archcomm-no-completion-signal.md

≈918 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Agent completion detected by heuristic instead of explicit signal

Orchestrator detects "done" by counting idle iterations, checking output files, or waiting for a timeout. A thinking pause looks like completion; slow API calls trigger premature termination. Violates Parity: the UI must receive an explicit signal, not guess.

What goes wrong

Agent researches a complex question. Makes 3 tool calls, then pauses 8 seconds composing a response. Orchestrator counts 2 idle iterations, hits maxIdleIterations: 2, terminates. User sees a truncated answer.

Detection

Surfaces: agent-tool-execution, agent-dashboard

Static signals:

  1. Find the orchestrator control loop that decides continue/stop.
  2. Check for idle-counting, timeout-based completion, or file-existence as termination.
  3. Check which terminal reasons the loop handles. end_turn alone is not enough: pause_turn means resend and continue, max_tokens means truncated, so a loop that returns on "anything but tool_use" presents a cut-off answer as complete. The full list is in references/framework-signals.md.
  4. Flag any heuristic used as the primary completion signal.

Concrete commands:

rg '(consecutiveIdle|noToolCall|idleCount|maxIdle)' --type=ts src/
rg '(setTimeout|setInterval)' --type=ts -A 5 src/ | rg '(done|complete|finish|terminate)'
rg '(stop_reason|stopReason|finishReason|end_turn|pause_turn|RUN_FINISHED|shouldContinue)' --type=ts src/

False-positive guards:

  • Skip files with // ax-audit-ignore:comm-no-completion-signal.
  • Skip timeout logic alongside an explicit signal (both stop_reason AND setTimeout).
  • Skip test files and fixtures.

Fix

// before: heuristic completion
let idle = 0;
while (idle < 3) {
  const res = await llm.chat(messages);
  if (!res.toolCalls.length) { idle++; continue; }
  idle = 0;
  await executeTools(res.toolCalls, messages);
}

// after: every terminal reason handled explicitly, none inferred
while (true) {
  const res = await llm.chat(messages);
  switch (res.stopReason) {
    case "tool_use":
      for (const tc of res.toolCalls) {
        if (tc.name === "task_complete") return { status: "complete", summary: tc.args.summary };
        messages.push({ role: "tool", content: await executeTool(tc) });
      }
      continue;
    case "pause_turn":  continue;                                        // server tools mid-run: resend, not done
    case "end_turn":    return { status: "complete", content: res.content };
    case "max_tokens":  return { status: "truncated", content: res.content };
    default:            return { status: "failed", reason: res.stopReason }; // refusal, context window
  }
}

Default tier and overrides

Defaults to: release-blocker

Surface Tier
Agent tool execution release-blocker
Agent dashboard release-blocker
Agent chat fix-this-sprint
Agent config backlog

Examples

Anti-pattern (fails): while (noToolCalls < 2): thinking pause triggers premature termination.

Applied (passes): if (res.stopReason === "end_turn") return res.content: explicit model signal.

Suppression

// ax-audit-ignore:comm-no-completion-signal, timeout is safety net, primary signal is stop_reason
const SAFETY_TIMEOUT = 120_000;

Source: SKILL.md on GitHub

No alerts13d3 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    The skill is a specialized auditing framework for AI agent products, focusing on architectural integrity and user trust. It uses standard shell tools for static analysis of codebases. The analysis found no malicious behavior, obfuscation, or data exfiltration risks.

  • Socket13d

    No alerts

  • Snyk13d

    Risk: LOW · No issues

Signed by skilld at 57eb304. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 2 weeks ago

README badge

README badge for mblode/agent-skills/ax-audit