All skills
mblode avatar

/ax-audit

@57eb304
by Matthew Blodemblode/agent-skills134 stars
12

Audits agentic products for tool parity, authority, approval payloads, recovery, and trust using 27 rules and a ship verdict. Use when asked for an "AX audit", to review an agent approval flow, or whether an agent can operate the product. For human-facing API ergonomics use dx-audit; for ordinary UI use ui-design.

Use this Skill: https://skilld.dev/gh/mblode/agent-skills/ax-audit

This session only. Nothing lands on disk.

rules-axtrust-no-uncertainty-markers.md

≈701 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Agent presents everything with equal certainty

Agent is 95% sure of one recommendation, 40% of another, but both render identically. When the 40% answer is wrong, the user distrusts not just it but everything. Confident wrong answers cause permanent trust damage.

What goes wrong

Two recommendations in one response: one well-supported, one a guess. Same font, weight, formatting. User treats both as equally reliable; the guess is wrong; now they second-guess every future response. Trust is binary when the interface gives no gradient.

Detection

Surfaces: agent-chat

Auditability: observational

Static signals:

  1. Find agent output containers.
  2. Check for confidence props (confidence, certainty, score) or uncertainty components.
  3. Absence of all = flag.

Concrete commands:

rg 'confidence|certainty|ConfidenceBadge|UncertaintyIndicator' --type=ts src/

Judgment signals:

  • Hedging in prompt instructions is weaker than structured indicators but better than nothing.
  • A badge always showing "high" is not meaningful: check for actual variation.

False-positive guards:

  • Skip // ax-audit-ignore:trust-no-uncertainty-markers, test, and Storybook files.
  • Skip trivial outputs (confirmations, acknowledgments) where confidence is always 100%.

Fix

Add confidence indicators: numeric score, visual badge (high/medium/low), hedging language, or expandable reasoning that shows uncertainty.

Examples

Anti-pattern (fails):

<ul>
  {recommendations.map((rec) => (
    <li key={rec.id}>{rec.text}</li>
  ))}
</ul>

Applied (passes):

<ul>
  {recommendations.map((rec) => (
    <li key={rec.id}>
      {rec.text}
      <ConfidenceBadge level={rec.confidence > 0.8 ? "high" : "low"} />
    </li>
  ))}
</ul>

Default tier and overrides

Defaults to: fix-this-sprint

Surface Tier
Agent tool execution fix-this-sprint
Agent chat fix-this-sprint
Agent config backlog
Agent dashboard fix-this-sprint

No tool-execution bump. Hedging in prose changes nothing about what a tool did, and an observational rule that can only return unknown on static evidence should not be the single finding that flips a verdict. The blockers on that surface are the gate, its payload, and the escape hatch.

Suppression

{/* ax-audit-ignore:trust-no-uncertainty-markers, deterministic lookups, no uncertainty */}
<AgentRecommendation text={result.text} />

Source: SKILL.md on GitHub

No alerts13d3 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    The skill is a specialized auditing framework for AI agent products, focusing on architectural integrity and user trust. It uses standard shell tools for static analysis of codebases. The analysis found no malicious behavior, obfuscation, or data exfiltration risks.

  • Socket13d

    No alerts

  • Snyk13d

    Risk: LOW · No issues

Signed by skilld at 57eb304. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 2 weeks ago

README badge

README badge for mblode/agent-skills/ax-audit