All skills
lukemurraynz avatar

/hitl-design-patterns

@2cc2455

Design human-in-the-loop (HITL) approval gates for AI agent workflows. Use when you need to decide "where do I put a human approval gate", "should this agent action need confirmation", "classify this action's reversibility", "design the confirmation UX", "set a timeout and fallback policy", "define an escalation path", or "wire HITL into MCP / MAF / CopilotKit / Copilot Studio". Covers maker-checker / four-eyes approval, confirmation dialogs, fail-closed timeouts, and audit. Do not use for general security review, runtime policy enforcement (use agent-governance-toolkit instead), or fully automated loops with no human decision point. Pairs with owasp-agentic for the threat-model perspective (ASI09).

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/hitl-design-patterns

This session only. Nothing lands on disk.

CHANGELOG.md

≈1.8k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Changelog — hitl-design-patterns

[1.1.0] - 2026-08-25

Three-tier approval chain pattern from production agent-harness source review (Claude Code snapshot, 2026-03).

Added

  • Tiered Decision Chain: Rules → Classifier → Human section: deterministic allowlist/denylist tier, LLM classifier tier (capped at Recoverable actions), human gate tier. Design rules: classifier may never approve Irreversible actions; run lower tiers before rendering prompts (approval-fatigue mitigation); feed denials back to the model as corrective context to prevent retry loops; atomic claim-and-resolve for concurrent decision handlers; audit every tier's decision with the deciding tier recorded.

Changed

  • Frontmatter version 1.0.2 → 1.1.0.

[1.0.2] - 2026-08-16

Recovery-escalation patterns from "When AI Agents Fail: Engineering Reliable Recovery with Microsoft Foundry" (Microsoft Foundry blog, 2026-08-13) deep review.

Added

  • Escalation as a recovery path (not only an approval gate) subsection: humans as recovery path of last resort for unverifiable unknown state, unavailable verification tooling, or created duplicates. Decision-grade context handoff checklist (request, operation/tool, side-effect + verification status, records found, recommended action, approve/reject consequences). Cross-links the microsoft-agent-framework tool failure semantics reference for action-ledger schema and recovery decision logic.
  • Two escalation triggers added to the table: unverifiable unknown state (no blind retry), conflicting data sources / authorization-policy conflict on side-effecting actions (escalate rather than let the model arbitrate).

[Unreleased] - 2026-07-18

AzureFeeds newsletter sweep — citation only.

Added

[1.0.1]

[1.0.1] - 2026-06-16

Quality-improvement pass (1 review round, 3 parallel reviewer angles + self-run content-accuracy and scorer/structural angles). Baseline deterministic scorer 50/100 (LAUNCH-FLOOR FAIL) → 100/100 (Excellent, LAUNCH-FLOOR PASS). Validator green with 0 warnings.

External validation baseline

  • Official source: MCP ToolAnnotations (spec rev 2025-03-26), CopilotKit useHumanInTheLoop, MAF HITL (Microsoft Learn) checked.
  • Official docs: OWASP Top 10 for Agentic Applications v1.0 (Dec 2025) — ASI09 confirmed.
  • Community: OWASP ASI mapping repos cross-checked the ASI09 code/title.

Pack A — correctness, structure, robustness (SKILL.md)

  • MCP annotations corrected: "destructive": true → standard destructiveHint: true, plus the full standard hint set (title, readOnlyHint, idempotentHint, openWorldHint). x-human-approval-required / x-approval-class now explicitly labelled custom, non-standard extensions.
  • MAF section corrected: named the real mechanism (RequestPort / RequestInfoExecutor, request/response handling); the prior type: human_in_the_loop YAML is now marked ILLUSTRATIVE — not a literal MAF schema.
  • CopilotKit corrected: replaced the imprecise "renderConfirmation must block until resolve" with the documented respond-callback pause/resume pattern.
  • Added ## Guardrails section: never fabricate framework APIs (mark illustrative), default fail-closed, no soft gate on irreversible, no agent-satisfied non-agentable gate, ask when reversibility is unclear, no invented thresholds/approvers.
  • Added ## When NOT to Use section delegating to agent-governance-toolkit, owasp-agentic, autonomous-agent-loops.
  • Rewrote description with explicit quoted triggers, a "Use when…" routing phrase, and a "Do not use for…" exclusion clause; added escalation path to argument-hint.
  • Fail-closed extended to audit-write failure in prose and in the requestApproval TypeScript example.
  • Added an end-to-end worked example stitching classification → gate → UX → timeout → audit.
  • Reference style standardised: agentic-security.instructions.md → owasp-agentic (matches the line-15 reference for the same ASI09 content).
  • Build26 session codes removed (stale-prone) in favour of surface-type descriptions.
  • Thresholds: added guidance to set explicit written numbers for "bulk"/"blast radius"; added a "one tier more dangerous when unknown" fallback; gave the 60-second re-gate rule a rationale.
  • Keyword discoverability: added maker-checker, four-eyes, confirmation dialog, function/tool-calling synonyms.
  • Copilot Studio: added caveat to prefer a signed token over a bare boolean gate variable.
  • Added cowork: block (category automation).

GitHub / live source validation

  • INCORRECT → fixed: MCP destructive field name (now destructiveHint).
  • INCORRECT → labelled: MAF human_in_the_loop step type (now illustrative; real mechanism named).
  • VERIFIED: OWASP ASI09 "Human-Agent Trust Exploitation"; CopilotKit useHumanInTheLoop hook exists.
  • UNVERIFIED (left illustrative): exact MAF/CopilotKit symbol signatures against any specific installed version — flagged in-text with last-checked date.

Token efficiency

  • Before: ~2,188 est tokens (1,683 words). After: ~3,641 est tokens (2,801 words).
  • Method: wc -w × 1.3 (±30%; tiktoken not available).
  • Net increase is added capability (Guardrails, When NOT to Use, anti-fabrication discipline, fail-closed audit logic, worked example), not bloat. One dedup: merged the overlapping "Use When" + "When to Use" tables.
  • Capability preserved: reversibility matrix, five-element confirmation UX, fail-closed timeout code, escalation/four-eyes, independent append-only audit.

Validator

  • validate_skill.py: PASS — 0 errors, 0 warnings.
  • Deterministic scorer: 100/100 (Excellent), LAUNCH-FLOOR PASS, faithfulness PASS, safety destructive=WARN (down from FAIL).
  • Security scan (bandit + prose): PASS.

Other validation

  • Markdown/link checks: not available (no link checker in environment); cross-references inspected manually.
  • Tests: not available (documentation/guidance skill — no executable test surface).
  • Build/compile: not available; code blocks are illustrative (tsx/ts/json/yaml), statically inspected.
  • Smoke test: validator + scorer serve as the static smoke test.

Honest unknowns

  • Exact MAF / CopilotKit symbol signatures against a specific installed package version (left illustrative with last-checked date).
  • Whether sibling skills owasp-agentic, copilotkit-agui, agent-governance-toolkit, autonomous-agent-loops exist in the consumer's bundle (out of scope — this bundle ships only SKILL.md; references degrade gracefully).

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    This skill provides defensive design patterns for implementing Human-in-the-Loop (HITL) approval gates. It focuses on security best practices such as action reversibility classification, fail-closed timeouts, and mandatory audit logging for AI agent workflows.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
cowork
{
  "category": "automation",
  "icon": "shield-check"
}
metadata
{
  "last_verified": "2026-08-16",
  "version": "1.1.0"
}
Other metadata
argument-hint
<agent action | approval UX | timeout policy | reversibility classification | escalation path | framework>

README badge

README badge for lukemurraynz/hve-agent-skills/hitl-design-patterns