Human-in-the-Loop Design Patterns
HITL approval gates are design components, not just security controls. This skill covers the positive design pattern: when to require human checkpoints (approval gates), what to show the user in a confirmation dialog, how to handle timing, and how to fail safely. For the threat-model perspective (ASI09 Human-Agent Trust Exploitation), load owasp-agentic. For CopilotKit's useHumanInTheLoop hook, load copilotkit-agui. Both are sibling skills; if they are not present in your bundle, the patterns below still stand on their own.
When to Use
| User intent | Output |
|---|---|
| Deciding where to place human approval gates | Reversibility classification matrix, gate placement rules |
| Designing the confirmation UX (confirmation dialog) for an agent action | Required elements checklist, component structure |
| Handling timeout when the human does not respond | Timeout policy patterns, fail-closed fallback |
| Wiring HITL into a specific framework | MAF, CopilotKit, MCP annotation, Copilot Studio patterns |
| Classifying an action's blast radius | Impact Γ reversibility decision table |
| Defining maker-checker / four-eyes approval | Escalation paths, multi-approver rules |
When NOT to Use
| Situation | Use instead |
|---|---|
| Runtime policy enforcement / guardrail execution (not gate design) | use agent-governance-toolkit |
| Threat-model framing of trust exploitation (ASI09) | use owasp-agentic |
| Bounded automation where the main challenge is loop control, not a human decision | use autonomous-agent-loops |
| General security review with no approval-gate question | a security-review skill |
| A fully automated workflow with no human decision point | no HITL gate is needed |
Do not steal these triggers: this skill owns gate design (placement, UX, timeout, escalation, audit), not runtime enforcement or threat modelling.
Action Reversibility Classification
Classify every agent-executable action before deciding on a gate. The classification drives the gate type and the confirmation elements required.
| Class | Reversibility | Scope | Gate required | Examples |
|---|---|---|---|---|
| Safe | Fully reversible | Single resource | None | Read, search, generate draft |
| Recoverable | Reversible with effort | Single resource | Soft gate (confirm + undo path) | Send email, create record, post comment |
| Hard to reverse | Partial rollback only | Multiple resources or pipeline | Hard gate (explicit confirm + preview) | Bulk update, publish, deploy to staging |
| Irreversible | Cannot be undone | Any | Non-agentable gate | Delete data, send payment, deploy to prod, revoke access |
When you cannot determine the class, treat the action as one tier more dangerous than your best guess (Recoverable β Hard to reverse). Never default an unknown action to Safe.
Non-agentable gate: Approval must come from a human outside the prompt stream - a real-time UI surface the agent does not render, a signed change ticket, a CI gate, or a CODEOWNERS review. The agent must not be able to satisfy this approval by generating text in the conversation, by rendering its own approval UI, or by parsing an "approved" string it produced. Verify the gate is wired to an external approval service with its own identity check, and prefer a signed approval token over a plain boolean. See owasp-agentic ASI09 (Human-Agent Trust Exploitation).
Gate Placement Rules
- Place the gate before the action, not after. An agent that acts first and asks forgiveness later violates the HITL contract.
- Gate every action classified as Hard to reverse or Irreversible.
- Gate Recoverable actions when the blast radius exceeds a single user's data, the action runs at bulk scale, or the context is production.
- Set explicit, written thresholds for "bulk" and "blast radius" - they are context-dependent, so pick numbers and record them (e.g. "bulk = >25 records or >1 recipient list"; "blast radius = touches more than the requesting user's own resources"). Do not leave them to interpretation.
- Do not gate Safe actions - unnecessary gates train users to click through without reading (approval fatigue).
- Combine gates for compound actions: A-then-B requires one gate covering both, not two separate confirmations.
- Re-gate a long-running multi-step operation when the approved context goes stale. A practical default is re-gate after ~60 seconds with no new human instruction, because the human's mental model of what they approved drifts as the agent takes further autonomous steps; tune the window to your step size and risk.
Tiered Decision Chain: Rules β Classifier β Human
Not every gated action needs a human on the first pass. Production agent harnesses (verified in the Claude Code source snapshot, 2026-03) resolve tool approvals through a three-tier chain, escalating to the next tier only when the previous one cannot decide safely:
| Tier | Decider | Decides | Escalates when |
|---|---|---|---|
| 1. Deterministic rules | Allowlist / denylist, path checks, static analysis | Clearly-safe reads and clearly-forbidden writes | Action matches neither list |
| 2. Classifier (LLM-as-judge) | A separate model call evaluating the action against policy | Ambiguous but low-risk actions | Confidence below threshold, or action class β₯ Hard-to-reverse |
| 3. Human gate | The confirmation UX from this skill | Everything the lower tiers could not clear | β |
Design rules for the chain:
- The classifier tier may never approve Irreversible actions. Cap it at Recoverable; Hard-to-reverse and above go straight to the human tier regardless of classifier confidence.
- Run tiers 1β2 before rendering any prompt so users only see gates that genuinely need judgment - this is also your approval-fatigue mitigation.
- Feed denials back to the model as corrective context ("denied: <reason>"), not as silent failures. A denial the model can learn from prevents repeat attempts; a silent denial produces retry loops.
- Resolve each decision exactly once. When multiple handlers (hook, classifier callback, user click, abort) can decide concurrently, use an atomic claim-and-resolve primitive - check-then-set leaves a window where two handlers both act on the same pending request.
- Audit every tier's decision, not just the human's: record which tier decided and why, so post-incident review can tell "classifier said no" from "user said no".
This chain is an approval-routing pattern; it does not change reversibility classification or gate placement rules above.
Confirmation UX: Required Elements
A confirmation component (confirmation dialog) must include all five elements. Missing elements cause users to approve blindly.
| Element | Content | Required for |
|---|---|---|
| Action summary | One sentence: what the agent is about to do | All gates |
| Target | The specific resource, record, or system (not a generic description) | All gates |
| Consequence | What changes and whether it can be undone | All gates |
| Preview | Diff, rendered preview, or sample of the output before application | Hard-to-reverse and Irreversible |
| Undo path | How to reverse if approved in error - or "This action cannot be undone" | All gates |
// Minimum confirmation component structure. On deny, the agent must STOP the
// gated action (not retry, not work around) and record the denial in the audit log.
<ConfirmationCard
summary="Send weekly report to all 847 subscribers"
target="Marketing list: newsletter-subscribers@example.com"
consequence="Sends immediately. Recipients cannot be recalled."
preview={<EmailPreview subject={draft.subject} body={draft.body} />}
undoPath="Cannot be unsent. Contact support within 5 minutes to attempt suppression."
onConfirm={handleConfirm}
onDeny={handleDeny} // halt + audit; do not proceed
/>Timeout and Fallback Design
Define timeout policy for every gate. A gate with no timeout becomes an indefinite block.
| Scenario | Recommended timeout | On timeout |
|---|---|---|
| Synchronous user-facing gate | 5β15 minutes | Fail closed (cancel) |
| Async approval (manager workflow) | 24β48 hours | First timeout β escalate; second timeout β cancel |
| Batch operation (multiple approvals) | Timeout per step; cancel remainder on first timeout | Cancel |
| Production deployment | No gate timeout - the CI/CD pipeline owns the timeout | Pipeline decides |
Fail closed, not open. When a timeout fires, the default is to cancel the action. Proceeding without approval after timeout is a silent authorisation bypass. The same rule applies if the audit write fails: if the gate cannot record its decision, treat it as a failed gate and cancel - never proceed un-audited.
async function requestApproval(
action: AgentAction,
timeoutMs = 10 * 60 * 1000,
): Promise<boolean> {
const result = await Promise.race([
waitForHumanDecision(action),
sleep(timeoutMs).then(() => ({ decision: "timeout" as const })),
]);
try {
await auditLog.record({ action, outcome: result.decision, timestamp: new Date() });
} catch {
return false; // Audit write failed -> fail closed, do not proceed un-audited
}
if (result.decision === "timeout") {
return false; // Fail closed
}
return result.decision === "approved";
}Escalation Paths
| Trigger | Escalation |
|---|---|
| Blast radius spans more than one team | Require manager approval in addition to requestor |
| Irreversible action in production | Require two independent approvers (four-eyes / maker-checker principle) |
| Approval not received within timeout | Route to on-call escalation queue |
| Same action repeated by the same agent in rapid succession | Flag as anomaly, pause for human review |
| Unverifiable unknown state (tool timed out after a side effect and state verification cannot resolve it) | Escalate as recovery path β do not blind-retry the action |
| Conflicting data sources or authorization/policy conflict on a side-effecting action | Escalate rather than retry or let the model arbitrate |
Escalation as a recovery path (not only an approval gate)
HITL is usually designed as a gate before an action. For side-effecting agent operations, humans are also the recovery path of last resort when autonomous recovery is unsafe: unknown state that cannot be verified, verification tooling unavailable, or a duplicate already created. Pattern source: "When AI Agents Fail: Engineering Reliable Recovery with Microsoft Foundry" (Microsoft Foundry blog, 2026-08-13).
A recovery escalation must hand the human decision-grade context, never a bare "the agent failed":
- Original user request
- Operation attempted and the tool called
- Whether a side effect was possible, and whether downstream state was verified
- What downstream records were found (e.g. "PR-10482 exists and PR-10483 also exists")
- Recommended recovery action, and what happens on approve vs reject
Pair this with an action ledger (durable record of attempted operations with idempotency keys) so the human (and post-incident forensics) can reconstruct what was attempted and what may have changed external state. See the microsoft-agent-framework skill's tool failure semantics reference for the ledger schema and the recovery decision logic that decides when to escalate.
Document the escalation path in the agent's capability manifest so the operations team can verify it without reading the source code. The manifest is the owner of record for who approves what; reconcile it against the audit log periodically so a manifest that says "manager approval" cannot silently diverge from logs that show otherwise.
Authoritative citation: Microsoft least-privilege for agents (July 2026)
The canonical scenarios for how identity + RBAC + scope + safe tool binding interact with approval gates live in the sibling agent-governance-toolkit skill β "Microsoft least-privilege guidance (July 2026)" section, citing "Least privilege for AI agents: Identity, access, and tool binding" (Microsoft Security blog, July 16, 2026). Cross-link rather than duplicate.
Key principles the HITL designer should carry from that section:
- Approval gates alone are insufficient. Pair every HITL gate with scoped identity and tool-binding constraints so a mistakenly-approved action cannot escape its capability ring.
- Time-boxed elevation for sensitive scopes (write, deploy, RBAC, billing) is preferred over permanent grants, the HITL gate exists to authorize the elevation, not to compensate for over-broad standing access.
- Separation of agent identity from the developer/operator identity that provisioned it, so approval audit trails name the agent that acted, not the human who built it.
Pairs with the owasp-agentic ASI09 (Human-Agent Trust Exploitation) and ASI03 (Identity and Privilege Abuse) reference entries.
Audit Requirements
Every gate must produce an audit record regardless of outcome.
Required fields: agent identity, action description, target resource, requestor, approvers, decision (approved / denied / timeout), timestamp, conversation or session ID.
Store audit records outside the agent's own storage - a separate, append-only store the agent cannot rewrite, so an agent can never modify its own audit trail. For async approvals that span sessions, the session ID links the original request to the eventual decision; persist the record across the wait, not just in memory.
Framework Integration
Framework-API discipline. Field names and step types below are pinned to the dates noted. Treat any framework symbol you cannot confirm against the installed version as illustrative and verify before shipping (see Guardrails). Frameworks rename things between versions.
Computer-using and multi-surface agents
As agents act through desktops, Cloud PCs, local sandboxes, and multi-agent workflows, HITL gates must be explicit on each surface.
| Surface | HITL implication |
|---|---|
| Foundry / MAF production agents | Gates must be part of the workflow contract, not only UI copy |
| Computer-using agents (e.g. agents driving Windows / legacy GUI apps) | Previews and non-agentable approvals are mandatory for writes |
| Local agent sandboxes / containers | Containment reduces blast radius but does not replace approval for irreversible actions (see agent-governance-toolkit) |
For computer-using agents, the confirmation UX must show the exact application/window, target record or file, and proposed UI action. A screenshot-only preview is insufficient for regulated actions: pair it with a structured action summary (the parsed intent and target, not just pixels) and an undo path, so the human reviews the meaning of the action, not only its appearance.
MCP Tool Annotations
Use the standard MCP ToolAnnotations hints so MCP-aware clients can surface a confirmation UI before invoking the tool. Standard fields (MCP spec rev 2025-03-26, verified 2026-06-16): title, readOnlyHint, destructiveHint, idempotentHint, openWorldHint. These are hints, not enforcement - the client decides what to do with them; do not rely on an annotation alone to stop a destructive call.
{
"name": "deleteResource",
"description": "Permanently deletes a cloud resource",
"annotations": {
"title": "Delete cloud resource",
"readOnlyHint": false,
"destructiveHint": true,
"idempotentHint": false,
"openWorldHint": true,
"x-human-approval-required": true,
"x-approval-class": "irreversible"
},
"inputSchema": {
"type": "object",
"properties": {
"resourceId": { "type": "string" },
"resourceType": { "type": "string" }
},
"required": ["resourceId", "resourceType"]
}
}x-human-approval-required and x-approval-class are custom (non-standard) extensions, not part of the MCP spec - only clients you control will honour them. The spec-standard signal for "this is destructive" is destructiveHint: true.
CopilotKit
Use the useHumanInTheLoop hook (@copilotkit/react-core) - see copilotkit-agui for the full pattern. It pauses agent execution and renders your component while the tool is in the executing state; the agent stays paused until your render component calls the supplied respond callback with the user's decision. Key rule: do not resolve the tool until a human has acted. (Hook verified to exist 2026-06-16; confirm the exact import path and callback name against your installed CopilotKit version.)
Microsoft Agent Framework (MAF)
MAF implements HITL through request/response handling: an executor sends a request out of the workflow via a RequestPort (RequestInfoExecutor / request-response events) and waits for the external response before proceeding (Microsoft Learn, verified 2026-06-16). Drive the gate with that mechanism, and own the timeout in your executor logic - fail closed on no response.
# ILLUSTRATIVE declarative sketch only - NOT a literal MAF schema.
# MAF does not ship a `human_in_the_loop` step type; model the gate with a
# RequestPort / RequestInfoExecutor in code. Verify against your MAF version.
- gate: confirm_bulk_delete
message: "Delete ${count} records from ${table}?"
timeout_seconds: 600
on_timeout: cancel # fail closed
approvers: [data-owner]Copilot Studio
Use an "Ask a question" action to present the confirmation before any Power Automate flow that performs writes. Gate the flow trigger on the confirmation response. Do not rely on the user "not continuing" - model the gate explicitly in the topic flow. A bare boolean variable is fragile (it can initialise truthy or be overwritten mid-flow); prefer a value the flow cannot set itself - a signed approval token or the literal response from the "Ask a question" node - as the gate condition.
End-to-End Worked Example
Classification β gate β UX β timeout β audit, stitched for one action:
- Action: "Agent deletes 1,200 stale customer records from the prod table."
- Classify: Irreversible (cannot undo) + multi-resource + production β Non-agentable gate, two approvers.
- Gate placement: before execution; one combined gate for the whole batch.
- Confirmation UX: summary ("Delete 1,200 records from
customers"), target (table + filter), consequence ("permanent, no undo"), preview (the 1,200-row sample/diff), undo path ("cannot be undone"). - Timeout: synchronous, 10 min, fail closed; if the second approver is async, first-timeout β escalate to on-call, second β cancel.
- Audit: record agent id, action, target filter, requestor, both approvers, decision, timestamp, session id - to the append-only store. If the audit write fails, cancel.
Pitfalls
- Approval fatigue: Too many gates train users to click through without reading. Gate only where the classification requires it.
- Soft gate on an irreversible action: Showing a warning but proceeding without blocking is not a gate - it is a notification. A gate must halt execution until a decision is recorded.
- Agent self-approval: Agent text ("I have confirmed this is safe") must not satisfy a gate. Approval requires a human action outside the prompt stream. Watch for the async race where the agent calls the action before the human responds - the action tool must be blocked until approval resolves.
- Missing preview: Users cannot make an informed decision about a bulk or destructive operation without seeing what will change.
- No timeout: A gate with no timeout creates an indefinite lock. Always pair a gate with a timeout and a fail-closed fallback.
- Timeout fail-open: Proceeding after timeout is an authorisation bypass. The default on timeout is always cancel.
- Missing audit record: A gate with no audit record is unverifiable. Log every outcome including denied and timed-out gates, and fail closed if the log write fails.
- Fabricated framework APIs: Asserting a gate "exists" in a framework when the symbol is unverified gives false assurance. Mark unconfirmed APIs illustrative.
Guardrails
- Never fabricate a framework API, field name, step type, or guarantee. If you cannot confirm a symbol against the installed framework version, label it
Illustrative β verify against <framework> versionand say so to the user. A labelled unknown beats a confident wrong answer. - Default to fail-closed everywhere: on timeout, on denial, and on audit-write failure, the action does not proceed.
- Never recommend a soft gate (warn-and-continue) for a Hard-to-reverse or Irreversible action. Those require a blocking gate.
- Never let agent-generated text satisfy a non-agentable gate. Approval must come from a human on a surface the agent does not control.
- Always require the five confirmation elements, and always show the preview before a destructive or bulk action.
- When reversibility is unclear, ask rather than assume Safe; if you must proceed without input, classify one tier more dangerous and gate accordingly.
- Do not invent thresholds, approver names, timeouts, or audit fields the user has not given - state the recommended default and mark it as a default to confirm, rather than presenting it as their policy.
- Privacy: keep only the audit fields listed; do not log full record contents or secrets in the approval trail.
Final Output Contract
Every engagement produces:
- A reversibility-based gate design for the scoped actions or workflow.
- Concrete approval UX requirements, timeout behavior, and escalation paths.
- Audit and evidence requirements for every approval outcome.
- Framework-specific integration notes when the user names MCP, MAF, CopilotKit, Copilot Studio, or a similar surface - with any unverified API clearly marked illustrative.
Quality Gate
Do not mark this engagement complete until:
- Irreversible or high-blast-radius actions are not left on a soft gate.
- Timeout behavior fails closed and is explicit, including on audit-write failure.
- The approver sees the action, target, consequence, and undo path or irreversibility statement.
- Every framework symbol is either verified against the installed version or marked illustrative.