Quality Review — hitl-design-patterns
Final score: 9/10 · Review date: 2026-06-16 · Version: 1.0.1
Validator green (0 errors, 0 warnings); deterministic scorer 100/100; all Critical/High findings closed; the two central framework claims (MCP, MAF) externally verified and corrected. Held at 9/10 (not 10) because exact framework symbol signatures could not be pinned to a specific installed package version and remain labelled illustrative — an honest, non-blocking constraint rather than a defect.
1.0.0 (baseline) → 1.0.1 Findings (Closed)
Summary
- Rounds completed: 1 (stopping condition met — zero open Critical, zero open High after fixes).
- Angles covered: content accuracy (external), builder's-eye, final correctness, cross-reference quality, keyword discoverability, robustness/error-handling, LLM-safety, operational completeness, plus deterministic scorer-structural (trigger clarity, scope boundaries, instruction specificity).
- Final score: 9/10. Validator: PASS, 0 warnings.
Must-fix items closed (Critical)
- MCP annotation field
destructive→destructiveHint(spec rev 2025-03-26); customx-*fields labelled non-standard. - MAF
type: human_in_the_loopYAML labelled illustrative; real mechanism (RequestPort/RequestInfoExecutor) named. - Missing
## Guardrailssection added — this flipped the scorer'sdestructivesafety signal from FAIL to WARN and faithfulness from WARN to PASS.
Should-fix items closed (High)
- Description rewritten with quoted triggers + "Use when" + exclusion clause (trigger_clarity 8→25).
## When NOT to Usesection added with named competing skills (scope_boundaries 4→25).- Audit-write-failure now fails closed (prose + code).
- Reference style standardised (
agentic-security.instructions.md→owasp-agentic). - CopilotKit wording aligned to the documented
respondcallback.
Medium / Low closed
- Build26 session codes removed; 60-second re-gate given a rationale; explicit-threshold guidance added; "one tier more dangerous when unknown" fallback; keyword synonyms (maker-checker, four-eyes, confirmation dialog, function/tool calling); Copilot Studio signed-token caveat; argument-hint gained "escalation path"; end-to-end worked example added.
External validation results
- GitHub / live source: MCP
ToolAnnotationsshape confirmed; CopilotKituseHumanInTheLoopconfirmed to exist; MAF HITL confirmed to use request/response (RequestPort), not a declarative step type. - Official docs: OWASP Top 10 for Agentic Applications v1.0 (Dec 2025) — ASI09 "Human-Agent Trust Exploitation" confirmed (code + title correct).
- Community / release notes: OWASP ASI mapping repositories cross-checked the ASI09 code/title; no contradicting regression found.
Example executability
- Total code examples: 4 (tsx confirmation component, TypeScript
requestApproval, MCP JSON, MAF YAML). - Executed: 0 (documentation/guidance skill — no project/runtime to execute against).
- Validated statically: 4 (syntax inspected; JSON well-formed; MAF YAML explicitly marked illustrative).
- Illustrative-only: MAF YAML (labelled), framework symbol signatures (labelled with last-checked date).
- End-to-end path: yes (the worked example stitches classification → gate → UX → timeout → audit).
Token efficiency
- Total estimated tokens: ~3,641 (method:
wc -w × 1.3, ±30%; tiktoken unavailable). Single file, under the 5,000-token review flag. - Largest file: SKILL.md (~3,641 tokens) — the only content file.
- Dedup win: merged overlapping "Use When" + "When to Use" tables.
- Capability signals preserved: reversibility classification matrix; five-element confirmation UX; fail-closed timeout (incl. audit-write failure); four-eyes/maker-checker escalation; independent append-only audit trail.
- Stable routing content (description, When to Use) anchored at top for cache alignment; structure ≤2 levels, fully reachable from SKILL.md.
Honest unknowns / non-blocking recommendations
- Pin MAF and CopilotKit symbols to a specific installed package version before relying on them in production (currently labelled illustrative with a 2026-06-16 last-checked date). Risk if ignored: low — the surrounding guidance is framework-agnostic and the symbols are flagged.
- Sibling skills (
owasp-agentic,copilotkit-agui,agent-governance-toolkit,autonomous-agent-loops) are referenced but not shipped in this single-file bundle; references degrade gracefully and the patterns stand alone. Risk if ignored: low. - No automated link checker in the environment; cross-references inspected manually. Risk if ignored: low.
Exact validator output
validate_skill.py: passed=True, non-PASS checks: 0
score_skill.py: hitl-design-patterns: 100/100 (Excellent) risk=high floor=80 LAUNCH-FLOOR=PASS
trigger_clarity 25/25 · instruction_specificity 25/25 · scope_boundaries 25/25 · robustness 25/25
faithfulness: PASS · safety: destructive=WARN
security_scan.py: PASS (bandit 0 findings; prose scan clean)