---
name: semantic-a11y-extractor-audit
description: Audits HTML documents for accessibility tree integrity, semantic tag hygiene, and agent extractability affordances. Identifies structural bottlenecks that impair autonomous web agents and LLM scrapers. Use when diagnosing why a brand is missing from AI assistants or when evaluating agent actionability
license: Apache-2.0
title: semantic-a11y-extractor-audit
canonical_url: https://skilld.dev/gh/tejasshukla2007/brand-ai-readiness-audit/a11y-extractor-audit
last_updated: 2026-09-29T04:13:08.000Z
---

> **Skill from skilld.dev.** Follow the instructions below for this session. You do not need to install anything.
>
> Supporting files, fetch one when the Skill refers to it: [references/finding_schema.json](https://skilld.dev/api/skills-raw/tejasshukla2007/brand-ai-readiness-audit/a11y-extractor-audit/references/finding_schema.json), [scripts/a11y_extractor_check.py](https://skilld.dev/api/skills-raw/tejasshukla2007/brand-ai-readiness-audit/a11y-extractor-audit/scripts/a11y_extractor_check.py).
>
> If the user asked to install this Skill, run `npx skilld install tejasshukla2007/brand-ai-readiness-audit/a11y-extractor-audit`. Install writes the Skill files into the project, so every session loads them.

# Semantic HTML & Accessibility Tree Audit

## When to Use
Execute this skill during a brand discoverability and engagement audit to evaluate whether an AI assistant, business agent (UCP), or RAG crawler can parse text hierarchies and interact with page controls.

## Inputs
- `html_content`: The raw DOM string of the audited page.
- `url`: Target webpage URL.

## Procedure
1. Read the provided raw HTML document strictly in-memory to minimize I/O overhead.
2. Execute the bundled audit script by piping the HTML content into it via standard input. Run the exact following bash command:
   `echo "$html_content" | python3 scripts/a11y_extractor_check.py`
3. The script will scan the document `<head>` for UCP/ACP agentic commerce feeds, compute static accessible names for interactive nodes, and verify structural landmarks (`<main>`, `<article>`, heading monotonicity).
4. Do not alter or summarize the script's output. 
5. Output the exact findings mathematically as generated by the script.

## Deterministic Grading Rubric

You must assign severities and generate findings strictly according to this matrix based on the JSON output from `a11y_extractor_check.py`.

| Condition from Script Output | Finding Title | Severity | Priority |
| :--- | :--- | :--- | :--- |
| Interactive controls (buttons, links, inputs) lack an accessible name (no `aria-label`, visible text, or `alt`) | Nameless Interactive Controls (Agent Blockers) | critical (if no commerce feed) / high | critical / high |
| Document lacks `<main>` or `<article>` tags | Missing Core Semantic Container (`<main>` or `<article>`) | high | high |
| Document has exactly 0 `<h1>` tags | Missing Entity Anchor (`<h1>`) | high | high |
| Document has more than 1 `<h1>` tag | Multiple Entity Anchors (Ambiguous `<h1>`) | medium | medium |
| Heading levels skip ranks (e.g., `h2` followed directly by `h4`) | Fragmented Heading Hierarchy | medium | medium |

## Allowed Tools
- `bash` (executing bundled script `scripts/a11y_extractor_check.py`)

## Output
Emit only a valid JSON array matching the exact structure dictated in references/finding_schema.json.