Semantic Structure Audit
When to use
- Trigger this skill when invoked by the
audit-orchestratorduring a comprehensive AI-readiness audit. - Use this skill to evaluate off-site discoverability factors related to explicit data structuring.
- Trigger to determine if a site provides direct, machine-readable pathways (
llms.txt, JSON-LD schema) for crawlers to extract facts without relying on visual DOM rendering.
Inputs
url: The target website URL or domain string to evaluate.
Procedure
- Execute Semantic Extraction: Run the command
python scripts/audit_semantics.py <url>to fetch the domain's/llms.txtfile and extract all<script type="application/ld+json">blocks from the target URL's raw HTML. - Evaluate
llms.txtPresence:- Check the script output to see if
/llms.txtor/llms-full.txtresolved successfully. - If missing, flag as a discoverability gap, as AI agents rely on this for sitemap and context navigation.
- Check the script output to see if
- Evaluate Schema.org Coverage & Quality:
- Analyze the extracted JSON-LD blocks.
- Check for the presence of core schemas (e.g.,
Organization,WebSite,Article,FAQPage, orProduct). - Assess entity disambiguation: Verify if
OrganizationorPersonschemas include asameAsarray pointing to authoritative external graphs (e.g., Wikidata, Crunchbase, LinkedIn, Wikipedia).
- Synthesize Findings: For every missing element or low-quality schema configuration, create a distinct finding. Assign a
severity(e.g., missingOrganizationschema ishigh, missingsameAslinks ismedium). - Format Actions: Ensure every finding includes concrete
evidence(e.g., "0/3 required schemas found in JSON-LD payload") and a prioritizedsuggested_actiondetailing exactly what to add.
Deterministic Grading Rubric
You must assign severities and generate findings strictly according to this matrix based on the JSON output from audit_semantics.py.
Note: For SEM-000, phrase the evidence and suggested_action as an informative warning. Inform the user that the audit was inconclusive due to a firewall, and gently recommend verifying that official AI crawlers (like GPTBot) are whitelisted.
| Condition from Script Output | Finding ID | Finding Title | Severity | Priority |
|---|---|---|---|---|
errors contains "HTTP error" or "Network error" fetching target URL |
SEM-000 | Semantic Audit Inconclusive: Bot Protection Active | medium | medium |
json_ld_schemas is empty (and SEM-000 is NOT triggered) |
SEM-001 | Complete Absence of JSON-LD Structured Data | critical | critical |
llms_txt_found: false |
SEM-002 | Missing llms.txt at Domain Root | high | high |
json_ld_schemas lacks Organization or WebSite type |
SEM-003 | Missing Core Entity Schema (Organization/WebSite) | high | high |
errors contains "malformed JSON-LD" |
SEM-004 | Malformed JSON-LD Syntax Detected | high | high |
Organization schema exists but lacks sameAs property |
SEM-005 | Missing Entity Disambiguation (sameAs Links) | medium | medium |
llms_txt_found: true but llms_full_txt_found: false |
SEM-006 | Missing llms-full.txt for Expanded Context | low | low |
missing_contextual_schemas: true |
SEM-007 | Missing Contextual Content Schemas (Product/FAQ/Article) | medium | medium |
Output
Emit a JSON array strictly conforming to the schema defined in references/finding_schema.json. Do not include top-level report metadata (such as site, timestamp, or summary metrics), as these will be handled by the audit-orchestrator.