All skills
tejasshukla2007 avatar

/semantic-structure-audit

@4af6077

Audits a website for machine-readable semantic context, specifically evaluating the presence and quality of an llms.txt file and schema.org JSON-LD structured data. Use when diagnosing why a brand's specific facts or entity definitions are missed by AI assistants.

  • 3 files
  • 9.2 KB
  • Updated 3 weeks ago
  • GitHub

Use this Skill: https://skilld.dev/gh/tejasshukla2007/brand-ai-readiness-audit/semantic-structure-audit

This session only. Nothing lands on disk.

SKILL.md

≈73 tokens always: the name and description. ≈913 when used: this file. ≈189 more on demand in 1 file.

Semantic Structure Audit

When to use

  • Trigger this skill when invoked by the audit-orchestrator during a comprehensive AI-readiness audit.
  • Use this skill to evaluate off-site discoverability factors related to explicit data structuring.
  • Trigger to determine if a site provides direct, machine-readable pathways (llms.txt, JSON-LD schema) for crawlers to extract facts without relying on visual DOM rendering.

Inputs

  • url: The target website URL or domain string to evaluate.

Procedure

  1. Execute Semantic Extraction: Run the command python scripts/audit_semantics.py <url> to fetch the domain's /llms.txt file and extract all <script type="application/ld+json"> blocks from the target URL's raw HTML.
  2. Evaluate llms.txt Presence:
    • Check the script output to see if /llms.txt or /llms-full.txt resolved successfully.
    • If missing, flag as a discoverability gap, as AI agents rely on this for sitemap and context navigation.
  3. Evaluate Schema.org Coverage & Quality:
    • Analyze the extracted JSON-LD blocks.
    • Check for the presence of core schemas (e.g., Organization, WebSite, Article, FAQPage, or Product).
    • Assess entity disambiguation: Verify if Organization or Person schemas include a sameAs array pointing to authoritative external graphs (e.g., Wikidata, Crunchbase, LinkedIn, Wikipedia).
  4. Synthesize Findings: For every missing element or low-quality schema configuration, create a distinct finding. Assign a severity (e.g., missing Organization schema is high, missing sameAs links is medium).
  5. Format Actions: Ensure every finding includes concrete evidence (e.g., "0/3 required schemas found in JSON-LD payload") and a prioritized suggested_action detailing exactly what to add.

Deterministic Grading Rubric

You must assign severities and generate findings strictly according to this matrix based on the JSON output from audit_semantics.py. Note: For SEM-000, phrase the evidence and suggested_action as an informative warning. Inform the user that the audit was inconclusive due to a firewall, and gently recommend verifying that official AI crawlers (like GPTBot) are whitelisted.

Condition from Script Output Finding ID Finding Title Severity Priority
errors contains "HTTP error" or "Network error" fetching target URL SEM-000 Semantic Audit Inconclusive: Bot Protection Active medium medium
json_ld_schemas is empty (and SEM-000 is NOT triggered) SEM-001 Complete Absence of JSON-LD Structured Data critical critical
llms_txt_found: false SEM-002 Missing llms.txt at Domain Root high high
json_ld_schemas lacks Organization or WebSite type SEM-003 Missing Core Entity Schema (Organization/WebSite) high high
errors contains "malformed JSON-LD" SEM-004 Malformed JSON-LD Syntax Detected high high
Organization schema exists but lacks sameAs property SEM-005 Missing Entity Disambiguation (sameAs Links) medium medium
llms_txt_found: true but llms_full_txt_found: false SEM-006 Missing llms-full.txt for Expanded Context low low
missing_contextual_schemas: true SEM-007 Missing Contextual Content Schemas (Product/FAQ/Article) medium medium

Output

Emit a JSON array strictly conforming to the schema defined in references/finding_schema.json. Do not include top-level report metadata (such as site, timestamp, or summary metrics), as these will be handled by the audit-orchestrator.

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 4af6077. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 3 weeks ago

README badge

README badge for tejasshukla2007/brand-ai-readiness-audit/semantic-structure-audit