Linguistic Clarity & Content Quality Audit
When to use
Execute this skill when the audit orchestrator requests an evaluation of a domain's content quality for AI discoverability and generative engine optimization.
Inputs
- target_url: The full URL or domain name to be audited.
Procedure
- Fetch the raw HTML content of the target_url homepage.
- Execute
python scripts/extract_text.pypassing the full HTML via stdin. The script extracts visible text, brand ground truth (from<title>and<meta>tags), and breaks content into numbered paragraphs[P1],[P2], etc. - Read the Tool Output Data. You are provided with:
- BRAND GROUND TRUTH: The page's
<title>and<meta>tags. Use this to understand the core domain/purpose of the brand. - VISIBLE TEXT: The cleaned text of the webpage, broken into numbered paragraphs.
- BRAND GROUND TRUTH: The page's
- Handling Bot Protection: If the visible text says
Error: No HTML provided.(meaning the fetch failed due to a 403 Forbidden firewall), you MUST skip the 3 rules below and emit exactly one finding (ID:CQA-000, Title:Content Quality Audit Inconclusive: Bot Protection Active, Severity:medium) advising the user to whitelist AI crawlers. - Evaluate the text against the Deterministic Grading Rubric below. For each finding, use the exact text snippet or paragraph number as evidence.
- Format all detected issues strictly into the JSON structure defined in
references/finding_schema.json.
Deterministic Grading Rubric
1. Factual Correctness (Tiered Verification)
Extract major claims and statistics.
- Step 1: Check your internal memory. If verified, go straight to Outcome.
- Step 2: Only if internal memory lacks the knowledge, use your web search tool (if available) to verify it against third-party sources.
- Outcome:
- If demonstrably FALSE: Emit a
criticalfinding. - If it cannot be verified internally or externally: Pass (assume true, as the structural citation checks in the freshness skill will catch unlinked claims).
- If TRUE: Pass.
- If demonstrably FALSE: Emit a
2. Contextual Drift (Semantic Isolation)
- Brand Alignment: Compare each paragraph against the Brand Ground Truth (derived from
<title>and<meta>tags). If a paragraph drifts completely off-topic, flag it. - Jargon Isolation: Assume the reader (the AI chunker) has ZERO human common sense. If a highly niche term, double-meaning, or metaphor (e.g., using the word "Apple" without surrounding tech keywords like "Mac" or "software", or using an undefined acronym) is used but not defined within the same paragraph, it is an AI parsing risk. The AI will misclassify the context.
3. Substance Assessment (Fluff vs. Fact)
Use your holistic semantic judgment to scan for "Information Density."
- If a paragraph's primary purpose is emotional persuasion (relying on marketing buzzwords like 'revolutionary', 'best-in-class') rather than providing concrete, verifiable nouns or quantifiable data, flag it as a "Low Information Density" risk. AI engines penalize abstract marketing and prioritize dense, factual text.
Evaluation Rubric
| Condition | Finding ID | Title | Severity | Priority |
|---|---|---|---|---|
| HTML fetch failed (bot protection) | CQA-000 | Content Quality Audit Inconclusive: Bot Protection Active | medium | medium |
| Demonstrably false claim detected | CQA-001 | Factually Incorrect Claim Detected | critical | critical |
| Paragraph drifts off-topic or contains undefined jargon/acronyms | CQA-002 | Contextual Drift or Undefined Jargon Risk | medium | medium |
| Paragraph relies on marketing fluff without concrete facts | CQA-003 | Low Information Density Paragraph | low | low |
Output
For every violation found, emit a JSON finding that strictly conforms to the schema in references/finding_schema.json.
evidence: MUST contain the exact text snippet or paragraph number (e.g., "[P4] Our revolutionary product...")suggested_action.summary: Provide the exact rewritten sentence or the exact noun/statistic needed to fix it.
IMPORTANT: Output ONLY the raw JSON array. Do not include markdown code blocks (```json).