---
name: seo-technical
description: >
  Audit technical SEO across crawlability, indexability, security, URLs, mobile,
  Core Web Vitals, rendering, structured data, and IndexNow. Exclude content
  strategy and backlinks.
user-invocable: true
argument-hint: "[url]"
license: MIT
metadata:
  author: AgriciDaniel
  version: "2.4.1"
  category: seo
---

# Technical SEO Audit

## Categories

### 1. Crawlability
- robots.txt: exists, valid, not blocking important resources
- XML sitemap: run `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run sitemap_discovery.py <url> --json`; require a
  valid entry in `found`, and report stale or unsafe robots.txt declarations
  separately from working fallback locations
- Noindex tags: intentional vs accidental
- Crawl depth: important pages within 3 clicks of homepage
- JavaScript rendering: check if critical content requires JS execution (method in section 8)
- Crawl budget: for large sites (>10k pages), efficiency matters
- Googlebot **fetch limits**: Googlebot fetches the first **2MB of HTML** and first **64MB of a PDF** (uncompressed; 15MB is the broader crawler-infra default). Long-standing, not a 2026 change, but inline base64 images, oversized inline CSS/JS, or bloated nav can push critical content/JSON-LD past the cap and out of the index. Keep key content + structured data within the first 2MB.
- Crawl rate **auto-adjusts** (backs off on 5xx/slow responses); there is **no manual crawl-rate control** (the legacy Search Console setting was removed Jan 2024). Influence crawling via sitemaps, server responsiveness, and robots controls.
- Google's canonical crawling/robots reference moved to **developers.google.com/crawling** (migrated 2025-11-20); IP-range files relocated to `/crawling/ipranges/` and `googlebot.json` was renamed `common-crawlers.json`.
- AMP has no separate ranking advantage. Since 2026-07-01, Google Search sends
  users directly to publisher-hosted AMP URLs, so do not recommend AMP Cache,
  AMP Viewer, or signed exchange maintenance. Audit AMP against the same content,
  action-parity, and quality requirements as other pages.

#### AI Crawler Management

As of 2025-2026, AI companies actively crawl the web to train models and power AI search. Managing these crawlers via robots.txt is a critical technical SEO consideration.

**Known AI crawlers** (the authoritative table, with robots.txt behaviour per crawler, is in `seo-geo`):

| Crawler | Company | robots.txt token | Purpose |
|---------|---------|-----------------|---------|
| GPTBot | OpenAI | `GPTBot` | Model training (NOT ChatGPT Search) |
| OAI-SearchBot | OpenAI | `OAI-SearchBot` | ChatGPT Search citability |
| ChatGPT-User | OpenAI | `ChatGPT-User` | Real-time browsing (user-triggered) |
| ClaudeBot | Anthropic | `ClaudeBot` | Model training (NOT Claude search citability) |
| Claude-SearchBot | Anthropic | `Claude-SearchBot` | Claude search-result citability |
| PerplexityBot | Perplexity | `PerplexityBot` | Perplexity search index (not model training) |
| Bytespider | ByteDance | `Bytespider` | Model training |
| Google-Extended | Google | `Google-Extended` | Gemini training and grounding, and training of the models behind Search gen-AI features (no effect on Search inclusion or ranking) |
| Applebot-Extended | Apple | `Applebot-Extended` | Apple Intelligence training opt-out (NOT Siri/Spotlight/Safari) |
| CCBot | Common Crawl | `CCBot` | Open dataset |

**Key distinctions:**
- Blocking `Google-Extended` prevents Gemini training and grounding use (and training of the models behind Search gen-AI features) but does NOT affect Google Search indexing or AI Overviews (those use `Googlebot`)
- Blocking `GPTBot` prevents OpenAI training but does NOT affect ChatGPT Search
  citability, which is governed by `OAI-SearchBot`, nor user-triggered browsing
  (`ChatGPT-User`). Check `OAI-SearchBot` for any citability claim; `GPTBot`
  status is evidence about training use only
- Blocking `ClaudeBot` prevents Anthropic model training but does NOT affect
  citability in Claude's own search features, which is governed by
  `Claude-SearchBot` (per Anthropic's crawler support article). Check
  `Claude-SearchBot` for any Claude-search citability claim; `ClaudeBot` status
  is evidence about training use only
- Blocking `Applebot-Extended` opts out of Apple Intelligence / generative-model
  training use but does NOT affect discoverability via Siri, Spotlight, or Safari,
  which follows `Applebot` (per Apple's support article); `Applebot-Extended` does
  not itself crawl

**Example, selective AI crawler blocking:**
```
# Allow search indexing, block AI training crawlers
User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Bytespider
Disallow: /

# Allow all other crawlers (including Googlebot for search)
User-agent: *
Allow: /
```

**Recommendation:** Consider your AI visibility strategy before blocking: blocking an AI search crawler removes the site from that engine's answers. Do not promise traffic from allowing one. Cross-reference the `seo-geo` skill for the full AI crawler/fetcher taxonomy.

> **Google's user-triggered fetchers generally ignore robots.txt rules** (other vendors differ: Anthropic's Claude-User honors it). Google now documents **Google-Agent** (user-triggered agentic browsing) plus **Google-GeminiNotebook** (formerly Google-NotebookLM) and **Google Messages** as *user-triggered* fetchers that **cannot be blocked via robots.txt**. Use server-side access controls instead. By contrast, `Google-Extended` and `Google-CloudVertexBot` obey robots.txt. Emerging: **Web Bot Auth** (RFC 9421) lets bots authenticate cryptographically via a `Signature-Agent` header + key directory at `agent.bot.goog` (used by Google-Agent); reverse-DNS verification remains the fallback.

### 2. Indexability
- Canonical tags: self-referencing, no conflicts with noindex
- Duplicate content: near-duplicates, parameter URLs, www vs non-www
- Canonicalization fixes can take time: Google may retain corrected pages in a
  duplicate cluster for **up to two weeks** while re-evaluating them. Do not
  interpret an unchanged canonical immediately after a fix as proof that the
  fix failed.
- Thin content: pages below minimum word counts per type
- Pagination: crawlable `<a href>` links to each page (Google no longer uses rel=next/prev; it announced this in 2019); give each page a self-referencing canonical; load-more and infinite scroll need paginated URLs behind them
- Hreflang: correct for multi-language/multi-region sites
- Index bloat: unnecessary pages consuming crawl budget

### 3. Security
- HTTPS: enforced, valid SSL certificate, no mixed content
- Security headers:
  - Content-Security-Policy (CSP)
  - Strict-Transport-Security (HSTS)
  - X-Frame-Options
  - X-Content-Type-Options
  - Referrer-Policy
- HSTS preload: check preload list inclusion for high-security sites
- **Back-button hijacking** (spam-policy violation, malicious practices): flag pages that defeat the Back button via `history.pushState`/`replaceState` (including scripts injected by third-party ad/library platforms). Added to Google's spam policies 2026-04-13; **enforcement live since 2026-06-15** (manual actions + automated demotions): treat as Critical.

### 4. URL Structure
- Clean URLs: descriptive, hyphenated, no query parameters for content
- Hierarchy: logical folder structure reflecting site architecture
- Redirects: no chains (max 1 hop), 301 for permanent moves
- URL length: flag >100 characters
- Trailing slashes: consistent usage

### 5. Mobile Optimization & Page Experience
- Responsive design: viewport meta tag, responsive CSS
- Touch targets: WCAG 2.2 AA requires at least 24x24 CSS px; 48x48px with spacing is the comfortable guideline (not a Google requirement)
- Font size: readable text without zooming (16px base is common practice, not a Google rule)
- No horizontal scroll
- Mobile-first indexing: Googlebot Smartphone is the primary crawler (rollout completed 2024). A mobile version is **not strictly required** (Google says "very strongly recommended"), sites that don't work on mobile can still be indexed, but the real risk is **content/parity loss**, not hard exclusion.
- **Mobile/desktop content parity** (highest-value mobile check): equivalent primary content, matching robots meta tags, matching titles/descriptions, equivalent structured data, crawlable resources; avoid lazy-loading primary content that requires user interaction.
- **Intrusive interstitials / ad density**: flag full-page interstitials, standalone consent-redirect pages, persistent blocking dialogs, and excessive/distracting ad density (a named page-experience aspect). Acceptable: small banners, standard CMS/legal dialogs.
- **"Read more" deep links**: keep key content **immediately visible on load** (not behind tabs/accordions), don't hijack scroll on load, and preserve URL hash fragments, content hidden behind expandable sections is less likely to qualify.

> **Page experience is guidance, not a single ranking system.** Only **Core Web Vitals** feeds ranking directly; **HTTPS** is a confirmed but lightweight signal (Google called it very lightweight when it was announced in 2014). Relevance can still win even when page experience is sub-par, so don't over-weight security headers. Note: the standalone **Page Experience report was removed** from Search Console (monitor via the Core Web Vitals + HTTPS reports).

### 6. Core Web Vitals
- **LCP** (Largest Contentful Paint): target <=2.5s
- **INP** (Interaction to Next Paint): target <=200ms
  - INP replaced FID on March 12, 2024. FID was removed from Chrome's field-data tools (CrUX API, PageSpeed Insights) on September 9, 2024 (Lighthouse is a lab tool that never reported FID). Do NOT reference FID anywhere.
- **CLS** (Cumulative Layout Shift): target <=0.1
- Evaluation uses 75th percentile of real user data
- Use PageSpeed Insights API or CrUX data if MCP available

### 7. Structured Data
- Detection: JSON-LD (preferred), Microdata, RDFa
- Validation against Google's supported types
- See seo-schema skill for full analysis

### 8. JavaScript Rendering
- Method: `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run agentic_check.py <url> --json` reports visible words in the raw HTML (`server-rendered`); compare with `render_page.py <url> --mode always --json` when Chromium is available. Without Chromium, report the raw-HTML result and say rendered content was not compared.
- Check if content visible in initial HTML vs requires JS
- Identify client-side rendered (CSR) vs server-side rendered (SSR)
- Flag SPA frameworks (React, Vue, Angular) that may cause indexing issues
- If dynamic rendering is detected, flag it as technical debt rather than a valid setup.
  Google documents it as "a workaround and not a recommended solution" because of the added
  complexity and resource cost.
  See https://developers.google.com/search/docs/crawling-indexing/javascript/dynamic-rendering

**Recommended rendering strategy:**

| Strategy | Use Case |
|----------|----------|
| **SSR** | Public SEO content, dynamic pages |
| **SSG** | Static content, blogs, docs |
| **CSR** | Authenticated / behind-login content only |

**Preferred frameworks:** Next.js, Astro, React Router v7 (Remix), SvelteKit

#### JavaScript SEO: Canonical & Indexing Guidance (December 2025)

Google updated its JavaScript SEO documentation in December 2025 with critical clarifications:

1. **Canonical conflicts:** If a canonical tag in raw HTML differs from one injected by JavaScript, Google may use EITHER one. Ensure canonical tags are identical between server-rendered HTML and JS-rendered output.
2. **noindex with JavaScript:** If raw HTML contains `<meta name="robots" content="noindex">` but JavaScript removes it, Google MAY still honor the noindex from raw HTML. Serve correct robots directives in the initial HTML response.
3. **Non-200 status codes:** Google does NOT render JavaScript on pages returning non-200 HTTP status codes. Any content or meta tags injected via JS on error pages will be invisible to Googlebot.
4. **Structured data in JavaScript:** Product, Article, and other structured data injected via JS may face delayed processing. For time-sensitive structured data (especially e-commerce Product markup), include it in the initial server-rendered HTML.

**Best practice:** Serve critical SEO elements (canonical, meta robots, structured data, title, meta description) in the initial server-rendered HTML rather than relying on JavaScript injection.

### 9. IndexNow Protocol
- Check if site supports IndexNow for Bing, Yandex, Naver
- Supported by search engines other than Google
- Recommend implementation for faster indexing on non-Google engines

## Agent-Friendly Pages & Agentic Browsing

Agent readiness has its own sub-skill: `/seo agentic <url>` (`seo-agentic`).
It owns the Lighthouse **Agentic Browsing** category (a fraction, X of N, not
a 0-100 score), the accessibility tree for agents, AI agent access policy,
llms.txt, Markdown delivery, ai-catalog.json, `/.well-known` discovery files,
and WebMCP. During a technical audit, record only these two signals and point
to `seo-agentic` for the rest:

- JS rendering: primary content missing from the raw HTML also hides it from
  agents that do not run JavaScript.
- A 5xx robots.txt, which compliant crawlers read as "disallow everything".

```bash
"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run agent_ux_check.py https://example.com --json
```

The Agent-UX 0-100 score above is a local heuristic. Keep it distinct from the
Lighthouse fraction, and surface its findings as opportunities, not failures.
A failing Lighthouse `agent-accessibility-tree` audit is different: `seo-agentic`
rates it P0, because it is Google's own measured check.

## Output

### Technical Score: XX/100

Score only what was measured. Each category score is the share of that
category's checks that passed, adjusted for severity; a category you could not
measure is reported as "not measured", never given a number. Show the checks
behind every score.

### Category Breakdown
| Category | Status | Score |
|----------|--------|-------|
| Crawlability | pass/warn/fail | XX/100 |
| Indexability | pass/warn/fail | XX/100 |
| Security | pass/warn/fail | XX/100 |
| URL Structure | pass/warn/fail | XX/100 |
| Mobile | pass/warn/fail | XX/100 |
| Core Web Vitals | pass/warn/fail | XX/100 |
| Structured Data | pass/warn/fail | XX/100 |
| JS Rendering | pass/warn/fail | XX/100 |
| IndexNow | pass/warn/fail | XX/100 |

### Critical Issues (fix immediately)
### High Priority (fix within 1 week)
### Medium Priority (fix within 1 month)
### Low Priority (backlog)

## DataForSEO Integration (Optional)

If DataForSEO MCP tools are available, use `on_page_instant_pages` for real page analysis (status codes, page timing, broken links, on-page checks), `on_page_lighthouse` for Lighthouse audits (performance, accessibility, SEO scores), and `domain_analytics_technologies_domain_technologies` for technology stack detection.

## Google API Integration (Optional)

If Google API credentials are configured, use `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run pagespeed_check.py <url> --json` for real PSI + CrUX field data (replaces lab-only CWV estimates), `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run crux_history.py <url> --form-factor PHONE --json` for 25-week CWV trends (use PHONE: the all-devices view can hide a mobile failure), and `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run gsc_inspect.py <url> --json` for real indexation status per URL.

## Auditing a Local or Private Host

`url_safety` refuses loopback and private addresses by default, so `http://localhost:3000` and a staging host on Tailscale fail with "Blocked hostname" or "Blocked IP literal". That default is deliberate: these scripts follow URLs found on the pages they crawl.

To audit a pre-deployment host, the operator names it in `CLAUDE_SEO_LOCAL_TARGETS`, a comma-separated list of `host` or `host:port` entries:

```bash
CLAUDE_SEO_LOCAL_TARGETS="localhost:3000,127.0.0.1:8080,100.101.102.103" \
  "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run fetch_page.py http://localhost:3000/
```

What it does and does not cover:

| Behaviour | Allowlisted host |
|-----------|------------------|
| First, top-level URL over raw HTTP | Allowed |
| Redirect target reached from that URL | Refused |
| Subresource fetched by a rendered page | Refused |
| Playwright renders (`--render`, screenshots) | Refused; use the raw-HTTP path |
| A host not named in the variable | Refused |
| Cloud metadata endpoints, even when listed | Refused |

`host:port` matches that port only; a bare `host` matches any port. With the variable unset the policy is unchanged. Never suggest setting it for a host the user does not control. See SECURITY.md.

## Error Handling

| Scenario | Action |
|----------|--------|
| URL unreachable | Report connection error with status code. Suggest verifying URL, checking DNS resolution, and confirming the site is publicly accessible. |
| robots.txt not found | Note that no robots.txt was detected at the root domain. Recommend creating one with appropriate directives. Continue audit on remaining categories. |
| HTTPS not configured | Flag as a critical issue. Report whether HTTP is served without redirect, mixed content exists, or SSL certificate is missing/expired. |
| Core Web Vitals data unavailable | Note that CrUX data is not available (common for low-traffic sites). Suggest using Lighthouse lab data as a proxy and recommend increasing traffic before re-testing. |
