All skills
jwynia avatar

/scraper-builder

@dbeae09
by J Wyniajwynia/agent-skills160 stars
20

Guide AI agents to generate complete PageObject pattern web scraper projects using Playwright and TypeScript with Docker deployment. Supports agent-browser site analysis for automated selector discovery. Keywords: scraper, playwright, pageobject, web scraping, docker, typescript, data extraction, automation.

Use this Skill: https://skilld.dev/gh/jwynia/agent-skills/scraper-builder

This session only. Nothing lands on disk.

assetstemplatesbase-page.ts.md

≈652 tokens on demand. Your agent reads this file only when SKILL.md points to it.

BasePage Template

Abstract base class for all page objects. Provides navigation, waiting, screenshot, and text extraction utilities.

Usage

Every page object in a scraper project extends BasePage. Do not instantiate it directly.

Template

import { Page } from 'playwright';

/**
 * Abstract base class for all page objects.
 *
 * Provides common navigation, waiting, and extraction methods.
 * All site-specific page objects extend this class.
 */
export abstract class BasePage {
  constructor(protected readonly page: Page) {}

  /**
   * Navigate to a URL and wait for the page to stabilize.
   */
  async navigate(url: string): Promise<void> {
    await this.page.goto(url, { waitUntil: 'networkidle' });
  }

  /**
   * Wait for the current page to finish loading.
   * Use after interactions that trigger navigation or content updates.
   */
  async waitForPageLoad(): Promise<void> {
    await this.page.waitForLoadState('networkidle');
  }

  /**
   * Save a full-page screenshot for debugging.
   * Screenshots are saved to the screenshots/ directory.
   */
  async screenshot(name: string): Promise<void> {
    await this.page.screenshot({
      path: `screenshots/${name}.png`,
      fullPage: true,
    });
  }

  /**
   * Extract text content from a single element.
   * Returns empty string if element not found.
   */
  async getText(selector: string): Promise<string> {
    const el = this.page.locator(selector);
    return (await el.textContent()) ?? '';
  }

  /**
   * Extract text content from all matching elements.
   */
  async getTexts(selector: string): Promise<string[]> {
    return this.page.locator(selector).allTextContents();
  }

  /**
   * Get an attribute value from a single element.
   */
  async getAttribute(selector: string, attr: string): Promise<string | null> {
    return this.page.locator(selector).getAttribute(attr);
  }

  /**
   * Check whether an element exists on the page.
   */
  async exists(selector: string): Promise<boolean> {
    return (await this.page.locator(selector).count()) > 0;
  }

  /**
   * Get the current page URL.
   */
  get currentUrl(): string {
    return this.page.url();
  }
}

Customization Notes

  • Wait strategy: networkidle works for most server-rendered sites. For SPAs with streaming data, consider domcontentloaded plus explicit element waits.
  • Screenshot path: Adjust the screenshots/ prefix if your project uses a different directory structure.
  • Additional helpers: Add methods like getNumber(), getHref(), or getImageSrc() if multiple page objects need them.

Source: SKILL.md on GitHub

2 warnings14d5 checks · Risk SAFE
  • Gen Agent Trust Hub14d

    This skill is a generator for web scraper projects using Playwright, TypeScript, and Docker. It follows security best practices by implementing data validation via Zod schemas, using non-root users in Docker containers, and referencing official images and tools from trusted organizations.

  • Socket14d

    No alerts

  • Snyk14d

    Risk: MEDIUM · 1 issue

  • Runlayer7mo

    22/22 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at dbeae09. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 months ago.

Dormantupdated 8 months ago
Other metadata
compatibility
Requires Deno for generator scripts. Generated projects use Node.js with Playwright.
metadata
{
  "author": "agent-skills",
  "version": "1.0",
  "domain": "development",
  "type": "generator",
  "mode": "generative"
}

README badge

README badge for jwynia/agent-skills/scraper-builder