All skills
jwynia avatar

/scraper-builder

@dbeae09
by J Wyniajwynia/agent-skills160 stars
20

Guide AI agents to generate complete PageObject pattern web scraper projects using Playwright and TypeScript with Docker deployment. Supports agent-browser site analysis for automated selector discovery. Keywords: scraper, playwright, pageobject, web scraping, docker, typescript, data extraction, automation.

Use this Skill: https://skilld.dev/gh/jwynia/agent-skills/scraper-builder

This session only. Nothing lands on disk.

assetsconfigsdocker-compose.yml.md

≈443 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Docker Compose Template

Docker Compose configuration for running the scraper with volume mounts for output data and debug screenshots.

Template

services:
  scraper:
    build: .
    environment:
      - NODE_ENV=production
      - BASE_URL=${BASE_URL:-https://example.com}
      - HEADLESS=true
      - MAX_PAGES=${MAX_PAGES:-10}
      - REQUEST_DELAY=${REQUEST_DELAY:-1000}
      - OUTPUT_DIR=/app/data
    volumes:
      # Persist scraped data on the host
      - ./data:/app/data
      # Persist debug screenshots on the host
      - ./screenshots:/app/screenshots
    # Prevent runaway browser processes from consuming all resources
    deploy:
      resources:
        limits:
          memory: 2G
          cpus: '2'

Usage

# Build and run
docker compose up --build

# Run with custom URL
BASE_URL=https://shop.example.com docker compose up

# Run in background
docker compose up -d

# View logs
docker compose logs -f scraper

# One-off run (remove container after)
docker compose run --rm scraper

Customization Notes

  • Environment variables: Override at runtime via shell exports or a .env file in the project root. Docker Compose automatically reads .env.
  • Resource limits: Adjust memory and cpus based on target site complexity. Heavy JavaScript sites need more memory.
  • Shared memory: If Chromium crashes with out-of-memory errors, add shm_size: '512mb' under the service.
  • Networking: Add a networks section if the scraper needs to communicate with other services (proxy, database).
  • Restart policy: Add restart: on-failure for scheduled/recurring scrapes.

See Also

  • dockerfile.md — Dockerfile template
  • ../../references/docker-setup.md — Full Docker setup guide

Source: SKILL.md on GitHub

2 warnings14d5 checks · Risk SAFE
  • Gen Agent Trust Hub14d

    This skill is a generator for web scraper projects using Playwright, TypeScript, and Docker. It follows security best practices by implementing data validation via Zod schemas, using non-root users in Docker containers, and referencing official images and tools from trusted organizations.

  • Socket14d

    No alerts

  • Snyk14d

    Risk: MEDIUM · 1 issue

  • Runlayer7mo

    22/22 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at dbeae09. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 months ago.

Dormantupdated 8 months ago
Other metadata
compatibility
Requires Deno for generator scripts. Generated projects use Node.js with Playwright.
metadata
{
  "author": "agent-skills",
  "version": "1.0",
  "domain": "development",
  "type": "generator",
  "mode": "generative"
}

README badge

README badge for jwynia/agent-skills/scraper-builder