All skills

Build an automated agent that collects public data, adds useful labels or summaries, and saves the results. Use it for jobs, prices, news, GitHub projects, sports, events, and other public sources. It can run on a set schedule with GitHub Actions and save data to Notion, Google Sheets, or Supabase.

  • 1 file
  • 8 KB
  • Updated 3 weeks ago
  • GitHub

Use this Skill: https://skilld.dev/gh/agenticluke/scheduled-data-scraper-plus/skill

This session only. Nothing lands on disk.

SKILL.md

≈77 tokens always: the name and description. ≈2k when used: this file.

Data Scraper Agent

Credit: This skill comes from the open-source community author. Keep clear credit to the original author when you share or change it.

Build an agent that collects public data on a schedule. It can clean, score, sum up, or sort the data. It then saves the results for later use.

Keep the tool simple. Only add AI when it gives clear value.

Main tools: Python, Gemini Flash, GitHub Actions, Notion, Google Sheets, or Supabase.

When to use this skill

Use this skill when the user wants to:

  • Watch a public website or API.
  • Collect jobs, prices, news, code projects, scores, events, or lists.
  • Find new or changed items on a schedule.
  • Build a bot that checks a source.
  • Run a small data task with low hosting cost.
  • Use past user choices to improve later scores.

Do not use this skill to:

  • Read private data without clear access.
  • Bypass a login, paywall, CAPTCHA, or access rule.
  • Ignore a site's terms or robots.txt.
  • Collect passwords, private messages, or secret keys.
  • Collect personal data that is not needed.
  • Send forms, place orders, post content, or change a source.
  • Copy large amounts of protected text when a link and short note will work.

Core design

Each agent has three main parts:

COLLECT -> ENRICH -> STORE
   |          |         |
 Website      Rules     Notion
 or API       or AI     Sheets
                         Supabase

A fourth part runs the job on a schedule:

GitHub Actions -> COLLECT -> ENRICH -> STORE
Part Tool Job
Collect Requests and BeautifulSoup Read simple pages or APIs
Collect Playwright Read pages that need JavaScript
Enrich Plain rules Clean, filter, and score clear facts
Enrich Gemini Flash Sum up or sort text when rules are not enough
Store Supabase Save rows with strong search and update rules
Store Google Sheets Save small lists that people edit
Store Notion Save items for simple review
Schedule GitHub Actions Run the job at set times

Free plans and limits can change. Check the current limits before you promise that a setup will stay free.

Build steps

1. Set the goal

Write down:

  • The source URL or API.
  • The fields to collect.
  • How often to check.
  • Where to save the data.
  • What counts as new or changed.
  • How the user will review bad results.

Ask for missing facts before you build.

2. Check the source

Before writing code:

  • Confirm that the data is public.
  • Read the site's terms and robots.txt.
  • Use an official API when one is easy to use.
  • Check if pages use JavaScript.
  • Find a stable item key, such as an ID or full URL.
  • Set a slow request rate.
  • Add a clear user agent when the site allows scraping.

Do not try to beat blocks or access rules.

3. Build the collector

Use the smallest tool that works:

  • Use requests for JSON APIs and plain HTML.
  • Add BeautifulSoup for simple HTML pages.
  • Use Playwright only when the data needs JavaScript.

The collector should:

  • Set a timeout.
  • Retry short network errors with a wait between tries.
  • Stop after a small retry count.
  • handle empty pages and missing fields.
  • Check response status and data type.
  • Limit pages and items per run.
  • log counts and errors without logging secrets.
  • Save the source URL and time for each item.

A bad page must not erase good stored data.

4. Clean and check data

For each item:

  • Trim extra space.
  • Turn dates into one time zone and format.
  • Turn prices into a number plus currency.
  • Keep the raw value when a clean value may be wrong.
  • Mark missing fields instead of making facts up.
  • Reject items that fail key checks.

Use a stable key to stop copies. Good keys include:

  • A source item ID.
  • A full item URL.
  • A hash of stable fields, such as title, group, and date.

5. Add rules or AI

Use plain rules first for facts such as dates, prices, words, and known tags.

Use an AI model only for tasks such as:

  • A short summary.
  • A topic label.
  • A fit score.
  • A reason for the score.

Ask the model for strict JSON. Check every field before saving it.

Store:

  • The model name.
  • The prompt version.
  • The raw item text or its source link.
  • The score or label.
  • A short reason.
  • Any model error.

If the model fails, save the clean item without AI fields. Do not lose the item.

Never send secrets or data that is not needed to the model.

6. Save results

Pick one store:

  • Use Google Sheets for a small list and easy hand edits.
  • Use Notion for a simple review board.
  • Use Supabase for larger data, fast search, and safe updates.

Use an upsert when possible. This means:

  • Add a new row when the key is new.
  • Update the old row when the same key changes.
  • Leave the old row alone when nothing changed.

Track these fields when useful:

source_id
source_url
title
raw_text
clean_text
score
label
first_seen_at
last_seen_at
content_hash
status
error

Do not remove missing items at once. A page may be broken. Mark them as unseen, then remove or close them only after several good runs.

7. Add the schedule

Create a GitHub Actions job that:

  • Runs by hand and on a set schedule.
  • Uses a locked Python version.
  • Installs pinned package versions.
  • Reads secrets from GitHub Secrets.
  • Sets a run timeout.
  • Stops two runs from changing the same data at once.
  • Keeps useful logs.
  • Returns a failed status when the main job fails.

GitHub schedule times use UTC. State the local time next to the UTC time.

Do not place keys in code, logs, files, or the workflow text.

8. Add a feedback loop

Store simple user choices, such as:

  • Keep.
  • Skip.
  • Good match.
  • Bad match.

Use this data to tune clear rules first. For example, raise the score for wanted tags and lower it for blocked tags.

Do not claim that the agent learns by itself unless the code truly reads old choices and changes later results.

Keep a way to reset bad rules.

Key edge cases

Plan for these cases:

  • The page layout changes.
  • The API rate limit is reached.
  • A request times out.
  • A page gives an error page with status 200.
  • A field is missing or has a new shape.
  • Dates have no time zone.
  • Prices use a new currency or comma style.
  • The same item has more than one URL.
  • An old item changes after it was saved.
  • An item goes away for one run and comes back.
  • Two scheduled runs start at the same time.
  • The AI returns bad JSON or false facts.
  • The data store is down.
  • A secret is missing.
  • A run gets no items.

A zero-item run should warn and stop updates. It should not clear old data.

Concrete example

User request:

Check a public job page each morning. Save new Python jobs to Google Sheets. Give each job a fit score from 1 to 5.

Build it like this:

  1. Confirm the page allows the planned checks.
  2. Collect the job ID, title, group, place, date, text, and URL.
  3. Use the job ID as the stable key.
  4. Keep jobs that mention Python.
  5. Ask Gemini Flash for strict JSON:
{
  "score": 4,
  "reason": "The job needs Python and data work."
}
  1. Check that score is a whole number from 1 to 5.
  2. Save new jobs to Google Sheets.
  3. Update a row if its text changes.
  4. Run each day at 12:00 UTC.
  5. If the page is empty or broken, log a warning and keep all old rows.
  6. Let the user mark each row as keep or skip.
  7. Use those marks to tune later rules.

Done checklist

The work is done when:

  • The collector can run twice without making copies.
  • One broken item does not stop the full run.
  • A zero-item run does not delete old data.
  • AI failure does not lose clean data.
  • Secrets stay out of code and logs.
  • The store can add and update items.
  • The schedule can also run by hand.
  • The README shows setup, fields, limits, and a test command.
  • The user knows where to review results and errors.

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 110404d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 weeks ago.

Activeupdated 3 weeks ago
origin
community

README badge

README badge for agenticluke/scheduled-data-scraper-plus