All skills
apify avatar

/apify-ultimate-scraper

@97047d3 official
by apifyapify/agent-skills2.4k stars
259

Universal AI-powered web scraper for any platform. Scrape data from Instagram, Facebook, TikTok, YouTube, LinkedIn, X/Twitter, Google Maps, Google Search, Google Trends, Reddit, Airbnb, Yelp, and 15+ more platforms. Use for lead generation, brand monitoring, competitor analysis, influencer discovery, trend research, content analytics, audience analysis, review analysis, SEO intelligence, recruitment, or any data extraction task.

Use this Skill: https://skilld.dev/gh/apify/agent-skills/apify-ultimate-scraper

This session only. Nothing lands on disk.

referencesworkflowsknowledge-base-and-rag.md

≈948 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Knowledge base and RAG pipeline workflows

Website to RAG knowledge base via sitemap crawl

When: User wants to ingest an entire website or documentation site into a vector database for AI retrieval (chatbots, search, AI agents).

Pipeline

  1. Extract sitemap -> apify/sitemap-extractor
    • Key input: sitemapUrl or domain
  2. Crawl and convert to markdown -> apify/website-content-crawler
    • Pipe: results[].url -> startUrls (or pass dataset ID)
    • Key input: startUrls, maxCrawlPages, htmlTransformer: "readableText", outputFormats: ["markdown"]
  3. Chunk and embed (n8n: Recursive Character Text Splitter -> OpenAI Embeddings node)
  4. Upsert to vector DB (n8n: Supabase / Qdrant node with document + metadata)

Output fields

text (clean markdown), url, metadata.title, metadata.description, crawledAt

Gotcha

apify/rag-web-browser is purpose-built for RAG use cases and returns pre-chunked, clean text without boilerplate - use it when you want simpler output and don't need full site coverage. For comprehensive crawls (full docs sites, 100+ pages), website-content-crawler gives more control over depth and URL filtering.


Deep research agent with web crawling

When: User or an AI agent submits a research question and wants a synthesized report drawn from live web sources.

Pipeline

  1. Generate search queries (n8n: AI node expands research question into 3-5 distinct queries)
  2. Search -> apify/google-search-scraper
    • Pipe: generated queries -> queries (array)
    • Key input: queries, maxResultsPerPage (5-10)
  3. Retrieve content -> apify/rag-web-browser
    • Pipe: results[].organicResults[].url -> query (RAG browser takes query + crawls most relevant result)
    • Key input: query, maxResults, requestTimeoutSecs
  4. Synthesize (n8n: OpenAI node assembles final report from per-source summaries)
  5. Output to n8n Data Table, Notion, or Google Docs

Output fields

Search: organicResults[].url, organicResults[].title, organicResults[].snippet RAG browser: text, url, metadata.title

Gotcha

apify/rag-web-browser fetches and summarizes a single URL per call. To process multiple search results in parallel, use n8n's Split In Batches node with a concurrency of 3-5 rather than running them sequentially. This cuts total runtime significantly for 10+ URLs.


Scheduled news monitoring to AI knowledge feed

When: User wants to track industry news sources daily, filter new articles, summarize them, and store in a searchable knowledge base (Notion, NocoDB, Supabase).

Pipeline

  1. Extract articles -> lukaskrivka/article-extractor-smart
    • Key input: startUrls (news site listing pages), maxCrawlPages, articleSelector (optional CSS hint)
  2. Filter new articles only (n8n: compare publishedAt or URL against stored records in DB)
  3. Full article content (optional) -> apify/website-content-crawler
    • Pipe: new article url values -> startUrls
    • Use when listing-page extract is too short for quality summarization
  4. AI summarize + tag (n8n: OpenAI node generates 3-sentence summary + keyword tags)
  5. Upsert to knowledge base (n8n: Notion / NocoDB / Supabase node)

Output fields

Step 1: title, text, publishedAt, author, url, tags Step 3 (WCC): full text, metadata.title, metadata.description

Gotcha

lukaskrivka/article-extractor-smart handles most news formats well, but paywalled sites return truncated content. Check text length - if consistently under 200 characters for a given source, that site is paywalled and should be removed from the list. Deduplicate by URL before summarizing to avoid re-processing old articles on re-runs.

Source: SKILL.md on GitHub

1 alert16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    The skill provides comprehensive instructions for orchestrating and interacting with web scrapers (Actors) via the Apify CLI. It provides structured playbooks for different data extraction workflows (B2B lead generation, brand monitoring, social media analytics, etc.) and contains standard guidelines on configuration, parameter passing, and error handling. No malicious behaviors, obfuscation techniques, or hidden actions were detected.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: MEDIUM · 1 issue

  • Runlayer7mo

    2/2 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 97047d3. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 3 months ago
  • SEO
  • apify
  • scraping
  • web-scraping
  • instagram
  • facebook
  • tiktok
  • youtube
  • linkedin
  • twitter
  • google-maps
  • reddit
  • lead-generation
  • competitor-analysis
  • brand-monitoring

README badge

README badge for apify/agent-skills/apify-ultimate-scraper

Scrapes data from 100+ Apify Actors covering Instagram, Facebook, TikTok, YouTube, LinkedIn, X/Twitter, Google Maps, Reddit, Airbnb, Yelp, and 15+ other platforms via the Apify CLI. Use for lead generation, competitor analysis, brand monitoring, influencer discovery, review analysis, or SEO intelligence on any public web platform.

Generated from the current SKILL.md.

Does this skill work with all platforms?
It covers ~100 Actors across 15+ platforms including Instagram, Facebook, TikTok, YouTube, LinkedIn, X/Twitter, Google Maps, Reddit, Airbnb, Yelp, and more. Not all platforms are equally supported; check the actor-index.md reference for your target platform.
What authentication is required?
You need an Apify account and CLI token. Authenticate via `apify login` (OAuth), set `APIFY_TOKEN` as an environment variable, or source from a .env file.
Do I need to install anything locally?
Yes, Apify CLI v1.5.0 or later must be installed via npm (`npm install -g apify-cli`).
Can I export results in different formats?
Yes. Results can be fetched as JSON or CSV using `apify datasets get-items` with the `--format` flag.
What should I read before running a scraper for the first time?
Check `references/actor-index.md` to find the right Actor for your platform, then read `references/gotchas.md` for common pitfalls specific to that Actor before running.

Generated from the current SKILL.md. These answers refresh after source changes.