All skills
apify avatar

/apify-ultimate-scraper

@97047d3 official
by apifyapify/agent-skills2.4k stars
259

Universal AI-powered web scraper for any platform. Scrape data from Instagram, Facebook, TikTok, YouTube, LinkedIn, X/Twitter, Google Maps, Google Search, Google Trends, Reddit, Airbnb, Yelp, and 15+ more platforms. Use for lead generation, brand monitoring, competitor analysis, influencer discovery, trend research, content analytics, audience analysis, review analysis, SEO intelligence, recruitment, or any data extraction task.

Use this Skill: https://skilld.dev/gh/apify/agent-skills/apify-ultimate-scraper

This session only. Nothing lands on disk.

referencesworkflowslead-generation.md

≈1.3k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Lead generation workflows

Local business leads with email enrichment

When: User wants business contacts, emails, or phone numbers for businesses in a specific location.

Pipeline

  1. Find businesses -> compass/crawler-google-places
    • Key input: searchStringsArray, locationQuery, maxCrawledPlaces
  2. Enrich with contacts -> compass/enrich-google-maps-dataset-with-contacts
    • Pipe: results[].url -> startUrls (or pass the dataset ID directly)
    • Key input: datasetId (from step 1), maxRequestsPerCrawl

Output fields

Step 1: title, address, phone, website, categoryName, totalScore, reviewsCount, url Step 2: emails[], phones[], socialLinks, linkedInUrl, twitterUrl

Gotcha

Google Maps results vary by language and location. Set language: "en" explicitly. Also set locationQuery to a specific city/region, not just a country.


B2B prospect discovery via LinkedIn

When: User wants to find professionals by role, company, or industry.

Pipeline

  1. Search profiles -> harvestapi/linkedin-profile-search
    • Key input: keyword, location, title, limit
  2. Enrich with details -> harvestapi/linkedin-profile-scraper
    • Pipe: results[].profileUrl -> urls
    • Key input: urls, includeEmail (set to true for email discovery)

Output fields

Step 1: fullName, headline, location, profileUrl, currentCompany Step 2: experience[], education[], skills[], email, phone

Cost estimate

Step 2 with includeEmail: true costs ~$0.01/profile. For 500 profiles, budget ~$5.

Gotcha

LinkedIn Actors are all PPE. Estimate and confirm with user before running at scale.


Sales Navigator bulk lead extraction

When: User wants daily 100-1,000 lead extraction from a Sales Navigator search for outbound sequences.

Pipeline

  1. Extract leads -> harvestapi/linkedin-profile-search
    • Key input: searchUrl (Sales Navigator search URL), maxResults, proxy settings
  2. Verify emails -> native n8n Hunter.io node or HTTP Request to ZeroBounce API
    • Pipe: results[].email -> email verification input

Output fields

Step 1: fullName, email, companyName, jobTitle, connectionDegree, profileUrl Step 2: result (valid/risky/invalid), score

Cost estimate

harvestapi/linkedin-profile-search is PPE. 1,000 leads at typical rates runs ~$5-10. Confirm before scheduling daily runs.

Gotcha

Sales Navigator URL must be a saved search URL, not a one-time results URL. The URL changes each session unless saved.


SERP-based B2B prospect discovery

When: User wants to find companies matching niche keywords via Google, AI-qualify them against ICP criteria, and push qualified leads to CRM.

Pipeline

  1. Find companies -> apify/google-search-scraper
    • Key input: queries (search terms array), maxResultsPerPage, countryCode
  2. Crawl company sites -> apify/website-content-crawler
    • Pipe: results[].organicResults[].url -> startUrls
    • Key input: startUrls, maxCrawlDepth (set to 2), maxCrawlPages (set to 5)

Output fields

Step 1: organicResults[].url, organicResults[].title, organicResults[].snippet Step 2: text (clean markdown), url, metadata.title, metadata.description

Gotcha

Pass only company root domains from SERP results into WCC - not individual blog post URLs. Filter organicResults[].url for root domains before piping.


Apollo leads + AI website icebreakers

When: User has an Apollo lead list with company websites and wants personalized cold email icebreakers generated from each company's web presence.

Pipeline

  1. Scrape company sites -> apify/website-content-crawler
    • Key input: startUrls (homepage URLs from Apollo export), maxCrawlDepth (set to 2), maxCrawlPages (set to 5)
  2. Generate icebreakers -> AI node (GPT-4o or Claude)
    • Pipe: results[].text -> prompt context per lead
    • Key input: company summary + lead name + role

Output fields

Step 1: text (clean markdown), metadata.title, metadata.description, url Step 2: AI-generated icebreaker string per lead

Gotcha

Some Apollo exports include LinkedIn URLs instead of company websites. Filter the list for http URLs before passing to WCC - LinkedIn blocks crawlers.


Reddit community lead mining

When: User wants to find prospects actively posting problems that their product or service solves in relevant subreddits.

Pipeline

  1. Mine subreddit posts -> trudax/reddit-scraper-lite
    • Key input: startUrls (subreddit URLs), searchTerms (problem keywords), maxItems, sort (hot/new/top)
  2. Qualify leads -> AI node
    • Pipe: results[].title, results[].body -> qualification prompt
    • Key input: ICP criteria, pain point keywords

Output fields

Step 1: title, body, subreddit, url, score, numberOfComments, createdAt, author Step 2: AI qualification score, extracted contact intent, suggested outreach angle

Gotcha

Reddit usernames are pseudonymous - there is no direct email enrichment path. The output is intent signals and post URLs for manual outreach via Reddit DM or to cross-reference against other platforms.

Source: SKILL.md on GitHub

1 alert16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    The skill provides comprehensive instructions for orchestrating and interacting with web scrapers (Actors) via the Apify CLI. It provides structured playbooks for different data extraction workflows (B2B lead generation, brand monitoring, social media analytics, etc.) and contains standard guidelines on configuration, parameter passing, and error handling. No malicious behaviors, obfuscation techniques, or hidden actions were detected.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: MEDIUM · 1 issue

  • Runlayer7mo

    2/2 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 97047d3. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 3 months ago
  • SEO
  • apify
  • scraping
  • web-scraping
  • instagram
  • facebook
  • tiktok
  • youtube
  • linkedin
  • twitter
  • google-maps
  • reddit
  • lead-generation
  • competitor-analysis
  • brand-monitoring

README badge

README badge for apify/agent-skills/apify-ultimate-scraper

Scrapes data from 100+ Apify Actors covering Instagram, Facebook, TikTok, YouTube, LinkedIn, X/Twitter, Google Maps, Reddit, Airbnb, Yelp, and 15+ other platforms via the Apify CLI. Use for lead generation, competitor analysis, brand monitoring, influencer discovery, review analysis, or SEO intelligence on any public web platform.

Generated from the current SKILL.md.

Does this skill work with all platforms?
It covers ~100 Actors across 15+ platforms including Instagram, Facebook, TikTok, YouTube, LinkedIn, X/Twitter, Google Maps, Reddit, Airbnb, Yelp, and more. Not all platforms are equally supported; check the actor-index.md reference for your target platform.
What authentication is required?
You need an Apify account and CLI token. Authenticate via `apify login` (OAuth), set `APIFY_TOKEN` as an environment variable, or source from a .env file.
Do I need to install anything locally?
Yes, Apify CLI v1.5.0 or later must be installed via npm (`npm install -g apify-cli`).
Can I export results in different formats?
Yes. Results can be fetched as JSON or CSV using `apify datasets get-items` with the `--format` flag.
What should I read before running a scraper for the first time?
Check `references/actor-index.md` to find the right Actor for your platform, then read `references/gotchas.md` for common pitfalls specific to that Actor before running.

Generated from the current SKILL.md. These answers refresh after source changes.