All skills
browserbase avatar

/event-prospecting

@592061a official
by browserbasebrowserbase/skills3.7k stars
240

Event prospecting skill. Takes a conference / event speakers URL, extracts the people, filters their companies against the user's ICP, then deep-researches only the speakers at ICP-fit companies. Outputs a person-first HTML report where each card answers "why should the AE talk to this person?" with all public links and a one-click DM opener. Use when the user wants to: (1) find leads at a specific conference, (2) prep for an event, (3) research event speakers, (4) build a target list from a sponsor/exhibitor page, (5) scrape conference speakers and rank by ICP fit. Triggers: "find leads at {event}", "research speakers at", "prospect this conference", "stripe sessions leads", "ai engineer summit prospects", "event prospecting", "scrape conference speakers", "who should I meet at".

Use this Skill: https://skilld.dev/gh/browserbase/skills/event-prospecting

This session only. Nothing lands on disk.

referencesevent-platforms.md

≈2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Event-Prospecting — Platform Reference

Contents


Detection Priority

recon.mjs probes the event URL once via browse open + browse eval, then chooses the first matching platform from this list:

  1. Next.js / __NEXT_DATA__ — document.getElementById('__NEXT_DATA__') returns a <script> tag.
  2. Sessionize — <meta name="generator"> content matches /sessionize/i.
  3. Lu.ma — location.hostname matches /lu\.ma/.
  4. Eventbrite — <meta property="og:site_name"> matches /eventbrite/i.
  5. JSON-LD Event — at least one <script type="application/ld+json"> block has @type === 'Event'.
  6. Custom / markdown fallback — none of the above; the extractor uses browse get markdown and parses speaker blocks heuristically.

The order matters: Next.js sites often also embed JSON-LD, so probing __NEXT_DATA__ first gives us the structured data path before falling back to JSON-LD heuristics.


Next.js / __NEXT_DATA__

Detection signature

!!document.getElementById('__NEXT_DATA__')

This script tag is emitted by every getServerSideProps / getStaticProps Next.js page. Its textContent is a JSON blob containing every prop hydrated into the React tree at build/request time — including the speaker list for an event microsite.

Extraction strategy — next-data-eval

recon.mjs walks the parsed JSON looking for arrays whose elements are objects with both a name-ish key AND a linkedin substring somewhere. It records the JSON path of every such array (e.g. .props.pageProps.featuredSpeakers.speakers.items) into recon.nextDataPaths. The Phase B extractor then runs ONE browse eval to harvest those arrays and union them into a single people list.

Sample output shape (Stripe Sessions 2026)

{
  "platform": "next-data",
  "strategy": "next-data-eval",
  "nextDataPaths": [
    ".props.pageProps.featuredSpeakers.speakers.items",
    ".props.pageProps.moreSpeakers.speakers.items"
  ]
}

A typical speaker object inside one of those arrays:

{
  "name": "Patrick Collison",
  "title": "CEO and Co-founder",
  "companyName": "Stripe",
  "linkedInProfile": "https://www.linkedin.com/in/patrickcollison/",
  "bio": "..."
}

Known gotchas

  • The walker also matches talks[N].speakers arrays (a denormalized re-listing of the same speakers per session). recon.mjs filters those out via regex so we don't double-count.
  • Some Next sites lazy-load speakers AFTER hydration. The 2.5s browse wait timeout is usually enough; if a site fails extraction, bump the wait to 5s before declaring it a different platform.
  • Field names vary across sites (companyName vs company vs org, linkedInProfile vs linkedinUrl). The Phase B extractor normalizes via fallback chains.

Sessionize

Detection signature

<meta name="generator" content="Sessionize.com">

Extraction strategy — sessionize-api (stub in v0.1; full implementation in a future phase)

Sessionize exposes a public read-only JSON API at https://sessionize.com/api/v2/{event_id}/view/Speakers. The event ID is in the page URL or embedded JS. v0.1 of recon.mjs only sets strategy: "sessionize-api" and emits the URL — no API discovery yet.

Sample output shape

{
  "platform": "sessionize",
  "strategy": "sessionize-api"
}

Known gotchas

  • Some Sessionize-hosted pages bury the event ID behind a custom domain. We may need to scan the page for sessionize.com/api/v2/... URLs and pull the ID from there.
  • The Sessionize API returns a flat speaker list — much cleaner than scraping Next.js, when it works.

Lu.ma

Detection signature

/lu\.ma/.test(location.hostname)

Lu.ma always serves on lu.ma (or rarely a custom CNAME); the hostname check is decisive.

Extraction strategy — json-ld (stub in v0.1)

Lu.ma embeds an Event JSON-LD block with attendee/speaker info, but the volume of structured data is event-dependent. v0.1 sets strategy: "json-ld" and defers actual extraction to Phase B+. The fallback markdown extractor handles Lu.ma pages reasonably well in the meantime.

Sample output shape

{
  "platform": "luma",
  "strategy": "json-ld"
}

Known gotchas

  • Lu.ma often gates the full attendee list behind a login. Public-facing pages usually only show "featured" speakers, not the full list.
  • Some Lu.ma events are private; recon.mjs should NOT crash if __NEXT_DATA__ isn't present and JSON-LD is empty — it falls through to the markdown strategy.

Eventbrite

Detection signature

<meta property="og:site_name" content="Eventbrite">

Extraction strategy — json-ld (stub in v0.1)

Eventbrite emits standard Event JSON-LD on every public event page. Speakers are usually only in the prose body, not the structured data — Eventbrite is more of an RSVP platform than a speaker-directory platform. We may need to combine JSON-LD parsing with markdown extraction of the description body.

Sample output shape

{
  "platform": "eventbrite",
  "strategy": "json-ld"
}

Known gotchas

  • Eventbrite event pages can be heavy. If browse cloud fetch fails or returns thin content, use browse get markdown instead.
  • Most Eventbrite events do NOT publish a speaker list at all. This platform is low-yield for prospecting.

Custom / Markdown Fallback

Detection signature

None of the above match. recon.mjs sets platform: "custom" and strategy: "markdown".

Extraction strategy — markdown

The Phase B extractor calls browse get markdown, splits the output on heading boundaries (####, ###, ##), and treats each block as a candidate speaker:

  • Line 1: name (must start with capital letter)
  • Line 2: title/role
  • Line 3: company
  • Anywhere in the block: linkedin.com/in/{handle}

This is a best-effort fallback. Coverage is typically 60-80% of the actual speaker list; field accuracy is lower than the structured paths.

Sample output shape

{
  "platform": "custom",
  "strategy": "markdown"
}

Known gotchas

  • Many static event sites use cards (not headings) for speakers. The heading-split heuristic misses those entirely.
  • "Section title" lines (## Speakers, ## Schedule) get parsed as candidate speakers and need to be filtered downstream by checking for the LinkedIn pattern.
  • If markdown extraction yields zero people, the pipeline should surface a "platform unsupported" error rather than emit an empty people.jsonl.

Adding a New Platform

  1. Add a detection branch in recon.mjs probe() after the existing ones. Pick a cheap signal (a meta tag, hostname, or a specific script tag) that's distinctive.
  2. Pick a strategy name — kebab-case verb-phrase like sessionize-api or json-ld-events.
  3. Add an extractor branch in extract_event.mjs that handles the new strategy.
  4. Add a section to this file with: detection signature, extraction strategy, sample output shape, known gotchas.
  5. Drop a fixture under scripts/__fixtures__/{platform}-snapshot.json capturing the expected recon.json shape so future refactors don't silently break extraction.

Keep the detection branches small. If a platform needs more than ~20 lines of detection logic, factor it out into scripts/detectors/{platform}.mjs and import.

Source: SKILL.md on GitHub

1 warning16d3 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    The event-prospecting skill automates lead discovery from conference websites by scraping speaker lists, performing ICP-based company triage, and enriching contact data. It generates a comprehensive local HTML report and CSV export.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: MEDIUM · 2 issues

Signed by skilld at 592061a. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Activeupdated 5 months ago
metadata
{
  "author": "browserbase",
  "version": "0.1.0"
}
All 1 allowed tools
Bash Agent AskUserQuestion
Other metadata
compatibility
Requires browse CLI (`npm install -g browse`) and BROWSERBASE_API_KEY env var. The same `browse` binary covers both API commands and JS-rendered page fallback.

README badge

README badge for browserbase/skills/event-prospecting