All skills
jimliu avatar

/baoyu-url-to-markdown

@d6442e1 official
by Jim Liu 宝玉jimliu/baoyu-skills26k stars
2,896

Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.

Use this Skill: https://skilld.dev/gh/jimliu/baoyu-skills/baoyu-url-to-markdown

This session only. Nothing lands on disk.

referencesadapters.md

≈808 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Adapters & Media

Read when choosing an adapter, handling media, or answering adapter-specific questions.

Built-in Adapters

Adapter URLs Key Features
x x.com, twitter.com Tweets, threads, X Articles, media, login detection
youtube youtube.com, youtu.be Transcript/captions, chapters, cover image, metadata
hn news.ycombinator.com Threaded comments, story metadata, nested replies
generic Any URL (fallback) Defuddle extraction, Readability fallback, auto-scroll, network idle detection

Adapter is auto-selected based on URL. Override with --adapter <name>.

YouTube

  • Extracts transcripts/captions when available
  • Transcript format: [MM:SS] Text segment with chapter headings
  • Availability depends on YouTube exposing a caption track; videos with captions disabled or restricted playback may produce description-only output
  • Use --wait-for force if the page needs time to finish loading player metadata

X/Twitter

  • Extracts single tweets, threads, and X Articles
  • Auto-detects login state; if logged out and content requires auth, JSON output shows "status": "needs_interaction"
  • Use --wait-for interaction for login-protected content

Hacker News

  • Parses threaded comments with proper nesting and reply hierarchy
  • Includes story metadata (title, URL, author, score, comment count)
  • Shows comment deletion/dead status

Media Download Workflow

Driven by download_media in EXTEND.md:

Setting Behavior
1 (always) Run CLI with --download-media --output <path>
0 (never) Run CLI with --output <path> (no media download)
ask (default) Follow the ask-each-time flow below

Ask-Each-Time Flow

  1. Run the CLI without --download-media with --output <path> → markdown saved
  2. Check the saved markdown for remote media URLs (https:// in image/video links)
  3. If no remote media found → done, no prompt needed
  4. If remote media found → ask via AskUserQuestion:
    • header: "Media", question: "Download N images/videos to local files?"
    • "Yes" — Download to local directories
    • "No" — Keep remote URLs
  5. If the user confirms → run the CLI again with --download-media --output <same-path> (overwrites markdown with localized links)

Media Layout

When --download-media is enabled:

  • Images → imgs/ next to the output file (or --media-dir)
  • Videos → videos/ next to the output file (or --media-dir)
  • Markdown media links are rewritten to local relative paths

Output Format

Markdown to stdout (or file with --output).

JSON output (--format json) returns structured data:

  • adapter — which adapter handled the URL
  • status — "ok" or "needs_interaction"
  • login — login state detection (logged_in, logged_out, unknown)
  • interaction — interaction gate details (kind, provider, prompt)
  • document — structured content (url, title, author, publishedAt, content blocks, metadata)
  • media — collected media assets with url, kind, role
  • markdown — converted markdown text
  • downloads — media download results (when --download-media used)

Source: SKILL.md on GitHub

2 warnings16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill fetches web pages and converts them to markdown using a browser-based extraction tool. It includes specific adapters for X (Twitter), YouTube, and Hacker News, and provides a generic fallback for other sites. The skill also handles automated media downloads and manages login sessions locally.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    4/15 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at d6442e1. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 5 months ago
version
1.61.0
Other metadata
metadata
{
  "openclaw": {
    "homepage": "https://github.com/JimLiu/baoyu-skills#baoyu-url-to-markdown",
    "requires": {
      "anyBins": [
        "bun"
      ]
    }
  }
}
  • url-to-markdown
  • web-scraping
  • chrome-cdp
  • markdown
  • baoyu-fetch
  • x-twitter
  • youtube
  • hacker-news
  • login-captcha

README badge

README badge for jimliu/baoyu-skills/baoyu-url-to-markdown

Fetches any URL and converts it to markdown using Chrome DevTools Protocol with site-specific adapters for X, YouTube, Hacker News, and generic pages. Handles login and CAPTCHA interactions via wait modes, and optionally downloads images and videos to local directories.

Generated from the current SKILL.md.

Does this skill require Chrome or Chromium installed?
Yes. The skill uses Chrome DevTools Protocol to fetch and render pages. You can specify a custom Chrome binary with `--browser-path` if the default is not found.
What happens if a page requires login or CAPTCHA?
Use `--wait-for interaction` to pause and let you complete login or CAPTCHA manually, then continue. The skill will wait up to 10 minutes by default (`--interaction-timeout`).
Which websites have built-in adapters?
X (Twitter), YouTube, Hacker News, and generic pages. The skill auto-detects the adapter or you can force one with `--adapter`.
Can this skill download images and videos from a page?
Yes, with `--download-media`. Media files are saved to `imgs/` and `videos/` subdirectories next to the markdown file, and links are rewritten to point to local paths.
Does this skill require Bun?
Yes. Bun is the runtime for the vendored `baoyu-fetch` CLI. If not installed, the skill will prompt you to install it.

Generated from the current SKILL.md. These answers refresh after source changes.