All skills
coreyhaines31 avatar

/read-book

@7c08527

When you want to read and extract structured notes from a book — PDF, EPUB, MOBI, markdown, .txt, pasted text, or URL to a public-domain work. Reads in chunks (by chapter when a TOC exists, by 50-page blocks otherwise), extracts per-chapter TL;DR + key concepts + quotes + action items + frameworks, and offers to capture to second-brain raw/ as a highlights- file. Four modes — notes (default, chapter-by-chapter), summary (whole-book TL;DR + 3–5 takeaways), quotes (pull-quote highlights only), study (notes + Q&A spaced-rep prep). Triggers on "/read-book," "read this book," "extract notes from this PDF," "what's in this book," "summarize this ebook," "pull quotes from this." Sibling to watch-video (same content-consumption pattern, different medium).

Use this Skill: https://skilld.dev/gh/coreyhaines31/makerskills/read-book

This session only. Nothing lands on disk.

referencessources.md

≈778 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Per-source ingestion

How to get usable text from each input type. Skill picks based on file extension or URL pattern.


PDF

Native path — Claude's Read tool handles PDFs directly:

Read tool with file_path="<pdf>" pages="1-10"

Max ~10 pages per call. For longer books, chunk per the chunking plan.

Get TOC + page count first:

# Page count
pdfinfo "<pdf>" 2>/dev/null | awk '/^Pages:/ {print $2}'

# Outline / bookmarks (PDF TOC)
pdftk "<pdf>" dump_data 2>/dev/null | grep -A 2 "BookmarkTitle:" | head -50

# Or with mutool (faster, comes with mupdf):
mutool show "<pdf>" outline 2>/dev/null | head -50

# Fallback: scan the first 5 pages for "Contents" or "Table of Contents"
Read tool with file_path="<pdf>" pages="1-5"

Scanned PDFs (no text layer):

# Check if text exists
pdftotext -layout "<pdf>" - 2>/dev/null | wc -c
# If <1000 chars for a 100-page book, it's scanned. Need OCR:
brew install ocrmypdf
ocrmypdf "<pdf>" "<pdf-ocr.pdf>"
# Then proceed with the OCR'd version

EPUB

Convert to markdown via pandoc (preserves chapter structure):

# Install once: brew install pandoc
pandoc "<book.epub>" -o "<workdir>/book.md" --wrap=none

The output has # Chapter X headers — easy to chunk.

For richer metadata + cover extraction, use calibre's ebook-convert:

# Install once: brew install calibre
ebook-convert "<book.epub>" "<workdir>/book.txt"
# Calibre also extracts metadata.opf alongside

MOBI / AZW3

Use ebook-convert (calibre) — pandoc doesn't handle MOBI well:

ebook-convert "<book.mobi>" "<workdir>/book.epub"
# Then pandoc as above:
pandoc "<workdir>/book.epub" -o "<workdir>/book.md" --wrap=none

Or directly to txt:

ebook-convert "<book.mobi>" "<workdir>/book.txt"

Markdown / .txt

Just Read it. No conversion. Chunk by character count (30K per chunk).


Pasted text

Use what was pasted. Treat like a single chunk if short (<30K chars), or chunk by character count if long.

If the user pastes only a section of a book ("read this chapter for me"), treat as 1-chunk and skip the per-chapter aggregation.


URL (public-domain text)

# WebFetch the URL
# Project Gutenberg pattern: https://www.gutenberg.org/files/<id>/<id>-0.txt
# Archive.org pattern: https://archive.org/stream/<id>/<id>_djvu.txt

For Project Gutenberg, prefer the .txt URL over HTML — clean text, easy to chunk.


Metadata extraction

For every source, extract these into the workdir's metadata.json:

{
  "title": "<book title>",
  "author": "<author>",
  "year": "<year if available>",
  "source": "<original path or URL>",
  "type": "pdf" | "epub" | "mobi" | "markdown" | "text" | "url",
  "page_count": 287,
  "word_count": 95000,
  "has_toc": true,
  "captured_at": "YYYY-MM-DD"
}

For PDFs: pdfinfo gives title, author, page count out of the box. For EPUB: unzip -p <file> META-INF/container.xml and the OPF file inside have full metadata. For markdown/text: ask the user for title + author if not in the filename.

Source: SKILL.md on GitHub

2 warnings2mo3 checks · Risk MEDIUM
  • Gen Agent Trust Hub2mo

    The 'read-book' skill automates the extraction of notes from various book formats and URLs. While useful, it contains security risks related to how it executes background commands with user-provided files, and it could be manipulated by malicious instructions hidden inside the books it processes.

  • Socket2mo

    No alerts

  • Snyk2mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 7c08527. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 4 weeks ago.

Activeupdated 3 months ago
metadata
{
  "version": "0.1.0"
}

README badge

README badge for coreyhaines31/makerskills/read-book