All skills
paulrberg avatar
by Paul Bergpaulrberg/agent-skills94 stars
7

Use when PDF files are the primary input or output: read, compare, reconcile, extract text/tables/images, OCR scans, fill forms, split, merge, rotate, rename, compress, or convert between PDF and images. Optimized for private financial, tax, legal, and health documents on macOS.

Use this Skill: https://skilld.dev/gh/paulrberg/agent-skills/pdf

This session only. Nothing lands on disk.

SKILL.md

โ‰ˆ71 tokens always: the name and description. โ‰ˆ1.3k when used: this file. โ‰ˆ1.9k more on demand in 3 files.

PDF

Process PDFs locally on macOS with exact extraction, source preservation, deliberate tool routing, and structural plus semantic validation.

Invariants

  1. Run extraction and transformations locally. Task-relevant document evidence in tool output and internal agent reports may be processed by the configured model provider. Require explicit user authorization and an external-disclosure review before uploading or sending document contents outside that agent workflow. Package and language-data downloads do not authorize document disclosure.
  2. Preserve every original PDF byte-for-byte. Write a sibling output, copy, or explicitly named destination unless the user authorizes destructive replacement.
  3. Preserve monetary values, identifiers, dates, signs, and displayed precision as strings. Use decimal.Decimal for arithmetic; never infer missing rows or silently discard headers, footnotes, continuation lines, or boundary pages.
  4. Inspect structure and representative renders before choosing a transformation. Use the smallest tool that preserves the required layout, forms, annotations, and image quality.
  5. Validate every written PDF structurally and against task semantics. A command exiting successfully is not evidence that extracted rows, totals, page boundaries, form appearances, or visual layout are correct.
  6. Keep reports concise for private financial, tax, legal, and health documents. Prefer counts, reconciliations, and file references over raw sensitive rows unless the rows materially support the task or the user asks for them.

Profile First

Resolve the skill directory from this SKILL.md, then profile every unknown input:

uv run "<skill-dir>/scripts/profile.py" "<input.pdf>"

The helper emits schema-versioned JSON with integrity, encryption, page geometry/rotation, image counts, and per-page text coverage without document text. Stop on password_required; password handling is outside this skill.

When layout, cropping, OCR quality, signatures, or form placement matters, render the first and last page, every structural boundary, and any page behind a discrepancy. For dense charts, tables, or technical drawings, render at higher resolution and crop or zoom the relevant region before reading values.

Route by Evidence

Need Preferred route
Quick reading or page-aware extraction Host PDF reader when available, then pdftotext -layout
Coordinates, columns, or difficult tables Poppler bounding boxes, then pdfplumber through uv run
Image-only or materially incomplete text OCRmyPDF with Tesseract; default languages eng+ron
Merge, split, rotate, or integrity checks qpdf
Render pages or extract embedded images pdftocairo or pdfimages
Convert ordered images into a PDF img2pdf
Reduce size qpdf lossless rewrite first; Ghostscript only for an accepted lossy pass
Inspect, fill, flatten, or overlay forms Read references/forms.md first

Read references/recipes.md only when exact commands for the selected extraction, transformation, OCR, image, comparison, or compression branch are needed.

Execute and Reconcile

  1. Profile inputs and identify whether each page is digital, scanned, mixed, rotated, or image-heavy.
  2. Extract or transform into a new path. For tabular documents, retain page provenance and parse continuations across page breaks before assigning rows.
  3. Reconcile financial and evidentiary output with every available invariant: page and row counts, opening/closing balances, inflows/outflows, subtotals, displayed totals, date coverage, and source hashes when provenance matters.
  4. For comparisons, extract both sources independently, enumerate overlapping and unique facts, and render the pages behind every material disagreement. Distinguish a real discrepancy from an extraction failure.
  5. For split or rename work, establish an old-to-new map from stable content identifiers. Copy by default, preserve contextual boundary pages when needed, and verify the first and last page of every result.
  6. Validate outputs with qpdf, expected page count/dimensions, text coverage, representative renders, and the task's semantic invariants. Retain OCR sidecars or extraction intermediates only when they are requested or useful evidence.

Completion requires preserved originals, intentional outputs, successful structural checks, semantic reconciliation, and a concise report of paths and evidence. Lead read-only reports with ### ๐Ÿ“„ PDF โ€” ๐Ÿ”Ž inspected, no files written; use ### ๐Ÿ“„ PDF โ€” โœ… updated only after all required validation passes, and ### ๐Ÿ“„ PDF โ€” โ›” not deliverable when a required check fails.

Source: SKILL.md on GitHub

No alerts5d3 checks ยท Risk SAFE
  • Gen Agent Trust Hub5d

    The skill provides local PDF processing capabilities on macOS using a suite of standard utilities and Python scripts. It involves the inherent risk of indirect prompt injection from untrusted document data and relies on local command execution, though the implementation follows security best practices to prevent direct command injection.

  • Socket5d

    No alerts

  • Snyk5d

    Risk: LOW ยท No issues

Signed by skilld at 85eb569. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last week
argument-hint
[file ...]
Other metadata
compatibility
Requires macOS, uv, Poppler, qpdf, Ghostscript, OCRmyPDF with Tesseract language data, and img2pdf.

README badge

README badge for paulrberg/agent-skills/pdf