All skills
dkyazzentwatwa avatar

/ocr-document-processor

@103b430

Extract text and structure from scans, images, and scanned PDFs. Use for OCR, searchable PDFs, table extraction, receipt parsing, and business card parsing.

Use this Skill: https://skilld.dev/gh/dkyazzentwatwa/chatgpt-skills/ocr-document-processor

This session only. Nothing lands on disk.

SKILL.md

≈45 tokens always: the name and description. ≈258 when used: this file. ≈46 more on demand in 1 file.

OCR Document Processor

Handle OCR-heavy inputs where text must be recovered from images or scanned pages.

Use This For

  • OCR on images and scanned PDFs
  • Searchable PDF export
  • Structured extraction to text, markdown, JSON, or HTML
  • Table extraction from scanned material
  • Receipt parsing and business card parsing

Workflow

  1. Decide whether plain OCR, structured extraction, or document-specific parsing is needed.
  2. Preprocess noisy inputs before extraction when skew, blur, or shadows are present.
  3. Use scripts/ocr_processor.py for core OCR tasks.
  4. Use the focused helpers when the input is specialized:
    • scripts/business_card_scanner.py
    • scripts/receipt_scanner.py
  5. Return confidence caveats when the source is low quality, rotated, handwritten, or multilingual.

Guardrails

  • Prefer explicit language selection when accuracy matters.
  • Do not claim fields are exact when OCR confidence is weak.
  • Route non-scanned digital PDFs to document-converter-suite instead of OCR by default.

Source: SKILL.md on GitHub

No alerts15d4 checks · Risk SAFE
  • Gen Agent Trust Hub15d

    This skill provides OCR capabilities to extract text from images and PDF files. While functional, it is vulnerable to indirect prompt injection because it reads content from external documents that may contain instructions intended to hijack the AI's behavior.

  • Socket15d

    No alerts

  • Snyk15d

    Risk: LOW · No issues

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 103b430. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 months ago.

Steadyupdated 6 months ago
  • Python
  • ocr
  • document-processing
  • pdf
  • image-extraction
  • receipt-parsing
  • business-card
  • table-extraction
  • text-recognition

README badge

README badge for dkyazzentwatwa/chatgpt-skills/ocr-document-processor

Extracts text and structure from scanned images and PDFs using OCR, with specialized helpers for receipts, business cards, and tables. Returns text, markdown, JSON, or HTML output with confidence caveats for low-quality or multilingual sources.

Generated from the current SKILL.md.

What input formats does this skill support?
Images and scanned PDFs. For non-scanned digital PDFs, the skill recommends routing to document-converter-suite instead.
Can this skill extract tables from scanned documents?
Yes. Table extraction is listed as a core use case alongside receipt and business card parsing.
What output formats are available?
Text, markdown, JSON, and HTML. The skill also supports searchable PDF export.
How does the skill handle low-quality or rotated scans?
It includes preprocessing steps for noisy inputs with skew, blur, or shadows. When confidence is weak, handwritten, or multilingual, it returns confidence caveats rather than claiming field exactness.
Are there specialized scripts for specific document types?
Yes. The skill provides focused helpers for business card scanning and receipt parsing alongside the core OCR processor.

Generated from the current SKILL.md. These answers refresh after source changes.