OCRing PDFs
ocrmypdf writes an invisible text layer over the original page images, so the
output file is both the scan you can look at and a document pdftotext, grep,
and pdfplumber can read. Rasterize-then-tesseract gives you a .txt divorced
from the pages; page numbers and coordinates are gone.
The toolchain is not in the base container. It installs in 18 seconds (measured 2026-09-12: apt 3s, pip 15s), so install it when a scan shows up rather than carrying it in a container layer.
Probe before installing
pdftotext in.pdf - | tr -d '\f \n' | wc -cNonzero means the PDF already has a text layer and is not a scan. Extract with
pdftotext or pdfplumber and stop. Running OCR on it wastes a minute, and with
--force-ocr it replaces exact embedded text with a lossy reading of a raster of
itself.
A small nonzero count (tens of characters across many pages) is the mixed case:
a born-digital cover page in front of scanned body pages, or a scan whose
producer stamped a header. --skip-text handles it.
Install
sh scripts/ensure_ocr.sh # English
sh scripts/ensure_ocr.sh nor deu # plus Norwegian and GermanIdempotent: 0.8s when everything is already present, 2.6s to add one more
language pack. Installs ghostscript, pngquant, poppler-utils, tesseract
and its language packs via apt, then ocrmypdf via pip.
Run
ocrmypdf --skip-text --deskew --rotate-pages --output-type pdf in.pdf out.pdf
pdftotext out.pdf - | wc -w # verify: zero words means it failed quietlyAbout 2s per page for a single dense page at 200 DPI on one core. A 300-page
scan is therefore a background job, not a single bash call — launch it detached
with a sentinel file per the external-call pattern in bash-tool-timeout.
Which text-layer mode
| flag | use it when |
|---|---|
--skip-text |
Default. Pages that already carry text are passed through untouched; image-only pages get OCR. The safe choice for anything mixed. |
--force-ocr |
Every page is rasterized and re-OCRed, discarding any existing text. Correct for a scan carrying a junk text layer, and for pages with text-over-image that --skip-text would skip. Destroys real embedded text, so probe first. |
--redo-ocr |
Replaces a previous OCR layer while leaving born-digital text alone. Narrower than --force-ocr and slower to fail on odd inputs. |
--output-type pdf skips PDF/A conversion. Drop it when the output is going into
an archive that requires PDF/A; ghostscript does the conversion either way.
Languages
-l eng+nor for a mixed-language document, -l nor for a monolingual one. Order
does not matter. Every code needs its tesseract-ocr-<code> pack installed.
Pass the codes to ensure_ocr.sh and it handles them. Accuracy drops noticeably
when the language is wrong, and tesseract will not tell you; it returns
confident garbage instead.
Container facts (measured 2026-09-12)
apt-get updateexits 100 here. A preconfigured nodesource repo is off the egress allowlist and returns 403, and the nonzero exit aborts any&&chain behind it. The Ubuntu mirrors are reachable without an update. Runapt-get installdirectly.unpaperis absent, so--cleanand--clean-finalfail. Don't pass them.- One core, so
--jobsbuys nothing on claude.ai. CCotw has four. ocrmypdf --versionprints to stderr. Capture with2>&1or a version check reads as empty.jbig2is absent; output uses CCITT/JPEG instead, which costs some file size and nothing else.
When to use transcribing-images instead
This skill produces glyphs. It does not read a chart, describe a diagram, or recover handwriting. Tesseract on those pages returns nothing useful and gives no sign that it lost anything.
Route to transcribing-images when the meaningful content is a picture, or when
the deliverable is a reading rather than a file. Both is a normal answer: OCR the
document so it is greppable, then send the pages that carry figures to a vision
model.
In an interactive session, native vision beats both for a handful of pages:
rasterize with pdftoppm -r 200 -png and view the images. Reach for OCR when
the document is longer than context will hold, or when the text has to outlive
the conversation as a file.