All skills
oaustegard avatar

/transcribing-images

@f5c056b

Reads the visual content of slides, pages, and images the way a human would, not just their embedded text. Use when a PPTX or PDF has image slides, screenshots, charts, scanned figures, or flattened-to-image layouts that the built-in pptx/pdf skills read as empty; when asked to transcribe, describe, OCR, or extract what is shown in an image, slide deck, or document page; or when embedded-text extraction returned little or nothing from a visually rich file. Triggers on 'read this deck', 'what's on these slides', 'transcribe', 'OCR', 'extract text from image', 'describe this chart/diagram', .pptx/.pdf/.png/.jpg with visual content.

Use this Skill: https://skilld.dev/gh/oaustegard/claude-skills/transcribing-images

This session only. Nothing lands on disk.

CHANGELOG.md

≈201 tokens on demand. Your agent reads this file only when SKILL.md points to it.

transcribing-images - Changelog

All notable changes to the transcribing-images skill are documented in this file. The format is based on Keep a Changelog.

[0.1.2] - 2026-09-12

Other

  • Add ocring-pdfs skill; route file-deliverable OCR out of transcribing-images (#796)

[0.1.2] - 2026-09-12

Changed

  • Route file-deliverable OCR to the new ocring-pdfs skill; the tesseract engine here stays for loose-text passes on pages already known to be plain scans.

[0.1.1] - 2026-09-09

Other

  • prompt-audit: dated prompting patterns across the skill catalogue (#791)
  • Deprecate mapping-codebases; adopt ruff 0.16.0 baseline (#747)

[0.1.0] - 2026-06-19

Added

  • transcribing-images skill (visual reading of slides/pages) (#703)

Source: SKILL.md on GitHub

No alerts13d3 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    The skill is safe and performs its stated task of transcribing visual content from document files using vision models or local OCR. It uses standard system utilities for file conversion and rasterization in a secure manner. The primary risk identified is the inherent possibility of indirect prompt injection from the content of the processed documents, which is a standard risk for vision-based transcription tools.

  • Socket13d

    No alerts

  • Snyk13d

    Risk: LOW · No issues

Signed by skilld at f5c056b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 3 weeks ago
metadata
{
  "version": "0.1.2"
}

README badge

README badge for oaustegard/claude-skills/transcribing-images