---
name: pdf-processing
description: Extract text and tables from PDF files, fill forms, merge documents. Use when working with PDF files or when the user mentions PDFs, forms, or document extraction.
title: pdf-processing
canonical_url: https://skilld.dev/gh/tencent/weknora/pdf-processing
last_updated: 2026-09-29T02:47:02.000Z
---

> **Skill from skilld.dev.** Follow the instructions below for this session. You do not need to install anything.
>
> Supporting files, fetch one when the Skill refers to it: [FORMS.md](https://skilld.dev/api/skills-raw/tencent/weknora/pdf-processing/FORMS.md), [scripts/analyze_form.py](https://skilld.dev/api/skills-raw/tencent/weknora/pdf-processing/scripts/analyze_form.py), [scripts/extract_text.py](https://skilld.dev/api/skills-raw/tencent/weknora/pdf-processing/scripts/extract_text.py).
>
> If the user asked to install this Skill, run `npx skilld install tencent/weknora/pdf-processing`. Install writes the Skill files into the project, so every session loads them.

# PDF Processing

This skill provides utilities for working with PDF documents.

## Quick Start

Use pdfplumber to extract text from PDFs:

```python
import pdfplumber

with pdfplumber.open("document.pdf") as pdf:
    text = pdf.pages[0].extract_text()
    print(text)
```

## Available Operations

1. **Text Extraction**: Extract text content from PDF pages
2. **Table Extraction**: Extract tabular data from PDFs
3. **Form Filling**: Fill PDF forms with provided data
4. **Document Merging**: Combine multiple PDFs into one

## Advanced Features

**Form filling**: See [FORMS.md](https://skilld.dev/api/skills-raw/tencent/weknora/pdf-processing/FORMS.md) for complete guide

**Utility scripts**: 
- Run `scripts/analyze_form.py` to extract form fields
- Run `scripts/extract_text.py` to extract text from a PDF

## Best Practices

1. Always validate PDF files before processing
2. Handle password-protected PDFs gracefully
3. Check for scanned PDFs that may require OCR
