---
title: "skill by agenticluke · skilld"
canonical_url: "https://skilld.dev/gh/agenticluke/document-processor-plus"
meta:
  description: "Process, convert, scan, read, edit, sign, and fill documents with the Nutrient DWS API. Supports PDF, DOCX, XLSX, PPTX, HTML, and image files. From agenticluke/document-processor-plus."
  "og:description": "Process, convert, scan, read, edit, sign, and fill documents with the Nutrient DWS API. Supports PDF, DOCX, XLSX, PPTX, HTML, and image files. From agenticluke/document-processor-plus."
  "og:title": "skill by agenticluke"
  "twitter:description": "Process, convert, scan, read, edit, sign, and fill documents with the Nutrient DWS API. Supports PDF, DOCX, XLSX, PPTX, HTML, and image files. From agenticluke/document-processor-plus."
  "twitter:title": "skill by agenticluke"
---

`

[All skills](https://skilld.dev/skills)

[![agenticluke avatar](https://skilld.dev/_img/avatar?url=https%3A%2F%2Fgithub.com%2Fagenticluke.png%3Fsize%3D96)](https://skilld.dev/gh/agenticluke)

# **/skill**

[@2265d4b](https://github.com/agenticluke/document-processor-plus/commit/2265d4b036287294b61e160acb6750cad76a9c19 "Your agent reads SKILL.md at commit 2265d4b")

by [agenticluke](https://skilld.dev/gh/agenticluke)· [agenticluke](https://skilld.dev/gh/agenticluke)/ [document-processor-plus](https://skilld.dev/gh/agenticluke/document-processor-plus)

Process, convert, scan, read, edit, sign, and fill documents with the Nutrient DWS API. Supports PDF, DOCX, XLSX, PPTX, HTML, and image files.

- 1 file
- 9.6 KB
- Updated 2 weeks ago
- [GitHub](https://github.com/agenticluke/document-processor-plus/blob/main/skill/SKILL.md "View SKILL.md on GitHub")

## SKILL.md

9.6 KB

**≈37** tokens always: the name and description. **≈2.4k** when used: this file.

## Nutrient Document Processing

Use the [Nutrient DWS Processor API](https://www.nutrient.io/api/) to work with documents.

Use this skill to:

- Change file types.
- Read text and tables.
- Run OCR on scans and images.
- Hide private data.
- Add watermarks.
- Sign PDF files.
- Fill PDF forms.

### Setup

Get an API key from [Nutrient](https://dashboard.nutrient.io/sign_up/?product=processor).

Set the key in your shell:

```
export NUTRIENT_API_KEY="pdf_live_..."
```

Never put the key in code, logs, or saved files.

Each request is a multipart `POST` request to:

```
https://api.nutrient.io/build
```

Each request must include an `instructions` JSON field.

### Basic Rules

- Check that `NUTRIENT_API_KEY` is set before a request.
- Check that each input file exists.
- Use the same file name in the upload field and the JSON.
- Put file names with spaces inside quotes.
- Pick an output name that will not replace the input.
- Do not send private files unless the user agrees.
- Treat all output as unsafe until the request succeeds.
- Open or test the output before deleting the source file.
- Do not retry a failed request in a loop.
- Do not guess form field names. Read them from the PDF first.
- Test redaction on a copy. Make sure the hidden text cannot be found or copied.

### Convert Documents

#### DOCX to PDF

```
curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.docx=@document.docx" \
  -F 'instructions={"parts":[{"file":"document.docx"}]}' \
  -o output.pdf
```

#### PDF to DOCX

```
curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"output":{"type":"docx"}}' \
  -o output.docx
```

#### HTML to PDF

```
curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "index.html=@index.html" \
  -F 'instructions={"parts":[{"html":"index.html"}]}' \
  -o output.pdf
```

Input types include PDF, DOCX, XLSX, PPTX, DOC, XLS, PPT, PPS, PPSX, ODT, RTF, HTML, JPG, PNG, TIFF, HEIC, GIF, WebP, SVG, TGA, and EPS.

Some files may lose fonts, links, layout, notes, or other parts during a change. Check the result by sight.

### Read Text and Tables

#### Read plain text

```
curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"output":{"type":"text"}}' \
  -o output.txt
```

#### Read tables into Excel

```
curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"output":{"type":"xlsx"}}' \
  -o tables.xlsx
```

A scan may have no text layer. Run OCR first if the text output is empty or poor.

Table output may need a manual check. Merged cells, small print, and page breaks can cause errors.

### Run OCR

Use OCR to add a text layer to a scan.

```
curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "scanned.pdf=@scanned.pdf" \
  -F 'instructions={"parts":[{"file":"scanned.pdf"}],"actions":[{"type":"ocr","language":"eng"}]}' \
  -o searchable.pdf
```

OCR supports more than 100 languages. Common codes include:

- `eng` for English
- `deu` for German
- `fra` for French
- `spa` for Spanish
- `jpn` for Japanese
- `kor` for Korean
- `chi_sim` for simple Chinese
- `chi_tra` for old-style Chinese
- `ara` for Arabic
- `hin` for Hindi
- `rus` for Russian

Full names such as `english` may also work. Codes are safer. See the [OCR language list](https://www.nutrient.io/guides/document-engine/ocr/language-support/).

OCR can make mistakes with handwriting, low light, tilted pages, or small text. Check names, dates, and numbers by hand.

### Hide Private Data

Redaction should remove private text, not just cover it with a dark box.

#### Use known data types

```
curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"actions":[{"type":"redaction","strategy":"preset","strategyOptions":{"preset":"social-security-number"}},{"type":"redaction","strategy":"preset","strategyOptions":{"preset":"email-address"}}]}' \
  -o redacted.pdf
```

#### Use a text pattern

```
curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"actions":[{"type":"redaction","strategy":"regex","strategyOptions":{"regex":"\\b[A-Z]{2}\\d{6}\\b"}}]}' \
  -o redacted.pdf
```

Preset names include:

- `social-security-number`
- `email-address`
- `credit-card-number`
- `international-phone-number`
- `north-american-phone-number`
- `date`
- `time`
- `url`
- `ipv4`
- `ipv6`
- `mac-address`
- `us-zip-code`
- `vin`

A preset can miss data or hide safe data. Check every page. Search the saved PDF for the old text before sharing it.

### Add a Watermark

```
curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"actions":[{"type":"watermark","text":"CONFIDENTIAL","fontSize":72,"opacity":0.3,"rotation":-45}]}' \
  -o watermarked.pdf
```

Use short text. Check that the mark does not hide key content.

### Sign a PDF

```
curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"actions":[{"type":"sign","signatureType":"cms"}]}' \
  -o signed.pdf
```

This adds a CMS signature. It may not meet every legal or trust rule. Confirm the needed type before signing a legal file. Do not say a signature is valid until a PDF reader has checked it.

### Fill a PDF Form

```
curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "form.pdf=@form.pdf" \
  -F 'instructions={"parts":[{"file":"form.pdf"}],"actions":[{"type":"fillForm","formFields":{"name":"Jane Smith","email":"jane@example.com","date":"2026-02-06"}}]}' \
  -o filled.pdf
```

The keys in `formFields` must match the field names inside the PDF. A label shown on the page may not be the real field name. Some PDFs only look like forms and have no form fields.

### Full Example

This example turns a scan into a PDF with text, then reads that text.

```
test -n "$NUTRIENT_API_KEY" || {
  echo "NUTRIENT_API_KEY is not set" >&2
  exit 1
}

test -f receipt-scan.pdf || {
  echo "receipt-scan.pdf was not found" >&2
  exit 1
}

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "receipt-scan.pdf=@receipt-scan.pdf" \
  -F 'instructions={"parts":[{"file":"receipt-scan.pdf"}],"actions":[{"type":"ocr","language":"eng"}]}' \
  -o receipt-searchable.pdf

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "receipt-searchable.pdf=@receipt-searchable.pdf" \
  -F 'instructions={"parts":[{"file":"receipt-searchable.pdf"}],"output":{"type":"text"}}' \
  -o receipt.txt
```

After it runs:

1. Open `receipt-searchable.pdf`.
2. Make sure the pages look right.
3. Read `receipt.txt`.
4. Check all prices, dates, and names against the scan.

### Handle Errors

`curl --fail-with-body` stops on an HTTP error and shows the error text.

If a request fails:

1. Check that the API key is set.
2. Check the input path and file name.
3. Check that the upload name matches the name in `instructions`.
4. Check that the JSON uses valid quotes and braces.
5. Check that the file type is supported.
6. Check the API error text.
7. Keep the source file and remove any bad output file.

A failed request can still create an empty output file. Check the file size before using it:

```
test -s output.pdf || {
  echo "The output file is empty" >&2
  exit 1
}
```

### MCP Server Option

Use the MCP server when a tool needs direct access without a hand-written `curl` command.

```
{
  "mcpServers": {
    "nutrient-dws": {
      "command": "npx",
      "args": ["-y", "@nutrient-sdk/dws-mcp-server"],
      "env": {
        "NUTRIENT_DWS_API_KEY": "YOUR_API_KEY",
        "SANDBOX_PATH": "/path/to/working/directory"
      }
    }
  }
}
```

Set `SANDBOX_PATH` to the smallest folder the tool needs. Do not use a broad path such as `/` or your full home folder. Keep the real API key out of files that may be shared or saved in source control.

### When to Use This Skill

Use this skill when the user asks to:

- Change a document to PDF, DOCX, XLSX, or another supported type.
- Read text, tables, or key-value data from a PDF.
- Run OCR on a scan or image.
- Remove private data before sharing a file.
- Add a watermark.
- Sign a PDF.
- Fill a PDF form with code.

Ask before sending a file if it may hold health, money, legal, work, or identity data.

### Links

- [API Playground](https://dashboard.nutrient.io/processor-api/playground/)
- [API Guide](https://www.nutrient.io/guides/dws-processor/)
- [MCP Server on npm](https://www.npmjs.com/package/@nutrient-sdk/dws-mcp-server)

Source: [SKILL.md on GitHub](https://github.com/agenticluke/document-processor-plus/blob/main/skill/SKILL.md)

## Third-party checks

No third-party reports yet.

## Provenance

[Signed by skilld at 2265d4b.](https://github.com/agenticluke/document-processor-plus/commit/2265d4b036287294b61e160acb6750cad76a9c19 "2265d4b036287294b61e160acb6750cad76a9c19") This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 weeks ago.

Activeupdated 2 weeks ago

## README badge

![README badge for agenticluke/document-processor-plus](https://skilld.dev/b/agenticluke/document-processor-plus?theme=light&label=0)

## Related skills

-
-
-
-
-
-