All skills

Process, convert, scan, read, edit, sign, and fill documents with the Nutrient DWS API. Supports PDF, DOCX, XLSX, PPTX, HTML, and image files.

  • 1 file
  • 9.6 KB
  • Updated 2 weeks ago
  • GitHub

Use this Skill: https://skilld.dev/gh/agenticluke/document-processor-plus/skill

This session only. Nothing lands on disk.

SKILL.md

≈37 tokens always: the name and description. ≈2.4k when used: this file.

Nutrient Document Processing

Use the Nutrient DWS Processor API to work with documents.

Use this skill to:

  • Change file types.
  • Read text and tables.
  • Run OCR on scans and images.
  • Hide private data.
  • Add watermarks.
  • Sign PDF files.
  • Fill PDF forms.

Setup

Get an API key from Nutrient.

Set the key in your shell:

export NUTRIENT_API_KEY="pdf_live_..."

Never put the key in code, logs, or saved files.

Each request is a multipart POST request to:

https://api.nutrient.io/build

Each request must include an instructions JSON field.

Basic Rules

  • Check that NUTRIENT_API_KEY is set before a request.
  • Check that each input file exists.
  • Use the same file name in the upload field and the JSON.
  • Put file names with spaces inside quotes.
  • Pick an output name that will not replace the input.
  • Do not send private files unless the user agrees.
  • Treat all output as unsafe until the request succeeds.
  • Open or test the output before deleting the source file.
  • Do not retry a failed request in a loop.
  • Do not guess form field names. Read them from the PDF first.
  • Test redaction on a copy. Make sure the hidden text cannot be found or copied.

Convert Documents

DOCX to PDF

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.docx=@document.docx" \
  -F 'instructions={"parts":[{"file":"document.docx"}]}' \
  -o output.pdf

PDF to DOCX

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"output":{"type":"docx"}}' \
  -o output.docx

HTML to PDF

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "index.html=@index.html" \
  -F 'instructions={"parts":[{"html":"index.html"}]}' \
  -o output.pdf

Input types include PDF, DOCX, XLSX, PPTX, DOC, XLS, PPT, PPS, PPSX, ODT, RTF, HTML, JPG, PNG, TIFF, HEIC, GIF, WebP, SVG, TGA, and EPS.

Some files may lose fonts, links, layout, notes, or other parts during a change. Check the result by sight.

Read Text and Tables

Read plain text

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"output":{"type":"text"}}' \
  -o output.txt

Read tables into Excel

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"output":{"type":"xlsx"}}' \
  -o tables.xlsx

A scan may have no text layer. Run OCR first if the text output is empty or poor.

Table output may need a manual check. Merged cells, small print, and page breaks can cause errors.

Run OCR

Use OCR to add a text layer to a scan.

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "scanned.pdf=@scanned.pdf" \
  -F 'instructions={"parts":[{"file":"scanned.pdf"}],"actions":[{"type":"ocr","language":"eng"}]}' \
  -o searchable.pdf

OCR supports more than 100 languages. Common codes include:

  • eng for English
  • deu for German
  • fra for French
  • spa for Spanish
  • jpn for Japanese
  • kor for Korean
  • chi_sim for simple Chinese
  • chi_tra for old-style Chinese
  • ara for Arabic
  • hin for Hindi
  • rus for Russian

Full names such as english may also work. Codes are safer. See the OCR language list.

OCR can make mistakes with handwriting, low light, tilted pages, or small text. Check names, dates, and numbers by hand.

Hide Private Data

Redaction should remove private text, not just cover it with a dark box.

Use known data types

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"actions":[{"type":"redaction","strategy":"preset","strategyOptions":{"preset":"social-security-number"}},{"type":"redaction","strategy":"preset","strategyOptions":{"preset":"email-address"}}]}' \
  -o redacted.pdf

Use a text pattern

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"actions":[{"type":"redaction","strategy":"regex","strategyOptions":{"regex":"\\b[A-Z]{2}\\d{6}\\b"}}]}' \
  -o redacted.pdf

Preset names include:

  • social-security-number
  • email-address
  • credit-card-number
  • international-phone-number
  • north-american-phone-number
  • date
  • time
  • url
  • ipv4
  • ipv6
  • mac-address
  • us-zip-code
  • vin

A preset can miss data or hide safe data. Check every page. Search the saved PDF for the old text before sharing it.

Add a Watermark

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"actions":[{"type":"watermark","text":"CONFIDENTIAL","fontSize":72,"opacity":0.3,"rotation":-45}]}' \
  -o watermarked.pdf

Use short text. Check that the mark does not hide key content.

Sign a PDF

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "document.pdf=@document.pdf" \
  -F 'instructions={"parts":[{"file":"document.pdf"}],"actions":[{"type":"sign","signatureType":"cms"}]}' \
  -o signed.pdf

This adds a CMS signature. It may not meet every legal or trust rule. Confirm the needed type before signing a legal file. Do not say a signature is valid until a PDF reader has checked it.

Fill a PDF Form

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "form.pdf=@form.pdf" \
  -F 'instructions={"parts":[{"file":"form.pdf"}],"actions":[{"type":"fillForm","formFields":{"name":"Jane Smith","email":"jane@example.com","date":"2026-02-06"}}]}' \
  -o filled.pdf

The keys in formFields must match the field names inside the PDF. A label shown on the page may not be the real field name. Some PDFs only look like forms and have no form fields.

Full Example

This example turns a scan into a PDF with text, then reads that text.

test -n "$NUTRIENT_API_KEY" || {
  echo "NUTRIENT_API_KEY is not set" >&2
  exit 1
}

test -f receipt-scan.pdf || {
  echo "receipt-scan.pdf was not found" >&2
  exit 1
}

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "receipt-scan.pdf=@receipt-scan.pdf" \
  -F 'instructions={"parts":[{"file":"receipt-scan.pdf"}],"actions":[{"type":"ocr","language":"eng"}]}' \
  -o receipt-searchable.pdf

curl --fail-with-body -X POST https://api.nutrient.io/build \
  -H "Authorization: Bearer $NUTRIENT_API_KEY" \
  -F "receipt-searchable.pdf=@receipt-searchable.pdf" \
  -F 'instructions={"parts":[{"file":"receipt-searchable.pdf"}],"output":{"type":"text"}}' \
  -o receipt.txt

After it runs:

  1. Open receipt-searchable.pdf.
  2. Make sure the pages look right.
  3. Read receipt.txt.
  4. Check all prices, dates, and names against the scan.

Handle Errors

curl --fail-with-body stops on an HTTP error and shows the error text.

If a request fails:

  1. Check that the API key is set.
  2. Check the input path and file name.
  3. Check that the upload name matches the name in instructions.
  4. Check that the JSON uses valid quotes and braces.
  5. Check that the file type is supported.
  6. Check the API error text.
  7. Keep the source file and remove any bad output file.

A failed request can still create an empty output file. Check the file size before using it:

test -s output.pdf || {
  echo "The output file is empty" >&2
  exit 1
}

MCP Server Option

Use the MCP server when a tool needs direct access without a hand-written curl command.

{
  "mcpServers": {
    "nutrient-dws": {
      "command": "npx",
      "args": ["-y", "@nutrient-sdk/dws-mcp-server"],
      "env": {
        "NUTRIENT_DWS_API_KEY": "YOUR_API_KEY",
        "SANDBOX_PATH": "/path/to/working/directory"
      }
    }
  }
}

Set SANDBOX_PATH to the smallest folder the tool needs. Do not use a broad path such as / or your full home folder. Keep the real API key out of files that may be shared or saved in source control.

When to Use This Skill

Use this skill when the user asks to:

  • Change a document to PDF, DOCX, XLSX, or another supported type.
  • Read text, tables, or key-value data from a PDF.
  • Run OCR on a scan or image.
  • Remove private data before sharing a file.
  • Add a watermark.
  • Sign a PDF.
  • Fill a PDF form with code.

Ask before sending a file if it may hold health, money, legal, work, or identity data.

Links

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 2265d4b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 weeks ago.

Activeupdated 2 weeks ago

README badge

README badge for agenticluke/document-processor-plus