All skills
massimodeluisa avatar

/recursive-decomposition

@e0a8b04

Decompose dense codebase-wide, multi-document, PDF, and aggregation work even when the input fits the context window, following Recursive Language Models (Zhang, Kraska, Khattab, 2025). Use when the user asks to analyse all files, a whole repo, all docs, large PDFs, or to aggregate or multi-hop across scattered sources. Skip one file, one function, a single needle, or a one-page PDF conversion. Triggers: long context, context rot, large codebase, many files, all files, big document, multi-document, PDF, aggregate, summarize everything, codebase-wide, multi-hop, recursive, sub-agents, map-reduce.

Use this Skill: https://skilld.dev/gh/massimodeluisa/recursive-decomposition-skill/recursive-decomposition

This session only. Nothing lands on disk.

evalsREADME.md

≈579 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Evals

Large PDFs live in a git submodule. They are not copied into this repository.

Corpus: tccao/mortgage-doc-rag (MIT). 131 public-domain mortgage PDFs, about 63 MB.

git submodule update --init --depth 1 skills/recursive-decomposition/evals/files/mortgage-doc-rag

Layout follows Evaluating skills. Runs go under evals-workspace/ at the repository root (gitignored).

Check (no agent)

bash .github/scripts/eval-skill.sh check

Validates trigger queries and, if the submodule is present, PDF count, byte floor, and the largest filename. CI does not clone the submodule. Missing PDFs print SKIP, not ERROR.

Agent runs (with vs without the skill)

Same prompt twice, clean context:

evals-workspace/iteration-1/eval-pdf-corpus/with_skill/outputs/result.json
evals-workspace/iteration-1/eval-pdf-corpus/without_skill/outputs/result.json
{
  "decomposed": true,
  "depth": 1,
  "subagents_spawned_subagents": false,
  "pdf_count": 131,
  "largest": [{"path": "data/degraded/appraisal/urar_form_1004_epa_scan.pdf", "bytes": 1639534}]
}
bash .github/scripts/eval-skill.sh score evals-workspace/iteration-1

Trigger queries live in trigger-queries.json. Score those by whether the agent loaded this skill.

Firecrawl document skills

Skill What it does
firecrawl/anydoc@convert-documents-to-markdown Local npx -y @firecrawl/anydoc FILE -o out.md. Exit 3 means OCR; then --ocr hosted.
firecrawl/cli@firecrawl-parse Cloud firecrawl parse. 50 MB cap, about 1 credit per page.

Prefer anydoc. Many files in this corpus are scans; pdftotext on the whole tree extracts nothing useful. Size first, then parse only what the question needs.

Live pre/post (same prompt, 2026-09-11): size-first 0.02 s. anydoc on two digital PDFs 1.64 s, titles recovered. The largest file is a 4-page scan: anydoc exit 3, firecrawl parse 15.45 s, title Uniform Residential Appraisal Report. Naive pdftotext on all 131 files: 1.65 s and 0 bytes; on that scan, 4 bytes and no title.

Source: SKILL.md on GitHub

1 warning13d4 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    The skill implements a recursive decomposition strategy for processing large inputs. It is generally safe and follows established research protocols, but it possesses an inherent risk of indirect prompt injection because it processes external files and documents. It also relies on the Firecrawl utility for document conversion, which involves executing remote packages.

  • Socket13d

    No alerts

  • Snyk13d

    Risk: LOW · No issues

  • Runlayer7mo

    5/5 files flagged

Signed by skilld at e0a8b04. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 weeks ago.

Activeupdated 3 weeks ago
Other metadata
metadata
{
  "author": "massimodeluisa",
  "version": "1.2.0",
  "paper": "https://arxiv.org/abs/2512.24601"
}

README badge

README badge for massimodeluisa/recursive-decomposition-skill