All skills
agents365-ai avatar

/scholar-deep-research

@72ee04b

Use when the user asks for a literature review, academic deep dive, research report, state-of-the-art survey, topic scoping, comparative analysis of methods/papers, grant background, or any request that needs multi-source scholarly evidence with citations. Also trigger proactively when a user question clearly requires academic grounding (e.g. "what's known about X", "compare approach A vs B in the literature", "summarize the field of Y"). Runs an 8-phase (Phase 0..7), script-driven research workflow across 7 federated sources (OpenAlex, arXiv, Crossref, PubMed, DBLP, bioRxiv, Exa) with optional Semantic Scholar / Brave MCP enrichment, with deduplication, transparent ranking, dual-backend citation chasing (OpenAlex + Semantic Scholar), self-critique, and structured report output with verifiable citations.

Use this Skill: https://skilld.dev/gh/agents365-ai/365-skills/scholar-deep-research

This session only. Nothing lands on disk.

referencessource_selection.md

≈804 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Source Selection

When you have four databases and limited rounds, where do you spend the queries? This is a one-page decision guide.

Default — always run

Source Why
OpenAlex Backbone. Free, no key, 240M+ works, citation counts, DOI, PDF URLs when OA. Run first on every cluster.

Add by domain

If the question is about... Add this source
Biology, medicine, public health, drugs, clinical trials PubMed — MeSH terms, clinical trial filters, abstracts
Computer science, ML, AI, statistics, physics, math arXiv — preprints with the latest unpublished work
A specific paper (have a DOI/title) Crossref — authoritative metadata, journal lookup, version-of-record
Cross-disciplinary topics All four (overlap is feature, not bug — dedupe handles it)

Add by question type

Question type Source priority
"What is known about X?" (overview) OpenAlex → +PubMed/arXiv by domain
"What's the latest in Y?" (recency) arXiv (CS/ML/physics) or PubMed (bio) → OpenAlex
"Compare A vs B" OpenAlex → cited-by graph from top results
"Who works on Z?" (people) OpenAlex → search by author
"What is the seminal paper on W?" OpenAlex → sort by citations descending
"Has X been replicated?" OpenAlex (forward snowball) → look for "replication" / "reproduction" terms
"Are there critiques of paper P?" Crossref + author search for the critic; OpenAlex cited-by

What about Semantic Scholar / Google Scholar / Web of Science?

  • Semantic Scholar is excellent. We expose it as enrichment via the asta MCP tools when available. If those time out, you lose some semantic ranking but the rest of the pipeline keeps going. Treat it as a bonus, not a dependency.
  • Google Scholar has no public API (scraping is against ToS and brittle). Skip.
  • Web of Science / Scopus require institutional access. Mention in the report appendix as "not consulted" if your user explicitly cares about indexing rigor.
  • Brave Search (web search MCP) is for non-academic sources — press releases, blog posts, community discussion of preprints. Only run when the question explicitly needs them.

What each source is BAD at

  • OpenAlex: occasional metadata errors (wrong author, wrong year). Verify the top-N before they enter the report.
  • arXiv: no citation counts, no peer review.
  • Crossref: only catalogs DOI-registered works; misses arXiv preprints and many gray-lit reports. Citation counts are incoming-only (not as semantic as OpenAlex).
  • PubMed: biomedical only, and some preprints/non-MEDLINE journals are missing.

Rate limits and politeness

Source Polite-pool ID Notes
OpenAlex --email <you@host> Higher rate limit, faster queue
Crossref --email <you@host> Same
PubMed --api-key <key> 10 req/s with key, 3 req/s without
arXiv User-Agent only ~1 req/3s; the script doesn't paginate aggressively

When in doubt, pass --email. It costs nothing and unblocks higher throughput.

Source: SKILL.md on GitHub

1 warning4mo3 checks · Risk SAFE
  • Gen Agent Trust Hub4mo

    This skill is a robust academic research tool that automates literature reviews across multiple scholarly databases. It uses strict input validation for paper identifiers and handles sensitive API keys through environment variables. While the skill executes a local script from a companion 'paper-fetch' skill to download PDFs, this is a standard design pattern within the author's ecosystem. It also processes external PDF text, which carries a minor risk of indirect prompt injection common to all document-reading agents.

  • Socket4mo

    No alerts

  • Snyk4mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 72ee04b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 11 hours ago.

Activeupdated 3 weeks ago
homepage
https://github.com/Agents365-ai/365-skills
platforms
[macos, linux, windows]
Other metadata
compatibility
Requires Python 3.9+ with httpx and pypdf (see requirements.txt). Optional: `pip install docling` to enable layout-aware markdown PDF extraction (`extract_pdf.py --engine docling`); auto-used as a fallback for scanned/sparse PDFs. Works offline-first (no MCP required) but enriches with Semantic Scholar / Brave MCP tools when available.
metadata
{
  "openclaw": {
    "requires": {
      "bins": [
        "python3"
      ]
    },
    "emoji": "🔬"
  },
  "hermes": {
    "tags": [
      "research",
      "literature-review",
      "academic",
      "papers",
      "citations",
      "survey"
    ],
    "category": "research"
  },
  "pimo": {
    "tags": [
      "research",
      "literature-review",
      "academic"
    ],
    "category": "research"
  },
  "author": "Agents365-ai",
  "version": "0.17.0"
}

README badge

README badge for agents365-ai/365-skills/scholar-deep-research