All skills
agents365-ai avatar

/scholar-deep-research

@72ee04b

Use when the user asks for a literature review, academic deep dive, research report, state-of-the-art survey, topic scoping, comparative analysis of methods/papers, grant background, or any request that needs multi-source scholarly evidence with citations. Also trigger proactively when a user question clearly requires academic grounding (e.g. "what's known about X", "compare approach A vs B in the literature", "summarize the field of Y"). Runs an 8-phase (Phase 0..7), script-driven research workflow across 7 federated sources (OpenAlex, arXiv, Crossref, PubMed, DBLP, bioRxiv, Exa) with optional Semantic Scholar / Brave MCP enrichment, with deduplication, transparent ranking, dual-backend citation chasing (OpenAlex + Semantic Scholar), self-critique, and structured report output with verifiable citations.

Use this Skill: https://skilld.dev/gh/agents365-ai/365-skills/scholar-deep-research

This session only. Nothing lands on disk.

referencessearch_strategies.md

≈1.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Search Strategies

Reference for Phase 1 (Discovery) and Phase 4 (Citation chasing). Load this when planning queries or when discovery feels unproductive.

Boolean clusters

A query is rarely a single keyword. Build 3-5 clusters of synonyms and join them with AND. Example for "CRISPR base editing for muscular dystrophy":

Cluster A (technique): "base editing" OR "adenine base editor" OR ABE OR
                       "cytosine base editor" OR CBE OR prime editing
Cluster B (disease):   "Duchenne muscular dystrophy" OR DMD OR
                       "Becker muscular dystrophy" OR dystrophinopathy OR
                       dystrophin
Cluster C (delivery):  AAV OR "adeno-associated virus" OR LNP OR
                       "lipid nanoparticle" OR "in vivo delivery"
Negative:              NOT review NOT editorial NOT comment

Each search script accepts one query at a time — run one cluster per call, then dedupe across all of them. Don't pre-AND clusters; that over-constrains and misses cross-references.

PICO (or PICO-style)

For comparative or systematic questions, decompose with PICO:

Letter What
P Population, problem, or phenomenon
I Intervention or independent variable
C Comparator (often baseline or alternative)
O Outcome — what is measured

Non-biomedical translations:

  • ML: Task, Method, Baseline, Metric (TMBM)
  • Engineering: System, Modification, Reference, Performance
  • Social science: Population, Treatment, Control, Effect

PICO matters because it forces you to articulate what counts as an answer — and that constrains your search.

Snowballing

Two flavors. Both are run by build_citation_graph.py:

  • Backward snowballing. Pull the references of high-quality seed papers. The best work cites foundational papers you'd otherwise miss with keywords.
  • Forward snowballing. Pull the cited-by list. The most recent critiques, replications, and extensions live here.

Special case — citation chasing for criticism. When a paper has very high citations but no critique appears in your corpus, search explicitly:

"<first author> <year>" (critique OR limitations OR reanalysis OR
                          replication OR failed OR flawed OR overstated)

If no critique exists, that's interesting. Note it. Don't assume "uncriticized" = "correct."

Saturation

Discovery ends when adding more search rounds stops adding signal, not when you get tired. The script research_state.py saturation formalizes this:

saturated when:
  (new_papers_in_round / total_in_round) < 20%   AND
  max(citations of papers first seen in round) < 100

The first condition catches "we've seen most of these before." The second catches "but there's still a high-impact paper we missed." Both must hold.

If one cluster saturates and another doesn't, run more rounds on the unsaturated cluster only.

Iteration patterns

A productive search sequence usually looks like:

Round 1: broad keywords on each cluster, limit=50
Round 2: tighter keywords (using terms learned from Round 1 abstracts)
Round 3: author search for repeat-appearing authors (search by author name)
Round 4: citation chase (build_citation_graph) on top 8 seeds
Round 5: targeted gap fills based on Phase 6 self-critique

A keyword you didn't know existed yesterday but appears in 6 abstracts today is a signal. Re-run Round 1 with that keyword added to the cluster.

Source-specific tips

  • OpenAlex is the best general-purpose first stop. Free, no key, citation counts included.
  • arXiv for CS/ML/physics preprints. Note: papers are unrefereed until cross-listed with a venue.
  • Crossref is best for known DOIs — verify metadata, find venue, get bibliographic detail.
  • PubMed for biomedical questions. Use MeSH terms when you know them: "CRISPR-Cas Systems"[MeSH].

Common failure modes

  • First-page fixation. The first 10 OpenAlex results aren't "the literature." Run multiple rounds and chase citations.
  • Acronym blindness. "ABE" means base editor, but also "average bit error" — disambiguate with co-occurring terms.
  • English-only. Some fields have important non-English literature. Note this in the report's limitations.
  • Recency collapse. Saturation can fire before the most recent year is well-covered. Always run a final query restricted to the last 18 months.

Source: SKILL.md on GitHub

1 warning4mo3 checks · Risk SAFE
  • Gen Agent Trust Hub4mo

    This skill is a robust academic research tool that automates literature reviews across multiple scholarly databases. It uses strict input validation for paper identifiers and handles sensitive API keys through environment variables. While the skill executes a local script from a companion 'paper-fetch' skill to download PDFs, this is a standard design pattern within the author's ecosystem. It also processes external PDF text, which carries a minor risk of indirect prompt injection common to all document-reading agents.

  • Socket4mo

    No alerts

  • Snyk4mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 72ee04b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 11 hours ago.

Activeupdated 3 weeks ago
homepage
https://github.com/Agents365-ai/365-skills
platforms
[macos, linux, windows]
Other metadata
compatibility
Requires Python 3.9+ with httpx and pypdf (see requirements.txt). Optional: `pip install docling` to enable layout-aware markdown PDF extraction (`extract_pdf.py --engine docling`); auto-used as a fallback for scanned/sparse PDFs. Works offline-first (no MCP required) but enriches with Semantic Scholar / Brave MCP tools when available.
metadata
{
  "openclaw": {
    "requires": {
      "bins": [
        "python3"
      ]
    },
    "emoji": "🔬"
  },
  "hermes": {
    "tags": [
      "research",
      "literature-review",
      "academic",
      "papers",
      "citations",
      "survey"
    ],
    "category": "research"
  },
  "pimo": {
    "tags": [
      "research",
      "literature-review",
      "academic"
    ],
    "category": "research"
  },
  "author": "Agents365-ai",
  "version": "0.17.0"
}

README badge

README badge for agents365-ai/365-skills/scholar-deep-research