All skills
agents365-ai avatar

/scholar-deep-research

@72ee04b

Use when the user asks for a literature review, academic deep dive, research report, state-of-the-art survey, topic scoping, comparative analysis of methods/papers, grant background, or any request that needs multi-source scholarly evidence with citations. Also trigger proactively when a user question clearly requires academic grounding (e.g. "what's known about X", "compare approach A vs B in the literature", "summarize the field of Y"). Runs an 8-phase (Phase 0..7), script-driven research workflow across 7 federated sources (OpenAlex, arXiv, Crossref, PubMed, DBLP, bioRxiv, Exa) with optional Semantic Scholar / Brave MCP enrichment, with deduplication, transparent ranking, dual-backend citation chasing (OpenAlex + Semantic Scholar), self-critique, and structured report output with verifiable citations.

Use this Skill: https://skilld.dev/gh/agents365-ai/365-skills/scholar-deep-research

This session only. Nothing lands on disk.

referencesquality_assessment.md

≈1.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Quality Assessment

Reference for Phase 3 (Deep read) and Phase 6 (Self-critique). Load this when judging whether a paper deserves the weight it would carry in the report.

CRAAP test (with adjustments)

The CRAAP framework is from library science but maps cleanly to academic work.

Letter Question What good looks like
Currency When was this published? Within field's relevance horizon (2-5 yr in fast fields, decades in others)
Relevance Does it match your question? Title + abstract directly speak to the PICO
Authority Who wrote it, where? Established lab, peer-reviewed venue, conflict-of-interest disclosed
Accuracy Is the methodology sound? Pre-registered if possible, sample size justified, code/data available
Purpose Why was it written? Primary research, not advocacy or marketing

CRAAP is necessary but not sufficient. A CRAAP-perfect paper can still be wrong.

Venue tiers (rough)

The rank_papers.py script uses a small built-in tier-1 list (Nature/Science/Cell/PNAS/NeurIPS/ICML/ICLR/...). Treat it as a starting prior, not a verdict.

Tier-1 signals (in approximate order of evidential weight):

  1. Replication by an independent group
  2. Multiple cited-by criticisms that failed to overturn the paper
  3. Inclusion in a Cochrane / NICE / FDA / consensus document
  4. Published in a top-tier venue
  5. High citation count (with attention to who is citing)

Citation count alone is a weak signal. A wrong paper can be highly cited (often cited as a counter-example).

Preprint handling

arXiv, bioRxiv, medRxiv, ChemRxiv, SSRN — none are peer-reviewed.

Rules:

  • Tag every preprint as preprint in the evidence section.
  • A claim in the report should not rest only on preprints unless the entire report is about pre-registered or in-flight work.
  • Check whether the preprint has since appeared in a journal — Crossref by author + title or OpenAlex by title is the fastest check.
  • Preprints with >100 citations and >12 months in the wild without journal publication deserve a note: "in pre-print since X, not yet peer reviewed."

Retraction check

Before deeply citing a paper:

  • Look up the DOI on Retraction Watch (https://retractionwatch.com) — it's slow but authoritative
  • Check OpenAlex is_retracted: true flag if available
  • Crossref returns a relation: is-corrected-by link when an erratum exists

The cost of citing a retracted paper as if it were valid is high. Spend the 30 seconds.

Conflicts and funding

Note in the evidence section when:

  • Funding came from an entity with a direct stake in the result (e.g., drug trial funded by manufacturer)
  • Senior author is on the advisory board of a directly-relevant company
  • The paper is part of a regulatory dossier, not an independent science paper

This isn't dismissal — it's calibration. A pharma-sponsored Phase 3 trial is still evidence; it's evidence that needs context.

Sample-size sanity

Quick smell tests by field:

Field Suspicious n Notes
Mouse studies n < 5 per arm Effect sizes inflate at small n
Human RCT n < 20 per arm Underpowered for most clinically meaningful effects
fMRI / neuroimaging n < 30 Voodoo correlations territory
Genome-wide association n < 1000 Won't survive multiple-testing correction
Survey research n < 200 Confidence intervals will be huge
ML benchmarks single seed reported Variance unknown — high risk of cherry-picking

These are smells, not failures. A well-controlled n=8 mouse experiment can be informative; a sloppy n=80 one isn't.

When the paper disagrees with the consensus

Two questions to ask:

  1. Did the author engage with the prior consensus? If they ignore it, they're either revolutionary or ignorant. Both are possible; the methodology section will tell you which.
  2. Did they explain the discrepancy? If they say "X claimed Y; we find Z because of W methodological choice," that's evidence. If they don't mention X at all, they may not know.

A reasonable report can include a heterodox paper, but the surrounding text should acknowledge that it's heterodox.

Quick decision tree for Phase 3

Have full text? ──no──> mark depth=shallow, fall back to abstract
       │
       yes
       │
Sample size sane? ──no──> mark "low_power" in evidence.limitations
       │
       yes
       │
Pre-registered or replicated? ──yes──> evidence weight = high
       │
       no
       │
Single lab, novel finding, hot field? ──yes──> evidence weight = medium
       │                                       (note as "needs replication")
       no
       │
Established methodology, multiple corroborating papers in corpus?
   ──yes──> evidence weight = high
   ──no──>  evidence weight = medium

Source: SKILL.md on GitHub

1 warning4mo3 checks · Risk SAFE
  • Gen Agent Trust Hub4mo

    This skill is a robust academic research tool that automates literature reviews across multiple scholarly databases. It uses strict input validation for paper identifiers and handles sensitive API keys through environment variables. While the skill executes a local script from a companion 'paper-fetch' skill to download PDFs, this is a standard design pattern within the author's ecosystem. It also processes external PDF text, which carries a minor risk of indirect prompt injection common to all document-reading agents.

  • Socket4mo

    No alerts

  • Snyk4mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 72ee04b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 11 hours ago.

Activeupdated 3 weeks ago
homepage
https://github.com/Agents365-ai/365-skills
platforms
[macos, linux, windows]
Other metadata
compatibility
Requires Python 3.9+ with httpx and pypdf (see requirements.txt). Optional: `pip install docling` to enable layout-aware markdown PDF extraction (`extract_pdf.py --engine docling`); auto-used as a fallback for scanned/sparse PDFs. Works offline-first (no MCP required) but enriches with Semantic Scholar / Brave MCP tools when available.
metadata
{
  "openclaw": {
    "requires": {
      "bins": [
        "python3"
      ]
    },
    "emoji": "🔬"
  },
  "hermes": {
    "tags": [
      "research",
      "literature-review",
      "academic",
      "papers",
      "citations",
      "survey"
    ],
    "category": "research"
  },
  "pimo": {
    "tags": [
      "research",
      "literature-review",
      "academic"
    ],
    "category": "research"
  },
  "author": "Agents365-ai",
  "version": "0.17.0"
}

README badge

README badge for agents365-ai/365-skills/scholar-deep-research