All skills
imbad0202 avatar

/academic-paper-reviewer

@a3f6569

Multi-perspective academic paper review with dynamic reviewer personas. Runs a 5-seat, role-separated review panel (Journal-Fit Reviewer + 3 peer-review roles + Devil's Advocate) with field-specific expertise; role separation is not a claim of independent error processes. Supports full review, re-review (verification), quick assessment, methodology focus, Socratic guided, and calibration modes. Triggers on: review paper, peer review, manuscript review, referee report, review my paper, critique paper, simulate review, editorial review, calibrate reviewer, reviewer calibration, measure reviewer accuracy, 審查論文, 論文審查, 模擬審查, 同儕審查, 幫我審這篇, 以審查人角度評估, 審查者校準, 논문 심사, 동료 심사, 모의 심사, 심사자 관점에서 평가, 심사자 보정, revisar artículo, revisión entre pares, revisión de manuscrito, informe de árbitro, revisa mi artículo, criticar artículo, simular revisión, revisión editorial, calibrar revisor.

Use this Skill: https://skilld.dev/gh/imbad0202/academic-research-skills/academic-paper-reviewer

This session only. Nothing lands on disk.

examplessubclaim_decomposition_example.md

≈2.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Example: Sub-Claim Decomposition in Editorial Synthesis (§F.3.2 partial-evidence trap)

This example shows the editorial synthesizer's Step 1 → Step 2 → Step 5 flow on a compound weakness — the case the §F.3.2 partial-evidence trap exists to catch — and contrasts the old bundle-level aggregation (which buries a minority sub-claim) with the sub-claim decomposition introduced for #214.

It is a focused illustration of one weakness bundle, not a full Phase 0–2 walkthrough (see hei_paper_review_example.md for that).

In the machine-shaped excerpts below, the stable serialized reviewer ID EIC maps to the public display role Journal-Fit Reviewer.


The input: what the 4 reviewers said about one weakness area

The paper reports a multi-site mixed-effects analysis. Three reviewers raised concerns that look like "one statistics weakness" but are actually two distinct claims:

  • R1 (Methodology), Severity: Major, Confidence 5 (anchors: absence: §3 Methods — expected the random-effects grouping factor; checked §3 Methods, §4 Results / table: Table 3 standardized betas vs Table 4 raw coefficients, no key; competence: core expertise — multi-site mixed-effects design): "Two problems. First, the statistical reporting is inconsistent — Table 3 gives standardized betas, Table 4 raw coefficients, with no key. Second, the mixed-model random-effects grouping is never specified: are sites or participants the grouping factor?"
  • R2 (Domain), Severity: Major, Confidence 4 (anchor: absence: §3 Methods — expected the grouping variable; checked §3 Methods; competence: core expertise — quantitative methods in the domain): "I agree the random-effects structure is unclear — the paper doesn't say what the grouping variable is."
  • R3 (Cross-disciplinary), Severity: Major, Confidence 3 (anchor: table: Tables 3–4 — units shift between tables, no key; competence: adjacent — reporting clarity, not mixed-model design): "The reporting tables were hard to follow; the units shift between tables."
  • EIC, Confidence 4: Did not comment on the statistics specifically.

Atomicity note (#574 A2/A3): R1's card bundles two sub-claims into ONE weakness — a fully-conformant current-format card would have emitted them as two separate atomic weaknesses, each with its own single evidence_anchor object (Schema 6). Step 1b decomposition is the synthesizer's guard for bundles that still arrive; the two parenthetical anchors above are the bundle's internal evidence, and the transport rows below carry each sub-claim's own anchor.


❌ Old behavior — bundle-level aggregation (the trap)

The Step 1 inventory collapses this into a single Key Weaknesses cell ("statistics problems"), and Step 2 reaches one consensus verdict on the bundle:

[CONSENSUS-3] Statistics problems — R1, R2, R3 agree; EIC silent. Author must fix.

What went wrong: the bundle holds two sub-claims with different support levels, and the verdict averages them into one. "Random-effects grouping unspecified" was raised by R1 and corroborated by R2 — a genuine 2-reviewer consensus. "Inconsistent reporting units" was raised by R1 and R3 but is a different defect. Folding both into one "[CONSENSUS-3]" item produces a single roadmap line that lets the author treat the whole thing as one fix, and the precise grouping-factor defect can be lost inside a vague "clean up the statistics" instruction. Partial support read as full resolution.


✅ New behavior — Step 1b sub-claim inventory, then per-sub-claim consensus

Step 1b — Weakness Sub-Claim Inventory

sub_claim_id parent_weakness reviewer_id position evidence_pointer severity confidence
SC-1 Statistics R1 raised absence: §3 Methods — expected the random-effects grouping factor; checked §3, §4 major 5
SC-1 Statistics R2 corroborated absence: §3 Methods — expected the grouping variable; checked §3 Methods major 4
SC-1 Statistics R3 not-mentioned — — —
SC-1 Statistics EIC not-mentioned — — —
SC-2 Statistics R1 raised table: Table 3 standardized betas vs Table 4 raw coefficients, no key major 5
SC-2 Statistics R3 corroborated table: Tables 3–4 — units shift between tables, no key major 3
SC-2 Statistics R2 not-mentioned — — —
SC-2 Statistics EIC not-mentioned — — —

Two atomic sub-claims fall out of the one "statistics" bundle. not-mentioned is recorded as silence, not opposition. Note the severity column: both R1 rows carry the SAME transported severity — they decompose from R1's one "Statistics" weakness, whose single per-finding Severity is inherited by every sub-claim split from it (#574 A3); a severity difference between sub-claims requires different parent weaknesses, never re-rating.

Step 2 — consensus per sub_claim (denominator = 4 non-DA reviewers)

  • SC-1 (random-effects grouping unspecified): agree = 2 (R1 raised, R2 corroborated), conflict = 0, silent = 2 (R3, EIC). Two agree, none conflict → a corroborated finding (below the CONSENSUS-3 bar of 3/4), action-bearing because two anchored reviewer rows identify the same major methods defect. Confidence remains uncertainty/scope metadata and does not prioritize the finding. Not a CONSENSUS-3/4 label; not a SPLIT.
  • SC-2 (inconsistent reporting units): agree = 2 (R1 raised, R3 corroborated), conflict = 0, silent = 2 (R2, EIC) → a corroborated finding because two exact table anchors identify the same reporting defect. Confidence remains uncertainty/scope metadata and does not prioritize the finding. Not a SPLIT.

Note the threshold discipline: even though every reviewer who spoke agreed, neither sub-claim is promoted to CONSENSUS-4 — the consensus denominator is the four non-DA scoring reviewers, not the 2 who spoke, so 2/4 is a corroborated finding, not unanimity. This is exactly the mislabel the absolute-count rule prevents.

Neither sub-claim triggers Journal-Fit Reviewer arbitration — there is no disputed position anywhere, so finer granularity did not manufacture arbitration load. (Had R2 instead written "the grouping is clearly by site, this is a non-issue," SC-1 would carry a disputed position against R1's raised, and then it would be a genuine SPLIT requiring Journal-Fit Reviewer arbitration.)

Step 5 — Revision Roadmap (sub-claim-keyed)

Instead of one blurred "fix the statistics" line, two separately-prioritized, traceable items:

Priority 1 — Structural (Must Fix)

  • [SC-1] Specify the mixed-model random-effects grouping factor (sites vs participants) and re-state the model. Severity: major (transported) | Anchor: absence: §3 Methods — expected the random-effects grouping factor; checked §3 Methods, §4 Results (transported) | Raised R1 (conf 5), corroborated R2 (conf 4).

Priority 2 — Content Supplementation (Should Fix)

  • [SC-2] Standardize coefficient reporting across Tables 3–4 (one convention + a key). Severity: major (transported) | Anchor: table: Table 3 standardized betas vs Table 4 raw coefficients, no key (transported) | Raised R1 (conf 5), corroborated R3 (conf 3).

Each item carries its sub_claim_id, traces back to the Step 1b inventory, and flows forward into academic-paper revision mode unchanged in format. The minority-risk sub-claim (the grouping factor) is now its own Priority-1 item rather than buried in a bundle.


What this example demonstrates (acceptance for #214)

  • Step 1b inventory with sub_claim as the primary key, alongside the retained Step 1a summary matrix.
  • Per-sub-claim consensus in Step 2, with not-mentioned correctly excluded from both consensus and SPLIT counting.
  • SPLIT bound respected — no arbitration is triggered absent an explicit conflicting position (the parenthetical shows what would trigger one).
  • Revision Roadmap at sub-claim granularity — the compound weakness yields two correctly-prioritized, traceable items instead of one.
  • DA-CRITICAL flow and the v3.6.2 sprint-contract arithmetic path are untouched (neither appears in this general-protocol example).

Source: SKILL.md on GitHub

1 warning6d5 checks · Risk SAFE
  • Gen Agent Trust Hub6d

    The skill is a multi-agent framework for academic paper review. It is well-architected with significant security defenses against prompt injection from the manuscripts it processes. The primary risk is the large attack surface provided by untrusted input data, though this is mitigated by explicit boundary instructions.

  • Socket6d

    No alerts

  • Snyk6d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    18 files scanned · No issues

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at a3f6569. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated last week
Other metadata
metadata
{
  "version": "1.11.1",
  "last_updated": "2026-08-15",
  "status": "active",
  "data_access_level": "raw",
  "task_type": "open-ended",
  "related_skills": [
    "academic-paper",
    "academic-pipeline"
  ]
}

README badge

README badge for imbad0202/academic-research-skills/academic-paper-reviewer