All skills
oaustegard avatar

/verifying-claims

@4b31e7d

Check that a document's claims about code are actually true by reading the prose, the code, and the tests and reporting (or fixing) where they disagree. Use whenever the user wants to verify a README, guide, spec, or docstring still matches the code; whenever they mention documentation drift, doc-code sync, "is this still accurate", stale docs, or keeping docs/tests/code consistent; before publishing or merging a docs change; or as a periodic doc-accuracy sweep. The agent reads the prose's meaning directly — there is no claim-comment DSL to maintain. Pairs with TDD — the test suite is the deterministic behavioral gate, this skill is the semantic prose-vs-reality review.

Use this Skill: https://skilld.dev/gh/oaustegard/claude-skills/verifying-claims

This session only. Nothing lands on disk.

referencesdrift-report-example.md

≈454 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Drift report — example

What a review produces. This is the report for a small parser package whose README.md claims parse(text) turns text into records and Reader(path).read() streams them, checked against the source and tests via gather_context.py.


Document: README.md · Sources: pkg/parser.py · Tests: tests/

Verdict Claim (prose) Reality
PASS parse(text) turns text into records parse(text, strict=False) exists; test_parse_empty exercises it
UNSUPPORTED Reader(path).read() streams records Reader.read(self, n=10) exists and matches, but test_reader_reads asserts True — it never calls read(). No test protects this claim.

Summary: 1 PASS, 1 UNSUPPORTED, 0 FAIL, 0 STALE.

Recommended actions:

  • The read() claim is accurate today but unprotected. Add a test that calls Reader(path).read() and asserts on its output, so a future change to read fails loudly instead of silently invalidating the README.

Notes on reading this report:

  • UNSUPPORTED is the interesting verdict. A dumb signature check would have marked read() green — the signature matches. Reading the test shows the claim rests on nothing. That gap is what an agent review adds over a declarative check, and it points at a missing test rather than a doc edit.
  • FAIL would mean the doc is wrong now (e.g., README says parse returns a dict but the code returns a list). Those get a prose fix.
  • STALE would mean the doc references something gone (e.g., a removed parse_strict function). Those get a prose fix or removal.
  • The report never claims the tests are correct — only that the prose agrees with code+tests as they stand. Bad tests upstream still produce a confident PASS.

Source: SKILL.md on GitHub

1 warning2mo3 checks · Risk SAFE
  • Gen Agent Trust Hub2mo

    The skill is well-implemented and uses safe techniques like AST parsing to inspect code without execution. However, it is susceptible to indirect prompt injection because it ingests untrusted documentation and source code into the agent's context without explicit security boundaries or sanitization. Mitigations include adding human review checkpoints and explicit instructions for the agent to ignore any commands found within the analyzed documentation.

  • Socket2mo

    No alerts

  • Snyk2mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 4b31e7d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago
metadata
{
  "version": "0.2.0"
}

README badge

README badge for oaustegard/claude-skills/verifying-claims