All skills
bitwarden avatar

/performing-multi-agent-code-review

@0a5f03a official
by bitwardenbitwarden/ai-plugins155 stars
20

Perform a rigorous, multi-agent code review with architecture-compliance, parallel quality/security analysis, finding validation, and severity audit. Use when the user asks for a structured, deep, thorough, multi-pass, or multi-agent code review — or a review that includes architecture/pattern compliance, confidence-scored findings, or a severity audit. Use when the user asks for a code review across a commit range, time window, or N most recent commits in a locally checked-out repo.

Use this Skill: https://skilld.dev/gh/bitwarden/ai-plugins/performing-multi-agent-code-review

This session only. Nothing lands on disk.

referencesevaluation-standards.md

≈760 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Evaluation Standards

Loaded by the orchestrator in Step 1. Severity Levels, Do Not Flag, and Confidence Scoring below are propagated verbatim into every Step 2–5 subagent prompt except Agent 5, which takes the carve-out subset in agent-5-skill-review.md. The Finding Shape schema lives in finding-shape.md and is propagated the same way.

Severity Levels

Every finding must be assigned one of the following. Do not guess — apply these definitions literally.

  • 🛑 Blocker — Will cause a production failure, data loss, or security breach.
  • ⚠️ Important — A real bug or significant risk that is likely to be hit in practice.
  • ♻️ Refactor — True technical debt being created that will cost more to maintain over time, even if it doesn't cause immediate problems. Must cite concrete evidence — duplication of an existing pattern, violation of a documented convention, or a measurable structural improvement. If the rationale can't be made concrete, it isn't a finding.

There is no "suggestion" or other lower tier. Findings that don't clear the Refactor bar are not findings.

Do Not Flag

The following are not valid findings under any tier. Subagents must not emit them, and Step 5 dismisses any that slip through.

  • Code style or quality concerns absent a rule explicitly documented in the repo's CLAUDE.md, README.md, or other project guidelines already loaded and forwarded by the orchestrator.
  • Subjective suggestions or improvements — "could be cleaner", "consider doing X", "this might be simpler".
  • Pedantic nit-picks a senior engineer would not raise in code review.
  • Issues a linter would catch.
  • Speculative issues that depend on specific inputs or runtime state without evidence those inputs occur in practice.
  • Pre-existing issues not introduced or worsened by this change. Exception: a CWE-1427 observation in a changed file stands whether or not the diff touched the line it sits on.

Confidence Scoring

Rate each potential finding on a 0–100 scale:

  • 0: Not confident — false positive or pre-existing issue.
  • 25: Somewhat confident — might be real, might be a false positive. Stylistic issues not called out in project guidelines land here.
  • 50: Moderately confident — real issue, but a nitpick, unlikely to hit in practice, or is a stylistic preference without project-rule backing.
  • 80: Highly confident — verified; very likely to hit in practice. Directly impacts functionality or violates a project guideline.
  • 100: Certain — evidence directly confirms it will happen frequently.

Only report findings with confidence ≥ 80. Findings rated 50–79 are dismissed silently; do not re-rate upward to clear the threshold.

Finding Shape

Every finding and every Step 4/5 return object follows the JSON schema in finding-shape.md. The main orchestrator loads that file in Step 1 and propagates its contents verbatim to every subagent except Agent 5, which returns prose the orchestrator translates.

Source: SKILL.md on GitHub

1 warning14d3 checks · Risk SAFE
  • Gen Agent Trust Hub14d

    The skill provides a rigorous multi-agent code review process with several built-in security safeguards. It implements a defensive boundary against indirect prompt injection by instructing subagents to treat instructions found within code changes as security findings. It also restricts tool usage (e.g., banning network tools for subagents) to prevent data exfiltration. All external functions utilized are internal or vendor-associated plugins.

  • Socket14d

    No alerts

  • Snyk14d

    Risk: MEDIUM · 1 issue

Signed by skilld at 0a5f03a. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last month
What it can do
Runs commands Reads files Edits files
All 12 allowed tools
Bash(gh pr diff:*)Bash(gh pr view:*)Bash(git diff:*)Bash(git status:*)Bash(git rev-parse:*)Bash(git log:*)ReadWriteGrepGlobSkillAskUserQuestion
Other metadata
argument-hint
[pr-number | commit-range] [--model <model>] [--model-analysis <model>] [--model-security <model>] [--model-validation <model>] [--model-audit <model>] [--output-dir <path>]

README badge

README badge for bitwarden/ai-plugins/performing-multi-agent-code-review

Orchestrates a multi-pass code review by spawning architecture, code-quality, bug-analysis, and security agents in parallel, each emitting confidence-scored findings against Bitwarden's zero-knowledge and threat-model principles. Use this skill when the user requests a deep, structured, or multi-agent review—or when reviewing commit ranges in a locally checked-out repo.

Generated from the current SKILL.md.

Does this skill work with local commits and PR branches?
Yes. The skill supports multiple modes: PR review (via `gh pr diff`), local HEAD changes, branch comparisons, and commit ranges. Pass the PR number or commit range as the first argument.
What happens if a prerequisite plugin is missing?
The skill aborts immediately with a clear error message identifying the missing plugin (bitwarden-tech-lead or bitwarden-security-engineer) and does not proceed.
Where does the review output go?
By default, reviews write to `${CLAUDE_PLUGIN_DATA}/code-reviews/` organized by project. You can override this with `--output-dir <path>` at invocation time.
Can I specify which model to use?
Yes. Pass `--model <model>` in the arguments; otherwise the skill defaults to the opus model.
Does this skill upload findings to GitHub?
No. All findings are written to a local markdown file only. The skill does not create pull request comments or push any data to GitHub.

Generated from the current SKILL.md. These answers refresh after source changes.