All skills
bitwarden avatar

/classifying-review-findings

@c9e5b81 official
by bitwardenbitwarden/ai-plugins155 stars
20

Use this skill when categorizing code review findings into severity levels. Apply when determining which emoji and label to use for PR comments, deciding if an issue should be flagged at all, or classifying findings as CRITICAL, IMPORTANT, DEBT, SUGGESTED, or QUESTION.

Use this Skill: https://skilld.dev/gh/bitwarden/ai-plugins/classifying-review-findings

This session only. Nothing lands on disk.

SKILL.md

≈75 tokens always: the name and description. ≈598 when used: this file.

Classifying Review Findings

Severity Categories

Emoji Category Criteria
❌ CRITICAL Will break, crash, expose data, or violate requirements
⚠️ IMPORTANT Missing error handling, unhandled edge cases, could cause bugs
♻️ DEBT Duplicates patterns, violates conventions, needs rework within 6 months
🎨 SUGGESTED Measurably improves security, reduces complexity by 3+, eliminates bug classes
❓ QUESTION Requires human knowledge - unclear requirements, intent, or system conflicts

ALWAYS use hybrid emoji + text format for each finding (if multiple severities apply, use the most severe: ❌ > ⚠️ > ♻️ > 🎨 > ❓):

Before Classifying

Verify ALL three:

  1. Can you trace the execution path showing incorrect behavior?
  2. Is this handled elsewhere (error boundaries, middleware, validators)?
  3. Are you certain about framework behavior and language semantics?

If any answer is "no" or "unsure" → DO NOT classify as a finding.

Not Valid Findings (Reject)

  • Praise ("great implementation")
  • Vague suggestions ("could be simpler")
  • Style preferences without enforced standard
  • Naming nitpicks unless actively misleading
  • PR metadata issues (title, description, test plan) - handled by summary skill, not classified here
  • Renovate/Dependabot minor/patch updates to existing dependencies with passing CI — these are routine Stage 5 monitoring, not reviewable findings

Suggested Improvements (🎨) Criteria

Only suggest improvements that provide measurable value:

  1. Security gain - Eliminates entire vulnerability class (SQL injection, XSS, etc.)
  2. Complexity reduction - Reduces cyclomatic complexity by 3+, eliminates nesting level
  3. Bug prevention - Makes entire category of bugs impossible (type safety, null safety)
  4. Performance gain - Reduces O(n²) to O(n), eliminates N+1 queries (provide evidence)

Provide concrete metrics:

  • ❌ "This could be simpler"
  • ✅ "This has cyclomatic complexity of 12; extracting validation logic would reduce to 6"

If you can't measure the improvement, don't suggest it.

Source: SKILL.md on GitHub

1 warning14d4 checks · Risk SAFE
  • Gen Agent Trust Hub14d

    No security issues detected. The skill provides guidelines for categorizing code review findings and does not contain any code, tool requests, or network operations.

  • Socket14d

    No alerts

  • Snyk14d

    Risk: LOW · No issues

  • Runlayer7mo

    1/1 file flagged

Signed by skilld at c9e5b81. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 18 hours ago.

Activeupdated 6 months ago
  • Security
  • code-review
  • severity-classification
  • pr-comments
  • quality-gates
  • bug-detection
  • technical-debt
  • workflow

README badge

README badge for bitwarden/ai-plugins/classifying-review-findings

Classifies code review findings into five severity levels (CRITICAL, IMPORTANT, DEBT, SUGGESTED, QUESTION) using emoji markers, with strict criteria for execution-path tracing, error-handling verification, and measurable improvement thresholds. Targets PR comment workflows where AI agents need to distinguish blocking issues from style preferences and route findings to appropriate remediation timelines.

Generated from the current SKILL.md.

What severity categories does this skill define?
Five categories: CRITICAL (breaks, crashes, exposes data), IMPORTANT (missing error handling, unhandled edge cases), DEBT (duplicates patterns, violates conventions), SUGGESTED (measurably improves security or reduces complexity by 3+), and QUESTION (requires human knowledge).
When should I reject a finding instead of classifying it?
Reject praise, vague suggestions, style preferences without enforced standards, naming nitpicks unless misleading, PR metadata issues, and routine Renovate/Dependabot updates with passing CI.
What must I verify before classifying a finding?
You must trace the execution path showing incorrect behavior, confirm the issue is not handled elsewhere (error boundaries, middleware, validators), and be certain about framework behavior and language semantics. If any answer is no or unsure, do not classify it as a finding.
How should I decide between multiple applicable severity levels?
Use the most severe category that applies, in order: CRITICAL > IMPORTANT > DEBT > SUGGESTED > QUESTION.
What makes a valid SUGGESTED improvement?
It must provide measurable value: eliminate an entire vulnerability class, reduce cyclomatic complexity by 3+, make a category of bugs impossible, or provide concrete performance gains. Do not suggest improvements you cannot measure.

Generated from the current SKILL.md. These answers refresh after source changes.