All skills
github avatar

/diagnose

@7f59e01 official
by githubgithub/awesome-copilot40k stars
5,040

Perform a systematic diagnostic scan of an AI workflow across 5 quality dimensions — prompt quality, context efficiency, tool health, architecture fitness, and safety — producing a scored report with prioritized remediation actions.

Use this Skill: https://skilld.dev/gh/github/awesome-copilot/diagnose

This session only. Nothing lands on disk.

SKILL.md

≈62 tokens always: the name and description. ≈1k when used: this file.

AI Workflow Diagnostics

You are a systematic AI workflow auditor. Perform a diagnostic scan across 5 dimensions. For each dimension, score 1–5 and provide specific findings.

Dimension 1: Prompt Quality (1–5)

Evaluate:

  • Structure (role, context, instructions, output zones)
  • Output schema definition (explicit vs. implicit)
  • Instruction clarity (specific vs. vague)
  • Edge case handling (addressed vs. ignored)
  • Anti-patterns (wall of text, contradictions, implicit format)

Dimension 2: Context Efficiency (1–5)

Evaluate:

  • Context budget allocation (planned vs. ad-hoc)
  • Attention gradient awareness (critical info at start/end)
  • Context window utilization (efficient vs. wasteful)
  • State management (explicit vs. implicit)
  • Memory strategy (appropriate for conversation length)

Dimension 3: Tool Health (1–5)

Evaluate:

  • Tool count (3–7 ideal, 13+ problematic)
  • Description quality (specific vs. vague)
  • Error handling (graceful vs. none)
  • Schema completeness (input/output/error defined)
  • Idempotency (safe to retry vs. side-effect prone)
  • Scope attribution: Distinguish project-configured tools (custom scripts, project MCP servers) from agent-level tools (built-in IDE tools, global MCP servers). Only flag tool overhead for tools the project can actually control.

Dimension 4: Architecture Fitness (1–5)

Evaluate:

  • Topology appropriateness (single vs. multi-agent justified)
  • Agent boundaries (clear vs. overlapping)
  • Handoff protocols (structured vs. ad-hoc)
  • Observability (decisions logged vs. black box)
  • Cost awareness (budgeted vs. unbounded)

Dimension 5: Safety & Reliability (1–5)

Evaluate:

  • Input validation (present vs. absent)
  • Output filtering (PII, content policy) — scope contextually: data between a user's own frontend and backend is lower risk than data exposed to external services
  • Cost controls (ceilings set vs. unbounded)
  • Error recovery (fallbacks vs. crash)
  • Evaluation strategy (golden tests vs. "it seems to work")

Diagnostic Report Format

╔══════════════════════════════════════╗
║          WORKFLOW DIAGNOSTIC        ║
╠══════════════════════════════════════╣
║ Prompt Quality      ████░  4/5      ║
║ Context Efficiency   ███░░  3/5      ║
║ Tool Health          ██░░░  2/5      ║
║ Architecture         ████░  4/5      ║
║ Safety & Reliability ██░░░  2/5      ║
╠══════════════════════════════════════╣
║ Overall Score:       15/25           ║
╚══════════════════════════════════════╝

CRITICAL FINDINGS:
1. [Most severe issue — immediate action needed]
2. [Second most severe]
3. [Third]

RECOMMENDED ACTIONS:
1. [Specific remediation for finding #1]
2. [Specific remediation for finding #2]
3. [Specific remediation for finding #3]

Scoring Guide

Score Meaning Recommended Action
5 Production-excellent No action needed
4 Good with minor gaps Polish prompt clarity or output schema
3 Functional but risky Add error handling or reduce complexity
2 Significant issues Immediate attention — add retries/guards
1 Broken or missing Rebuild from scratch with clear structure

Usage

Invoke this skill when you want to:

  • Find hidden problems before a workflow goes to production
  • Audit an existing agent for quality and reliability
  • Get a prioritized remediation plan with concrete next steps
  • Health-check a workflow after significant changes

Provide the workflow description, prompt text, tool list, or agent configuration as context. The more detail you provide, the more precise the findings.

Source: SKILL.md on GitHub

No alerts14d3 checks · Risk SAFE
  • Gen Agent Trust Hub14d

    This skill provides a systematic diagnostic framework for auditing AI workflows. It includes evaluation criteria for prompt quality, context efficiency, tool health, architecture, and safety. The analysis found no security concerns, as the skill is purely informational and contains no executable code or dangerous instructions.

  • Socket14d

    No alerts

  • Snyk14d

    Risk: LOW · No issues

Signed by skilld at 7f59e01. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 19 hours ago.

Activeupdated 5 months ago
  • diagnostic
  • audit
  • prompt-quality
  • ai-workflow
  • context-efficiency
  • tool-health
  • safety
  • reliability
  • evaluation

README badge

README badge for github/awesome-copilot/diagnose

Performs a systematic audit of an AI workflow across five dimensions—prompt quality, context efficiency, tool health, architecture fitness, and safety—and returns a scored diagnostic report with prioritized remediation actions. Use this skill to identify hidden problems in agent prompts, tool configurations, and handoff protocols before deploying to production.

Generated from the current SKILL.md.

What should I provide to get an accurate diagnostic?
The skill works best with the full workflow context: the system prompt or role definition, tool list with descriptions, any multi-agent handoff logic, and the conversation or use case you're diagnosing. More detail yields more precise findings.
Does this evaluate Claude-specific features like vision or artifacts?
No. The skill diagnoses workflows across 5 generic dimensions (prompt quality, context efficiency, tool health, architecture, safety) and does not assume a particular AI model or platform.
Can this flag tools that come with my IDE or agent by default?
The skill distinguishes between project-configured tools (custom scripts, project MCP servers) and agent-level built-in tools; it only flags tool overhead for tools your project can actually control.
What does the final score mean?
Scores range 1–5 per dimension, summed to a 0–25 overall score. A 4 or 5 on any dimension means no immediate action needed; 2 or below requires urgent attention and specific remediation steps.
Does this skill help optimize token usage?
Yes. The Context Efficiency dimension evaluates context budget allocation, attention gradient awareness, and window utilization, and returns specific feedback on how to reduce wasteful token consumption.

Generated from the current SKILL.md. These answers refresh after source changes.