All skills
shipshitdev avatar

/advanced-evaluation

@400068d
by Ship Shitshipshitdev/library37 stars
3

Design and operate LLM-as-a-Judge evaluation systems using direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated quality assessment. Use when building LLM-as-judge systems, comparing model responses, calibrating rubrics, debugging inconsistent evaluations, or designing A/B tests for prompt or model changes.

Use this Skill: https://skilld.dev/gh/shipshitdev/library/advanced-evaluation

This session only. Nothing lands on disk.

referencesevaluation-pipeline.md

≈723 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Evaluation Pipeline Diagram

Visual layout of a production evaluation pipeline.

┌─────────────────────────────────────────────────┐
│                 Evaluation Pipeline              │
├─────────────────────────────────────────────────┤
│                                                   │
│  Input: Response + Prompt + Context               │
│           │                                       │
│           ▼                                       │
│  ┌─────────────────────┐                         │
│  │   Criteria Loader   │ ◄── Rubrics, weights    │
│  └──────────┬──────────┘                         │
│             │                                     │
│             ▼                                     │
│  ┌─────────────────────┐                         │
│  │   Primary Scorer    │ ◄── Direct or Pairwise  │
│  └──────────┬──────────┘                         │
│             │                                     │
│             ▼                                     │
│  ┌─────────────────────┐                         │
│  │   Bias Mitigation   │ ◄── Position swap, etc. │
│  └──────────┬──────────┘                         │
│             │                                     │
│             ▼                                     │
│  ┌─────────────────────┐                         │
│  │ Confidence Scoring  │ ◄── Calibration         │
│  └──────────┬──────────┘                         │
│             │                                     │
│             ▼                                     │
│  Output: Scores + Justifications + Confidence     │
│                                                   │
└─────────────────────────────────────────────────┘

Pipeline Stages

  1. Criteria Loader: Loads rubrics and criterion weights from configuration
  2. Primary Scorer: Applies direct scoring or pairwise comparison
  3. Bias Mitigation: Runs position swaps, length normalization, and other debiasing
  4. Confidence Scoring: Calibrates confidence based on position consistency and evidence strength

Source: SKILL.md on GitHub

1 warning16d4 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    No security issues or malicious patterns were detected. The skill contains standard documentation, guidelines, and benign example code for implementing LLM-as-a-Judge evaluation pipelines.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer7mo

    7/7 files flagged

Signed by skilld at 400068d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 3 months ago
Other metadata
metadata
{
  "version": "2.1.1",
  "source": "https://github.com/muratcankoylan/Agent-Skills-for-Context-Engineering/blob/main/skills/advanced-evaluation/SKILL.md",
  "upstream_repo": "muratcankoylan/Agent-Skills-for-Context-Engineering",
  "upstream_ref": "main",
  "upstream_commit": "25e1fa79a33f",
  "last_synced": "2026-06-13",
  "license": "MIT",
  "tags": "evaluation, llm-as-judge, quality, bias-mitigation"
}

README badge

README badge for shipshitdev/library/advanced-evaluation