All skills
shipshitdev avatar

/advanced-evaluation

@400068d
by Ship Shitshipshitdev/library37 stars
3

Design and operate LLM-as-a-Judge evaluation systems using direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated quality assessment. Use when building LLM-as-judge systems, comparing model responses, calibrating rubrics, debugging inconsistent evaluations, or designing A/B tests for prompt or model changes.

Use this Skill: https://skilld.dev/gh/shipshitdev/library/advanced-evaluation

This session only. Nothing lands on disk.

README.md

≈541 tokens on demand. Your agent reads this file only when SKILL.md points to it.

advanced-evaluation

Master LLM-as-a-Judge techniques — direct scoring, pairwise comparison, rubric generation, and bias mitigation (position, length, verbosity, authority).

Upstream

Derived from muratcankoylan/Agent-Skills-for-Context-Engineering (MIT).

Field Value
Source skills/advanced-evaluation/SKILL.md
Upstream ref main
Synced at commit 25e1fa79a33f
Last synced 2026-06-13
License MIT

Local modifications: Imported 2026-01-20 (this repo's commit ef42a98) from muratcankoylan/Agent-Skills-for-Context-Engineering at v1.0.0-era content (then pinned to creation commit 0b9a3b81bfea). Ported forward 2026-06-13 to upstream HEAD (commit 25e1fa79a33f); local body now tracks upstream v2.1.0 — carried full direct-scoring and pairwise prompt templates, the Metric Selection Framework table, three worked JSON examples, a 10-item Guidelines section, 8-entry Gotchas, a Scaling Evaluation section, the claim-advanced-evaluation-position-swap ID, and a fully rewritten scripts/evaluation_example.py. references/full-guide.md was renamed to references/evaluation-pipeline.md to match upstream. Reference to a sibling not vendored here (harness-engineering) was stripped; cross-links to tool-design (vendored) are retained. Local divergence: concrete vendor model names in references were genericized. To diff: compare the upstream path on main since commit 25e1fa79a33f; restructured 2026-07-10: long examples moved to references/.

Checking for upstream changes: when upstream has moved ahead of the synced marker above, diff skills/advanced-evaluation/SKILL.md on main since commit 25e1fa79a33f, port anything worth bringing home, then bump metadata.upstream_commit (or metadata.upstream_version) and metadata.last_synced in SKILL.md and this table.

Source: SKILL.md on GitHub

1 warning16d4 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    No security issues or malicious patterns were detected. The skill contains standard documentation, guidelines, and benign example code for implementing LLM-as-a-Judge evaluation pipelines.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer7mo

    7/7 files flagged

Signed by skilld at 400068d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 3 months ago
Other metadata
metadata
{
  "version": "2.1.1",
  "source": "https://github.com/muratcankoylan/Agent-Skills-for-Context-Engineering/blob/main/skills/advanced-evaluation/SKILL.md",
  "upstream_repo": "muratcankoylan/Agent-Skills-for-Context-Engineering",
  "upstream_ref": "main",
  "upstream_commit": "25e1fa79a33f",
  "last_synced": "2026-06-13",
  "license": "MIT",
  "tags": "evaluation, llm-as-judge, quality, bias-mitigation"
}

README badge

README badge for shipshitdev/library/advanced-evaluation