All skills
launchdarkly avatar

/online-evals

@ef54971 official

Attach judges to config variations for automatic LLM-as-a-judge evaluation. Create custom judges, configure sampling rates, and monitor quality scores.

Use this Skill: https://skilld.dev/gh/launchdarkly/agent-skills/online-evals

This session only. Nothing lands on disk.

README.md

≈323 tokens on demand. Your agent reads this file only when SKILL.md points to it.

LaunchDarkly Config Online Evaluations Skill

An Agent Skill for attaching judges to config variations for automatic LLM-as-a-judge evaluation.

Overview

This skill teaches agents how to:

  • Create custom judge configs with evaluation criteria
  • Attach judges to config variations via API
  • Configure sampling rates for cost control
  • Monitor evaluation results in the dashboard

Installation (Local)

Copy skills/agentcontrol/online-evals/ into your agent client's skills path.

Prerequisites

  • LaunchDarkly API access token with ai-configs:write permission
  • Existing config with variations (use configs-create skill)
  • For custom judges: understanding of LLM-as-a-judge methodology

Usage

Attach security and API contract judges to the model-selector config at 100% sampling
Create a custom judge that checks for scope creep in code changes

Structure

online-evals/
├── SKILL.md
└── README.md

Related

License

Apache-2.0

Source: SKILL.md on GitHub

No alerts2d3 checks · Risk SAFE
  • Gen Agent Trust Hub2d

    This skill provides instructions and code for integrating LaunchDarkly's LLM-as-a-judge evaluation features into AI agents. It uses official LaunchDarkly SDKs and interacts exclusively with the vendor's API endpoints.

  • Socket2d

    No alerts

  • Snyk2d

    Risk: LOW · No issues

Signed by skilld at ef54971. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 3 days ago
metadata
{
  "author": "launchdarkly",
  "version": "0.1.0"
}
Other metadata
compatibility
Requires LaunchDarkly API access token with ai-configs:write permission. SDK versions Python v0.20.0+ or Node.js v0.20.0+ for automatic metric recording and the consolidated `track_judge_result` / `trackJudgeResult` API.
  • Python
  • launchdarkly
  • evaluation
  • llm-as-a-judge
  • quality-scoring
  • ai-config
  • node.js
  • sdk

README badge

README badge for launchdarkly/agent-skills/online-evals

Attaches LLM-as-a-judge evaluators to LaunchDarkly config variations for automatic quality scoring. Supports built-in judges (accuracy, relevance, toxicity) and custom judges via the LaunchDarkly API, with configurable sampling rates and automatic metric recording in Python v0.20.0+ and Node.js v0.20.0+ SDKs.

Generated from the current SKILL.md.

What SDK versions are required for automatic judge evaluation?
Python v0.20.0+ or Node.js v0.20.0+ for automatic metric recording and the consolidated track_judge_result / trackJudgeResult API.
Can judges be attached to agent mode configs?
No. Judges can only be attached to completion mode configs in the UI. For agent mode or custom pipelines, use programmatic evaluation via the SDK instead.
What are the built-in judges LaunchDarkly provides?
Three pre-configured judges: Accuracy (measures correctness and grounding), Relevance (measures how well it addresses the request), and Toxicity (measures harmful phrasing, where lower scores are safer).
Can multiple judges with the same metric key be attached to one variation?
No. You cannot attach multiple judges with the same metric key to a single variation.
What API permissions are required?
Your LaunchDarkly API access token must have ai-configs:write permission.

Generated from the current SKILL.md. These answers refresh after source changes.