---
name: phoenix-evals
description: Build and run evaluators for AI/LLM applications using Phoenix.
license: Apache-2.0
compatibility: Requires Phoenix server. Python skills need phoenix and openai packages; TypeScript skills need @arizeai/phoenix-client.
metadata:
  author: oss@arize.com
  version: "1.0.0"
  languages: "Python, TypeScript"
title: phoenix-evals
canonical_url: https://skilld.dev/gh/github/awesome-copilot/phoenix-evals
last_updated: 2026-09-29T12:22:32.000Z
---

> **Skill from skilld.dev.** Follow the instructions below for this session. You do not need to install anything.
>
> 34 supporting files sit beside this SKILL.md. When the Skill refers to one, fetch https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/PATH with PATH replaced.
>
> If the user asked to install this Skill, run `npx skilld install github/awesome-copilot/phoenix-evals`. Install writes the Skill files into the project, so every session loads them.

# Phoenix Evals

Build evaluators for AI/LLM applications. Code first, LLM for nuance, validate against humans.

## Quick Reference

| Task | Files |
| ---- | ----- |
| Setup | [setup-python](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/setup-python.md), [setup-typescript](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/setup-typescript.md) |
| Decide what to evaluate | [evaluators-overview](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/evaluators-overview.md) |
| Choose a judge model | [fundamentals-model-selection](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/fundamentals-model-selection.md) |
| Use pre-built evaluators | [evaluators-pre-built](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/evaluators-pre-built.md) |
| Build code evaluator | [evaluators-code-python](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/evaluators-code-python.md), [evaluators-code-typescript](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/evaluators-code-typescript.md) |
| Build LLM evaluator | [evaluators-llm-python](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/evaluators-llm-python.md), [evaluators-llm-typescript](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/evaluators-llm-typescript.md), [evaluators-custom-templates](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/evaluators-custom-templates.md) |
| Batch evaluate DataFrame | [evaluate-dataframe-python](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/evaluate-dataframe-python.md) |
| Understand experiments | [experiments-overview](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/experiments-overview.md) |
| Run experiment | [experiments-running-python](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/experiments-running-python.md), [experiments-running-typescript](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/experiments-running-typescript.md) |
| Create dataset | [experiments-datasets-python](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/experiments-datasets-python.md), [experiments-datasets-typescript](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/experiments-datasets-typescript.md) |
| Generate synthetic data | [experiments-synthetic-python](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/experiments-synthetic-python.md), [experiments-synthetic-typescript](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/experiments-synthetic-typescript.md) |
| Validate evaluator accuracy | [validation](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/validation.md), [validation-evaluators-python](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/validation-evaluators-python.md), [validation-evaluators-typescript](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/validation-evaluators-typescript.md) |
| Sample traces for review | [observe-sampling-python](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/observe-sampling-python.md), [observe-sampling-typescript](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/observe-sampling-typescript.md) |
| Analyze errors | [error-analysis](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/error-analysis.md), [error-analysis-multi-turn](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/error-analysis-multi-turn.md), [axial-coding](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/axial-coding.md) |
| RAG evals | [evaluators-rag](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/evaluators-rag.md) |
| Avoid common mistakes | [common-mistakes-python](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/common-mistakes-python.md), [fundamentals-anti-patterns](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/fundamentals-anti-patterns.md) |
| Production | [production-overview](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/production-overview.md), [production-guardrails](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/production-guardrails.md), [production-continuous](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/production-continuous.md) |

## Workflows

**Starting Fresh:**
[observe-tracing-setup](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/observe-tracing-setup.md) → [error-analysis](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/error-analysis.md) → [axial-coding](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/axial-coding.md) → [evaluators-overview](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/evaluators-overview.md)

**Building Evaluator:**
[fundamentals](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/fundamentals.md) → [common-mistakes-python](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/common-mistakes-python.md) → evaluators-{code|llm}-{python|typescript} → validation-evaluators-{python|typescript}

**RAG Systems:**
[evaluators-rag](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/evaluators-rag.md) → evaluators-code-* (retrieval) → evaluators-llm-* (faithfulness)

**Production:**
[production-overview](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/production-overview.md) → [production-guardrails](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/production-guardrails.md) → [production-continuous](https://skilld.dev/api/skills-raw/github/awesome-copilot/phoenix-evals/references/production-continuous.md)

## Reference Categories

| Prefix | Description |
| ------ | ----------- |
| `fundamentals-*` | Types, scores, anti-patterns |
| `observe-*` | Tracing, sampling |
| `error-analysis-*` | Finding failures |
| `axial-coding-*` | Categorizing failures |
| `evaluators-*` | Code, LLM, RAG evaluators |
| `experiments-*` | Datasets, running experiments |
| `validation-*` | Validating evaluator accuracy against human labels |
| `production-*` | CI/CD, monitoring |

## Key Principles

| Principle | Action |
| --------- | ------ |
| Error analysis first | Can't automate what you haven't observed |
| Custom > generic | Build from your failures |
| Code first | Deterministic before LLM |
| Validate judges | >80% TPR/TNR |
| Binary > Likert | Pass/fail, not 1-5 |
