All skills
github avatar

/arize-evaluator

@4e136f3 official
by githubgithub/awesome-copilot40k stars
5,040

Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and continuous monitoring. Use when the user mentions create evaluator, LLM judge, hallucination, faithfulness, correctness, relevance, run eval, score spans, score experiment, trigger-run, column mapping, continuous monitoring, or improve evaluator prompt.

Use this Skill: https://skilld.dev/gh/github/awesome-copilot/arize-evaluator

This session only. Nothing lands on disk.

referencesax-setup.md

≈383 tokens on demand. Your agent reads this file only when SKILL.md points to it.

ax CLI — Troubleshooting

Consult this only when an ax command fails. Do NOT run these checks proactively.

Check version first

If ax is installed (not command not found), always run ax --version before investigating further. The version must be 0.14.0 or higher — many errors are caused by an outdated install. If the version is too old, see Version too old below.

ax: command not found

macOS/Linux:

  1. Check common locations: ~/.local/bin/ax, ~/Library/Python/*/bin/ax
  2. Install: uv tool install arize-ax-cli (preferred), pipx install arize-ax-cli, or pip install arize-ax-cli
  3. Add to PATH if needed: export PATH="$HOME/.local/bin:$PATH"

Windows (PowerShell):

  1. Check: Get-Command ax or where.exe ax
  2. Common locations: %APPDATA%\Python\Scripts\ax.exe, %LOCALAPPDATA%\Programs\Python\Python*\Scripts\ax.exe
  3. Install: pip install arize-ax-cli
  4. Add to PATH: $env:PATH = "$env:APPDATA\Python\Scripts;$env:PATH"

Version too old (below 0.14.0)

Upgrade: uv tool install --force --reinstall arize-ax-cli, pipx upgrade arize-ax-cli, or pip install --upgrade arize-ax-cli

SSL/certificate error

  • macOS: export SSL_CERT_FILE=/etc/ssl/cert.pem
  • Linux: export SSL_CERT_FILE=/etc/ssl/certs/ca-certificates.crt
  • Fallback: export SSL_CERT_FILE=$(python -c "import certifi; print(certifi.where())")

Subcommand not recognized

Upgrade ax (see above) or use the closest available alternative.

Still failing

Stop and ask the user for help.

Source: SKILL.md on GitHub

1 warning16d4 checks · Risk MEDIUM
  • Gen Agent Trust Hub16d

    This skill interacts with the Arize 'ax' CLI to manage LLM evaluation workflows. It is generally safe but includes persistence mechanisms to save configuration to shell profiles and uses runtime Python execution for data formatting. It also creates a potential surface for indirect prompt injection as it processes external trace and experiment data using LLM-as-judge templates.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • ZeroLeaks5mo

    1 finding · Score: 86/100

Signed by skilld at 4e136f3. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 5 months ago
metadata
{
  "author": "arize",
  "version": "1.0"
}
compatibility
Requires the ax CLI and a configured Arize profile with an AI integration.
  • arize
  • llm-as-judge
  • evaluation
  • monitoring
  • span-scoring
  • prompt-engineering
  • ai-integration
  • experiment-analysis

README badge

README badge for github/awesome-copilot/arize-evaluator

Manages LLM-as-judge evaluators on Arize, including creating evaluators with custom prompts, running evaluations against spans or experiments, column mapping, and continuous monitoring of new traces. Requires the ax CLI and an Arize profile with LLM provider credentials configured via AI integrations.

Generated from the current SKILL.md.

What LLM providers does this skill support?
OpenAI, Anthropic, Azure, Bedrock, Vertex, Gemini, NVIDIA NIM, and custom providers. Credentials are managed via the arize-ai-provider-integration skill or the ax ai-integrations command.
Can I run evaluators on experiment data or only live project traces?
Both. Tasks can attach evaluators to a project for continuous or one-time scoring of live spans, or to a dataset/experiment for backfill evaluation of experiment runs.
What happens if an evaluation fails or is cancelled?
The skill reports the failure and explains what went wrong — it never fabricates or estimates evaluation scores. You should fix the identified issue and retry, or verify integration credentials with ax ai-integrations list.
Does this skill handle evaluator versioning?
Yes. Creating a new version with ax evaluators create-version preserves the old version as immutable while the new version becomes active. This lets you update the prompt or model without affecting past evaluation runs.
What is the difference between span, trace, and session granularity?
Span evaluates individual LLM calls; trace evaluates all spans in a call chain grouped by trace_id; session evaluates all traces in a conversation grouped by session.id. The {conversation} template variable is only available at session granularity.

Generated from the current SKILL.md. These answers refresh after source changes.