Capability
Other metadata
- compatibility
- Requires Phoenix server. Python skills need phoenix and openai packages; TypeScript skills need @arizeai/phoenix-client.
- metadata
{
"author": "oss@arize.com",
"version": "1.0.0",
"languages": "Python, TypeScript"
}
What it does
Builds and runs evaluators for LLM applications using Phoenix, supporting code-based checks, LLM-as-judge approaches, and human validation workflows. Includes pre-built evaluators for RAG systems, error analysis, experiment tracking, and production monitoring across Python and TypeScript.
Generated from the current SKILL.md.
Frequently asked
Does this skill work with Python and TypeScript?
Yes. Python skills require the phoenix and openai packages; TypeScript skills require @arizeai/phoenix-client. Both require a running Phoenix server.
Can I build custom evaluators or only use pre-built ones?
You can build both code-based evaluators (deterministic logic) and LLM-based evaluators (using an LLM as a judge), with templates and validation support for both.
Does this cover RAG system evaluation?
Yes. The skill includes a dedicated RAG evaluators workflow covering retrieval quality and answer faithfulness.
Can I validate that my evaluators are accurate?
Yes. The skill provides validation references to test your evaluators' accuracy (TPR/TNR targets) against human labels.
What should I do before building evaluators?
The skill recommends starting with error analysis and tracing to observe actual failures, then categorizing them before automating evaluation.
Generated from the current SKILL.md. These answers refresh after source changes.