All skills
github avatar

/eval-driven-dev

@2860790 official
by githubgithub/awesome-copilot40k stars
5,040

Improve AI application with evaluation-driven development. Define eval criteria, instrument the application, build golden datasets, observe and evaluate application runs, analyze results, and produce a concrete action plan for improvements. ALWAYS USE THIS SKILL when the user asks to set up QA, add tests, add evals, evaluate, benchmark, fix wrong behaviors, improve quality, or do quality assurance for any Python project that calls an LLM model.

Use this Skill: https://skilld.dev/gh/github/awesome-copilot/eval-driven-dev

This session only. Nothing lands on disk.

references2c-capture-and-verify-trace.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Step 2c: Capture and verify a reference trace

Goal: Run the app through the Runnable, capture a trace, and verify that instrumentation and the Runnable are working correctly. The trace proves everything is wired up and provides the exact data shapes needed for dataset creation in Step 4.


Choose the trace input

The trace input determines what code paths are captured. A trivial input produces a trivial trace that misses the app's real behavior.

The input must reflect the "Realistic input characteristics" section, according to pixie_qa/00-project-analysis.md you've read in step 2b.

The input has two parts — understand the boundary between them:

  • User-provided parameters (you author): What a real user types or configures — prompts, queries, configuration flags, URLs, schema definitions. Write these to be representative of real usage.
  • World data (captured from production code, not fabricated): Content the app fetches from external sources during execution — database records, API responses, files, etc. Run the production code once to capture this data into the trace. Only resort to synthetic data generation when:
    • The user explicitly instructs you to use synthetic data, OR
    • Fetching from real sources is impractical (too many fetches, incurs real monetary cost, or takes unreasonably long — more than ~30 minutes)

Quick check before writing input: "Would a real user create this data, or would the app get it from somewhere else?" If the app gets it, let the production code run and capture it.

App type User provides (you author) World provides (you source)
Web scraper URL + prompt + schema definition The HTML page content
Research agent Research question + scope constraints Source documents, search results
Customer support bot Customer's spoken message Customer profile from CRM, conversation history from session store
Code review tool PR URL + review criteria The actual diff, file contents, CI results

Capture multiple traces

Capture at least 2 traces with different input characteristics before building the dataset:

  • Different complexity (simple case vs. complex case)
  • Different capabilities (see 00-project-analysis.md capability inventory)
  • Different edge conditions (missing optional data, unusually large input)

This calibration prevents dataset homogeneity — you see what the app actually does with varied inputs.


Run pixie trace

First, verify the app can be imported: python -c "from <module> import <class>". Catch missing packages before entering a trace-install-retry loop.

# Create a JSON file with input data
echo '{"user_message": "a realistic sample input"}' > pixie_qa/sample-input.json

uv run pixie trace --runnable pixie_qa/run_app.py:AppRunnable \
  --input pixie_qa/sample-input.json \
  --output pixie_qa/reference-trace.jsonl

The --input flag takes a file path to a JSON file (not inline JSON). The JSON keys become kwargs for the Pydantic model.

For additional traces:

uv run pixie trace --runnable pixie_qa/run_app.py:AppRunnable \
  --input pixie_qa/sample-input-complex.json \
  --output pixie_qa/trace-complex.jsonl

Verify the trace

Quick inspection

The trace JSONL contains one line per wrap() event and one line per LLM span:

{"type": "kwargs", "value": {"user_message": "What are your hours?"}}
{"type": "wrap", "name": "customer_profile", "purpose": "input", "data": {...}, ...}
{"type": "llm_span", "request_model": "gpt-4o", "input_messages": [...], ...}
{"type": "wrap", "name": "response", "purpose": "output", "data": "Our hours are...", ...}

Check that:

  • Expected wrap entries appear (one per wrap() call in the code)
  • At least one llm_span entry appears (confirms real LLM calls were made)
  • Missing entries indicate the execution path was different than expected — fix before continuing

Format and verify coverage

Run pixie format to see the data in dataset-entry format:

pixie format --input trace.jsonl --output dataset_entry.json

The output shows:

  • input_data: the exact keys/values for runnable arguments
  • eval_input: data from wrap(purpose="input") calls
  • eval_output: the actual app output (from wrap(purpose="output"))

For each eval criterion from pixie_qa/02-eval-criteria.md, verify the format output contains the data needed. If a data point is missing, go back to Step 2a and add the wrap() call.

Trace audit

Before proceeding to Step 3, audit every trace:

  1. World data check: For each wrap(purpose="input") field, is the data realistically complex? Compare against 00-project-analysis.md "Realistic input characteristics." If the analysis says inputs are 5KB–500KB and yours is under 5KB, it's not representative.

  2. LLM span check: Do llm_span entries appear? If not, the app's LLM calls didn't fire — the Runnable may be misconfigured or the LLM may be mocked/faked. Fix this before continuing.

  3. Complexity check: Does the trace exercise the hard problems from 00-project-analysis.md? If it only exercises the happy path, capture an additional trace with harder inputs.

If any check fails, go back and fix the input or Runnable, then re-capture.


Output

  • pixie_qa/reference-trace.jsonl — reference trace with all expected wrap events and LLM spans
  • Additional trace files for varied inputs

Source: SKILL.md on GitHub

3 warnings16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill facilitates evaluation-driven development for Python AI applications using the pixie-qa framework. It automates the setup of an evaluation pipeline, including package installation, instrumentation, and result analysis. Security considerations include the installation of third-party packages, a self-updating mechanism, and the processing of potentially untrusted project specifications.

  • Socket16d

    1 alert: gptAnomaly

  • Snyk16d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    1/2 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 2860790. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 5 months ago
compatibility
Python 3.10+
Other metadata
metadata
{
  "version": "0.8.4",
  "pixie-qa-version": ">=0.8.4,<0.9.0",
  "pixie-qa-source": "https://github.com/yiouli/pixie-qa/"
}
  • Python
  • Testing
  • llm
  • evaluation
  • pixie-qa
  • quality-assurance
  • benchmarking
  • agent
  • observability

README badge

README badge for github/awesome-copilot/eval-driven-dev

Defines evaluation criteria, instruments a Python LLM application to inject controlled test data, builds datasets, and runs automated quality checks via the pixie-qa framework. Targets Python projects that call LLMs and need end-to-end testing of application code (routing, prompt assembly, response formatting) with real LLM calls — not mocking — scored by evaluators instead of deterministic assertions.

Generated from the current SKILL.md.

Does this skill work with any Python LLM application or only specific frameworks?
This skill works with any Python application that calls an LLM, regardless of framework. It instruments the app's data boundaries and runs the full application code end-to-end, exercising real routing, prompt assembly, and LLM calls.
Can I mock or stub the LLM during evaluation?
No. The skill requires real LLM calls — mocking or stubbing the LLM makes eval scores meaningless because you control both inputs and outputs. The app's own unit tests may mock the LLM; evals must not.
What is pixie-qa and do I need to install it separately?
pixie-qa is the Python package that powers the eval pipeline. The skill's setup.sh installs it automatically. If installation fails, the workflow cannot proceed and you must ask for help.
Does this skill create evals from scratch or does it assume I already have test data?
The skill builds evals from scratch. It guides you through defining eval criteria, instrumenting the app, capturing reference traces, defining evaluators, and building a golden dataset. If your prompt specifies a dataset or evaluation spec, the skill will use that instead.
What Python versions does this support?
Python 3.10 and later. The skill also requires pixie-qa version 0.8.4 or later.

Generated from the current SKILL.md. These answers refresh after source changes.