All skills
github avatar

/eval-driven-dev

@2860790 official
by githubgithub/awesome-copilot40k stars
5,040

Improve AI application with evaluation-driven development. Define eval criteria, instrument the application, build golden datasets, observe and evaluate application runs, analyze results, and produce a concrete action plan for improvements. ALWAYS USE THIS SKILL when the user asks to set up QA, add tests, add evals, evaluate, benchmark, fix wrong behaviors, improve quality, or do quality assurance for any Python project that calls an LLM model.

Use this Skill: https://skilld.dev/gh/github/awesome-copilot/eval-driven-dev

This session only. Nothing lands on disk.

references1-b-entry-point.md

≈587 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Step 1b: Entry Point & Execution Flow

Identify how the application starts and how a real user invokes it. Use the capability inventory from pixie_qa/00-project-analysis.md to prioritize — focus on the entry point(s) that exercise the most valuable and frequently-used capabilities, not just the first one you find.


What to investigate

1. How the software runs

What is the entry point? How do you start it? Is it a CLI, a server, a library function? What are the required arguments, config files, or environment variables?

Look for:

  • if __name__ == "__main__" blocks
  • Framework entry points (FastAPI app, Flask app, Django manage.py)
  • CLI entry points in pyproject.toml ([project.scripts])
  • Docker/compose configs that reveal startup commands

2. The real user entry point

How does a real user or client invoke the app? This is what the eval must exercise — not an inner function that bypasses the request pipeline.

  • Web server: Which HTTP endpoints accept user input? What methods (GET/POST)? What request body shape?
  • CLI: What command-line arguments does the user provide?
  • Library/function: What function does the caller import and call? What arguments?

3. Environment and configuration

  • What env vars does the app require? (service endpoints, database URLs, feature flags)
  • What config files does it read?
  • What has sensible defaults vs. what must be explicitly set?

Output: pixie_qa/01-entry-point.md

Write your findings to this file. Keep it focused — only entry point and execution flow.

Template

# Entry Point & Execution Flow

## How to run

<Command to start the app, required env vars, config files>

## Entry point

- **File**: <e.g., app.py, main.py>
- **Type**: <FastAPI server / CLI / standalone function / etc.>
- **Framework**: <FastAPI, Flask, Django, none>

## User-facing endpoints / interface

<For each way a user interacts with the app:>

- **Endpoint / command**: <e.g., POST /chat, python main.py --query "...">
- **Input format**: <request body shape, CLI args, function params>
- **Output format**: <response shape, stdout format, return type>

## Environment requirements

| Variable | Purpose | Required? | Default |
| -------- | ------- | --------- | ------- |
| ...      | ...     | ...       | ...     |

Source: SKILL.md on GitHub

3 warnings16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill facilitates evaluation-driven development for Python AI applications using the pixie-qa framework. It automates the setup of an evaluation pipeline, including package installation, instrumentation, and result analysis. Security considerations include the installation of third-party packages, a self-updating mechanism, and the processing of potentially untrusted project specifications.

  • Socket16d

    1 alert: gptAnomaly

  • Snyk16d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    1/2 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 2860790. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated 5 months ago
compatibility
Python 3.10+
Other metadata
metadata
{
  "version": "0.8.4",
  "pixie-qa-version": ">=0.8.4,<0.9.0",
  "pixie-qa-source": "https://github.com/yiouli/pixie-qa/"
}
  • Python
  • Testing
  • llm
  • evaluation
  • pixie-qa
  • quality-assurance
  • benchmarking
  • agent
  • observability

README badge

README badge for github/awesome-copilot/eval-driven-dev

Defines evaluation criteria, instruments a Python LLM application to inject controlled test data, builds datasets, and runs automated quality checks via the pixie-qa framework. Targets Python projects that call LLMs and need end-to-end testing of application code (routing, prompt assembly, response formatting) with real LLM calls — not mocking — scored by evaluators instead of deterministic assertions.

Generated from the current SKILL.md.

Does this skill work with any Python LLM application or only specific frameworks?
This skill works with any Python application that calls an LLM, regardless of framework. It instruments the app's data boundaries and runs the full application code end-to-end, exercising real routing, prompt assembly, and LLM calls.
Can I mock or stub the LLM during evaluation?
No. The skill requires real LLM calls — mocking or stubbing the LLM makes eval scores meaningless because you control both inputs and outputs. The app's own unit tests may mock the LLM; evals must not.
What is pixie-qa and do I need to install it separately?
pixie-qa is the Python package that powers the eval pipeline. The skill's setup.sh installs it automatically. If installation fails, the workflow cannot proceed and you must ask for help.
Does this skill create evals from scratch or does it assume I already have test data?
The skill builds evals from scratch. It guides you through defining eval criteria, instrumenting the app, capturing reference traces, defining evaluators, and building a golden dataset. If your prompt specifies a dataset or evaluation spec, the skill will use that instead.
What Python versions does this support?
Python 3.10 and later. The skill also requires pixie-qa version 0.8.4 or later.

Generated from the current SKILL.md. These answers refresh after source changes.