All skills
github avatar

/eval-driven-dev

@2860790 official
by githubgithub/awesome-copilot40k stars
5,040

Improve AI application with evaluation-driven development. Define eval criteria, instrument the application, build golden datasets, observe and evaluate application runs, analyze results, and produce a concrete action plan for improvements. ALWAYS USE THIS SKILL when the user asks to set up QA, add tests, add evals, evaluate, benchmark, fix wrong behaviors, improve quality, or do quality assurance for any Python project that calls an LLM model.

Use this Skill: https://skilld.dev/gh/github/awesome-copilot/eval-driven-dev

This session only. Nothing lands on disk.

referencesrunnable-examplesfastapi-web-server.md

≈1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Runnable Example: FastAPI / Web Server

When the app is a web server (FastAPI, Flask, Starlette) and you need to exercise the full HTTP request pipeline.

Approach: Use httpx.AsyncClient with ASGITransport to run the ASGI app in-process. This is the fastest and most reliable approach — no subprocess, no port management.

# pixie_qa/run_app.py
import httpx
from pydantic import BaseModel
import pixie


class AppArgs(BaseModel):
    user_message: str


class AppRunnable(pixie.Runnable[AppArgs]):
    """Drives a FastAPI app via in-process ASGI transport."""

    _client: httpx.AsyncClient

    @classmethod
    def create(cls) -> "AppRunnable":
        return cls()

    async def setup(self) -> None:
        from myapp.main import app  # your FastAPI/Starlette app instance

        transport = httpx.ASGITransport(app=app)
        self._client = httpx.AsyncClient(transport=transport, base_url="http://test")

    async def run(self, args: AppArgs) -> None:
        await self._client.post("/chat", json={"message": args.user_message})

    async def teardown(self) -> None:
        await self._client.aclose()

ASGITransport skips lifespan events

httpx.ASGITransport does not trigger ASGI lifespan events (startup / shutdown). If the app initializes resources in its lifespan (database connections, caches, service clients), you must replicate that initialization manually in setup():

async def setup(self) -> None:
    # Manually replicate what the app's lifespan does
    from myapp.db import get_connection, init_db, seed_data
    import myapp.main as app_module

    conn = get_connection()
    init_db(conn)
    seed_data(conn)
    app_module.db_conn = conn  # set the module-level global the app expects

    transport = httpx.ASGITransport(app=app_module.app)
    self._client = httpx.AsyncClient(transport=transport, base_url="http://test")

async def teardown(self) -> None:
    await self._client.aclose()
    # Clean up the manually-initialized resources
    import myapp.main as app_module
    if hasattr(app_module, "db_conn") and app_module.db_conn:
        app_module.db_conn.close()

Concurrency with shared mutable state

If the app uses shared mutable state (in-memory SQLite, file-based DB, global caches), add a semaphore to serialise access:

import asyncio

class AppRunnable(pixie.Runnable[AppArgs]):
    _client: httpx.AsyncClient
    _sem: asyncio.Semaphore

    @classmethod
    def create(cls) -> "AppRunnable":
        inst = cls()
        inst._sem = asyncio.Semaphore(1)
        return inst

    async def setup(self) -> None:
        from myapp.main import app
        transport = httpx.ASGITransport(app=app)
        self._client = httpx.AsyncClient(transport=transport, base_url="http://test")

    async def run(self, args: AppArgs) -> None:
        async with self._sem:
            await self._client.post("/chat", json={"message": args.user_message})

    async def teardown(self) -> None:
        await self._client.aclose()

Only use the semaphore when needed — if the app uses per-session state keyed by unique IDs (call_sid, session_id), concurrent calls are naturally isolated and no lock is needed.

Alternative: External server with httpx

When the app can't be imported directly (complex startup, uvicorn.run() in __main__), start it as a subprocess and hit it with HTTP:

class AppRunnable(pixie.Runnable[AppArgs]):
    _client: httpx.AsyncClient

    @classmethod
    def create(cls) -> "AppRunnable":
        return cls()

    async def setup(self) -> None:
        # Assumes the server is already running (started via run-with-timeout.sh)
        self._client = httpx.AsyncClient(base_url="http://localhost:8000")

    async def run(self, args: AppArgs) -> None:
        await self._client.post("/chat", json={"message": args.user_message})

    async def teardown(self) -> None:
        await self._client.aclose()

Start the server before running pixie trace or pixie test:

bash resources/run-with-timeout.sh 120 uv run python -m myapp.server
sleep 3  # wait for readiness

Source: SKILL.md on GitHub

3 warnings16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill facilitates evaluation-driven development for Python AI applications using the pixie-qa framework. It automates the setup of an evaluation pipeline, including package installation, instrumentation, and result analysis. Security considerations include the installation of third-party packages, a self-updating mechanism, and the processing of potentially untrusted project specifications.

  • Socket16d

    1 alert: gptAnomaly

  • Snyk16d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    1/2 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 2860790. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 5 months ago
compatibility
Python 3.10+
Other metadata
metadata
{
  "version": "0.8.4",
  "pixie-qa-version": ">=0.8.4,<0.9.0",
  "pixie-qa-source": "https://github.com/yiouli/pixie-qa/"
}
  • Python
  • Testing
  • llm
  • evaluation
  • pixie-qa
  • quality-assurance
  • benchmarking
  • agent
  • observability

README badge

README badge for github/awesome-copilot/eval-driven-dev

Defines evaluation criteria, instruments a Python LLM application to inject controlled test data, builds datasets, and runs automated quality checks via the pixie-qa framework. Targets Python projects that call LLMs and need end-to-end testing of application code (routing, prompt assembly, response formatting) with real LLM calls — not mocking — scored by evaluators instead of deterministic assertions.

Generated from the current SKILL.md.

Does this skill work with any Python LLM application or only specific frameworks?
This skill works with any Python application that calls an LLM, regardless of framework. It instruments the app's data boundaries and runs the full application code end-to-end, exercising real routing, prompt assembly, and LLM calls.
Can I mock or stub the LLM during evaluation?
No. The skill requires real LLM calls — mocking or stubbing the LLM makes eval scores meaningless because you control both inputs and outputs. The app's own unit tests may mock the LLM; evals must not.
What is pixie-qa and do I need to install it separately?
pixie-qa is the Python package that powers the eval pipeline. The skill's setup.sh installs it automatically. If installation fails, the workflow cannot proceed and you must ask for help.
Does this skill create evals from scratch or does it assume I already have test data?
The skill builds evals from scratch. It guides you through defining eval criteria, instrumenting the app, capturing reference traces, defining evaluators, and building a golden dataset. If your prompt specifies a dataset or evaluation spec, the skill will use that instead.
What Python versions does this support?
Python 3.10 and later. The skill also requires pixie-qa version 0.8.4 or later.

Generated from the current SKILL.md. These answers refresh after source changes.