All skills
google-gemini avatar

/behavioral-evals

@db14cdf
by google-geminigoogle-gemini/gemini-cli107k stars
14,675

Guidance for creating, running, fixing, and promoting behavioral evaluations. Use when verifying agent decision logic, debugging failures, debugging prompt steering, or adding workspace regression tests.

Use this Skill: https://skilld.dev/gh/google-gemini/gemini-cli/behavioral-evals

This session only. Nothing lands on disk.

referencespromoting.md

≈480 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Promoting Behavioral Evals

Use this guide when asked to analyze nightly results and promote incubated tests to stable suites.


1. 🔍 Investigate candidates

  1. Audit Nightly Logs: Use the gh CLI to fetch results from evals-nightly.yml (Direct URL: https://github.com/google-gemini/gemini-cli/actions/workflows/evals-nightly.yml).
    • Tip: The aggregate summary from the most recent run integrates the last 7 runs of history automatically.
    • Safety: DO NOT push changes or start remote runs. All verification is local.
  2. Assess Stability: Identify tests that pass 100% of the time across ALL enabled models over the last 7 nightly runs in a row.
    • 100% means the test passed 3/3 times for every model and run.
  3. Promotion Targets: Tests meeting this criteria are candidates for promotion from USUALLY_PASSES to ALWAYS_PASSES.

2. 🚥 Promotion Steps

  1. Locate File: Locate the eval file in the evals/ directory.
  2. Update Policy: Modify the policy argument to ALWAYS_PASSES.
    evalTest('ALWAYS_PASSES', { ... })
  3. Targeting: Follow guidelines in evals/README.md regarding stable suite organization.
  4. Constraint: Your final change must be minimal and targeted strictly to promoting the test status. Do not refactor the test or setup fixtures.

3. ✅ Verify

  1. Run Prompted Tests: Run the promoted test locally using non-interactive Vitest to confirm structure validity.
  2. Verify Suite Inclusion: Check that the test is successfully picked up by standard runnable ranges.

4. 📊 Report

Provide a summary of:

  • Which tests were promoted.
  • Provide the success rate evidence (e.g., 7/7 runs passed for all models).
  • If no candidates qualified, list the next closest candidates and their current pass rate.

Source: SKILL.md on GitHub

No alerts16d3 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides a structured framework for behavioral evaluations, authored by a trusted vendor. It includes templates and procedural guides for testing agent decision logic. No security issues were detected.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

Signed by skilld at db14cdf. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 1 hour ago.

Activeupdated 6 months ago

README badge

README badge for google-gemini/gemini-cli/behavioral-evals