All skills
google avatar

/google-agents-cli-eval

@2c39459
by googlegoogle/agents-cli6k stars
686

This skill should be used when the user wants to "run an evaluation", "evaluate my agent", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs guidance on the Agent Platform eval methodology and the Quality Flywheel. Covers eval metrics, dataset schema, LLM-as-judge scoring, and common failure causes. Applies to any agents-cli project, whatever framework the agent is written in. Do NOT use for agent API code patterns (ADK: use google-agents-cli-adk-code), deployment (use google-agents-cli-deploy), or project scaffolding (use google-agents-cli-scaffold).

Use this Skill: https://skilld.dev/gh/google/agents-cli/google-agents-cli-eval

This session only. Nothing lands on disk.

referencesadvanced-commands.md

≈598 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Advanced Eval Commands

Opt-in commands from the Quality Flywheel. The core loop (eval run, and eval generate / eval grade) lives in SKILL.md.

eval analyze

Runs LLM-based failure clustering and root-cause analysis over a results_*.json produced by an eval run. Use when you have 10+ failing cases and want categorized failure modes instead of reading the HTML case-by-case. Supported --metric values: multi_turn_task_success, multi_turn_tool_use_quality.

# Basic: analyze a results file with default settings
agents-cli eval analyze --eval-result artifacts/grade_results/results_<ts>.json

# Advanced: restrict to a specific metric and cap loss clusters
agents-cli eval analyze \
  --eval-result artifacts/grade_results/results_<ts>.json \
  --metric multi_turn_tool_use_quality \
  --top-k 5 \
  --output artifacts/analysis_<ts>.json

eval optimize

ADK Python projects. It wraps adk optimize via uv and loads the agent through ADK.

Runs GEPA prompt optimization against a target metric. Suitable after an eval run identifies prompt-only failures (wording, not tool/orchestration logic). --dataset and --target-metric override values in --config when both are passed. Long-running and expensive, see Stage 4 of the Quality Flywheel for usage guidance.

# Basic: optimize against a single metric on a dataset
agents-cli eval optimize --dataset tests/eval/datasets/basic-dataset.json --target-metric final_response_quality

# Advanced: drive multi-metric / multi-dataset optimization from a config file
agents-cli eval optimize --config tests/eval/optimization_config.json

eval submit / eval results (cloud-side)

The managed, asynchronous counterpart to the local path, for large or CI-driven runs: eval submit hands the dataset and metrics to the Agent Platform Eval Service, and eval results polls and downloads the scores. Pass --resource-name <agent> to also run inference server-side (managed generate + grade); omit it to grade an existing trace (managed grade).

# Grade an existing trace server-side; returns a run resource name to poll
agents-cli eval submit --dataset tests/eval/datasets/basic-dataset.json --dest gs://my-bucket
# Add --resource-name projects/<p>/locations/<l>/reasoningEngines/<id> to run inference too

agents-cli eval results --run-id <run-resource-name>

Source: SKILL.md on GitHub

No alertstoday3 checks · Risk SAFE
  • Gen Agent Trust Hubtoday

    This skill provides comprehensive guidance on using the Agent Platform evaluation framework. It includes security considerations regarding local code execution for custom metrics and data processing, which are managed through documented best practices and platform guardrails.

  • Sockettoday

    No alerts

  • Snyktoday

    Risk: LOW · No issues

Signed by skilld at 2c39459. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 2 days ago
Other metadata
metadata
{
  "author": "Google",
  "license": "Apache-2.0",
  "version": "1.8.0",
  "requires": {
    "bins": [
      "agents-cli"
    ],
    "install": "uv tool install google-agents-cli"
  }
}

README badge

README badge for google/agents-cli/google-agents-cli-eval