All skills
microsoft avatar

/m365-agent-evaluator

@a43d2c6
by microsoftmicrosoft/skills3.1k stars
351

Use this skill when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI. Trigger on intents such as "evaluate my agent", "test my agent", "run my evals", "create eval prompts", "add multi-turn tests", "tune evaluator thresholds", "why is my agent failing", or "set up eval environment variables".

Use this Skill: https://skilld.dev/gh/microsoft/skills/m365-agent-evaluator

This session only. Nothing lands on disk.

examplesrun-and-analyze.md

≈415 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Example: run evals and analyze results

User intent: "Run my evals and tell me why the agent is failing."

Safe preflight

node --version
npx -y --package @microsoft/m365-copilot-eval@latest runevals --version
npx -y --package @microsoft/m365-copilot-eval@latest runevals --help

Confirm env files exist without printing values. For first-time setup:

npx -y --package @microsoft/m365-copilot-eval@latest runevals accept-eula
npx -y --package @microsoft/m365-copilot-eval@latest runevals --init-only

Run with JSON output

npx -y --package @microsoft/m365-copilot-eval@latest runevals --prompts-file evals\evals.json --concurrency 1 --output .evals\latest.json

Use --concurrency 1 for debugging. Increase up to 5 only after setup is stable.

Optional human report

npx -y --package @microsoft/m365-copilot-eval@latest runevals --prompts-file evals\evals.json --output .evals\latest.html

Analysis approach

  1. Load references\result-analysis.md.
  2. Parse items from the JSON output.
  3. Check only score keys that exist.
  4. Separate setup/auth/model/schema failures from quality failures.
  5. Group quality failures by likely fix: instructions, grounding, citations, expected response, or capability gap.

Example response:

The main issue is grounding: two prompts passed relevance/coherence but failed groundedness. The agent answered with plausible project facts that were not present in the provided sources. Recommended change: add an instruction to answer only from retrieved workplace sources and say what is missing when evidence is insufficient.

Source: SKILL.md on GitHub

1 warning3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    This skill facilitates the evaluation of Microsoft 365 Copilot declarative agents using official Microsoft tooling. It includes robust security guardrails, such as guidance on managing environment secrets and avoiding the exposure of sensitive data in logs or chat. The skill's operations are aligned with its intended purpose and follow established development best practices.

  • Socket3mo

    No alerts

  • Snyk3mo

    Risk: MEDIUM · 1 issue

Signed by skilld at a43d2c6. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago

README badge

README badge for microsoft/skills/m365-agent-evaluator