All skills
microsoft avatar

/m365-agent-evaluator

@a43d2c6
by microsoftmicrosoft/skills3.1k stars
351

Use this skill when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI. Trigger on intents such as "evaluate my agent", "test my agent", "run my evals", "create eval prompts", "add multi-turn tests", "tune evaluator thresholds", "why is my agent failing", or "set up eval environment variables".

Use this Skill: https://skilld.dev/gh/microsoft/skills/m365-agent-evaluator

This session only. Nothing lands on disk.

referencesoutput-schema.md

≈678 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Output schema reference

The current runevals --output <file> JSON output is schema-compatible with the eval document format. It is not the older { "summary": ..., "results": [...] } shape.

Use references\output-schema.json for a compact validation-oriented schema and references\prompts-schema.json for the full package schema.

JSON output

Typical JSON output:

{
  "schemaVersion": "1.2.0",
  "metadata": {
    "evaluatedAt": "2025-01-01T00:00:00Z",
    "agentId": "00000000-0000-0000-0000-000000000000",
    "agentName": "Contoso Agent",
    "cliVersion": "1.5.0-preview.1"
  },
  "default_evaluators": {},
  "items": [
    {
      "prompt": "What can this agent help me with?",
      "response": "The agent response.",
      "expected_response": "The expected behavior.",
      "scores": {
        "relevance": {
          "score": 4,
          "result": "pass",
          "threshold": 3
        },
        "coherence": {
          "score": 5,
          "result": "pass",
          "threshold": 3
        }
      }
    }
  ]
}

Multi-turn output uses an item with turns and may include summary.

Score object

Scores are sparse. Keys appear only for evaluators that ran.

Score key Value shape
relevance `{ "score": 1-5, "result": "pass"
coherence `{ "score": 1-5, "result": "pass"
groundedness `{ "score": 1-5, "result": "pass"
similarity `{ "score": 1-5, "result": "pass"
citations `{ "count": number, "result": "pass"
exactMatch `{ "score": 0
partialMatch `{ "score": 0.0-1.0, "result": "pass"

Do not treat missing score keys as failures.

CSV output

CSV reports include:

  • metadata comments at the top,
  • aggregate statistics,
  • single-turn sections,
  • multi-turn sections,
  • serialized score JSON.

Use CSV when the user wants spreadsheet review or lightweight automation without parsing full JSON.

HTML output

HTML reports are for human review and may open in the default browser. Treat them as sensitive because they can contain prompts, retrieved data, and agent responses.

Automation guidance

When parsing JSON results:

  1. Read items.
  2. For each item, handle either single-turn prompt or multi-turn turns.
  3. Check only scores keys that exist.
  4. Treat setup/auth/model/schema errors separately from quality scores.
  5. Compare runs only when datasets, evaluators, thresholds, and model configuration are stable.

Source: SKILL.md on GitHub

1 warning3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    This skill facilitates the evaluation of Microsoft 365 Copilot declarative agents using official Microsoft tooling. It includes robust security guardrails, such as guidance on managing environment secrets and avoiding the exposure of sensitive data in logs or chat. The skill's operations are aligned with its intended purpose and follow established development best practices.

  • Socket3mo

    No alerts

  • Snyk3mo

    Risk: MEDIUM · 1 issue

Signed by skilld at a43d2c6. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago

README badge

README badge for microsoft/skills/m365-agent-evaluator