All skills
microsoft avatar

/m365-agent-evaluator

@a43d2c6
by microsoftmicrosoft/skills3.1k stars
351

Use this skill when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI. Trigger on intents such as "evaluate my agent", "test my agent", "run my evals", "create eval prompts", "add multi-turn tests", "tune evaluator thresholds", "why is my agent failing", or "set up eval environment variables".

Use this Skill: https://skilld.dev/gh/microsoft/skills/m365-agent-evaluator

This session only. Nothing lands on disk.

examplesbasic-generation.md

≈456 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Example: create a starter eval dataset

User intent: "Create eval prompts for my M365 Copilot agent."

Steps

  1. Confirm the target agent scenario and whether the repo already has evals\evals.json.
  2. Load references\eval-templates.md and references\pra-framework.md.
  3. Create a schema 1.2.0 dataset with root items.
  4. Save it under evals\evals.json unless the user asks for another path.

Starter file

{
  "schemaVersion": "1.2.0",
  "metadata": {
    "name": "Starter M365 Copilot agent evals",
    "description": "Core smoke tests for agent scope, grounding, and response quality.",
    "tags": ["starter", "regression"]
  },
  "default_evaluators": {
    "Relevance": {},
    "Coherence": {}
  },
  "items": [
    {
      "prompt": "What can this agent help me with?",
      "expected_response": "The agent explains its supported scope without claiming unsupported capabilities."
    },
    {
      "prompt": "Summarize the latest status for the project using available sources.",
      "expected_response": "The agent summarizes available status, distinguishes known facts from missing data, and avoids unsupported claims.",
      "evaluators": {
        "Groundedness": {
          "threshold": 3
        }
      },
      "evaluators_mode": "extend"
    },
    {
      "prompt": "List the open action items with owners.",
      "expected_response": "The agent lists action items and owners only when source data supports them.",
      "evaluators": {
        "Citations": {
          "threshold": 1
        }
      },
      "evaluators_mode": "extend"
    }
  ]
}

First safe command

npx -y --package @microsoft/m365-copilot-eval@latest runevals --init-only

Run real evals only after the user confirms tenant, agent, and Azure OpenAI configuration is ready.

Source: SKILL.md on GitHub

1 warning3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    This skill facilitates the evaluation of Microsoft 365 Copilot declarative agents using official Microsoft tooling. It includes robust security guardrails, such as guidance on managing environment secrets and avoiding the exposure of sensitive data in logs or chat. The skill's operations are aligned with its intended purpose and follow established development best practices.

  • Socket3mo

    No alerts

  • Snyk3mo

    Risk: MEDIUM · 1 issue

Signed by skilld at a43d2c6. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago

README badge

README badge for microsoft/skills/m365-agent-evaluator