All skills
microsoft avatar

/m365-agent-evaluator

@a43d2c6
by microsoftmicrosoft/skills3.1k stars
351

Use this skill when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI. Trigger on intents such as "evaluate my agent", "test my agent", "run my evals", "create eval prompts", "add multi-turn tests", "tune evaluator thresholds", "why is my agent failing", or "set up eval environment variables".

Use this Skill: https://skilld.dev/gh/microsoft/skills/m365-agent-evaluator

This session only. Nothing lands on disk.

examplesmissing-instructions.md

≈396 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Example: missing or weak instructions

User intent: "My agent gives vague answers in evals."

Symptoms

  • relevance passes but coherence fails.
  • groundedness fails because the agent invents missing facts.
  • similarity fails because the answer omits required structure or decisions.
  • Follow-up turns fail after a successful first turn.

Diagnosis

Check whether the agent instructions specify:

  1. supported scenarios and boundaries,
  2. source-grounding expectations,
  3. citation expectations,
  4. response format,
  5. behavior when data is missing,
  6. follow-up context handling.

Suggested instruction additions

Use only available workplace sources when answering source-backed questions. If the sources do not contain enough evidence, say what is missing instead of guessing.
For project-status answers, use this structure: Summary, Evidence, Risks, Next actions. Keep the answer concise and include citations when available.
For follow-up questions, preserve the project, customer, and time window from the prior turn unless the user changes them.

Matching eval update

Add or keep regression prompts that test the new instruction:

{
  "prompt": "Summarize the latest status for the project and include citations.",
  "expected_response": "The agent summarizes only source-backed status, cites available sources, and states when evidence is missing.",
  "evaluators": {
    "Groundedness": {
      "threshold": 4
    },
    "Citations": {
      "threshold": 1
    }
  },
  "evaluators_mode": "extend"
}

Source: SKILL.md on GitHub

1 warning3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    This skill facilitates the evaluation of Microsoft 365 Copilot declarative agents using official Microsoft tooling. It includes robust security guardrails, such as guidance on managing environment secrets and avoiding the exposure of sensitive data in logs or chat. The skill's operations are aligned with its intended purpose and follow established development best practices.

  • Socket3mo

    No alerts

  • Snyk3mo

    Risk: MEDIUM · 1 issue

Signed by skilld at a43d2c6. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago

README badge

README badge for microsoft/skills/m365-agent-evaluator