Example: run evals and analyze results
User intent: "Run my evals and tell me why the agent is failing."
Safe preflight
node --version
npx -y --package @microsoft/m365-copilot-eval@latest runevals --version
npx -y --package @microsoft/m365-copilot-eval@latest runevals --helpConfirm env files exist without printing values. For first-time setup:
npx -y --package @microsoft/m365-copilot-eval@latest runevals accept-eula
npx -y --package @microsoft/m365-copilot-eval@latest runevals --init-onlyRun with JSON output
npx -y --package @microsoft/m365-copilot-eval@latest runevals --prompts-file evals\evals.json --concurrency 1 --output .evals\latest.jsonUse --concurrency 1 for debugging. Increase up to 5 only after setup is stable.
Optional human report
npx -y --package @microsoft/m365-copilot-eval@latest runevals --prompts-file evals\evals.json --output .evals\latest.htmlAnalysis approach
- Load
references\result-analysis.md. - Parse
itemsfrom the JSON output. - Check only score keys that exist.
- Separate setup/auth/model/schema failures from quality failures.
- Group quality failures by likely fix: instructions, grounding, citations, expected response, or capability gap.
Example response:
The main issue is grounding: two prompts passed relevance/coherence but failed groundedness. The agent answered with plausible project facts that were not present in the provided sources. Recommended change: add an instruction to answer only from retrieved workplace sources and say what is missing when evidence is insufficient.