All skills
microsoft avatar

/microsoft-foundry

@04110d9
by microsoftmicrosoft/skills3.1k stars
351

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. USE FOR: foundry, azd ai agent, azd provision/deploy, hosted agent scaffold/develop/run/deploy/troubleshoot, prompt agent create, create agent, update agent, add tool to agent, invoke agent, agent.yaml, agent insights, pull agent insights, evaluate agent, batch eval, continuous eval, continuous monitoring, agent CI/CD, optimize prompt, improve prompt, prompt optimizer, optimize agent instructions, Agent Optimizer scaffold, dataset curation from traces, deploy model, model fine-tuning (SFT/DPO/RFT), Foundry project, RBAC, role assignment, permissions, quota, capacity, region, deployment failure, AI Services, create Foundry resource, knowledge index, customize deployment, onboard, availability, training-data, grader, distillation, large file upload. DO NOT USE FOR: Azure Functions, App Service, general Azure deploy (use azure-deploy), general Azure prep (use azure-prepare).

Use this Skill: https://skilld.dev/gh/microsoft/skills/microsoft-foundry

This session only. Nothing lands on disk.

foundry-agenteval-datasetsreferenceseval-trending.md

≈1.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Eval Trending — Metrics Over Time

Track evaluation metrics across multiple runs and versions to visualize improvement trends and detect regressions. This addresses the gap of understanding how agent quality changes over time.

Prerequisites

  • At least 2 evaluation runs in the same evaluation group (same evaluationId when created)
  • Project endpoint and selected environment resolved from azd or the selected metadata overlay

⚠️ Eval-group immutability: Trend a group only when its evaluator set and thresholds stayed fixed across runs. If either changed, start a new evaluation group and track that history separately.

Step 1 — Retrieve Evaluation History

Use evaluation_get to list all evaluation groups:

Parameter Required Description
projectEndpoint ✅ Azure AI Project endpoint
isRequestForRuns false (default) to list evaluation groups

Then retrieve all runs within the target evaluation group:

Parameter Required Description
projectEndpoint ✅ Azure AI Project endpoint
evalId ✅ Evaluation group ID
isRequestForRuns ✅ true to list runs

⚠️ Parameter guardrail: evaluation_get expects evalId, not evaluationId, even if the runs were grouped earlier with evaluationId.

Step 2 — Build Metrics Timeline

For each run, extract per-evaluator scores and build a timeline:

Run Agent Version Date Coherence Fluency Relevance Intent Resolution Task Adherence Safety
run-001 v1 2025-01-15 3.2 4.1 2.8 3.0 2.5 0.95
run-002 v2 2025-01-22 3.8 4.3 3.5 3.7 3.2 0.97
run-003 v3 2025-02-01 4.1 4.4 4.0 4.2 3.8 0.96
run-004 v4 2025-02-08 4.0 4.5 3.6 4.1 3.9 0.98

Step 3 — Trend Analysis

Calculate trends for each evaluator:

Evaluator v1 → v4 Change Trend Status
Coherence +0.8 (+25%) ↑ Improving ✅
Fluency +0.4 (+10%) ↑ Improving ✅
Relevance +0.8 (+29%) ↑ Improving (dip at v4) ⚠️
Intent Resolution +1.1 (+37%) ↑ Improving ✅
Task Adherence +1.4 (+56%) ↑ Improving ✅
Safety +0.03 (+3%) → Stable ✅

Detecting Regressions

Flag any evaluator where the latest run scored lower than the previous run:

Evaluator Previous (v3) Latest (v4) Delta Alert
Relevance 4.0 3.6 -0.4 (-10%) ⚠️ REGRESSION

⚠️ Regression detected: Relevance dropped 10% from v3 to v4. Investigate prompt changes or dataset drift. See Eval Regression for automated analysis.

Trend Visualization (Text-based)

Coherence   ████████████████████████████████░░░░░░ 4.0/5.0  ↑ +25%
Fluency     █████████████████████████████████████░░ 4.5/5.0  ↑ +10%
Relevance   ████████████████████████████░░░░░░░░░░ 3.6/5.0  ↑ +29% ⚠️ dip
Intent Res. █████████████████████████████████░░░░░░ 4.1/5.0  ↑ +37%
Task Adh.   ████████████████████████████████░░░░░░░ 3.9/5.0  ↑ +56%
Safety      ████████████████████████████████████████ 0.98     → Stable

Step 4 — Cross-Version Summary

Present an executive summary:

"Over 4 agent versions (v1→v4), your agent has improved significantly across all quality metrics. The biggest gain is Task Adherence (+56%). However, Relevance showed a 10% regression from v3 to v4 — recommend investigating recent prompt changes. Safety remains stable at 98%."

Recommended Thresholds

Severity Threshold Action
✅ Healthy ≤ 2% drop from previous run No action needed
⚠️ Warning 2–5% drop from previous run Review recent changes
🔴 Regression > 5% drop from previous run Block deployment, investigate
🔴 Critical Below baseline (v1) on any metric Rollback to last known good version

Next Steps

Source: SKILL.md on GitHub

2 warnings3d4 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    This skill provides a comprehensive environment for managing the end-to-end lifecycle of AI agents, models, and infrastructure on Microsoft Foundry. It includes sub-skills for deployment, evaluation, fine-tuning, and troubleshooting. The skill utilizes dynamic code execution and shell command wrappers, which are used within the context of local development and cloud orchestration. All external resources and dependencies originate from trusted organizations and well-known services.

  • Socket3d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk3d

    Risk: LOW · No issues

  • Runlayer7mo

    36/36 files flagged

Signed by skilld at 04110d9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated last week
metadata
{
  "author": "Microsoft",
  "version": "1.2.26"
}

README badge

README badge for microsoft/skills/microsoft-foundry