All skills
microsoft avatar

/microsoft-foundry

@04110d9
by microsoftmicrosoft/skills3.1k stars
351

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. USE FOR: foundry, azd ai agent, azd provision/deploy, hosted agent scaffold/develop/run/deploy/troubleshoot, prompt agent create, create agent, update agent, add tool to agent, invoke agent, agent.yaml, agent insights, pull agent insights, evaluate agent, batch eval, continuous eval, continuous monitoring, agent CI/CD, optimize prompt, improve prompt, prompt optimizer, optimize agent instructions, Agent Optimizer scaffold, dataset curation from traces, deploy model, model fine-tuning (SFT/DPO/RFT), Foundry project, RBAC, role assignment, permissions, quota, capacity, region, deployment failure, AI Services, create Foundry resource, knowledge index, customize deployment, onboard, availability, training-data, grader, distillation, large file upload. DO NOT USE FOR: Azure Functions, App Service, general Azure deploy (use azure-deploy), general Azure prep (use azure-prepare).

Use this Skill: https://skilld.dev/gh/microsoft/skills/microsoft-foundry

This session only. Nothing lands on disk.

foundry-agentobservereferencescompare-iterate.md

≈838 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Steps 8–10 — Re-Evaluate, Compare Versions, Iterate

Step 8 — Re-Evaluate

Use evaluation_agent_batch_eval_create for re-evaluation, even when the selected evaluation suite has suiteName. The generated suite preserves the reviewed dataset/evaluator bundle for selection and lineage, but the run should target the agent directly. Reuse the same evaluationId as the baseline run when the evaluator set and thresholds are unchanged. Use the same local or registered test dataset (from the selected agent root's .foundry/datasets/ and suite metadata) and evaluator bundle from the selected environment/evaluation suite. Update agentVersion to the new version.

⚠️ Parameter switch reminder: Agent-target batch re-evaluation creation uses evaluationId, but follow-up calls to evaluation_get and evaluation_comparison_create must use evalId. Do not call evaluation_suite_run for batch eval.

⚠️ Eval-group immutability: Reuse the same evaluationId only when evaluatorNames and thresholds are unchanged. If you add/remove evaluators or change thresholds, create a new evaluation group first, then compare runs within that new group.

Auto-poll for completion in a background terminal (same as Step 2).

Step 9 — Compare Versions

Critical: displayName is required in the insightRequest. Despite the MCP tool schema showing displayName as optional (type: ["string", "null"]), the API will reject requests without it with a BadRequest error. state must be "NotStarted".

Required Parameters for evaluation_comparison_create

Parameter Required Description
insightRequest.displayName ✅ Human-readable name. Omitting causes BadRequest.
insightRequest.state ✅ Must be "NotStarted"
insightRequest.request.evalId ✅ Eval group ID containing both runs
insightRequest.request.baselineRunId ✅ Run ID of the baseline
insightRequest.request.treatmentRunIds ✅ Array of treatment run IDs

Use evaluation_comparison_create with a nested insightRequest:

{
  "insightRequest": {
    "displayName": "V1 vs V2 Comparison",
    "state": "NotStarted",
    "request": {
      "type": "EvaluationComparison",
      "evalId": "<eval-group-id>",
      "baselineRunId": "<baseline-run-id>",
      "treatmentRunIds": ["<new-run-id>"]
    }
  }
}

Important: Both runs must be in the same eval group (same evaluationId in Steps 2 and 8), but comparison requests and lookups use evalId for that same group identifier. That shared group assumes the evaluator bundle is fixed for all runs in the group.

Then use evaluation_comparison_get (with the returned insightId) to retrieve comparison results. Present a summary showing which version performed better per evaluator, and recommend which version to keep.

Step 10 — Iterate or Finish

If more categories remain in the prioritized action table (from Step 4), loop back to Step 5 (dive into next category) → Step 6 (optimize) → Step 7 (deploy) → Step 8 (re-evaluate) → Step 9 (compare).

Otherwise, confirm the final agent version with the user, then prompt for CI/CD evals & monitoring.

Source: SKILL.md on GitHub

2 warnings3d4 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    This skill provides a comprehensive environment for managing the end-to-end lifecycle of AI agents, models, and infrastructure on Microsoft Foundry. It includes sub-skills for deployment, evaluation, fine-tuning, and troubleshooting. The skill utilizes dynamic code execution and shell command wrappers, which are used within the context of local development and cloud orchestration. All external resources and dependencies originate from trusted organizations and well-known services.

  • Socket3d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk3d

    Risk: LOW · No issues

  • Runlayer7mo

    36/36 files flagged

Signed by skilld at 04110d9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated last week
metadata
{
  "author": "Microsoft",
  "version": "1.2.26"
}

README badge

README badge for microsoft/skills/microsoft-foundry