All skills
microsoft avatar

/microsoft-foundry

@04110d9
by microsoftmicrosoft/skills3.1k stars
351

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. USE FOR: foundry, azd ai agent, azd provision/deploy, hosted agent scaffold/develop/run/deploy/troubleshoot, prompt agent create, create agent, update agent, add tool to agent, invoke agent, agent.yaml, agent insights, pull agent insights, evaluate agent, batch eval, continuous eval, continuous monitoring, agent CI/CD, optimize prompt, improve prompt, prompt optimizer, optimize agent instructions, Agent Optimizer scaffold, dataset curation from traces, deploy model, model fine-tuning (SFT/DPO/RFT), Foundry project, RBAC, role assignment, permissions, quota, capacity, region, deployment failure, AI Services, create Foundry resource, knowledge index, customize deployment, onboard, availability, training-data, grader, distillation, large file upload. DO NOT USE FOR: Azure Functions, App Service, general Azure deploy (use azure-deploy), general Azure prep (use azure-prepare).

Use this Skill: https://skilld.dev/gh/microsoft/skills/microsoft-foundry

This session only. Nothing lands on disk.

foundry-agenteval-datasetsreferencesdataset-curation.md

≈1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Dataset Curation — Human-in-the-Loop Review

Review, annotate, and approve harvested trace candidates before including them in evaluation datasets. This ensures dataset quality by adding a human review gate between raw trace extraction and finalized test cases.

Workflow Overview

Raw Traces (from KQL harvest)
    │
    ▼
[1] Candidate File (unreviewed)
    │
    ▼
[2] Human Review (approve/edit/reject each)
    │
    ▼
[3] Approved Dataset (versioned, ready for eval)

Step 1 — Generate Candidate File

After running a trace harvest, save candidates with a status field:

.foundry/datasets/<agent-name>-traces-candidates-<date>.jsonl

Each line includes a review status:

{"query": "How do I reset my password?", "response": "...", "status": "pending", "metadata": {"source": "trace", "conversationId": "conv-abc-123", "harvestRule": "error", "errorType": "TimeoutError", "duration": 12300}}
{"query": "What's the refund policy?", "response": "...", "status": "pending", "metadata": {"source": "trace", "conversationId": "conv-def-456", "harvestRule": "latency", "duration": 8700}}

Step 2 — Present for Review

Show candidates in a review table:

# Status Query (preview) Source Error Duration Eval Score
1 ⏳ pending "How do I reset my..." error harvest TimeoutError 12.3s —
2 ⏳ pending "What's the refund..." latency harvest — 8.7s —
3 ⏳ pending "Can you help me..." low-eval harvest — 0.4s 2.0

Review Actions

For each candidate, the user can:

Action Result
Approve Include in dataset as-is
Approve + Edit Include with modified query/response/ground_truth
Add Ground Truth Approve and add the expected correct answer
Reject Exclude from dataset
Flag Mark for later review

Batch Operations

  • "Approve all" — include all pending candidates
  • "Approve all errors" — include all candidates from error harvest
  • "Reject duplicates" — exclude candidates with similar queries to existing dataset entries
  • "Approve #1, #3, #5; reject #2, #4" — selective approval by number

Step 3 — Finalize Dataset

After review, filter approved candidates and save to a versioned dataset:

  1. Read .foundry/datasets/manifest.json to find the latest version number
  2. Filter candidates where status == "approved"
  3. Remove the status field from the output
  4. Save to .foundry/datasets/<agent-name>-<source>-v<N>.jsonl
  5. Update .foundry/datasets/manifest.json with metadata

Update Candidate Status

Mark the candidate file with final statuses:

{"query": "How do I reset my password?", "status": "approved", "ground_truth": "Navigate to Settings > Security > Reset Password", "metadata": {...}}
{"query": "What's the refund policy?", "status": "rejected", "rejectReason": "duplicate of existing test case", "metadata": {...}}
{"query": "Can you help me...", "status": "approved", "metadata": {...}}

💡 Tip: Keep candidate files as an audit trail. They document what was reviewed, when, and why items were accepted or rejected.

Quality Checks

Before finalizing, verify dataset quality:

Check Criteria
No duplicates Ensure no query appears in both the new dataset and existing datasets
Balanced categories Verify reasonable distribution across categories (not all edge-cases)
Ground truth coverage Flag examples without ground_truth that may benefit from one
Minimum size Warn if dataset has fewer than 20 examples (may not be statistically meaningful)
Safety coverage Ensure safety-related test cases are included if the agent handles sensitive topics

Next Steps

Source: SKILL.md on GitHub

2 warnings3d4 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    This skill provides a comprehensive environment for managing the end-to-end lifecycle of AI agents, models, and infrastructure on Microsoft Foundry. It includes sub-skills for deployment, evaluation, fine-tuning, and troubleshooting. The skill utilizes dynamic code execution and shell command wrappers, which are used within the context of local development and cloud orchestration. All external resources and dependencies originate from trusted organizations and well-known services.

  • Socket3d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk3d

    Risk: LOW · No issues

  • Runlayer7mo

    36/36 files flagged

Signed by skilld at 04110d9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated last week
metadata
{
  "author": "Microsoft",
  "version": "1.2.26"
}

README badge

README badge for microsoft/skills/microsoft-foundry