All skills
microsoft avatar

/microsoft-foundry

@04110d9
by microsoftmicrosoft/skills3.1k stars
351

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. USE FOR: foundry, azd ai agent, azd provision/deploy, hosted agent scaffold/develop/run/deploy/troubleshoot, prompt agent create, create agent, update agent, add tool to agent, invoke agent, agent.yaml, agent insights, pull agent insights, evaluate agent, batch eval, continuous eval, continuous monitoring, agent CI/CD, optimize prompt, improve prompt, prompt optimizer, optimize agent instructions, Agent Optimizer scaffold, dataset curation from traces, deploy model, model fine-tuning (SFT/DPO/RFT), Foundry project, RBAC, role assignment, permissions, quota, capacity, region, deployment failure, AI Services, create Foundry resource, knowledge index, customize deployment, onboard, availability, training-data, grader, distillation, large file upload. DO NOT USE FOR: Azure Functions, App Service, general Azure deploy (use azure-deploy), general Azure prep (use azure-prepare).

Use this Skill: https://skilld.dev/gh/microsoft/skills/microsoft-foundry

This session only. Nothing lands on disk.

finetuningworkflowsiterative-training.md

≈795 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Iterative Training Workflow

Systematically improve a fine-tuned model through successive experiments.

The Core Loop

1. Train with current config
2. Analyze training curves
3. Evaluate on held-out set
4. Diagnose what to change
5. Plan next experiment
→ Better than baseline? → Good enough? → Ship it (or loop back to 4)

Rule: Change ONE variable per experiment.

Experiment Tracking

Run Base model Dataset Epochs LR Batch Best val_loss Combined eval
R1 gpt-4.1-mini v1 (335 ex) 2 1.0 default 0.320 8.05
R2 gpt-4.1-mini v1 (335 ex) 2 0.5 default 0.310 9.15
... ... ... ... ... ... ... ...

What to Try (Priority Order)

Priority 1: Data Quality (highest leverage)

  • Fix inconsistencies: Contradicting examples confuse the model
  • Add diversity: Add examples for input types the model fails on
  • Reduce noise: Remove "correct but not ideal" outputs

Priority 2: Hyperparameters

See references/hyperparameters.md for full guide.

Quick sweep strategy:

  1. Baseline: epochs=2, lr=1.0
  2. Overfitting → lr=0.5 or epochs=1
  3. Underfitting → lr=1.5 or epochs=3
  4. Good LR found → try batch_size=16 or 32

Priority 3: Base Model

Model Best for
gpt-4.1-mini Best quality-per-dollar, most tasks
gpt-4.1-nano Fastest inference, simple tasks
gpt-oss-20b Large datasets, lowest absolute loss
Ministral-3B Lightweight, fast inference
Qwen-3-32B, Llama-3.3-70B Multilingual or specialized tasks

Priority 4: Training Type

  • SFT plateaued + need better reasoning → RFT (if model supports it)
  • Need style alignment → DPO
  • See references/training-types.md before switching

Diagnostic Decision Tree

Training curves healthy (no overfitting)?
├─ Yes
│  ├─ Eval improved? → Refine further
│  └─ Eval same/worse? → Data quality issue — filter or augment
└─ No (overfitting)
   ├─ Earlier checkpoint evals well? → Deploy that checkpoint
   ├─ Not severe → Reduce epochs or lower LR
   └─ Severe (ratio > 2.0)
      ├─ Dataset too small → Add more data
      └─ Dataset large → Lower LR dramatically (0.1-0.3)

When to Stop

  1. Beaten baseline by meaningful margin (>5%) and last 3 experiments didn't improve
  2. Diminishing returns: each experiment improves < 0.1 points
  3. Model is "good enough" for production
  4. Budget exhausted (time or money)

Multi-Model Strategy

Run the same dataset through 2-3 base models:

  1. gpt-4.1-mini — primary candidate
  2. gpt-oss-20b — large-dataset specialist (500+ examples)
  3. gpt-4.1-nano — fast inference option

Common Mistakes

  1. Not establishing a baseline first
  2. Changing multiple variables at once
  3. Overfitting to the eval set (keep a separate final test set)
  4. Ignoring training curves (they tell you what to change next)
  5. More data without quality check (lower-quality data often makes things worse)
  6. Not cleaning up old deployments (wastes quota and money)

Source: SKILL.md on GitHub

2 warnings3d4 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    This skill provides a comprehensive environment for managing the end-to-end lifecycle of AI agents, models, and infrastructure on Microsoft Foundry. It includes sub-skills for deployment, evaluation, fine-tuning, and troubleshooting. The skill utilizes dynamic code execution and shell command wrappers, which are used within the context of local development and cloud orchestration. All external resources and dependencies originate from trusted organizations and well-known services.

  • Socket3d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk3d

    Risk: LOW · No issues

  • Runlayer7mo

    36/36 files flagged

Signed by skilld at 04110d9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated last week
metadata
{
  "author": "Microsoft",
  "version": "1.2.26"
}

README badge

README badge for microsoft/skills/microsoft-foundry