All skills
microsoft avatar

/microsoft-foundry

@04110d9
by microsoftmicrosoft/skills3.1k stars
351

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. USE FOR: foundry, azd ai agent, azd provision/deploy, hosted agent scaffold/develop/run/deploy/troubleshoot, prompt agent create, create agent, update agent, add tool to agent, invoke agent, agent.yaml, agent insights, pull agent insights, evaluate agent, batch eval, continuous eval, continuous monitoring, agent CI/CD, optimize prompt, improve prompt, prompt optimizer, optimize agent instructions, Agent Optimizer scaffold, dataset curation from traces, deploy model, model fine-tuning (SFT/DPO/RFT), Foundry project, RBAC, role assignment, permissions, quota, capacity, region, deployment failure, AI Services, create Foundry resource, knowledge index, customize deployment, onboard, availability, training-data, grader, distillation, large file upload. DO NOT USE FOR: Azure Functions, App Service, general Azure deploy (use azure-deploy), general Azure prep (use azure-prepare).

Use this Skill: https://skilld.dev/gh/microsoft/skills/microsoft-foundry

This session only. Nothing lands on disk.

finetuningreferenceshyperparameters.md

≈809 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Hyperparameter Guide

SFT / DPO Core Parameters

Parameter What it controls Default Typical range
Epochs Passes through data 2 1–5
Learning rate multiplier Weight change aggressiveness 1.0 0.1–2.0
Batch size Examples per gradient step Model-dependent 4–32

Dataset Size vs Epochs

Dataset size Recommended epochs
< 100 examples 3–5
100–500 examples 2–3
500–2,000 examples 1–2
> 2,000 examples 1

Learning Rate Guidelines

  • Higher LR (1.5–2.0): Large/diverse datasets, task very different from pre-training
  • Lower LR (0.1–0.5): Small datasets (<200), refining not overwriting base behavior
  • For 1,000+ examples, LR 0.2–0.5 often beats default 1.0

DPO-Specific Parameters

  • beta (default 0.1): Alignment strength. Lower = more conservative.
  • l2_multiplier (default 0.1): Regularization to prevent drift from base model.

HP Sweep Strategy

Run Epochs LR Why
1 2 1.0 Baseline
2 2 0.5 Conservative
3 2 1.5 Aggressive
4 3 1.0 More training
5 1 1.0 Minimal intervention

Checkpoint Trick

When overfitting (val loss rises after epoch 2): deploy the epoch-2 checkpoint directly instead of retraining. Azure saves checkpoints at each epoch boundary.

checkpoints = client.fine_tuning.jobs.checkpoints.list(job_id)
for cp in checkpoints.data:
    print(f"Step {cp.step_number}: val_loss={cp.metrics.valid_loss}")

Model-Specific Recommendations

Model Recommended Start Notes
gpt-4.1-mini 2ep, lr=0.5–1.0 Very capable base; small nudges work
gpt-4.1-nano 2–3ep, lr=1.0–1.5 Smaller capacity, needs more epochs
gpt-oss-20b 2ep, lr=0.2–0.5 Lower LR critical; deployment may need capacity=100
o4-mini (RFT) Grader quality > HPs Focus on grader, not HP sweep

OSS Model Parameters

All OSS models require trainingType: "globalStandard" in the API request.

Model Recommended Start Best Found Notes
Ministral-3B 5ep, lr=1.0 10ep, lr=0.5 Small model, slow convergence
gpt-oss-20b 2ep, lr=0.3 2ep, lr=0.3 lr=1.0 overfits quickly
Llama-3.3-70B 3ep, lr=0.3 5ep, lr=0.5 lr=2.0 causes catastrophic degradation
Qwen-3-32B 3ep, lr=0.3 3ep, lr=0.3 Most fragile — more data can hurt

Key patterns: OSS models need 2–5× more epochs than nano. Lower LR (0.3–0.5) is safer. More data doesn't always help.

RFT Hyperparameters

Parameter Description Recommended Start
reasoning_effort "low", "medium", "high" "medium"
compute_multiplier Scales rollouts per step 1.5
learning_rate_multiplier Scales LR 1.0
n_epochs Data passes 2–3
eval_interval Eval every N steps 5
eval_samples Validation examples per eval 10
max_episode_steps Max tool calls + reasoning steps 5–10

Source: SKILL.md on GitHub

2 warnings3d4 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    This skill provides a comprehensive environment for managing the end-to-end lifecycle of AI agents, models, and infrastructure on Microsoft Foundry. It includes sub-skills for deployment, evaluation, fine-tuning, and troubleshooting. The skill utilizes dynamic code execution and shell command wrappers, which are used within the context of local development and cloud orchestration. All external resources and dependencies originate from trusted organizations and well-known services.

  • Socket3d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk3d

    Risk: LOW · No issues

  • Runlayer7mo

    36/36 files flagged

Signed by skilld at 04110d9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated last week
metadata
{
  "author": "Microsoft",
  "version": "1.2.26"
}

README badge

README badge for microsoft/skills/microsoft-foundry