All skills
microsoft avatar

/microsoft-foundry

@04110d9
by microsoftmicrosoft/skills3.1k stars
351

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. USE FOR: foundry, azd ai agent, azd provision/deploy, hosted agent scaffold/develop/run/deploy/troubleshoot, prompt agent create, create agent, update agent, add tool to agent, invoke agent, agent.yaml, agent insights, pull agent insights, evaluate agent, batch eval, continuous eval, continuous monitoring, agent CI/CD, optimize prompt, improve prompt, prompt optimizer, optimize agent instructions, Agent Optimizer scaffold, dataset curation from traces, deploy model, model fine-tuning (SFT/DPO/RFT), Foundry project, RBAC, role assignment, permissions, quota, capacity, region, deployment failure, AI Services, create Foundry resource, knowledge index, customize deployment, onboard, availability, training-data, grader, distillation, large file upload. DO NOT USE FOR: Azure Functions, App Service, general Azure deploy (use azure-deploy), general Azure prep (use azure-prepare).

Use this Skill: https://skilld.dev/gh/microsoft/skills/microsoft-foundry

This session only. Nothing lands on disk.

finetuningworkflowsdiagnose-poor-results.md

≈601 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Diagnosing Poor Results

When your fine-tuned model performs worse than expected, work through this checklist top-down (most common causes first).

Diagnostic Table

# Symptom Likely Cause Fix
1 Training loss → 0, validation loss rises Overfitting 1) Deploy earlier checkpoint. 2) Reduce epochs. 3) Lower LR. 4) Add more diverse data. Overfitting ratio > 1.5 is concerning.
2 High correctness, low conciseness (or reverse) Dataset style mismatch Verbose: Add concise examples, use "Be concise" system prompt, filter to shortest correct examples. Terse: Add detailed examples, increase dataset with quality-filtered data.
3 Model seems good on spot-check but auto-eval is low Evaluation rubric issue Manually grade 10 examples vs. LLM judge. Check: Is judge model strong enough? Is rubric clear? Do reference answers match desired output?
4 Garbage, empty outputs, or errors Deployment/client bug Check: wrong model format (→ HTTP 500), AzureOpenAI on project endpoint (→ "api-version not allowed"), low capacity (→ timeouts), wrong deployment name. Test with curl.
5 RFT model scores below base model RFT-specific issue See RFT section below.

RFT-Specific Diagnosis

Signal Meaning Fix
Train-val grader gap > 0.2 Model gaming the grader Use stricter/more deterministic grader (Python execution > LLM judge)
Grader too easy High grader scores but bad outputs Add multi-criteria grading (syntax + semantic)
Grader too noisy Random signal, no learning Use deterministic grader or increase val set size
All of the above fail RFT may not suit this task Switch back to SFT

Escalation Path

If nothing above helps:

  1. Try a different base model — some fine-tune better for certain tasks
  2. Increase dataset 2x-5x with synthetic data
  3. Simplify the task — fine-tune for a narrower sub-task first
  4. Try prompt engineering instead — sometimes a well-crafted system prompt beats fine-tuning
  5. Combine approaches — prompt engineering + fine-tuning together

Red Flags: Don't Fine-Tune

  • Base model already scores > 9.0 (minimal headroom)
  • Task changes frequently (constant retraining needed)
  • < 50 examples and can't generate synthetic data
  • "Correct" output is highly subjective

Source: SKILL.md on GitHub

2 warnings3d4 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    This skill provides a comprehensive environment for managing the end-to-end lifecycle of AI agents, models, and infrastructure on Microsoft Foundry. It includes sub-skills for deployment, evaluation, fine-tuning, and troubleshooting. The skill utilizes dynamic code execution and shell command wrappers, which are used within the context of local development and cloud orchestration. All external resources and dependencies originate from trusted organizations and well-known services.

  • Socket3d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk3d

    Risk: LOW · No issues

  • Runlayer7mo

    36/36 files flagged

Signed by skilld at 04110d9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last week
metadata
{
  "author": "Microsoft",
  "version": "1.2.26"
}

README badge

README badge for microsoft/skills/microsoft-foundry