All skills
microsoft avatar

/microsoft-foundry

@04110d9
by microsoftmicrosoft/skills3.1k stars
351

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. USE FOR: foundry, azd ai agent, azd provision/deploy, hosted agent scaffold/develop/run/deploy/troubleshoot, prompt agent create, create agent, update agent, add tool to agent, invoke agent, agent.yaml, agent insights, pull agent insights, evaluate agent, batch eval, continuous eval, continuous monitoring, agent CI/CD, optimize prompt, improve prompt, prompt optimizer, optimize agent instructions, Agent Optimizer scaffold, dataset curation from traces, deploy model, model fine-tuning (SFT/DPO/RFT), Foundry project, RBAC, role assignment, permissions, quota, capacity, region, deployment failure, AI Services, create Foundry resource, knowledge index, customize deployment, onboard, availability, training-data, grader, distillation, large file upload. DO NOT USE FOR: Azure Functions, App Service, general Azure deploy (use azure-deploy), general Azure prep (use azure-prepare).

Use this Skill: https://skilld.dev/gh/microsoft/skills/microsoft-foundry

This session only. Nothing lands on disk.

finetuningreferencesreward-hacking.md

≈737 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Reward Hacking Prevention in RFT

What Is Reward Hacking?

The model optimizes for the grader's scoring function rather than the actual task. The training grader becomes a proxy reward that diverges from true quality — the model games the proxy instead of improving.

Core rule: Your training grader MUST produce the same ranking as your evaluation methodology.

If you evaluate with… Then train with… NOT with…
LLM judge (semantic) LLM judge AST / regex / structural matching
Exact match Exact match Fuzzy or partial matching
Unit tests Unit tests Static analysis alone

Misaligned graders are the #1 cause of reward hacking.

Train-Val Gap Thresholds

Train-Val Gap Status Action
≤ 0.05 ✅ Healthy Continue training
0.05–0.10 ⚠️ Warning Monitor closely, check outputs qualitatively
> 0.10 🛑 Stop Stop training — reward hacking is likely

Pre-Training Checklist

  1. Baseline the grader: Run training grader on base model outputs. Record scores as your floor.
  2. Cross-validate graders: If training grader ≠ eval grader, generate 50 outputs, score with both, compute Spearman ρ. Proceed only if ρ ≥ 0.8. If ρ < 0.6, fix alignment first.
  3. Test hackability: Generate 5 intentionally bad outputs that might score well. If grader scores any > 5/10, redesign it.
  4. Set gap threshold: Monitor train-val gap every eval_interval. Stop if > 0.10.

Grader Iteration Loop

When reward hacking is detected:

1. STOP the training run
        ↓
2. COLLECT "hacked" outputs (high train score, low eval score)
        ↓
3. ANALYZE what pattern the model exploited
   (structural mimicry? verbosity? keyword stuffing?)
        ↓
4. UPDATE the grader to penalize that pattern
        ↓
5. RE-BASELINE the updated grader on base model outputs
        ↓
6. RESTART training with the improved grader

Red Flags Checklist

Investigate immediately if any are true:

  • Train-val gap > 0.10
  • Training reward increasing but eval quality stable or declining
  • Model outputs are longer/more verbose than base model
  • Outputs structurally match references but are semantically wrong
  • Different LLM judges disagree on quality
  • Conciseness/style scores dropping while correctness climbs
  • Model produces "template" responses

Key Principles

Principle Action
Align graders Training grader must rank outputs same as eval
Cross-validate first Spearman ρ ≥ 0.8 between training and eval graders
Monitor train-val gap ≤ 0.05 healthy, > 0.10 stop
Test hackability Bad outputs should score < 5/10
Prefer SFT when possible Use RFT only for verifiable-answer tasks
Iterate graders, not models Fix grader before restarting training

Source: SKILL.md on GitHub

2 warnings3d4 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    This skill provides a comprehensive environment for managing the end-to-end lifecycle of AI agents, models, and infrastructure on Microsoft Foundry. It includes sub-skills for deployment, evaluation, fine-tuning, and troubleshooting. The skill utilizes dynamic code execution and shell command wrappers, which are used within the context of local development and cloud orchestration. All external resources and dependencies originate from trusted organizations and well-known services.

  • Socket3d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk3d

    Risk: LOW · No issues

  • Runlayer7mo

    36/36 files flagged

Signed by skilld at 04110d9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last week
metadata
{
  "author": "Microsoft",
  "version": "1.2.26"
}

README badge

README badge for microsoft/skills/microsoft-foundry