All skills
microsoft avatar

/microsoft-foundry

@04110d9
by microsoftmicrosoft/skills3.1k stars
351

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. USE FOR: foundry, azd ai agent, azd provision/deploy, hosted agent scaffold/develop/run/deploy/troubleshoot, prompt agent create, create agent, update agent, add tool to agent, invoke agent, agent.yaml, agent insights, pull agent insights, evaluate agent, batch eval, continuous eval, continuous monitoring, agent CI/CD, optimize prompt, improve prompt, prompt optimizer, optimize agent instructions, Agent Optimizer scaffold, dataset curation from traces, deploy model, model fine-tuning (SFT/DPO/RFT), Foundry project, RBAC, role assignment, permissions, quota, capacity, region, deployment failure, AI Services, create Foundry resource, knowledge index, customize deployment, onboard, availability, training-data, grader, distillation, large file upload. DO NOT USE FOR: Azure Functions, App Service, general Azure deploy (use azure-deploy), general Azure prep (use azure-prepare).

Use this Skill: https://skilld.dev/gh/microsoft/skills/microsoft-foundry

This session only. Nothing lands on disk.

finetuningreferencesgrader-design.md

≈860 tokens on demand. Your agent reads this file only when SKILL.md points to it.

RFT Grader Design Guide

Grader Type Selection

Grader Type Best For Tradeoffs
Python grader (default) Most tasks incl. tool-calling. Accesses output_text and output_tools. Can't call external APIs or execute code.
Multi grader Combining multiple scoring dimensions. score_model component adds LLM cost per rollout.
Endpoint grader Tasks requiring external API calls (test suites, DB queries). HTTP latency, scaling risk. Under-provisioned endpoints can hang jobs.
String check Exact-match tasks (classification, yes/no, numeric). Binary 0/1 only — no partial credit.

Start with Python grader unless you need external API calls. Python graders are fast, deterministic, reliable, and tool-aware (sample.output_tools provides tool call metadata).

Partial Credit Pattern

Binary pass/fail gives sparse reward. Decompose into 2–4 scored dimensions:

def grade(sample, item):
    output_text = sample.get("output_text", "") or ""
    expected = item.get("expected_answer", "")
    
    score = 0.0
    
    # Core correctness (highest weight)
    if correct_action(output_text, expected):
        score += 0.4
    
    # Precision (exact amounts, specific values)
    score += 0.3 * precision_score(output_text, expected)
    
    # Reasoning quality (cited correct rules/facts)
    score += 0.2 * reasoning_score(output_text, expected)
    
    # Process quality (used the right tools)
    if used_correct_tools(sample.get("output_tools", [])):
        score += 0.1
    
    return round(min(score, 1.0), 3)

Weight Guidelines

Dimension Typical Weight Examples
Core correctness 0.3–0.5 Right action/answer/classification
Precision 0.2–0.3 Exact amounts, correct format
Reasoning 0.1–0.2 Cited correct rules, justified decision
Process quality 0.05–0.1 Used right tools, followed steps

Threshold Calibration Workflow

The pass_threshold determines what score counts as pass vs fail — the most important RFT hyperparameter.

  1. Run the base model on your training/validation set
  2. Score every output with your grader
  3. Compute pass rates at multiple thresholds:
for threshold in [0.5, 0.6, 0.7, 0.8, 0.85, 0.9, 0.95]:
    pass_rate = sum(1 for s in scores if s >= threshold) / len(scores)
    print(f"  @{threshold}: pass={pass_rate:.0%}, fail={1 - pass_rate:.0%}")
  1. Choose where 25–50% of base model rollouts fail:
Failure Rate Signal Quality
< 10% ❌ Too easy — no learning signal
10–25% ⚠️ Weak signal
25–50% ✅ Good — enough failures to learn from
50–70% ⚠️ Harsh — mostly negative reward
> 70% ❌ Too hard — training may diverge

Always re-run calibration when you change your dataset.

Consistency Rules

When using multiple graders (Python for training, endpoint for debugging, local script for eval):

  1. Identical scoring logic — same weights, keywords, dimension breakdown
  2. Identical default scores — same behavior when no action found, no amounts expected
  3. Test with same examples — run 10 samples through all graders and verify scores match

Mismatched scoring causes the model to learn different behavior than what your evaluation measures.

Source: SKILL.md on GitHub

2 warnings3d4 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    This skill provides a comprehensive environment for managing the end-to-end lifecycle of AI agents, models, and infrastructure on Microsoft Foundry. It includes sub-skills for deployment, evaluation, fine-tuning, and troubleshooting. The skill utilizes dynamic code execution and shell command wrappers, which are used within the context of local development and cloud orchestration. All external resources and dependencies originate from trusted organizations and well-known services.

  • Socket3d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk3d

    Risk: LOW · No issues

  • Runlayer7mo

    36/36 files flagged

Signed by skilld at 04110d9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated last week
metadata
{
  "author": "Microsoft",
  "version": "1.2.26"
}

README badge

README badge for microsoft/skills/microsoft-foundry