All skills
microsoft avatar

/finetuning

@04d245b
by microsoftmicrosoft/skills3.1k stars
351

Fine-tune models on Microsoft Foundry using SFT (supervised), DPO (preference), or RFT (reinforcement with graders). Covers dataset preparation, training job submission, deployment, and evaluation. USE FOR: fine-tune, SFT, DPO, RFT, training data, grader, distillation, fine-tuned model, training job, large file upload, calibrate grader, deploy fine-tuned model, evaluate fine-tuned model. DO NOT USE FOR: general model deployment without fine-tuning (use deploy-model), agent creation (use agents), prompt optimization without training (use prompt-optimizer).

Use this Skill: https://skilld.dev/gh/microsoft/skills/finetuning

This session only. Nothing lands on disk.

referencesagentic-rft.md

≈809 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Agentic RFT — Tool Calling

Train reasoning models (o4-mini) for agentic scenarios where the model invokes external tools during chain-of-thought reasoning.

⚠️ Access required: Agentic RFT with tool calling and GPT-5 RFT are behind feature flags. You must request access through the Microsoft Foundry portal or your Microsoft account team. o4-mini RFT without tools is generally available.

Tool Definition Format

tools = [
    {
        "name": "search",
        "server_url": "https://your-function-app.azurewebsites.net/api/tools",
        "headers": {
            "Authorization": "Bearer <your-key>"
        }
    },
    {
        "name": "get_by_id",
        "server_url": "https://your-function-app.azurewebsites.net/api/tools",
        "headers": {
            "Authorization": "Bearer <your-key>"
        }
    }
]

Submitting an Agentic RFT Job

job = client.fine_tuning.jobs.create(
    model="o4-mini-2025-04-16",
    training_file=train.id,
    validation_file=valid.id,
    method={
        "type": "reinforcement",
        "reinforcement": {
            "grader": grader,
            "tools": tools,
            "max_episode_steps": 10,
            "hyperparameters": {
                "eval_interval": 5,
                "eval_samples": 10,
                "compute_multiplier": 1.5,
                "reasoning_effort": "medium"
            }
        }
    }
)

Tool Response Format

Your tool endpoint must return:

{
    "type": "function_call_output",
    "call_id": "call_12345xyz",
    "output": "The result of the tool call...",
    "id": "fc_12345xyz"
}

Tool Endpoint Requirements

Constraint Limit
Recommended throughput 50 QPS
Max input payload 1 MB
Max return payload 1 MB (413 error if exceeded)
Timeout 10 minutes
Parallel calls Supported — handle race conditions
Retry on 5xx 3 attempts, then rollout discarded
On 4xx Error serialized and shown to model

Infrastructure: Use Always On, sufficient compute (S2+), multiple instances. Under-provisioned endpoints can cause jobs to hang during post-training eval.

RFT Hyperparameters

Parameter Description Recommended Start
reasoning_effort "low", "medium", "high" "medium"
compute_multiplier Scales rollouts per step 1.5
learning_rate_multiplier Scales the learning rate 1.0
n_epochs Data passes 2–3
eval_interval Eval every N steps 5
eval_samples Validation examples per eval 10
max_episode_steps Max tool calls + reasoning steps per rollout 5–10

Notes: Higher LR increases output verbosity without improving accuracy. Compute multiplier 1.5 balances rollout quality and training time. Platform may early-stop before all epochs.

When to Use Agentic RFT

  • Model needs to decide when to call tools (not just follow instructions)
  • Task involves multi-step reasoning with external data lookups
  • Model needs to learn tool selection — choosing the right tool for the job
  • Standard RFT (without tools) can't capture the agentic behavior

Source: SKILL.md on GitHub

2 warnings1mo3 checks · Risk SAFE
  • Gen Agent Trust Hub1mo

    This skill provides a robust toolkit for fine-tuning and evaluating models on Microsoft Foundry. It includes administrative scripts for job management and data processing. There are security considerations regarding the dynamic execution of user-supplied scripts and the invocation of the Azure CLI, which are standard for the skill's intended developer use-case.

  • Socket1mo

    1 alert: gptSecurity

  • Snyk1mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 04d245b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated 2 months ago
metadata
{
  "author": "Microsoft",
  "version": "0.0.0-placeholder"
}

README badge

README badge for microsoft/skills/finetuning