All skills
microsoft avatar

/finetuning

@04d245b
by microsoftmicrosoft/skills3.1k stars
351

Fine-tune models on Microsoft Foundry using SFT (supervised), DPO (preference), or RFT (reinforcement with graders). Covers dataset preparation, training job submission, deployment, and evaluation. USE FOR: fine-tune, SFT, DPO, RFT, training data, grader, distillation, fine-tuned model, training job, large file upload, calibrate grader, deploy fine-tuned model, evaluate fine-tuned model. DO NOT USE FOR: general model deployment without fine-tuning (use deploy-model), agent creation (use agents), prompt optimization without training (use prompt-optimizer).

Use this Skill: https://skilld.dev/gh/microsoft/skills/finetuning

This session only. Nothing lands on disk.

referencestraining-types.md

≈791 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Training Types: SFT vs DPO vs RFT

Decision Matrix

Factor SFT DPO RFT
Best for Teaching a new skill or format Aligning preferences/style Improving reasoning chains
Data needed Input–output pairs Chosen/rejected pairs Prompts + grading function
Data volume 50–5,000 examples 500–5,000 pairs 200–2,000 prompts
Effort to prepare data Low High (need contrasting pairs) Medium (need grader, not outputs)
Risk of regression Low Medium High (sensitive to grader quality)
Typical improvement 5–30% on task metrics Subtle style/safety shifts 0–15% on reasoning tasks
Supported models Most models Select models o4-mini

When to Use Each

SFT (Supervised Fine-Tuning)

  • You have high-quality input–output pairs
  • Task is well-defined (code generation, classification, extraction, summarization)
  • You want reliable, repeatable outputs in a specific format or style
  • Key insight: 300–500 high-quality examples often outperforms 1,500+ lower-quality ones

DPO (Direct Preference Optimization)

  • You want to adjust tone, verbosity, safety, or style
  • You have examples of "good" and "bad" outputs for the same input
  • SFT already works but outputs need refinement
  • DPO-specific params: beta (default 0.1), l2_multiplier (default 0.1)

RFT (Reinforcement Fine-Tuning)

  • Task has objectively verifiable answers (code execution, math, logic)
  • You can write a programmatic or LLM-based grader
  • You want to improve the model's reasoning, not just its outputs
  • Critical: RFT is extremely sensitive to grader quality. Train–val gap should be ≤ 0.05.

Choosing a Path

├─ Do you have labeled input–output pairs?
│  ├─ Yes → SFT
│  └─ No
│     ├─ Can you write a grading function? → RFT
│     └─ Can you rank "good" vs "bad" outputs? → DPO
│
After SFT:
├─ Results good enough? → Ship it
├─ Need style refinement? → DPO on top of SFT model
└─ Reasoning needs improvement? → RFT (if model supports it)

Model Compatibility (Microsoft Foundry)

Model SFT DPO RFT Vision FT
gpt-4.1 ✅ ✅ ❌ ✅
gpt-4.1-mini ✅ ❌ ❌ ❌
gpt-4.1-nano ✅ ❌ ❌ ❌
gpt-4o (2024-08-06) ✅ ✅ ❌ ✅
gpt-4o-mini ✅ ❌ ❌ ❌
o4-mini ❌ ❌ ✅ ❌
gpt-5 ❌ ❌ ✅ ⚠️ ❌
gpt-oss-20b ✅ ❌ ❌ ❌
Ministral-3B ✅ ❌ ❌ ❌
Llama-3.3-70B ✅ ❌ ❌ ❌
Qwen-3-32B ✅ ❌ ❌ ❌

DPO can be applied on top of an already SFT-fine-tuned model. Vision fine-tuning follows the same SFT workflow but with image data in messages.

⚠️ Feature flags: GPT-5 RFT and agentic RFT with tool calling require access requests. Contact your Microsoft account team or request access through the Microsoft Foundry portal. o4-mini RFT without tools is generally available.

Check Microsoft Foundry docs for the latest model availability.

Source: SKILL.md on GitHub

2 warnings1mo3 checks · Risk SAFE
  • Gen Agent Trust Hub1mo

    This skill provides a robust toolkit for fine-tuning and evaluating models on Microsoft Foundry. It includes administrative scripts for job management and data processing. There are security considerations regarding the dynamic execution of user-supplied scripts and the invocation of the Azure CLI, which are standard for the skill's intended developer use-case.

  • Socket1mo

    1 alert: gptSecurity

  • Snyk1mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 04d245b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated 2 months ago
metadata
{
  "author": "Microsoft",
  "version": "0.0.0-placeholder"
}

README badge

README badge for microsoft/skills/finetuning