All skills
microsoft avatar

/microsoft-foundry

@04110d9
by microsoftmicrosoft/skills3.1k stars
351

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. USE FOR: foundry, azd ai agent, azd provision/deploy, hosted agent scaffold/develop/run/deploy/troubleshoot, prompt agent create, create agent, update agent, add tool to agent, invoke agent, agent.yaml, agent insights, pull agent insights, evaluate agent, batch eval, continuous eval, continuous monitoring, agent CI/CD, optimize prompt, improve prompt, prompt optimizer, optimize agent instructions, Agent Optimizer scaffold, dataset curation from traces, deploy model, model fine-tuning (SFT/DPO/RFT), Foundry project, RBAC, role assignment, permissions, quota, capacity, region, deployment failure, AI Services, create Foundry resource, knowledge index, customize deployment, onboard, availability, training-data, grader, distillation, large file upload. DO NOT USE FOR: Azure Functions, App Service, general Azure deploy (use azure-deploy), general Azure prep (use azure-prepare).

Use this Skill: https://skilld.dev/gh/microsoft/skills/microsoft-foundry

This session only. Nothing lands on disk.

finetuningreferencesdeployment.md

≈876 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Deployment Formats

Model Format and SKU Mapping

Base model family model.format sku.name Endpoint type
gpt-4.1-mini "OpenAI" "Standard" Project
gpt-4.1-nano "OpenAI" "Standard" Project
o4-mini (RFT) "OpenAI" "Standard" Project
gpt-oss-20b "Microsoft" "GlobalStandard" Cognitive Services
Ministral-3B "Mistral AI" "GlobalStandard" Cognitive Services
Llama-3.3-70B "Meta" "GlobalStandard" Cognitive Services
Qwen-3-32B "Alibaba" "GlobalStandard" Cognitive Services

Format strings are case-sensitive. "Mistral AI" works; "mistral" does not.

Two Endpoint Types

Project Endpoint (OpenAI models): https://<resource>.services.ai.azure.com/api/projects/<project>/openai/v1/

  • Use openai.OpenAI(base_url=..., api_key=...) — NOT AzureOpenAI

Cognitive Services Endpoint (OSS models): https://<resource>.cognitiveservices.azure.com/openai/deployments/<name>/chat/completions?api-version=2025-04-01-preview

  • Use openai.AzureOpenAI(azure_endpoint=..., api_key=..., api_version=...)

CLI Deployment (az cognitiveservices)

The CLI uses different format strings than the ARM REST API for OSS models:

az cognitiveservices account deployment create \
  --name <resource> \
  --resource-group <rg> \
  --deployment-name <name> \
  --model-name <model> \
  --model-version "1" \
  --model-format "OpenAI-OSS" \
  --sku-capacity 100 \
  --sku-name "GlobalStandard"
Base model family ARM REST model.format CLI --model-format
gpt-4.1-mini/nano "OpenAI" "OpenAI"
gpt-oss-20b "Microsoft" "OpenAI-OSS"
Ministral-3B "Mistral AI" "OpenAI-OSS"
Llama-3.3-70B "Meta" "OpenAI-OSS"
Qwen-3-32B "Alibaba" "OpenAI-OSS"

⚠️ Using "OpenAI-OSS" in ARM REST or "Microsoft" in CLI will fail with HTTP 500.

ARM REST API Deployment

PUT https://management.azure.com/subscriptions/{sub_id}/resourceGroups/{rg}/providers/Microsoft.CognitiveServices/accounts/{account}/deployments/{deploy_name}?api-version=2024-10-01
{
  "sku": { "name": "GlobalStandard", "capacity": 100 },
  "properties": {
    "model": {
      "format": "Microsoft",
      "name": "gpt-oss-20b.ft-{jobid}-suffix",
      "version": "1"
    }
  }
}

ARM token: az account get-access-token --query accessToken -o tsv (expires ~60min).

Capacity Notes

  • Capacity = tokens-per-minute in thousands. 100 = 100K TPM.
  • Set capacity ≥ 100 for eval workloads. At capacity=1, OSS FT models hit "Failed to load LoRA" errors.
  • Quota is per-resource. After deleting a deployment, wait 15–20s before creating a new one.
  • Deployment names: max 64 chars, alphanumeric + hyphens, unique within resource.

Common Deployment Errors

Error Cause Fix
HTTP 500, no message Wrong model.format Check format table above
HTTP 409, deployment exists Name collision Use unique deployment name
HTTP 403 ARM token expired Refresh token
HTTP 400, "api-version not allowed" AzureOpenAI client on /v1/ endpoint Switch to openai.OpenAI
HTTP 429, quota exceeded Too many deployments Delete unused, wait 20s
ProvisioningState: Failed Model not available in region Try different region

Source: SKILL.md on GitHub

2 warnings3d4 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    This skill provides a comprehensive environment for managing the end-to-end lifecycle of AI agents, models, and infrastructure on Microsoft Foundry. It includes sub-skills for deployment, evaluation, fine-tuning, and troubleshooting. The skill utilizes dynamic code execution and shell command wrappers, which are used within the context of local development and cloud orchestration. All external resources and dependencies originate from trusted organizations and well-known services.

  • Socket3d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk3d

    Risk: LOW · No issues

  • Runlayer7mo

    36/36 files flagged

Signed by skilld at 04110d9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated last week
metadata
{
  "author": "Microsoft",
  "version": "1.2.26"
}

README badge

README badge for microsoft/skills/microsoft-foundry