All skills
microsoft avatar

/microsoft-foundry

@04110d9
by microsoftmicrosoft/skills3.1k stars
351

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. USE FOR: foundry, azd ai agent, azd provision/deploy, hosted agent scaffold/develop/run/deploy/troubleshoot, prompt agent create, create agent, update agent, add tool to agent, invoke agent, agent.yaml, agent insights, pull agent insights, evaluate agent, batch eval, continuous eval, continuous monitoring, agent CI/CD, optimize prompt, improve prompt, prompt optimizer, optimize agent instructions, Agent Optimizer scaffold, dataset curation from traces, deploy model, model fine-tuning (SFT/DPO/RFT), Foundry project, RBAC, role assignment, permissions, quota, capacity, region, deployment failure, AI Services, create Foundry resource, knowledge index, customize deployment, onboard, availability, training-data, grader, distillation, large file upload. DO NOT USE FOR: Azure Functions, App Service, general Azure deploy (use azure-deploy), general Azure prep (use azure-prepare).

Use this Skill: https://skilld.dev/gh/microsoft/skills/microsoft-foundry

This session only. Nothing lands on disk.

finetuningreferencesplatform-gotchas.md

≈518 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Platform Gotchas — Top 10

  1. OSS models require "trainingType": "globalStandard" in the request body — undocumented, and all OSS FT jobs fail without it.

  2. Model catalog fine_tune flag is wrong for OSS models — API returns fine_tune = false for all OSS models despite being FT-supported. Hardcode the supported list.

  3. Older SDK versions may fail on /v1/ project endpoints — client.files.create() throws "API version not supported" with older openai package versions. Upgrade to openai>=1.0 and use the /v1/ project endpoint (preferred). If you must use an older SDK, fall back to REST API with the non-project /openai/ endpoint.

  4. ARM "Succeeded" doesn't mean deployment is ready — provisioningState: Succeeded but data plane returns DeploymentNotReady indefinitely. Delete and recreate the deployment, then wait ~5 minutes.

  5. OSS FT deployments may fail with InternalServerError — use the correct provider-specific model.format (e.g., "Mistral AI" not "OpenAI") and try capacity=100.

  6. OSS FT inference hits "Failed to load LoRA" intermittently — deploy with capacity ≥ 100, use 8+ retries with exponential backoff, and wait 2+ minutes after deployment before first call.

  7. ARM REST and az cognitiveservices use different format strings for OSS models — ARM uses provider names ("Microsoft", "Meta"), CLI uses "OpenAI-OSS" for all OSS. Mixing them produces HTTP 500.

  8. Content safety false positives on entity extraction data — PII-dense data (medical records, legal docs, resumes) can trigger "Hate/Fairness" blocks at deployment time. Remove problematic document types.

  9. FT deployments at capacity=1 are severely rate-limited (~1 RPM) — evaluating 10 samples takes ~10 minutes. Use capacity ≥ 100 for eval workloads and exponential backoff.

  10. Wrong resource endpoint is a silent killer — jobs submitted to the wrong Foundry resource succeed via API but don't appear in the portal. Always verify the endpoint matches your Foundry project.

Source: SKILL.md on GitHub

2 warnings3d4 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    This skill provides a comprehensive environment for managing the end-to-end lifecycle of AI agents, models, and infrastructure on Microsoft Foundry. It includes sub-skills for deployment, evaluation, fine-tuning, and troubleshooting. The skill utilizes dynamic code execution and shell command wrappers, which are used within the context of local development and cloud orchestration. All external resources and dependencies originate from trusted organizations and well-known services.

  • Socket3d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk3d

    Risk: LOW · No issues

  • Runlayer7mo

    36/36 files flagged

Signed by skilld at 04110d9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated last week
metadata
{
  "author": "Microsoft",
  "version": "1.2.26"
}

README badge

README badge for microsoft/skills/microsoft-foundry