All skills
microsoft avatar

/microsoft-foundry

@04110d9
by microsoftmicrosoft/skills3.1k stars
351

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. USE FOR: foundry, azd ai agent, azd provision/deploy, hosted agent scaffold/develop/run/deploy/troubleshoot, prompt agent create, create agent, update agent, add tool to agent, invoke agent, agent.yaml, agent insights, pull agent insights, evaluate agent, batch eval, continuous eval, continuous monitoring, agent CI/CD, optimize prompt, improve prompt, prompt optimizer, optimize agent instructions, Agent Optimizer scaffold, dataset curation from traces, deploy model, model fine-tuning (SFT/DPO/RFT), Foundry project, RBAC, role assignment, permissions, quota, capacity, region, deployment failure, AI Services, create Foundry resource, knowledge index, customize deployment, onboard, availability, training-data, grader, distillation, large file upload. DO NOT USE FOR: Azure Functions, App Service, general Azure deploy (use azure-deploy), general Azure prep (use azure-prepare).

Use this Skill: https://skilld.dev/gh/microsoft/skills/microsoft-foundry

This session only. Nothing lands on disk.

modelsdeploy-modelcustomizereferencescustomize-guides.md

≈872 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Customize Guides — Selection Guides & Advanced Topics

Reference for: models/deploy-model/customize/SKILL.md

Table of Contents: Selection Guides · Advanced Topics

Selection Guides

How to Choose SKU

SKU Best For Cost Availability
GlobalStandard Production, high availability Medium Multi-region
Standard Development, testing Low Single region
ProvisionedManaged High-volume, predictable workloads Fixed (PTU) Reserved capacity
DataZoneStandard Data residency requirements Medium Specific zones

Decision Tree:

Do you need guaranteed throughput?
├─ Yes → ProvisionedManaged (PTU)
└─ No → Do you need high availability?
        ├─ Yes → GlobalStandard
        └─ No → Standard

How to Choose Capacity

For TPM-based SKUs (GlobalStandard, Standard):

Workload Recommended Capacity
Development/Testing 1K - 5K TPM
Small Production 5K - 20K TPM
Medium Production 20K - 100K TPM
Large Production 100K+ TPM

For PTU-based SKUs (ProvisionedManaged):

Use the PTU calculator based on:

  • Input tokens per minute
  • Output tokens per minute
  • Requests per minute

Capacity Planning Tips:

  • Start with recommended capacity
  • Monitor usage and adjust
  • Enable dynamic quota for flexibility
  • Consider spillover for peak loads

How to Choose RAI Policy

Policy Filtering Level Use Case
Microsoft.DefaultV2 Balanced Most applications
Microsoft.Prompt-Shield Enhanced Security-sensitive apps
Custom Configurable Specific requirements

Recommendation: Start with Microsoft.DefaultV2 and adjust based on application needs.


Advanced Topics

PTU (Provisioned Throughput Units) Deployments

What is PTU?

  • Reserved capacity with guaranteed throughput
  • Measured in PTU units, not TPM
  • Fixed cost regardless of usage
  • Best for high-volume, predictable workloads

PTU Calculator:

Estimated PTU = (Input TPM × 0.001) + (Output TPM × 0.002) + (Requests/min × 0.1)

Example:
- Input: 10,000 tokens/min
- Output: 5,000 tokens/min
- Requests: 100/min

PTU = (10,000 × 0.001) + (5,000 × 0.002) + (100 × 0.1)
    = 10 + 10 + 10
    = 30 PTU

PTU Deployment:

az cognitiveservices account deployment create \
  --name <account-name> \
  --resource-group <resource-group> \
  --deployment-name <deployment-name> \
  --model-name <model-name> \
  --model-version <version> \
  --model-format "OpenAI" \
  --sku-name "ProvisionedManaged" \
  --sku-capacity 100  # PTU units

Spillover Configuration

Spillover Workflow:

  1. Primary deployment receives requests
  2. When capacity reached, requests overflow to spillover target
  3. Spillover target must be same model or compatible
  4. Configure via deployment properties

Best Practices:

  • Use spillover for peak load handling
  • Spillover target should have sufficient capacity
  • Monitor both deployments
  • Test failover behavior

Priority Processing

What is Priority Processing?

  • Prioritizes your requests during high load
  • Available for ProvisionedManaged SKU
  • Additional charges apply
  • Ensures consistent performance

When to Use:

  • Mission-critical applications
  • SLA requirements
  • High-concurrency scenarios

Source: SKILL.md on GitHub

2 warnings3d4 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    This skill provides a comprehensive environment for managing the end-to-end lifecycle of AI agents, models, and infrastructure on Microsoft Foundry. It includes sub-skills for deployment, evaluation, fine-tuning, and troubleshooting. The skill utilizes dynamic code execution and shell command wrappers, which are used within the context of local development and cloud orchestration. All external resources and dependencies originate from trusted organizations and well-known services.

  • Socket3d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk3d

    Risk: LOW · No issues

  • Runlayer7mo

    36/36 files flagged

Signed by skilld at 04110d9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 19 hours ago.

Activeupdated last week
metadata
{
  "author": "Microsoft",
  "version": "1.2.26"
}

README badge

README badge for microsoft/skills/microsoft-foundry