All skills
microsoft avatar

/microsoft-foundry

@04110d9
by microsoftmicrosoft/skills3.1k stars
351

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. USE FOR: foundry, azd ai agent, azd provision/deploy, hosted agent scaffold/develop/run/deploy/troubleshoot, prompt agent create, create agent, update agent, add tool to agent, invoke agent, agent.yaml, agent insights, pull agent insights, evaluate agent, batch eval, continuous eval, continuous monitoring, agent CI/CD, optimize prompt, improve prompt, prompt optimizer, optimize agent instructions, Agent Optimizer scaffold, dataset curation from traces, deploy model, model fine-tuning (SFT/DPO/RFT), Foundry project, RBAC, role assignment, permissions, quota, capacity, region, deployment failure, AI Services, create Foundry resource, knowledge index, customize deployment, onboard, availability, training-data, grader, distillation, large file upload. DO NOT USE FOR: Azure Functions, App Service, general Azure deploy (use azure-deploy), general Azure prep (use azure-prepare).

Use this Skill: https://skilld.dev/gh/microsoft/skills/microsoft-foundry

This session only. Nothing lands on disk.

quotaquota.md

≈2.3k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Microsoft Foundry Quota Management

Quota and capacity management for Microsoft Foundry. Quotas are subscription + region level.

⚠️ Important: This is the authoritative skill for all Foundry quota operations. When a user asks about quota, capacity, TPM, PTU, quota errors, or deployment limits, always invoke this skill rather than using MCP tools (azure-quota, azure-documentation, azure-foundry) directly. This skill provides structured workflows and error handling that direct tool calls lack.

Important: All quota operations are control plane (management) operations. Use Azure CLI commands (az cognitiveservices, az rest, az ai) as the primary method.

Quota Types

Type Description
TPM Tokens Per Minute, pay-per-token, subject to rate limits
PTU Provisioned Throughput Units, monthly commitment, no rate limits
Region Max capacity per region, shared across subscription
Slots 10-20 deployment slots per resource

When to use PTU: Consistent high-volume production workloads where monthly commitment is cost-effective.


Use this sub-skill when the user needs to:

  • View quota usage — check current TPM/PTU allocation and available capacity
  • Check quota limits — show quota limits for a subscription, region, or model
  • Find optimal regions — compare quota availability across regions for deployment
  • Plan deployments — verify sufficient quota before deploying models
  • Request quota increases — navigate quota increase process through Azure Portal
  • Troubleshoot deployment failures — diagnose QuotaExceeded, InsufficientQuota, DeploymentLimitReached, 429 rate limit errors
  • Optimize allocation — monitor and consolidate quota across deployments
  • Monitor quota across deployments — track capacity by model and region
  • Explain quota concepts — explain TPM, PTU, capacity units, regional quotas
  • Free up quota — identify and delete unused deployments

Key Points:

  1. Isolated by region (East US ≠ West US)
  2. Regional capacity varies by model
  3. Multi-region enables failover and load distribution
  4. Quota requests specify target region

See detailed guide.


Core Workflows

1. Check Regional Quota

subId=$(az account show --query id -o tsv)
az rest --method get \
  --url "https://management.azure.com/subscriptions/$subId/providers/Microsoft.CognitiveServices/locations/eastus/usages?api-version=2023-05-01" \
  --query "value[?contains(name.value,'OpenAI')].{Model:name.value, Used:currentValue, Limit:limit}" -o table

Output interpretation:

  • Used: Current TPM consumed (10000 = 10K TPM)
  • Limit: Maximum TPM quota (15000 = 15K TPM)
  • Available: Limit - Used (5K TPM available)

Change region: eastus, eastus2, westus, westus2, swedencentral, uksouth.


2. Find Best Region for Deployment

Check specific regions for available quota:

subId=$(az account show --query id -o tsv)
region="eastus"
az rest --method get \
  --url "https://management.azure.com/subscriptions/$subId/providers/Microsoft.CognitiveServices/locations/$region/usages?api-version=2023-05-01" \
  --query "value[?name.value=='OpenAI.Standard.gpt-4o'].{Model:name.value, Used:currentValue, Limit:limit, Available:(limit-currentValue)}" -o table

See workflows reference for multi-region comparison.


3. Check Quota Before Deployment

Verify available quota for your target model:

subId=$(az account show --query id -o tsv)
region="eastus"
model="OpenAI.Standard.gpt-4o"

az rest --method get \
  --url "https://management.azure.com/subscriptions/$subId/providers/Microsoft.CognitiveServices/locations/$region/usages?api-version=2023-05-01" \
  --query "value[?name.value=='$model'].{Model:name.value, Used:currentValue, Limit:limit, Available:(limit-currentValue)}" -o table
  • Available > 0: Yes, you have quota
  • Available = 0: Delete unused deployments or try different region

4. Monitor Quota by Model

Show quota allocation grouped by model:

subId=$(az account show --query id -o tsv)
region="eastus"
az rest --method get \
  --url "https://management.azure.com/subscriptions/$subId/providers/Microsoft.CognitiveServices/locations/$region/usages?api-version=2023-05-01" \
  --query "value[?contains(name.value,'OpenAI')].{Model:name.value, Used:currentValue, Limit:limit, Available:(limit-currentValue)}" -o table

Shows aggregate usage across ALL deployments by model type.

Optional: List individual deployments:

  • Azure MCP tool: Use model_deployment_get to query deployments in a Foundry project
  • Azure CLI:
az cognitiveservices account list --query "[?kind=='AIServices'].{Name:name,RG:resourceGroup}" -o table

az cognitiveservices account deployment list --name <resource> --resource-group <rg> \
  --query "[].{Name:name,Model:properties.model.name,Capacity:sku.capacity}" -o table

5. Delete Deployment (Free Quota)

az cognitiveservices account deployment delete --name <resource> --resource-group <rg> \
  --deployment-name <deployment>

Quota freed immediately. Re-run Workflow #1 to verify.


6. Request Quota Increase

Azure Portal Process:

  1. Navigate to Azure Portal - All Resources → Filter "AI Services" → Click resource
  2. Select Quotas in left navigation
  3. Click Request quota increase
  4. Fill form: Model, Current Limit, Requested Limit, Region, Business Justification (required field)
  5. Wait for approval: 3-5 business days typically, up to 10 business days (source)

Business Justification is a mandatory field that explains why you need more quota. Azure reviews each request to ensure resources are allocated based on legitimate business needs. A strong justification includes:

  • Workload details: What you're building and which model you need
  • Data-driven estimates: Expected traffic volume and token usage calculations
  • Clear need: Why current quota is insufficient and what capacity you require
  • Timeline: When you need the increased quota (e.g., production launch date)

Business Justification template:

Production [workload type] using [model] in [region].
Expected traffic: [X requests/day] with [Y tokens/request].
Calculated required TPM: [Z TPM]. Current [N TPM] insufficient.
Request increase to [M TPM]. Deployment target: [date].

See detailed quota request guide for complete steps.


Quick Troubleshooting

Error Quick Fix Detailed Guide
QuotaExceeded Delete unused deployments or request increase Error Resolution
InsufficientQuota Reduce capacity or try different region Error Resolution
DeploymentLimitReached Delete unused deployments (10-20 slot limit) Error Resolution
429 Rate Limit Increase TPM or migrate to PTU Error Resolution

References

Detailed Guides:

Official Microsoft Documentation:

Calculators:

  • Azure Pricing Calculator - Official pricing estimator
  • Microsoft Foundry PTU calculator (Microsoft Foundry → Operate → Quota → Provisioned Throughput Unit tab) - PTU capacity sizing

Source: SKILL.md on GitHub

2 warnings3d4 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    This skill provides a comprehensive environment for managing the end-to-end lifecycle of AI agents, models, and infrastructure on Microsoft Foundry. It includes sub-skills for deployment, evaluation, fine-tuning, and troubleshooting. The skill utilizes dynamic code execution and shell command wrappers, which are used within the context of local development and cloud orchestration. All external resources and dependencies originate from trusted organizations and well-known services.

  • Socket3d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk3d

    Risk: LOW · No issues

  • Runlayer7mo

    36/36 files flagged

Signed by skilld at 04110d9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last week
metadata
{
  "author": "Microsoft",
  "version": "1.2.26"
}

README badge

README badge for microsoft/skills/microsoft-foundry