All skills
microsoft avatar

/deploy-model

@54c2a28 official

Unified Azure OpenAI model deployment skill with intelligent intent-based routing. Handles quick preset deployments, fully customized deployments (version/SKU/capacity/RAI policy), and capacity discovery across regions and projects. USE FOR: deploy model, deploy gpt, create deployment, model deployment, deploy openai model, set up model, provision model, find capacity, check model availability, where can I deploy, best region for model, capacity analysis. DO NOT USE FOR: listing existing deployments (use foundry_models_deployments_list MCP tool), deleting deployments, agent creation (use agent/create), project creation (use project/create).

Use this Skill: https://skilld.dev/gh/microsoft/github-copilot-for-azure/deploy-model

This session only. Nothing lands on disk.

customizeEXAMPLES.md

≈1.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

customize Examples

Example 1: Basic Deployment with Defaults

Scenario: Deploy gpt-4o accepting all defaults for quick setup. Config: gpt-4o / GlobalStandard / 10K TPM / Dynamic Quota enabled Result: Deployment gpt-4o created in ~2-3 min with auto-upgrade enabled.

Example 2: Production Deployment with Custom Capacity

Scenario: Deploy gpt-4o for production with high throughput. Config: gpt-4o / GlobalStandard / 50K TPM / Dynamic Quota / Name: gpt-4o-production Result: 50K TPM (500 req/10s). Suitable for moderate-to-high traffic production apps.

Example 3: PTU Deployment for High-Volume Workload

Scenario: Deploy gpt-4o with reserved capacity (PTU) for predictable workload. Config: gpt-4o / ProvisionedManaged / 200 PTU (min 50, max 1000) / Priority Processing enabled PTU sizing: 40K input + 20K output tokens/min → ~100 PTU estimated → 200 PTU recommended (2x headroom) Result: Guaranteed throughput, fixed monthly cost. Use case: customer service bots, document pipelines.

Example 4: Development Deployment with Standard SKU

Scenario: Deploy gpt-4o-mini for dev/testing with minimal cost. Config: gpt-4o-mini / Standard / 1K TPM / Name: gpt-4o-mini-dev Result: 1K TPM, 10 req/10s. Minimal pay-per-use cost for development and prototyping.

Example 5: Spillover Configuration

Scenario: Deploy gpt-4o with spillover to handle peak load overflow. Config: gpt-4o / GlobalStandard / 20K TPM / Dynamic Quota / Spillover → gpt-4o-backup Result: Primary handles up to 20K TPM; overflow auto-redirects to backup deployment.

Example 6: Anthropic Model Deployment (claude-sonnet-4-6)

Scenario: Deploy claude-sonnet-4-6 with customized settings. Config: claude-sonnet-4-6 / GlobalStandard / capacity 1 (MaaS) / Industry: Healthcare / No RAI policy (Anthropic manages content filtering) Result: User selected "Healthcare" as industry → tenant country code (US) and org name fetched automatically → deployed via ARM REST API with modelProviderData in ~2 min.


Comparison Matrix

Scenario Model SKU Capacity Dynamic Quota Priority Spillover Use Case
Ex 1 gpt-4o GlobalStandard 10K TPM ✓ - - Quick setup
Ex 2 gpt-4o GlobalStandard 50K TPM ✓ - - Production
Ex 3 gpt-4o ProvisionedManaged 200 PTU - ✓ - Predictable workload
Ex 4 gpt-4o-mini Standard 1K TPM - - - Dev/testing
Ex 5 gpt-4o GlobalStandard 20K TPM ✓ - ✓ Peak load
Ex 6 claude-sonnet-4-6 GlobalStandard 1 (MaaS) - - - Anthropic model

Common Patterns

Dev → Staging → Production

Stage Model SKU Capacity Extras
Dev gpt-4o-mini Standard 1K TPM —
Staging gpt-4o GlobalStandard 10K TPM —
Production gpt-4o GlobalStandard 50K TPM Dynamic Quota + Spillover

Cost Optimization

  • High priority: gpt-4o, ProvisionedManaged, 100 PTU, Priority Processing
  • Low priority: gpt-4o-mini, Standard, 5K TPM

Tips and Best Practices

Capacity: Start conservative → monitor with Azure Monitor → scale gradually → use spillover for peaks.

SKU Selection: Standard for dev → GlobalStandard + dynamic quota for variable production → ProvisionedManaged (PTU) for predictable load.

Cost: Right-size capacity; use gpt-4o-mini where possible (80-90% accuracy at lower cost); enable dynamic quota; consider PTU for consistent high-volume.

Versions: Auto-upgrade recommended; test new versions in staging first; pin only if compatibility requires it.

Content Filtering: Start with DefaultV2; use custom policies only for specific needs; monitor filtered requests.


Troubleshooting

Problem Solution
QuotaExceeded Check usage with az cognitiveservices usage list, reduce capacity, try different SKU, check other regions, or use the quota skill to request an increase
Version not available for SKU Check az cognitiveservices account list-models --query "[?name=='gpt-4o'].version", use latest
Deployment name exists Skill auto-generates unique name (e.g., gpt-4o-2), or specify custom name

Source: SKILL.md on GitHub

No alerts6mo3 checks · Risk SAFE
  • Gen Agent Trust Hub6mo

    This skill provides a unified workflow for deploying Azure OpenAI models using the Azure CLI. It includes intelligent routing, capacity discovery across regions, and support for both standard OpenAI models and Anthropic models on Azure while maintaining security best practices like mandatory project confirmation.

  • Socket6mo

    No alerts

  • Snyk6mo

    Risk: LOW · No issues

Signed by skilld at 54c2a28. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 17 hours ago.

Activeupdated 2 months ago
metadata
{
  "author": "Microsoft",
  "version": "1.0.0"
}
  • azure
  • openai
  • deployment
  • model
  • capacity
  • sku
  • foundry
  • azure-cli
  • provisioning

README badge

README badge for microsoft/github-copilot-for-azure/deploy-model

Deploys Azure OpenAI models with intent-based routing to preset deployment, customized configuration, or capacity discovery workflows. Routes user requests to the appropriate mode based on whether they want quick defaults, custom SKU/capacity/RAI settings, or to find available capacity across regions and projects.

Generated from the current SKILL.md.

Does this skill handle deployments in azd-managed Foundry projects?
No. For azd projects (those scaffolded from azd-ai-starter-basic or via azd ai agent init), declare deployments in azure.yaml instead — azd provision will create them through Bicep. Use this skill only for standalone Foundry projects or ad-hoc deployments outside the azd lifecycle.
What happens if I don't specify a project?
The skill checks the PROJECT_RESOURCE_ID environment variable first, then looks for clues in your prompt. If neither exists, it queries your projects and suggests the current one, with a confirmation step before deploying.
Can this skill list or delete existing deployments?
No. This skill creates deployments only. Use the foundry_models_deployments_list MCP tool to list existing deployments, or the Azure portal to delete them.
Does this skill validate quota and SKU support before deploying?
Yes. It queries the model catalog to confirm the model supports your chosen SKU, and checks your subscription's available quota via Azure CLI before presenting any deployment options.
What should I do if I hit a quota limit?
Defer to the quota skill (quota/quota.md) for quota increase requests, usage monitoring, and troubleshooting quota errors.

Generated from the current SKILL.md. These answers refresh after source changes.