All skills
microsoft avatar

/deploy-model

@54c2a28 official

Unified Azure OpenAI model deployment skill with intelligent intent-based routing. Handles quick preset deployments, fully customized deployments (version/SKU/capacity/RAI policy), and capacity discovery across regions and projects. USE FOR: deploy model, deploy gpt, create deployment, model deployment, deploy openai model, set up model, provision model, find capacity, check model availability, where can I deploy, best region for model, capacity analysis. DO NOT USE FOR: listing existing deployments (use foundry_models_deployments_list MCP tool), deleting deployments, agent creation (use agent/create), project creation (use project/create).

Use this Skill: https://skilld.dev/gh/microsoft/github-copilot-for-azure/deploy-model

This session only. Nothing lands on disk.

customizeSKILL.md

≈2.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Customize Model Deployment

Interactive guided workflow for deploying Azure OpenAI models with full customization control over version, SKU, capacity, content filtering, and advanced options.

Quick Reference

Property Description
Flow Interactive step-by-step guided deployment
Customization Version, SKU, Capacity, RAI Policy, Advanced Options
SKU Support GlobalStandard, Standard, ProvisionedManaged, DataZoneStandard
Best For Precise control over deployment configuration
Authentication Azure CLI (az login)
Tools Azure CLI, MCP tools (optional)

When to Use This Skill

Use this skill when you need precise control over deployment configuration:

  • ✅ Choose specific model version (not just latest)
  • ✅ Select deployment SKU (GlobalStandard vs Standard vs PTU)
  • ✅ Set exact capacity within available range
  • ✅ Configure content filtering (RAI policy selection)
  • ✅ Enable advanced features (dynamic quota, priority processing, spillover)
  • ✅ PTU deployments (Provisioned Throughput Units)

Alternative: Use preset for quick deployment to the best available region with automatic configuration.

Comparison: customize vs preset

Feature customize preset
Focus Full customization control Optimal region selection
Version Selection User chooses from available Uses latest automatically
SKU Selection User chooses (GlobalStandard/Standard/PTU) GlobalStandard only
Capacity User specifies exact value Auto-calculated (50% of available)
RAI Policy User selects from options Default policy only
Region Current region first, falls back to all regions if no capacity Checks capacity across all regions upfront
Use Case Precise deployment requirements Quick deployment to best region

Prerequisites

  • Azure subscription with Cognitive Services Contributor or Owner role
  • Microsoft Foundry project resource ID (format: /subscriptions/{sub}/resourceGroups/{rg}/providers/Microsoft.CognitiveServices/accounts/{account}/projects/{project})
  • Azure CLI installed and authenticated (az login)
  • Optional: Set PROJECT_RESOURCE_ID environment variable

Workflow Overview

Complete Flow (14 Phases)

1. Verify Authentication
2. Get Project Resource ID
3. Verify Project Exists
4. Get Model Name (if not provided)
5. List Model Versions → User Selects
6. List SKUs for Version → User Selects
7. Get Capacity Range → User Configures
   7b. If no capacity: Cross-Region Fallback → Query all regions → User selects region/project
8. List RAI Policies → User Selects
9. Configure Advanced Options (if applicable)
10. Configure Version Upgrade Policy
11. Generate Deployment Name
12. Review Configuration
13. Execute Deployment & Monitor

Fast Path (Defaults)

If user accepts all defaults (latest version, GlobalStandard SKU, recommended capacity, default RAI policy, standard upgrade policy), deployment completes in ~5 interactions.


Phase Summaries

⚠️ MUST READ: Before executing any phase, load references/customize-workflow.md for the full scripts and implementation details. The summaries below describe what each phase does — the reference file contains the how (CLI commands, quota patterns, capacity formulas, cross-region fallback logic).

Phase Action Key Details
1. Verify Auth Check az account show; prompt az login if needed Verify correct subscription is active
2. Get Project ID Read PROJECT_RESOURCE_ID env var or prompt user ARM resource ID format required
3. Verify Project Parse resource ID, call az cognitiveservices account show Extracts subscription, RG, account, project, region
4. Get Model List models via az cognitiveservices account list-models User selects from available or enters custom name
5. Select Version Query versions for chosen model Recommend latest; user picks from list
6. Select SKU Query model catalog + subscription quota, show only deployable SKUs ⚠️ Never hardcode SKU lists — always query live data
7. Configure Capacity Query capacity API, validate min/max/step, user enters value Cross-region fallback if no capacity in current region
8. Select RAI Policy Present content filter options Default: Microsoft.DefaultV2
9. Advanced Options Dynamic quota (GlobalStandard), priority processing (PTU), spillover SKU-dependent availability
10. Upgrade Policy Choose: OnceNewDefaultVersionAvailable / OnceCurrentVersionExpired / NoAutoUpgrade Default: auto-upgrade on new default
11. Deployment Name Auto-generate unique name, allow custom override Validates format: ^[\w.-]{2,64}$
12. Review Display full config summary, confirm before proceeding User approves or cancels
13. Deploy & Monitor az cognitiveservices account deployment create, poll status Timeout after 5 min; show endpoint + portal link

Error Handling

Common Issues and Resolutions

Error Cause Resolution
Model not found Invalid model name List available models with az cognitiveservices account list-models
Version not available Version not supported for SKU Select different version or SKU
Insufficient quota Capacity > available quota Skill auto-searches all regions; fails only if no region has quota
SKU not supported SKU not available in region Cross-region fallback searches other regions automatically
Capacity out of range Invalid capacity value PREVENTED: Skill validates min/max/step at input (Phase 7)
Deployment name exists Name conflict Auto-incremented name generation
Authentication failed Not logged in Run az login
Permission denied Insufficient permissions Assign Cognitive Services Contributor role
Capacity query fails API/permissions/network error DEPLOYMENT BLOCKED: Will not proceed without valid quota data

Troubleshooting Commands

# Check deployment status
az cognitiveservices account deployment show --name <account> --resource-group <rg> --deployment-name <name>

# List all deployments
az cognitiveservices account deployment list --name <account> --resource-group <rg> -o table

# Check quota usage
az cognitiveservices usage list --name <account> --resource-group <rg>

# Delete failed deployment
az cognitiveservices account deployment delete --name <account> --resource-group <rg> --deployment-name <name>

Selection Guides & Advanced Topics

For SKU comparison tables, PTU sizing formulas, and advanced option details, load references/customize-guides.md.

SKU selection: GlobalStandard (production/HA) → Standard (dev/test) → ProvisionedManaged (high-volume/guaranteed throughput) → DataZoneStandard (data residency).

Capacity: TPM-based SKUs range from 1K (dev) to 100K+ (large production). PTU-based use formula: (Input TPM × 0.001) + (Output TPM × 0.002) + (Requests/min × 0.1).

Advanced options: Dynamic quota (GlobalStandard only), priority processing (PTU only, extra cost), spillover (overflow to backup deployment).


Related Skills

  • preset - Quick deployment to best region with automatic configuration
  • microsoft-foundry - Parent skill for all Microsoft Foundry operations
  • quota — For quota viewing, increase requests, and troubleshooting quota errors, defer to this skill instead of duplicating guidance
  • rbac - Manage permissions and access control

Notes

  • Set PROJECT_RESOURCE_ID environment variable to skip prompt
  • Not all SKUs available in all regions; capacity varies by subscription/region/model
  • Custom RAI policies can be configured in Azure Portal
  • Automatic version upgrades occur during maintenance windows
  • Use Azure Monitor and Application Insights for production deployments

Source: SKILL.md on GitHub

No alerts6mo3 checks · Risk SAFE
  • Gen Agent Trust Hub6mo

    This skill provides a unified workflow for deploying Azure OpenAI models using the Azure CLI. It includes intelligent routing, capacity discovery across regions, and support for both standard OpenAI models and Anthropic models on Azure while maintaining security best practices like mandatory project confirmation.

  • Socket6mo

    No alerts

  • Snyk6mo

    Risk: LOW · No issues

Signed by skilld at 54c2a28. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated 2 months ago
metadata
{
  "author": "Microsoft",
  "version": "1.0.0"
}
  • azure
  • openai
  • deployment
  • model
  • capacity
  • sku
  • foundry
  • azure-cli
  • provisioning

README badge

README badge for microsoft/github-copilot-for-azure/deploy-model

Deploys Azure OpenAI models with intent-based routing to preset deployment, customized configuration, or capacity discovery workflows. Routes user requests to the appropriate mode based on whether they want quick defaults, custom SKU/capacity/RAI settings, or to find available capacity across regions and projects.

Generated from the current SKILL.md.

Does this skill handle deployments in azd-managed Foundry projects?
No. For azd projects (those scaffolded from azd-ai-starter-basic or via azd ai agent init), declare deployments in azure.yaml instead — azd provision will create them through Bicep. Use this skill only for standalone Foundry projects or ad-hoc deployments outside the azd lifecycle.
What happens if I don't specify a project?
The skill checks the PROJECT_RESOURCE_ID environment variable first, then looks for clues in your prompt. If neither exists, it queries your projects and suggests the current one, with a confirmation step before deploying.
Can this skill list or delete existing deployments?
No. This skill creates deployments only. Use the foundry_models_deployments_list MCP tool to list existing deployments, or the Azure portal to delete them.
Does this skill validate quota and SKU support before deploying?
Yes. It queries the model catalog to confirm the model supports your chosen SKU, and checks your subscription's available quota via Azure CLI before presenting any deployment options.
What should I do if I hit a quota limit?
Defer to the quota skill (quota/quota.md) for quota increase requests, usage monitoring, and troubleshooting quota errors.

Generated from the current SKILL.md. These answers refresh after source changes.