All skills
microsoft avatar

/azure-aigateway

@d0c21c0 official

Configure Azure API Management as an AI Gateway for AI models, MCP tools, and agents. WHEN: semantic caching, token limit, content safety, load balancing, AI model governance, MCP rate limiting, jailbreak detection, add Azure OpenAI backend, add AI Foundry model, test AI gateway, LLM policies, configure AI backend, token metrics, AI cost control, convert API to MCP, import OpenAPI to gateway.

Use this Skill: https://skilld.dev/gh/microsoft/github-copilot-for-azure/azure-aigateway

This session only. Nothing lands on disk.

SKILL.md

≈103 tokens always: the name and description. ≈1.1k when used: this file. ≈8.9k more on demand in 9 files.

Azure AI Gateway

Configure Azure API Management (APIM) as an AI Gateway for governing AI models, MCP tools, and agents.

To deploy APIM, use the azure-prepare skill. See APIM deployment guide.

When to Use This Skill

Category Triggers
Model Governance "semantic caching", "token limits", "load balance AI", "track token usage"
Tool Governance "rate limit MCP", "protect my tools", "configure my tool", "convert API to MCP"
Agent Governance "content safety", "jailbreak detection", "filter harmful content"
Configuration "add Azure OpenAI backend", "configure my model", "add AI Foundry model"
Testing "test AI gateway", "call OpenAI through gateway"

Quick Reference

Policy Purpose Details
azure-openai-token-limit Cost control Model Policies
azure-openai-semantic-cache-lookup/store 60-80% cost savings Model Policies
azure-openai-emit-token-metric Observability Model Policies
llm-content-safety Safety & compliance Agent Policies
rate-limit-by-key MCP/tool protection Tool Policies

Get Gateway Details

# Get gateway URL
az apim show --name <apim-name> --resource-group <rg> --query "gatewayUrl" -o tsv

# List backends (AI models)
az apim backend list --service-name <apim-name> --resource-group <rg> \
  --query "[].{id:name, url:url}" -o table

# Get subscription key
az apim subscription keys list \
  --service-name <apim-name> --resource-group <rg> --subscription-id <sub-id>

Test AI Endpoint

GATEWAY_URL=$(az apim show --name <apim-name> --resource-group <rg> --query "gatewayUrl" -o tsv)

curl -X POST "${GATEWAY_URL}/openai/deployments/<deployment>/chat/completions?api-version=2024-02-01" \
  -H "Content-Type: application/json" \
  -H "Ocp-Apim-Subscription-Key: <key>" \
  -d '{"messages": [{"role": "user", "content": "Hello"}], "max_tokens": 100}'

Common Tasks

Add AI Backend

See references/patterns.md for full steps.

# Discover AI resources
az cognitiveservices account list --query "[?kind=='OpenAI']" -o table

# Create backend
az apim backend create --service-name <apim> --resource-group <rg> \
  --backend-id openai-backend --protocol http --url "https://<aoai>.openai.azure.com/openai"

# Grant access (managed identity)
az role assignment create --assignee <apim-principal-id> \
  --role "Cognitive Services User" --scope <aoai-resource-id>

Apply AI Governance Policy

Recommended policy order in <inbound>:

  1. Authentication - Managed identity to backend
  2. Semantic Cache Lookup - Check cache before calling AI
  3. Token Limits - Cost control
  4. Content Safety - Filter harmful content
  5. Backend Selection - Load balancing
  6. Metrics - Token usage tracking

See references/policies.md for complete example.


Troubleshooting

Issue Solution
Token limit 429 Increase tokens-per-minute or add load balancing
No cache hits Lower score-threshold to 0.7
Content false positives Increase category thresholds (5-6)
Backend auth 401 Grant APIM "Cognitive Services User" role

See references/troubleshooting.md for details.


References

SDK Quick References

Source: SKILL.md on GitHub

1 warning16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides configuration guidance for Azure API Management as an AI Gateway. It incorporates security best practices such as Managed Identity authentication and content safety policies. All external resources originate from trusted Microsoft sources, and no security risks were identified.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer6mo

    7/9 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at d0c21c0. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 19 hours ago.

Activeupdated 5 months ago
metadata
{
  "author": "Microsoft",
  "version": "0.0.0-placeholder"
}
compatibility
Requires Azure CLI (az) for configuration and testing
  • MCP
  • azure
  • api-management
  • ai-gateway
  • llm
  • semantic-caching
  • rate-limiting
  • content-safety
  • token-limiting
  • load-balancing

README badge

README badge for microsoft/github-copilot-for-azure/azure-aigateway

Configures Azure API Management as a gateway to enforce semantic caching, token limits, content safety, and rate limiting across AI models, MCP tools, and agents. Use this skill to add Azure OpenAI or AI Foundry backends, apply LLM governance policies, and test the gateway with curl or Azure CLI.

Generated from the current SKILL.md.

Does this skill work with models other than Azure OpenAI?
Yes. The skill configures Azure API Management to govern any AI model backend, including AI Foundry models. You add backends via the `az apim backend create` command.
Can I use this skill to rate-limit MCP tools?
Yes. The skill includes the `rate-limit-by-key` policy for protecting MCP tools and APIs from overuse.
What do I need installed to use this skill?
You need the Azure CLI (az) installed and configured. The skill also assumes Azure API Management is already deployed; use the azure-prepare skill to deploy APIM first if needed.
Does this skill provide content safety and jailbreak detection?
Yes. The skill includes the `llm-content-safety` policy for filtering harmful content and detecting jailbreak attempts on AI agents.
Can semantic caching really save that much on API costs?
The skill documents 60-80% cost savings using the `azure-openai-semantic-cache-lookup` and `azure-openai-semantic-cache-store` policies, though actual savings depend on request patterns and cache hit rates.

Generated from the current SKILL.md. These answers refresh after source changes.