All skills
microsoft avatar

/azure-aigateway

@d0c21c0 official

Configure Azure API Management as an AI Gateway for AI models, MCP tools, and agents. WHEN: semantic caching, token limit, content safety, load balancing, AI model governance, MCP rate limiting, jailbreak detection, add Azure OpenAI backend, add AI Foundry model, test AI gateway, LLM policies, configure AI backend, token metrics, AI cost control, convert API to MCP, import OpenAPI to gateway.

Use this Skill: https://skilld.dev/gh/microsoft/github-copilot-for-azure/azure-aigateway

This session only. Nothing lands on disk.

referencestroubleshooting.md

≈2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

AI Gateway Troubleshooting

Common issues when using Azure API Management as an AI Gateway.


Authentication Issues

401 Unauthorized from Backend

Symptom: APIM returns 401 when calling Azure OpenAI.

Causes & Solutions:

Cause Fix
Managed identity not enabled on APIM az apim update --name <apim> --resource-group <rg> --set identity.type=SystemAssigned
Missing RBAC role az role assignment create --assignee <apim-principal-id> --role "Cognitive Services User" --scope <aoai-resource-id>
Wrong auth resource Ensure resource="https://cognitiveservices.azure.com" (not the endpoint URL)
RBAC propagation delay Wait 5-10 minutes after role assignment

Diagnostic:

# Verify identity is enabled
az apim show --name <apim> --resource-group <rg> --query "identity" -o json

# Check role assignments
AOAI_ID=$(az cognitiveservices account show --name <aoai> --resource-group <rg> --query id -o tsv)
az role assignment list --scope "$AOAI_ID" --query "[?principalType=='ServicePrincipal'].{role:roleDefinitionName, principal:principalId}" -o table

Rate Limiting Issues

429 Token Limit Exceeded

Symptom: Requests blocked with 429 Too Many Requests from azure-openai-token-limit policy.

Solutions:

  1. Increase limit: Raise tokens-per-minute value
  2. Add more backends: Load balance across regions for higher aggregate TPM
  3. Enable semantic caching: Reduce actual token consumption by serving cached responses
  4. Switch counter-key: Use per-user instead of global to prevent one user from exhausting the pool
<!-- Per-user instead of global -->
<azure-openai-token-limit
    tokens-per-minute="50000"
    counter-key="@(context.Request.Headers.GetValueOrDefault("X-User-Id", context.Subscription.Id))"
    estimate-prompt-tokens="true" />

429 from Azure OpenAI (Not APIM)

Symptom: Backend returns 429 even though APIM token limits are not exceeded.

Cause: Azure OpenAI's own TPM quota is exhausted.

Solutions:

  1. Increase Azure OpenAI deployment TPM quota in the portal
  2. Add load balancing across multiple Azure OpenAI instances
  3. Use retry with backoff:
<retry condition="@(context.Response.StatusCode == 429)" count="3" interval="10">
    <forward-request />
</retry>

Semantic Caching Issues

No Cache Hits

Symptom: Semantic cache is configured but cache hit rate is 0%.

Causes & Solutions:

Cause Fix
score-threshold too high Lower from 0.9 to 0.7 (more matches)
Embeddings backend misconfigured Verify backend URL and auth
Redis not configured Deploy Azure Cache for Redis Enterprise with RediSearch
Streaming requests Semantic caching doesn't work with "stream": true

Verify caching is working:

# Check cache-related headers in response
curl -v -X POST "${GATEWAY_URL}/openai/deployments/<deployment>/chat/completions?api-version=2024-02-01" \
  -H "Content-Type: application/json" \
  -H "Ocp-Apim-Subscription-Key: <key>" \
  -d '{"messages": [{"role": "user", "content": "What is Azure?"}], "max_tokens": 100}'

# Look for: x-cache-status header in response

Cache Returns Stale Data

Solution: Reduce duration in azure-openai-semantic-cache-store:

<!-- Shorter TTL for frequently changing knowledge -->
<azure-openai-semantic-cache-store duration="300" />  <!-- 5 minutes -->

Content Safety Issues

False Positives (Legitimate Content Blocked)

Symptom: Normal business content is being blocked by content safety policy.

Solutions:

  1. Increase thresholds (less strict):
<llm-content-safety backend-id="contentsafety-backend">
    <category name="Hate" threshold="5" />      <!-- Was 4, now less strict -->
    <category name="Sexual" threshold="5" />
    <category name="SelfHarm" threshold="5" />
    <category name="Violence" threshold="5" />
</llm-content-safety>
  1. Log blocked content for review:
<on-error>
    <choose>
        <when condition="@(context.LastError.Source == "llm-content-safety")">
            <trace source="content-safety" severity="warning">
                @{
                    return new JObject(
                        new JProperty("blocked", true),
                        new JProperty("subscription", context.Subscription.Id),
                        new JProperty("timestamp", DateTime.UtcNow)
                    ).ToString();
                }
            </trace>
            <return-response>
                <set-status code="400" reason="Content Filtered" />
                <set-body>{"error": "Content filtered by safety policy"}</set-body>
            </return-response>
        </when>
    </choose>
</on-error>

Content Safety Backend Error

Symptom: 500 error from llm-content-safety policy.

Causes:

Cause Fix
Content Safety resource not deployed Deploy Azure AI Content Safety resource
Backend URL wrong Check contentsafety-backend URL matches resource endpoint
Missing RBAC Grant APIM "Cognitive Services User" on the Content Safety resource
Region mismatch Content Safety must be in a supported region

Backend Configuration Issues

Backend Not Found

Symptom: 500 error with "Backend not found" message.

# Verify backend exists
az apim backend list --service-name <apim> --resource-group <rg> \
  --query "[].{id:name, url:url}" -o table

# Check backend ID matches policy reference

Timeout on AI Requests

Symptom: Requests timeout, especially for large context windows or complex prompts.

Solution: Increase timeout in <backend>:

<backend>
    <!-- Default is 30s, increase for large AI requests -->
    <forward-request timeout="120" />
</backend>

Diagnostic Tools

APIM Tracing

Enable request tracing for debugging policy flow:

# Get tracing subscription key
az apim subscription list --service-name <apim> --resource-group <rg> \
  --query "[?displayName=='Built-in all-access subscription'].primaryKey" -o tsv

# Send request with tracing
curl -X POST "${GATEWAY_URL}/..." \
  -H "Ocp-Apim-Trace: true" \
  -H "Ocp-Apim-Subscription-Key: <built-in-key>"

Application Insights

If APIM is connected to Application Insights:

// Failed AI gateway requests
requests
| where success == false
| where url contains "openai"
| project timestamp, resultCode, duration, url
| order by timestamp desc
| take 20

// Token metrics over time
customMetrics
| where name == "Total Tokens"
| summarize TotalTokens = sum(value) by bin(timestamp, 1h)
| render timechart

// Content safety blocks
traces
| where message contains "content-safety"
| project timestamp, message, customDimensions
| order by timestamp desc

Health Check

Quick validation that the AI Gateway is functioning:

# 1. Check APIM is running
az apim show --name <apim> --resource-group <rg> --query "provisioningState" -o tsv
# Expected: Succeeded

# 2. Check backends
az apim backend list --service-name <apim> --resource-group <rg> -o table

# 3. Test endpoint
curl -s -o /dev/null -w "%{http_code}" "${GATEWAY_URL}/openai/deployments/<deployment>/chat/completions?api-version=2024-02-01" \
  -H "Ocp-Apim-Subscription-Key: <key>" \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "ping"}], "max_tokens": 5}'
# Expected: 200

References

Source: SKILL.md on GitHub

1 warning16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides configuration guidance for Azure API Management as an AI Gateway. It incorporates security best practices such as Managed Identity authentication and content safety policies. All external resources originate from trusted Microsoft sources, and no security risks were identified.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer6mo

    7/9 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at d0c21c0. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 5 months ago
metadata
{
  "author": "Microsoft",
  "version": "0.0.0-placeholder"
}
compatibility
Requires Azure CLI (az) for configuration and testing
  • MCP
  • azure
  • api-management
  • ai-gateway
  • llm
  • semantic-caching
  • rate-limiting
  • content-safety
  • token-limiting
  • load-balancing

README badge

README badge for microsoft/github-copilot-for-azure/azure-aigateway

Configures Azure API Management as a gateway to enforce semantic caching, token limits, content safety, and rate limiting across AI models, MCP tools, and agents. Use this skill to add Azure OpenAI or AI Foundry backends, apply LLM governance policies, and test the gateway with curl or Azure CLI.

Generated from the current SKILL.md.

Does this skill work with models other than Azure OpenAI?
Yes. The skill configures Azure API Management to govern any AI model backend, including AI Foundry models. You add backends via the `az apim backend create` command.
Can I use this skill to rate-limit MCP tools?
Yes. The skill includes the `rate-limit-by-key` policy for protecting MCP tools and APIs from overuse.
What do I need installed to use this skill?
You need the Azure CLI (az) installed and configured. The skill also assumes Azure API Management is already deployed; use the azure-prepare skill to deploy APIM first if needed.
Does this skill provide content safety and jailbreak detection?
Yes. The skill includes the `llm-content-safety` policy for filtering harmful content and detecting jailbreak attempts on AI agents.
Can semantic caching really save that much on API costs?
The skill documents 60-80% cost savings using the `azure-openai-semantic-cache-lookup` and `azure-openai-semantic-cache-store` policies, though actual savings depend on request patterns and cache hit rates.

Generated from the current SKILL.md. These answers refresh after source changes.