All skills
microsoft avatar

/azure-aigateway

@d0c21c0 official

Configure Azure API Management as an AI Gateway for AI models, MCP tools, and agents. WHEN: semantic caching, token limit, content safety, load balancing, AI model governance, MCP rate limiting, jailbreak detection, add Azure OpenAI backend, add AI Foundry model, test AI gateway, LLM policies, configure AI backend, token metrics, AI cost control, convert API to MCP, import OpenAPI to gateway.

Use this Skill: https://skilld.dev/gh/microsoft/github-copilot-for-azure/azure-aigateway

This session only. Nothing lands on disk.

referencespatterns.md

≈1.7k tokens on demand. Your agent reads this file only when SKILL.md points to it.

AI Gateway Configuration Patterns

Step-by-step patterns for configuring Azure API Management as an AI Gateway.


Pattern 1: Add AI Model Backend

Connect Azure OpenAI or AI Foundry models to your APIM instance.

Prerequisites

  • APIM instance deployed (use azure-prepare skill to deploy APIM — see APIM deployment guide)
  • Azure OpenAI or AI Foundry resource provisioned
  • System-assigned or user-assigned managed identity enabled on APIM

Steps

1. Discover AI Resources
# Find Azure OpenAI resources
az cognitiveservices account list --query "[?kind=='OpenAI'].{name:name, rg:resourceGroup, endpoint:properties.endpoint}" -o table

# Find AI Foundry resources (if using)
az cognitiveservices account list --query "[?kind=='AIServices'].{name:name, rg:resourceGroup}" -o table
2. Enable Managed Identity on APIM
# Enable system-assigned identity
az apim update --name <apim-name> --resource-group <rg> --set identity.type=SystemAssigned

# Get principal ID
PRINCIPAL_ID=$(az apim show --name <apim-name> --resource-group <rg> --query "identity.principalId" -o tsv)
3. Grant RBAC Access
AOAI_ID=$(az cognitiveservices account show --name <aoai-name> --resource-group <rg> --query id -o tsv)

az role assignment create \
  --assignee "$PRINCIPAL_ID" \
  --role "Cognitive Services User" \
  --scope "$AOAI_ID"
4. Create Backend
az apim backend create \
  --service-name <apim-name> \
  --resource-group <rg> \
  --backend-id openai-backend \
  --protocol http \
  --url "https://<aoai-name>.openai.azure.com/openai"
5. Import API (OpenAPI Spec)
# Import the Azure OpenAI API specification
az apim api import \
  --service-name <apim-name> \
  --resource-group <rg> \
  --api-id azure-openai-api \
  --path "openai" \
  --specification-format OpenApi \
  --specification-url "https://raw.githubusercontent.com/Azure/azure-rest-api-specs/main/specification/cognitiveservices/data-plane/AzureOpenAI/inference/stable/2024-02-01/inference.json" \
  --service-url "https://<aoai-name>.openai.azure.com/openai"
6. Set Backend Policy

Add managed identity authentication in <inbound>:

<inbound>
    <base />
    <set-backend-service backend-id="openai-backend" />
    <authentication-managed-identity resource="https://cognitiveservices.azure.com" />
</inbound>

Pattern 2: Load Balance Across Multiple AI Backends

Distribute requests across multiple Azure OpenAI instances for higher throughput.

Steps

1. Create Multiple Backends
# Primary region
az apim backend create --service-name <apim> --resource-group <rg> \
  --backend-id openai-eastus --protocol http \
  --url "https://<aoai-eastus>.openai.azure.com/openai"

# Secondary region
az apim backend create --service-name <apim> --resource-group <rg> \
  --backend-id openai-westus --protocol http \
  --url "https://<aoai-westus>.openai.azure.com/openai"
2. Create Backend Pool

Using APIM backend pool (preview) or policy-based load balancing:

<inbound>
    <base />
    <set-variable name="backendUrl" value="@{
        var backends = new [] {
            "https://aoai-eastus.openai.azure.com",
            "https://aoai-westus.openai.azure.com"
        };
        var hash = Math.Abs(context.RequestId.GetHashCode());
        var index = hash % backends.Length;
        return backends[index];
    }" />
    <set-backend-service base-url="@((string)context.Variables["backendUrl"] + "/openai")" />
    <authentication-managed-identity resource="https://cognitiveservices.azure.com" />
</inbound>
3. Add Circuit Breaker (Retry on 429)
<retry condition="@(context.Response.StatusCode == 429)" count="3" interval="10" delta="5" max-interval="30" first-fast-retry="false">
    <set-variable name="backendUrl" value="@{
        var backends = new [] {
            "https://aoai-eastus.openai.azure.com",
            "https://aoai-westus.openai.azure.com"
        };
        var currentIndex = Array.IndexOf(backends, (string)context.Variables["backendUrl"]);
        return backends[(currentIndex + 1) % backends.Length];
    }" />
    <set-backend-service base-url="@((string)context.Variables["backendUrl"] + "/openai")" />
    <forward-request />
</retry>

Pattern 3: Convert API to MCP Tool

Expose an existing API through APIM as an MCP-compatible tool for AI agents.

Steps

  1. Import API into APIM using OpenAPI spec
  2. Add rate limiting to protect the tool endpoint
  3. Add content safety to filter harmful inputs
  4. Generate MCP manifest pointing to the APIM endpoint
<!-- Rate limit MCP tool calls -->
<inbound>
    <base />
    <rate-limit-by-key calls="10" renewal-period="60"
        counter-key="@(context.Request.Headers.GetValueOrDefault("X-Agent-Id", "anonymous"))" />
</inbound>

Pattern 4: Add Streaming Support

Configure APIM to properly handle Server-Sent Events (SSE) for streaming AI responses.

<inbound>
    <base />
    <set-backend-service backend-id="openai-backend" />
    <authentication-managed-identity resource="https://cognitiveservices.azure.com" />
</inbound>
<outbound>
    <base />
    <set-header name="Content-Type" exists-action="override">
        <value>@(context.Request.Body.As<JObject>()["stream"]?.Value<bool>() == true
            ? "text/event-stream" : "application/json")</value>
    </set-header>
</outbound>

Note: Semantic caching and token metrics policies are NOT compatible with streaming responses. Use non-streaming for cost control scenarios.


Pattern 5: Multi-Tenant AI Gateway

Isolate tenants with per-client rate limiting and tracking.

<inbound>
    <base />
    <!-- Extract tenant from subscription or header -->
    <set-variable name="tenantId" value="@(context.Subscription.Id)" />

    <!-- Per-tenant token limit -->
    <azure-openai-token-limit
        tokens-per-minute="10000"
        counter-key="@((string)context.Variables["tenantId"])"
        estimate-prompt-tokens="true" />

    <!-- Per-tenant metrics -->
    <azure-openai-emit-token-metric namespace="ai-gateway">
        <dimension name="Tenant" value="@((string)context.Variables["tenantId"])" />
        <dimension name="API" value="@(context.Api.Name)" />
    </azure-openai-emit-token-metric>

    <set-backend-service backend-id="openai-backend" />
    <authentication-managed-identity resource="https://cognitiveservices.azure.com" />
</inbound>

Next Steps

Source: SKILL.md on GitHub

1 warning16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides configuration guidance for Azure API Management as an AI Gateway. It incorporates security best practices such as Managed Identity authentication and content safety policies. All external resources originate from trusted Microsoft sources, and no security risks were identified.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer6mo

    7/9 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at d0c21c0. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 5 months ago
metadata
{
  "author": "Microsoft",
  "version": "0.0.0-placeholder"
}
compatibility
Requires Azure CLI (az) for configuration and testing
  • MCP
  • azure
  • api-management
  • ai-gateway
  • llm
  • semantic-caching
  • rate-limiting
  • content-safety
  • token-limiting
  • load-balancing

README badge

README badge for microsoft/github-copilot-for-azure/azure-aigateway

Configures Azure API Management as a gateway to enforce semantic caching, token limits, content safety, and rate limiting across AI models, MCP tools, and agents. Use this skill to add Azure OpenAI or AI Foundry backends, apply LLM governance policies, and test the gateway with curl or Azure CLI.

Generated from the current SKILL.md.

Does this skill work with models other than Azure OpenAI?
Yes. The skill configures Azure API Management to govern any AI model backend, including AI Foundry models. You add backends via the `az apim backend create` command.
Can I use this skill to rate-limit MCP tools?
Yes. The skill includes the `rate-limit-by-key` policy for protecting MCP tools and APIs from overuse.
What do I need installed to use this skill?
You need the Azure CLI (az) installed and configured. The skill also assumes Azure API Management is already deployed; use the azure-prepare skill to deploy APIM first if needed.
Does this skill provide content safety and jailbreak detection?
Yes. The skill includes the `llm-content-safety` policy for filtering harmful content and detecting jailbreak attempts on AI agents.
Can semantic caching really save that much on API costs?
The skill documents 60-80% cost savings using the `azure-openai-semantic-cache-lookup` and `azure-openai-semantic-cache-store` policies, though actual savings depend on request patterns and cache hit rates.

Generated from the current SKILL.md. These answers refresh after source changes.