All skills
microsoft avatar

/azure-reliability

@7804b4d official

Assess and improve the reliability posture of PaaS Applications (Azure Functions and Azure App Service). Scans deployed resources for zone redundancy, ZRS storage, health probes, and multi-region failover. Presents a feature-pivoted checklist, then drives staged remediation (CLI or IaC patches) end-to-end with user confirmation. WHEN: "assess reliability", "check reliability", "zone redundant", "multi-region failover", "high availability", "disaster recovery", "single points of failure", "reliability posture", "resiliency".

Use this Skill: https://skilld.dev/gh/microsoft/github-copilot-for-azure/azure-reliability

This session only. Nothing lands on disk.

referencesservicesfunctionsreliability.md

≈1.6k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Azure Functions — Reliability Reference

Supported Plans & Zone Redundancy

Plan Zone Redundancy Min Instances Health Check
Flex Consumption (FC1) ✅ zoneRedundant: true Auto-managed ❌ Platform health check not supported
Premium (EP1/EP2/EP3) ✅ zoneRedundant: true + sku.capacity: 2 minimumElasticInstanceCount: 2 per app ✅ healthCheckPath
Consumption (Y1) ❌ Not supported N/A ❌ Not supported
Dedicated (P1v2+) ✅ (treated as App Service) sku.capacity: 2 ✅ healthCheckPath

Assessment Queries

Zone Redundancy Check

az graph query -q "
resources
| where resourceGroup =~ '<rg>'
| where type =~ 'microsoft.web/serverfarms'
| where kind contains 'functionapp' or kind =~ 'linux' or kind =~ 'elastic'
| project name, sku=sku.name, zoneRedundant=properties.zoneRedundant, location
" --subscriptions <sub-id>

Function App Instance Count (Premium)

az functionapp show --name <app> --resource-group <rg> \
  --query "{minInstances:siteConfig.minimumElasticInstanceCount}" -o table

Configure: Zone Redundancy

Flex Consumption (FC1)

# Enable zone redundancy on plan
az resource update \
  --resource-group <rg> \
  --name <plan-name> \
  --resource-type "Microsoft.Web/serverfarms" \
  --set properties.zoneRedundant=true

Premium (EP1/EP2/EP3)

# Enable zone redundancy + set min capacity
az appservice plan update \
  --name <plan-name> \
  --resource-group <rg> \
  --number-of-workers 2

az resource update \
  --resource-group <rg> \
  --name <plan-name> \
  --resource-type "Microsoft.Web/serverfarms" \
  --set properties.zoneRedundant=true

# Set minimum elastic instances per app
az resource update \
  --resource-group <rg> \
  --name <app-name> \
  --resource-type "Microsoft.Web/sites" \
  --set properties.siteConfig.minimumElasticInstanceCount=2

Consumption (Y1) — upgrade path required

Consumption (Y1) plans do not support zone redundancy. The user must upgrade the plan first:

  • Recommended: Upgrade to Flex Consumption — similar serverless model, supports ZR, no per-app minimum cost.
  • Alternative: Upgrade to Premium (EP1+) — more control, higher base cost (always-ready instances charged 24/7).

⚠️ Inform the user of cost implications before initiating any plan change.

Configure: Health Endpoint

Flex Consumption does NOT support platform health check (healthCheckPath). Instead, add an HTTP endpoint in code:

TypeScript (v4 programming model)

import { app } from "@azure/functions";

app.http('health', {
  methods: ['GET'],
  authLevel: 'anonymous',
  route: 'health',
  handler: async () => ({ status: 200, body: 'OK' })
});

Python (v2 programming model)

import azure.functions as func

app = func.FunctionApp()

@app.route(route="health", methods=["GET"], auth_level=func.AuthLevel.ANONYMOUS)
def health(req: func.HttpRequest) -> func.HttpResponse:
    return func.HttpResponse("OK", status_code=200)

C# (isolated worker)

[Function("Health")]
public IActionResult Health([HttpTrigger(AuthorizationLevel.Anonymous, "get", Route = "health")] HttpRequest req)
{
    return new OkObjectResult("OK");
}

Premium Functions — Platform Health Check

az webapp config set \
  --name <app-name> \
  --resource-group <rg> \
  --generic-configurations '{"healthCheckPath": "/api/health"}'

⚠️ Enabling health check causes an app restart.

IaC Patching: Bicep

App Service Plan (AVM module)

module appServicePlan 'br/public:avm/res/web/serverfarm:0.5.0' = {
  params: {
    skuName: 'FC1'
    reserved: true
    zoneRedundant: true  // ← ADD
  }
}

Premium Plan — extra settings

module appServicePlan 'br/public:avm/res/web/serverfarm:0.5.0' = {
  params: {
    skuName: 'EP1'
    reserved: true
    zoneRedundant: true      // ← ADD
    skuCapacity: 2           // ← ADD (min 2 for ZR)
  }
}

// On the function app resource:
resource functionApp 'Microsoft.Web/sites@2023-12-01' = {
  properties: {
    siteConfig: {
      minimumElasticInstanceCount: 2  // ← ADD
    }
  }
}

IaC Patching: Terraform

resource "azurerm_service_plan" "plan" {
  sku_name               = "FC1"
  os_type                = "Linux"
  zone_balancing_enabled = true  # ← ADD
}

# Premium plan:
resource "azurerm_service_plan" "plan" {
  sku_name               = "EP1"
  os_type                = "Linux"
  zone_balancing_enabled = true  # ← ADD
  worker_count           = 2     # ← ADD
}

resource "azurerm_linux_function_app" "func" {
  site_config {
    minimum_elastic_instance_count = 2  # ← ADD (Premium only)
  }
}

Multi-Region Notes

  • Flex Consumption standby costs ~$0 (pay-per-execution) — ideal for active-passive
  • Code must be deployed to both regions separately
  • Event Hub checkpoints are per-app — secondary starts from its own checkpoint on failover
  • Consider Event Hubs Geo-DR for true event replication

Reporting (for the assessment table)

When the parent skill builds the feature-pivoted assessment table, report each Functions resource on the relevant rows:

Feature row What to report
Zone redundancy — compute 🟢 ON if the plan has zoneRedundant: true. For Premium plans, also requires sku.capacity ≥ 2 AND each Function App has minimumElasticInstanceCount ≥ 2. 🔴 OFF if the plan tier doesn't support ZR (Consumption Y1) — annotate (needs plan upgrade to Flex / Premium).
Health probes For Premium / Dedicated: 🟢 ON if siteConfig.healthCheckPath is set, 🔴 OFF otherwise. For Flex Consumption (FC1) / Consumption (Y1): always annotate 🔴 OFF (code-only fix) — healthCheckPath is not supported on these plans, so an HTTP-triggered /api/health function must be added in app code (gated by user consent — see configure-health-probes.md).
Multi-region failover 🟢 ON if the same Function App is deployed in ≥2 regions behind Front Door / Traffic Manager; otherwise 🔴 OFF.

Additional References

Source: SKILL.md on GitHub

No alerts4mo3 checks · Risk SAFE
  • Gen Agent Trust Hub4mo

    This skill provides a structured framework for assessing and enhancing the reliability of Azure Functions deployments. It utilizes standard Azure management tools to provide a checklist of improvements and automate remediation steps. The skill maintains a high level of security by requesting explicit user consent for all operations and focusing on official infrastructure tools.

  • Socket4mo

    No alerts

  • Snyk4mo

    Risk: LOW · No issues

Signed by skilld at 7804b4d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated 2 days ago
metadata
{
  "author": "Microsoft",
  "version": "0.0.0-placeholder"
}
  • azure
  • reliability
  • zone-redundancy
  • multi-region
  • failover
  • azure-functions
  • app-service
  • disaster-recovery
  • high-availability
  • iac

README badge

README badge for microsoft/github-copilot-for-azure/azure-reliability

Scans Azure Functions and App Service for zone redundancy, ZRS storage, health probes, and multi-region failover, then applies fixes via Azure CLI or infrastructure-as-code patches. Presents a feature-pivoted reliability checklist and walks through staged remediation with user confirmation at each step.

Generated from the current SKILL.md.

What Azure services does this skill assess?
Currently Azure Functions and Azure App Service only. Azure Container Apps support is planned but not yet available. The skill will identify Container Apps resources in your scope but mark them as not yet assessed.
Does this skill make changes to my resources automatically?
No. The skill always asks for confirmation before executing any changes. It presents a fix plan showing what will change, cost implications, and breaking changes, then lets you choose between applying fixes via CLI or patching your IaC (Bicep/Terraform).
What permissions do I need to run this skill?
Reader access on your subscription or resource group for assessment, and Contributor access if you want to make configuration changes. You must also have the Azure Resource Graph extension installed via `az extension add --name resource-graph`.
Does this skill help with multi-region failover setup?
Yes. After assessing and fixing core reliability features (zone redundancy, storage, health probes), the skill can help you set up multi-region failover using Front Door or Traffic Manager, but only with explicit user consent.

Generated from the current SKILL.md. These answers refresh after source changes.