All skills
microsoft avatar

/azure-reliability

@7804b4d official

Assess and improve the reliability posture of PaaS Applications (Azure Functions and Azure App Service). Scans deployed resources for zone redundancy, ZRS storage, health probes, and multi-region failover. Presents a feature-pivoted checklist, then drives staged remediation (CLI or IaC patches) end-to-end with user confirmation. WHEN: "assess reliability", "check reliability", "zone redundant", "multi-region failover", "high availability", "disaster recovery", "single points of failure", "reliability posture", "resiliency".

Use this Skill: https://skilld.dev/gh/microsoft/github-copilot-for-azure/azure-reliability

This session only. Nothing lands on disk.

referenceshealth-probe-checks.md

≈886 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Health Probe & Monitoring — Platform-Level Checks

Overview

Health probes enable automated failover and recovery. Without them, load balancers and platform services cannot detect failures automatically.

This file covers global / platform-level probe checks (Azure Front Door, Traffic Manager, Application Insights connectivity). For service-specific health-probe checks, configuration commands, and IaC patches, see:

Service Reference
Azure Functions services/functions/reliability.md

Azure App Service and Azure Container Apps per-service references are planned but not yet shipped in this skill version.

⚠️ Output format: Use --query "data[]" -o json for az graph query. Standard az afd / az network traffic-manager commands work fine with -o table.

Check Front Door Health Probe Configuration

az afd origin-group list \
  --profile-name <front-door-name> \
  --resource-group <rg> \
  --query "[].{name:name, probePath:healthProbeSettings.probePath, probeProtocol:healthProbeSettings.probeProtocol, intervalSeconds:healthProbeSettings.probeIntervalInSeconds}" -o table

Interpretation:

  • probePath empty / null → ❌ No active health probing → no automatic failover
  • probePath = /api/health (or similar) → ✅ Probe configured

Check Traffic Manager Endpoint Monitoring

az graph query -q "
Resources
| where type =~ 'microsoft.network/trafficmanagerprofiles'
| extend monitorPath = tostring(properties.monitorConfig.path)
| extend monitorProtocol = tostring(properties.monitorConfig.protocol)
| extend monitorPort = tostring(properties.monitorConfig.port)
| project name, resourceGroup, monitorProtocol, monitorPort, monitorPath
" --query "data[]" -o json

Check Application Insights Connectivity

App settings are not reliably queryable via Resource Graph. Use Azure CLI directly:

az webapp config appsettings list \
  --name <app-name> \
  --resource-group <rg> \
  --query "[?contains(name, 'APPINSIGHTS') || contains(name, 'APPLICATIONINSIGHTS')].{name:name}" -o table

For Function Apps:

az functionapp config appsettings list \
  --name <app-name> \
  --resource-group <rg> \
  --query "[?contains(name, 'APPINSIGHTS') || contains(name, 'APPLICATIONINSIGHTS')].{name:name}" -o table

Best Practices for Health Endpoints

These apply across all services:

  1. Keep health endpoints lightweight — return 200 quickly, no heavy DB/dependency queries on every probe.
  2. Use anonymous auth — health probes can't pass auth tokens.
  3. Two endpoints, not one — fast /health for the load balancer, optional /health/deep for on-call diagnostics.
  4. For Container Apps, both liveness AND readiness — liveness alone restarts the container without taking it out of rotation.
  5. Test the endpoint before relying on it: curl https://<app-url>/api/health.

Reporting (for the Multi-Region row)

For the Multi-region failover row of the assessment table:

  • ✅ — Front Door (or Traffic Manager) exists AND has a non-empty probePath / monitorConfig.path
  • ⚠️ Partial — global load balancer exists but has no health probe configured (manual failover only)
  • ❌ — no global load balancer

Per-service Health probes row reporting for Azure Functions is documented in services/functions/reliability.md. App Service and Container Apps per-service reporting is planned but not yet available.

Source: SKILL.md on GitHub

No alerts4mo3 checks · Risk SAFE
  • Gen Agent Trust Hub4mo

    This skill provides a structured framework for assessing and enhancing the reliability of Azure Functions deployments. It utilizes standard Azure management tools to provide a checklist of improvements and automate remediation steps. The skill maintains a high level of security by requesting explicit user consent for all operations and focusing on official infrastructure tools.

  • Socket4mo

    No alerts

  • Snyk4mo

    Risk: LOW · No issues

Signed by skilld at 7804b4d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated 2 days ago
metadata
{
  "author": "Microsoft",
  "version": "0.0.0-placeholder"
}
  • azure
  • reliability
  • zone-redundancy
  • multi-region
  • failover
  • azure-functions
  • app-service
  • disaster-recovery
  • high-availability
  • iac

README badge

README badge for microsoft/github-copilot-for-azure/azure-reliability

Scans Azure Functions and App Service for zone redundancy, ZRS storage, health probes, and multi-region failover, then applies fixes via Azure CLI or infrastructure-as-code patches. Presents a feature-pivoted reliability checklist and walks through staged remediation with user confirmation at each step.

Generated from the current SKILL.md.

What Azure services does this skill assess?
Currently Azure Functions and Azure App Service only. Azure Container Apps support is planned but not yet available. The skill will identify Container Apps resources in your scope but mark them as not yet assessed.
Does this skill make changes to my resources automatically?
No. The skill always asks for confirmation before executing any changes. It presents a fix plan showing what will change, cost implications, and breaking changes, then lets you choose between applying fixes via CLI or patching your IaC (Bicep/Terraform).
What permissions do I need to run this skill?
Reader access on your subscription or resource group for assessment, and Contributor access if you want to make configuration changes. You must also have the Azure Resource Graph extension installed via `az extension add --name resource-graph`.
Does this skill help with multi-region failover setup?
Yes. After assessing and fixing core reliability features (zone redundancy, storage, health probes), the skill can help you set up multi-region failover using Front Door or Traffic Manager, but only with explicit user consent.

Generated from the current SKILL.md. These answers refresh after source changes.