All skills
microsoft avatar

/azure-reliability

@7804b4d official

Assess and improve the reliability posture of PaaS Applications (Azure Functions and Azure App Service). Scans deployed resources for zone redundancy, ZRS storage, health probes, and multi-region failover. Presents a feature-pivoted checklist, then drives staged remediation (CLI or IaC patches) end-to-end with user confirmation. WHEN: "assess reliability", "check reliability", "zone redundant", "multi-region failover", "high availability", "disaster recovery", "single points of failure", "reliability posture", "resiliency".

Use this Skill: https://skilld.dev/gh/microsoft/github-copilot-for-azure/azure-reliability

This session only. Nothing lands on disk.

referencesiac-patching-terraform.md

≈1.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

IaC Patching — Terraform

When to Use

Use this reference when the user chooses "Patch my IaC" instead of "Fix now" (CLI). This patches Terraform files in the project's infra/ folder so reliability settings persist across terraform apply / azd up.

Detection

  1. Look for infra/ folder in the project root
  2. Check for *.tf files (especially main.tf, variables.tf)
  3. Confirm with user: "I found Terraform files in infra/. Want me to patch them for reliability?"

File Discovery

# Find all Terraform files
Get-ChildItem -Path infra -Recurse -Filter *.tf

# Common file structure:
# infra/main.tf              — main resources
# infra/variables.tf          — input variables
# infra/terraform.tfvars      — variable values
# infra/modules/              — reusable modules

Resource definitions may be in module files. Search all .tf files for the resource type.

Per-service Terraform patches

The patches for compute (zone redundancy on the App Service Plans / environments, Function App plan, health check path) live in the per-service references because the SKU rules and resource types differ:

Service Reference
Azure App Service services/app-service/reliability.md
Azure Functions services/functions/reliability.md

Azure App Service and Azure Container Apps per-service Terraform patches are planned for a future version of this skill.

The one truly cross-service patch — storage — lives below.


Patch: Storage Account — LRS / GRS → ZRS / GZRS

Find: azurerm_storage_account

Search pattern: resource "azurerm_storage_account"

Before:

resource "azurerm_storage_account" "storage" {
  name                     = var.storage_account_name
  resource_group_name      = azurerm_resource_group.rg.name
  location                 = azurerm_resource_group.rg.location
  account_tier             = "Standard"
  account_replication_type = "LRS"
}

After — change to ZRS:

resource "azurerm_storage_account" "storage" {
  name                     = var.storage_account_name
  resource_group_name      = azurerm_resource_group.rg.name
  location                 = azurerm_resource_group.rg.location
  account_tier             = "Standard"
  account_replication_type = "ZRS"
}

Parameterized Replication Type

If parameterized, update the default:

variables.tf — Before:

variable "storage_replication_type" {
  default = "LRS"
}

After:

variable "storage_replication_type" {
  default = "ZRS"
}

Also check terraform.tfvars for overrides.

⚠️ Existing Deployed Storage

Changing account_replication_type in Terraform expresses the desired end state, but LRS→ZRS is a storage redundancy conversion, not a simple property change. Terraform may attempt an in-place update that fails, or worse, plan a destroy+recreate (data loss risk).

Always follow this order for existing storage:

  1. Patch Terraform to account_replication_type = "ZRS" (desired end state)
  2. Run az storage account migration start to initiate the live conversion
  3. Wait for migration to complete (az storage account migration show)
  4. Run terraform plan — confirm it shows no changes (state now matches desired)
  5. If plan still shows changes, run terraform refresh to sync state, then re-plan

⛔ Do NOT run terraform apply before the migration completes. It may fail or attempt to recreate the storage account.


Deploy Plan (Skill executes this directly)

After patching, the skill executes the deploys itself — do not stop and tell the user to run commands. Confirm once with the user before each deploy, then run it.

Summarize the plan for the user:

✅ Terraform files patched for reliability.

Deploy plan (the skill will run these for you after your confirmation):
  1. `terraform plan -out tfplan` (skill will show the plan summary)
  2. Deploy 1 — `terraform apply tfplan` for the safe patches.
  3. Storage migration (only if upgrading LRS → ZRS).
     Command: `az storage account migration start ...`, then poll until `sku.name = Standard_ZRS`.
  4. Deploy 2 — second `terraform plan` + `apply` for the storage SKU patch (no-op confirmation).

Do NOT bundle the storage SKU change with the safe patches — a failed storage redundancy update can fail the whole apply.

⚠️ Note: If you have an existing Container Apps environment without zone redundancy,
   the environment name was changed to force recreation. The skill will surface the
   `terraform plan` summary before applying so you can confirm — apps will be recreated
   in the new environment.

Ready to run `terraform plan`? (yes / no)

Source: SKILL.md on GitHub

No alerts4mo3 checks · Risk SAFE
  • Gen Agent Trust Hub4mo

    This skill provides a structured framework for assessing and enhancing the reliability of Azure Functions deployments. It utilizes standard Azure management tools to provide a checklist of improvements and automate remediation steps. The skill maintains a high level of security by requesting explicit user consent for all operations and focusing on official infrastructure tools.

  • Socket4mo

    No alerts

  • Snyk4mo

    Risk: LOW · No issues

Signed by skilld at 7804b4d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated 2 days ago
metadata
{
  "author": "Microsoft",
  "version": "0.0.0-placeholder"
}
  • azure
  • reliability
  • zone-redundancy
  • multi-region
  • failover
  • azure-functions
  • app-service
  • disaster-recovery
  • high-availability
  • iac

README badge

README badge for microsoft/github-copilot-for-azure/azure-reliability

Scans Azure Functions and App Service for zone redundancy, ZRS storage, health probes, and multi-region failover, then applies fixes via Azure CLI or infrastructure-as-code patches. Presents a feature-pivoted reliability checklist and walks through staged remediation with user confirmation at each step.

Generated from the current SKILL.md.

What Azure services does this skill assess?
Currently Azure Functions and Azure App Service only. Azure Container Apps support is planned but not yet available. The skill will identify Container Apps resources in your scope but mark them as not yet assessed.
Does this skill make changes to my resources automatically?
No. The skill always asks for confirmation before executing any changes. It presents a fix plan showing what will change, cost implications, and breaking changes, then lets you choose between applying fixes via CLI or patching your IaC (Bicep/Terraform).
What permissions do I need to run this skill?
Reader access on your subscription or resource group for assessment, and Contributor access if you want to make configuration changes. You must also have the Azure Resource Graph extension installed via `az extension add --name resource-graph`.
Does this skill help with multi-region failover setup?
Yes. After assessing and fixing core reliability features (zone redundancy, storage, health probes), the skill can help you set up multi-region failover using Front Door or Traffic Manager, but only with explicit user consent.

Generated from the current SKILL.md. These answers refresh after source changes.