Official agent skill

Azure Reliability

by microsoft in microsoft/GitHub-Copilot-for-Azure

Assess and improve the reliability posture of PaaS Applications (Azure Functions and Azure App Service).

OfficialMITAuto-check passedDevOps & Cloud

Install Azure Reliability

skills CLI
$ npx skills add microsoft/GitHub-Copilot-for-Azure --skill azure-reliability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/GitHub-Copilot-for-Azure azure-reliability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/GitHub-Copilot-for-Azure.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/azure-skills/skills/azure-reliability .claude/skills/azure-reliability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
azure-reliability
GitHub stars
255
Used in
1 other repo
Token cost
~5.9k tokens
SKILL.md length
2,292 words
Files
14 (incl. references)
Skills in repo
56
Repo updated
First seen
Licence
MIT

At a glance

Assess and improve the reliability posture of PaaS Applications (Azure Functions and Azure App Service).

  • Works in 6 steps: Discover Resources → Assess Reliability → Generate Reliability Checklist → …
  • Tasks that involve Backup and disaster recovery
  • SKILL.md covers Quick Reference, When to Use This Skill, Prerequisites and MCP Tools, plus 6 more sections
  • Calls az, terraform and kind

What it does

Azure Reliability is an agent skill from microsoft/GitHub-Copilot-for-Azure, published by the product's own GitHub organization. Assess and improve the reliability posture of PaaS Applications (Azure Functions and Azure App Service). Scans deployed resources for zone redundancy, ZRS storage, health probes, and multi-region failover. Presents a feature-pivoted checklist, then drives staged remediation (CLI or IaC patches) end-to-end with user confirmation. WHEN: "assess reliability", "check reliability", "zone redundant", "multi-region failover", "high availability", "disaster recovery", "single points of failure", "reliability posture"…

Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files, including reference files (for example `references/configure-health-probes.md`, `references/configure-multi-region.md` and `references/configure-storage.md`).

It sits in DevOps & Cloud, covering Backup and disaster recovery. It works with Microsoft Azure, Azure Functions and Model Context Protocol. The repository describes itself as: GitHub Copilot for Azure. The licence is MIT.

When your agent uses it

  • Tasks that involve Backup and disaster recovery

Example prompts

  • “assess reliability”
  • “check reliability”
  • “zone redundant”
  • “/azure-reliability”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Discover Resources
  2. Assess Reliability
  3. Generate Reliability Checklist
  4. Present Fix Plan + Choose Path
  5. (both paths): Re-Assess
  6. (both paths): Multi-region follow-up — ASK and WAIT

What it can do on your machine

Read from SKILL.md and the folder at commit fcf2f3b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • az
    • terraform
    • kind

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use az, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Azure Reliability loads about 5.9k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 137 tokens; SKILL.md has 2,292 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~137
When it runs · the whole SKILL.md, loaded when a task matches
~5.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~25k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/GitHub-Copilot-for-Azure at commit fcf2f3b, republished under its MIT licence (© microsoft). 2,292 words, ~5,889 tokens.

Download SKILL.mdSave it as .claude/skills/azure-reliability/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
azure-reliability
description
Assess and improve the reliability posture of PaaS Applications (Azure Functions and Azure App Service). Scans deployed resources for zone redundancy, ZRS storage, health probes, and multi-region failover. Presents a feature-pivoted checklist, then drives staged remediation (CLI or IaC patches) end-to-end with user confirmation. WHEN: "assess reliability", "check reliability", "zone redundant", "multi-region failover", "high availability", "disaster recovery", "single points of failure", "reliability posture", "resiliency".
license
MIT
metadata.author
Microsoft
metadata.version
0.0.0-placeholder

Azure Reliability Assessment & Configuration

Quick Reference

PropertyDetails
Best forReliability posture assessment, zone redundancy enablement, multi-region failover setup
Primary capabilitiesReliability assessment table, Zone Redundancy Configuration, Multi-Region IaC Generation
Supported servicesAzure Functions, App Service (Container Apps planned for a future version)
MCP toolsAzure Resource Graph queries, Azure CLI commands

When to Use This Skill

Activate this skill when user wants to:

  • "Assess my Function app's reliability"
  • "Assess my Web app's reliability"
  • "Check the reliability of my resource group" (App Service and Functions resources only)
  • "Is my app zone redundant?" (App Service and Functions resources only)
  • "Is my app service plan zone redundant?"
  • "Make my app zone redundant" (App Service and Functions resources only)
  • "Make my app service plan zone redundant"
  • "Set up multi-region failover for my app" (App Service and Functions resources only)
  • "Check my reliability posture"
  • "Find single points of failure" (App Service and Functions resources only)
  • "Enable high availability for my app" (App Service and Functions resources only)
  • "Check disaster recovery readiness"
  • "Improve my app's resilience" (App Service and Functions resources only)

Scope note: This skill currently covers Azure Functions and Azure App Service only. If the user asks about Azure Container Apps reliability, acknowledge that support is planned but not yet available, and only proceed with the parts that apply to App Service and Functions resources in scope.

Prerequisites

  • Authentication: user is logged in to Azure via az login
  • Permissions: Reader access on target subscription/resource group (for assessment)
  • Permissions: Contributor access (for configuration changes)
  • Azure Resource Graph extension: az extension add --name resource-graph

MCP Tools

ToolPurpose
mcp_azure_mcp_extension_cli_generateGenerate az CLI commands for resource queries and configuration
mcp_azure_mcp_subscription_listList available subscriptions
mcp_azure_mcp_group_listList resource groups

Primary query method: Azure Resource Graph via az graph query (requires az extension add --name resource-graph).

Assessment Workflow

Phase 1: Discover Resources
  1. Identify scope — Ask user for resource group, subscription, or app name
  2. Query Azure Resource Graph to discover all resources in scope
  3. Classify resources by service type (Functions, Storage, etc.). If Container Apps is found, note it but do not deep-dive.

Important: Always scope queries to the user's specified resource group or subscription. Add these filters to every Resource Graph query:

  • Resource group: | where resourceGroup =~ '<rg-name>'
  • Subscription: Use --subscriptions <sub-id> flag on az graph query
  • App name: | where name =~ '<app-name>'
Phase 2: Assess Reliability

Two-step assessment: platform-level discovery first, then per-service deep dive.

Step 1 — Platform discovery (find what's there). Use these to enumerate resources in scope and detect cross-cutting reliability gaps:

Platform checkReference
Zone redundancy — discoveryreferences/zone-redundancy-checks.md
Storage redundancy (cross-service)references/storage-redundancy-checks.md
Multi-region & global load balancersreferences/multi-region-checks.md
Front Door / Traffic Manager / App Insights probesreferences/health-probe-checks.md

Step 2 — Per-service deep dive. For each compute resource discovered in Step 1, load the matching service reference. The service reference is the single source of truth for that service's plan/SKU rules, assessment queries, CLI commands, IaC patches (Bicep + Terraform + AVM), and reporting hints.

This skill version ships only the Azure Functions and App Service per-service references. Other compute services are listed below explicitly so the dispatch logic is unambiguous: if a resource matches an unsupported row, do not attempt to load a reference, fabricate CLI commands, or generate IaC patches for it.

Service detectedReference
Azure Functions (microsoft.web/serverfarms with kind contains 'functionapp')references/services/functions/reliability.md
Azure App Service (non-Functions sites: microsoft.web/sites without kind contains 'functionapp', microsoft.web/serverfarms without kind contains 'functionapp')references/services/app-service/reliability.md
Azure Container Apps (microsoft.app/containerapps, microsoft.app/managedenvironments)⚪ Not yet shipped — planned for a future version

Handling unsupported services: If a resource matches an unsupported row above, surface it in the discovery summary, mark it as ⚪ not assessed (planned) in the Phase 3 table, and skip the per-service remediation steps for it. Do not attempt to fabricate CLI commands or IaC patches for those services.

Phase 3: Generate Reliability Checklist

Present findings as a feature-pivoted table: one row per reliability feature (Zone redundancy on compute, Zone-redundant storage, Health probes, Multi-region failover), with a single status indicator and the specific resources that are relevant to that feature. This avoids the noise of one-row-per-resource with mostly n/a cells. Do not assign numeric scores or grades.

🔍 Reliability Assessment — {scope}
─────────────────────────────────────────────────────────────────────────────────────────────
Reliability Feature              Status      Resources
─────────────────────────────────────────────────────────────────────────────────────────────
Zone redundancy — compute        🔴 OFF      • plan-web-ii5trxva2ark4 (P1v3)
                                              • plan-ii5trxva2ark4 (FC1)

Zone-redundant storage           🔴 GRS      • stii5trxva2ark4 (defaulted; no SKU set in IaC)

Health probes                    🔴 OFF      • func-api-ii5trxva2ark4 — needs code change (FC1)
                                              • app-web-ii5trxva2ark4 — no health check path

Multi-region failover            🔴 OFF      • Single region (eastus) only — Front Door not configured
─────────────────────────────────────────────────────────────────────────────────────────────

Want me to fix the 🔴 items? I'll do the quick wins first (App
plan zone redundancy + health checks on supported plans), then ask before
storage migration and multi-region setup. (yes/no)

Rules for the table:

  • Four feature rows, in this order: Zone redundancy — compute · Zone-redundant storage · Health probes · Multi-region failover. Omit a row entirely only if no resource in scope could ever apply to it.
  • Status column is one symbol + one short word, no other characters:
    • 🟢 ON — feature is fully enabled across all relevant resources in scope
    • 🟡 PARTIAL — some resources have it, some don't (or partial config like liveness-only)
    • 🔴 OFF — feature is missing on all relevant resources
    • For storage, replace OFF with the current SKU when relevant (🔴 LRS, 🔴 GRS, 🟢 ZRS, 🟢 GZRS). When no SKU is set in IaC, label as 🔴 GRS (ARM/AVM default) and note that in the resource line.
  • Resources column lists only what's relevant to that feature, one bullet per resource:
    • For "needs fixing" resources, include a short inline reason ((FC1), (defaulted; no SKU set), liveness only, needs code change (FC1)).
    • For resources that are already ON for that feature, mention them on the same row with — already ON so the user sees credit for what's right.
  • Do not include n/a, —, or empty cells. If a feature doesn't apply to any resource in scope, drop the row.
  • Do not include numeric scores, grades, or point totals.
  • End the assessment with a single yes/no question that kicks off the staged remediation flow. Do not enumerate the per-resource fix list here — the user will see it after they say yes (Configuration Workflow Step 1).

UX Note: If the assessment finds the app already has all core reliability features (zone redundancy, ZRS/GZRS storage, health probes), skip the fix-it question and jump straight to Configuration Workflow Step 3 (Multi-region follow-up). Do NOT start any multi-region work without explicit consent.

Configuration Workflow

When user wants to fix findings from the assessment:

⛔ ALWAYS confirm with user before executing changes. Show what will change, any cost implications, and any destructive actions (e.g., environment recreation).

Step 1: Present Fix Plan + Choose Path

After assessment, if user says "fix it" / "improve my reliability" / "enable zone redundancy":

  1. List each fixable finding with the specific action
  2. Flag any cost implications or breaking changes
  3. Ask user which path they want:
I'll start with the quick wins (no downtime, fast):

1. ✏️  Enable zone redundancy on plan-ii5trxva2ark4 (Flex Consumption — no cost change)
2. ✏️  Set health check path to /api/health on func-api-ii5trxva2ark4

Then, separately, I'll ask if you want to upgrade storage:

3. 🕒  Upgrade stii5trxva2ark4 from LRS → ZRS (small cost increase, migration takes hours)
   — Required for full zone redundancy, but I'll confirm with you before starting.

How would you like to apply these changes?

  A) Fix now — Run az CLI commands against your live resources (immediate, one-time)
  B) Patch my IaC — Update your Bicep/Terraform files so changes persist across deploys

(If you use azd or Terraform, option B is recommended so `azd up` won't overwrite changes.)
Path A: Fix Now (CLI)

Run fixes against live resources using az CLI commands. Quick wins first, then ask before the slow storage migration.

The exact CLI commands per service live in the per-service references — pick the one(s) matching the resources discovered in Phase 2:

FixReference
Enable zone redundancy / configure health probes (Functions)references/services/functions/reliability.md
Enable zone redundancy / configure health probes (App Service)references/services/app-service/reliability.md
Upgrade storage replication (cross-service)references/configure-storage.md
Set up multi-region (cross-service)references/configure-multi-region.md
Platform overview / verificationreferences/configure-zone-redundancy.md, references/configure-health-probes.md

Execution order — always quick wins first:

  1. Zone redundancy on compute (fast, in-place property update on the App's plan).

  2. Health probes (Premium / Dedicated only — in-place; for FC1 / Consumption, follow the consent gate in configure-health-probes.md).

  3. Verify the compute changes succeeded before doing anything else.

  4. ⛔ STOP — Ask about storage upgrade. Compute is now zone-redundant, but storage may still be LRS or GRS. Ask the user explicitly:

    ✅ Compute is now zone-redundant.
    
    To be **fully zone-redundant**, your storage account also needs to be upgraded:
      • stii5trxva2ark4: currently `Standard_LRS` → needs `Standard_ZRS`
    
    ⚠️  This is a live storage redundancy conversion:
       • Takes hours to days depending on data volume
       • Small ongoing cost increase (~$0.01/GB/month more)
       • Only supported for Standard general-purpose v2 accounts
    
    Do you want me to start the storage migration now? (yes / no / later)
    • yes → run az storage account update --sku Standard_ZRS (or migration start if needed); poll az storage account show --query sku.name until it reports Standard_ZRS.
    • no / later → leave storage as-is; note in the re-assessment that ZR storage remains a gap.
  5. Multi-region — do NOT auto-run. Handled in Step 3 below as an explicit follow-up after re-assessment.

⚠️ Warning: If the user uses azd up or terraform apply later, CLI-only changes may be overwritten by the IaC definitions. Recommend also patching IaC after CLI fixes.

Show full SKILL.md (1,031 more words)Show less
Path B: Patch IaC

Update the user's Bicep or Terraform files so reliability settings are persistent.

Step 1: Detect IaC type

  1. Look for infra/ folder in project root
  2. If not found, check project root for *.bicep or *.tf files
  3. If still not found, ask user: "Where are your IaC files located?"
  4. Check for *.bicep files → use Bicep patching
  5. Check for *.tf files → use Terraform patching
  6. If both exist, ask user which to patch
  7. If no IaC exists, fall back to Path A (CLI) and inform user

Step 2: Classify each fix by risk level

FixRisk LevelWhat Happens
Zone redundancy (App plan)🟢 Safe patchIn-place property update on next deploy
Storage LRS → ZRS🟡 Pre-migration requiredLive storage migration must complete before the IaC SKU change can deploy. Never bundle with safe patches — use the two-deploy flow in Steps 3–5.
Health check path (Basic/Standard/Premium / Dedicated)🟢 Safe patchIn-place update, but causes app restart
Health check path (FC1 / Consumption)⚪ Code-only — ask firsthealthCheckPath is unsupported. Adding a health endpoint requires adding an HTTP-triggered /api/health function to app code. Always ask the user for explicit consent before touching source code. Do not patch IaC.

Step 3: Apply patches in two deploys (quick wins first)

The IaC patching framework (detection, AVM-module guidance, deploy-order rule, storage SKU patch) lives in:

The actual per-service compute patches (Function App plan ZR, App Service Plan ZR, etc.) live in the per-service references — load the matching service file from Phase 2 for the exact Bicep / Terraform / AVM snippets. Only Azure Functions and App Service have per-service references in this skill version; Container Apps is out of scope.

Deploy 1 — Quick wins only. Patch the 🟢 Safe items (zone redundancy on the App Service/Function App plan, health probes on Basic/Standard/Premium / Dedicated). Do NOT include the storage SKU patch in this deploy.

After patching, the skill runs the deploy itself (do not stop and tell the user to run it). Detect the deployment tool and confirm once before executing:

📦 Patches applied to your IaC. Ready to deploy:
   Tool detected: azd (found azure.yaml)
   Command:       azd up

Proceed with deployment? (yes / no)

On yes, run the appropriate command, stream output back to the user, and continue to the next step on success:

  • AZD project (has azure.yaml): azd up
  • Bicep-only: az deployment group create --resource-group <rg> --template-file infra/main.bicep --parameters @infra/main.parameters.json
  • Terraform: terraform plan -out tfplan → (show plan summary) → terraform apply tfplan

On no, stop and report the patched files; do not proceed to Step 4 / Re-Assess.

If deployment fails, surface the error and stop — do not continue to the storage step.

⛔ STOP — Ask about storage upgrade before Deploy 2. After Deploy 1 succeeds, ask the user explicitly:

✅ Quick-win patches deployed. Compute is now zone-redundant.

To be **fully zone-redundant**, your storage account also needs to be upgraded:
  • stii5trxva2ark4: currently `Standard_LRS` → needs `Standard_ZRS`

⚠️  This is a two-part change:
   1. Live storage migration (`az storage account migration start`) — takes hours to days
   2. A second deploy to update your IaC's storage SKU to match

Do you want me to start the storage migration now? (yes / no / later)
  • yes → the skill runs the migration command itself, polls until complete, then patches the storage SKU in IaC and runs Deploy 2 (now a no-op confirmation). The user does not need to run anything manually.
  • no / later → leave the storage SKU patch unapplied. Note in the re-assessment that ZR storage remains a gap; suggest revisiting later.

Step 4: Storage migration (only if user said yes in Step 3)

The skill runs these commands itself — do not ask the user to run them. Show progress as you go:

🔄 Starting storage migration (this can take up to 72 hours)...

   az storage account migration start --name stii5trxva2ark4 \
     --resource-group rg-example --sku Standard_ZRS --no-wait

   Polling: az storage account show --name stii5trxva2ark4 --query sku.name
   ...
   ✅ Migration complete: sku.name = Standard_ZRS

For very long migrations, you may surface a checkpoint to the user ("this is still running, check back later") rather than blocking the entire conversation.

Step 5: Deploy 2 — storage SKU patch

After the migration completes, the skill patches the storage SKU in IaC and runs the same deploy command as Step 3 (e.g. azd up). This deploy is a no-op confirmation that the IaC matches the live state. Confirm once with the user before executing, then run it directly.

Step 2 (both paths): Re-Assess

After changes are applied (CLI) or deployed (IaC), automatically re-run the assessment and show the same feature-pivoted table as Phase 3, with each feature row's status updated to reflect the new state. Briefly call out what changed since the previous run.

🔄 Reliability Re-Assessment — rg-eventhubs-python-jan13 (eastus)
───────────────────────────────────────────────────────────────────────────────────────
Reliability Feature              Status      Resources
───────────────────────────────────────────────────────────────────────────────────────
Zone redundancy — compute        🟢 ON       • plan-ii5trxva2ark4 (FC1)              — now ON
                                             • plan-web-ii5trxva2ark4 (P1v3)         — now ON

Zone-redundant storage           🟢 ZRS      • stii5trxva2ark4                       — GRS → ZRS

Health probes                    🟡 PARTIAL  • func-api-ii5trxva2ark4                — still off (FC1, code change declined)
                                             • app-web-ii5trxva2ark4                 — now ON

Multi-region failover            🔴 OFF      • Single region (eastus) only
───────────────────────────────────────────────────────────────────────────────────────

What changed: Function App and App Service plan zone redundancy, storage replication and health probes on App Service.
(Multi-region offered next — see Step 3.)
Step 3 (both paths): Multi-region follow-up — ASK and WAIT

Multi-region is a significant cost/complexity step. Do NOT start it automatically. After re-assessment, only if all core single-region reliability features are 🟢 ON (zone-redundant compute, ZRS/GZRS storage, health probes), explicitly ask the user and wait for their response before doing anything:

🟢 Your app is now fully zone-redundant in {region}.

The next step (optional) is multi-region failover with Azure Front Door:
   • Deploys compute + storage in a second region (paired region recommended)
   • Adds Azure Front Door for global load balancing with health-probe-driven failover
   • Protects against full region outages
   • Estimated additional cost: ~2x compute (active-passive); Front Door ~$35/month base

Do you want me to set up multi-region failover now? (yes / no / later)
  • yes → proceed with references/configure-multi-region.md. Confirm secondary region choice with the user, then:
    1. Generate the multi-region IaC (Bicep / Terraform additions for the secondary region + Front Door).
    2. Confirm once with the user: 📦 Multi-region IaC generated. Ready to deploy with \azd up`. Proceed? (yes / no)`
    3. On yes, the skill runs the deploy itself (azd up / az deployment group create / terraform apply) and streams output. Do not stop and tell the user to run it.
    4. After successful deploy, run a final re-assessment so the user sees Multi-region failover flip to 🟢 ON.
  • no / later → leave the deployment as-is. Note that single-region zone-redundant is a reliable end state; multi-region can be revisited anytime.

⛔ Do not skip the wait. Do not generate multi-region IaC, deploy a Front Door, or modify any files until the user has explicitly said yes. If core reliability is not yet all 🟢, do not ask about multi-region — finish the core gaps first.

Priority Classification

PriorityCriteriaAction
CriticalNo zone redundancy AND production workloadFix immediately
HighLRS storage on zone-redundant computeFix within days
MediumNo multi-region (single region but zone-redundant)Plan for next sprint
LowMissing health probes or monitoring gapsTrack and fix

Error Handling

ErrorMessageRemediation
Authentication required"Please login"Run az login and retry
Access denied"Forbidden"Confirm Reader/Contributor role assignment
Plan doesn't support ZR"Upgrade required"Inform user of plan upgrade path + cost delta
Region doesn't support AZ"Region limitation"Suggest supported regions

Best Practices

  • Run reliability assessments after every significant infrastructure change
  • Test failover scenarios periodically (at least quarterly)

Skill Boundaries

ActionThis skill doesHand off to
Assess reliability posture✅ Yes—
Recommend improvements✅ Yes—
Enable zone redundancy (CLI commands)✅ Yes—
Patch Bicep/Terraform for reliability✅ Yes—
Generate multi-region IaC✅ Yes (additions for the secondary region + Front Door)azure-prepare for full new-app IaC scaffolding
Deploy IaC for reliability changes✅ Yes (runs azd up / terraform apply / az deployment itself, after user confirmation)azure-deploy for general/non-reliability deploys
Validate pre-deploymentReliability checks onlyazure-validate for full validation

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (references) in plugins/azure-skills/skills/azure-reliability of microsoft/GitHub-Copilot-for-Azure.

  • SKILL.md
  • references/configure-health-probes.md
  • references/configure-multi-region.md
  • references/configure-storage.md
  • references/configure-zone-redundancy.md
  • references/health-probe-checks.md
  • references/iac-patching-bicep.md
  • references/iac-patching-terraform.md
  • references/multi-region-checks.md
  • references/services/app-service/reliability.md
  • references/services/functions/reliability.md
  • references/storage-redundancy-checks.md
  • references/zone-redundancy-checks.md
  • version.json

Open the folder on GitHubat commit fcf2f3b

Used in 3 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in microsoft/GitHub-Copilot-for-Azure, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Azure Reliability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Azure Reliability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Azure Reliability this skillmicrosoft/GitHub-Copilot-for-Azure2551 repos~5.9kAutomated safety check: PassMIT
Apex Azure Reliabilityjonathan-vella/apex217—~2.4kAutomated safety check: PassMIT
Azure Data API BuilderMicrosoftDocs/Agent-Skills777—~4.3kAutomated safety check: PassCC-BY-4.0
Terravision Cloud Diagramspatrickchugh/terravision1.6k—~5.6kAutomated safety check: NotesAGPL-3.0-only
Spotinfoalexei-led/spotinfo164—~1.8kAutomated safety check: PassApache-2.0
Install Boltmcpboltmcp/boltmcp371—~2.3kAutomated safety check: PassNone

Similar skills

  • Apex Azure Reliability

    jonathan-vella/apex

    ANALYSIS SKILL — Read-only reliability assessment for App Service and Azure Functions: zone redundancy, ZRS storage, health probes and multi-region failover.

    217 GitHub stars~2.4k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Azure Data API Builder

    MicrosoftDocs/Agent-Skills

    Official

    Expert knowledge for Azure Data Api Builder development including troubleshooting, best practices, decision making, limits & quotas, security, configuration, integrations & coding patterns, and…

    777 GitHub stars~4.3k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Terravision Cloud Diagrams

    patrickchugh/terravision

    Draw cloud architecture diagrams for AWS, Azure or GCP with the official provider icon sets, using TerraVision.

    1.6k GitHub stars~5.6k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Spotinfo

    alexei-led/spotinfo

    Query Spot/preemptible VM prices, savings and interruption risk across AWS, GCP and Azure with the spotinfo CLI.

    164 GitHub stars~1.8k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Install Boltmcp

    boltmcp/boltmcp

    A skill your agent uses when asked to help install or uninstall BoltMCP

    371 GitHub stars~2.3k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Apex Azure Diagnostics

    jonathan-vella/apex

    WORKFLOW SKILL — Debug Azure production issues: Container Apps, Functions, App Service, AKS, VMs and messaging, with KQL log analysis.

    217 GitHub stars~2.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from microsoft/GitHub-Copilot-for-Azure

All 56 skills in this repo
  • Capacity

    microsoft/GitHub-Copilot-for-Azure

    Official

    Discovers available Azure OpenAI model capacity across regions and projects.

    255 GitHub starsUsed in 2 repos~1.7k tokens
    Auto-check passed
  • Deploy Model

    microsoft/GitHub-Copilot-for-Azure

    Official

    Unified Azure OpenAI model deployment skill with intelligent intent-based routing.

    255 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Entra Agent Id

    microsoft/GitHub-Copilot-for-Azure

    Official

    Provision Microsoft Entra Agent Identity Blueprints, BlueprintPrincipals, and per-instance Agent Identities via Microsoft Graph, and configure OAuth 2.0 token exchange (fmipath, OBO, cross-tenant)…

    255 GitHub starsUsed in 3 repos~4k tokens
    Auto-check passed
  • Microsoft Foundry

    microsoft/GitHub-Copilot-for-Azure

    Official

    Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end.

    255 GitHub starsUsed in 1 repo~6.7k tokens
    Auto-check passed
  • Azure Storage

    microsoft/GitHub-Copilot-for-Azure

    Official

    Azure Storage Services including Blob Storage, File Shares, Queue Storage, Table Storage, and Data Lake.

    255 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • Preset

    microsoft/GitHub-Copilot-for-Azure

    Official

    Intelligently deploys Azure OpenAI models to optimal regions by analyzing capacity across all available regions.

    255 GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed

Categories

Questions about Azure Reliability

What does Azure Reliability do?

Assess and improve the reliability posture of PaaS Applications (Azure Functions and Azure App Service). Azure Reliability is an agent skill from microsoft/GitHub-Copilot-for-Azure, published by the product's own GitHub organization. Assess and improve the reliability posture of PaaS Applications (Azure Functions and Azure App Service).

When should I use Azure Reliability?

Azure Reliability fits situations like: tasks that involve Backup and disaster recovery.

How do I install Azure Reliability in Claude Code?

Run `npx skills add microsoft/GitHub-Copilot-for-Azure --skill azure-reliability -a claude-code`. Or copy the skill folder (plugins/azure-skills/skills/azure-reliability in microsoft/GitHub-Copilot-for-Azure) into .claude/skills/azure-reliability in your project. Claude Code loads it when a task matches its description.

How do I install Azure Reliability in Codex?

Run `npx skills add microsoft/GitHub-Copilot-for-Azure --skill azure-reliability -a codex`. Or copy the skill folder (plugins/azure-skills/skills/azure-reliability in microsoft/GitHub-Copilot-for-Azure) into .agents/skills/azure-reliability in your project. Codex loads it when a task matches its description.

Can I use Azure Reliability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/GitHub-Copilot-for-Azure --skill azure-reliability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/azure-reliability, .gemini/skills/azure-reliability, .github/skills/azure-reliability and .opencode/skills/azure-reliability in your project.

What does Azure Reliability need to run?

Going by SKILL.md and its folder, Azure Reliability needs the command-line tools its instructions call (az, terraform and kind).

Does Azure Reliability access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Azure Reliability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Azure Reliability use?

Azure Reliability is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Azure Reliability use?

About 5.9k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 19k tokens, read only when the agent opens those files.

What are the alternatives to Azure Reliability?

Skills that share tags, products or a category with Azure Reliability: Apex Azure Reliability (jonathan-vella/apex, 217 stars), Azure Data API Builder (MicrosoftDocs/Agent-Skills, 777 stars), Terravision Cloud Diagrams (patrickchugh/terravision, 1.6k stars) and Spotinfo (alexei-led/spotinfo, 164 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Azure Reliability?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/GitHub-Copilot-for-Azure, which has 255 GitHub stars. The repository holds 56 skills in this directory. The repository was last updated on October 8, 2026.

Source: microsoft/GitHub-Copilot-for-Azure on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.