Official agent skill

Sdaf Failure Triage

by Azure in Azure/sap-automation

Triage a failed or suspicious SDAF run: distinguish a real failure from a green no-op or a clean plan reported as failure, map the observed symptom to a documented cause in…

OfficialMITAuto-check passedDevOps & Cloud

Install Sdaf Failure Triage

skills CLI
$ npx skills add Azure/sap-automation --skill sdaf-failure-triage -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Azure/sap-automation sdaf-failure-triage --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Azure/sap-automation.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/sdaf-failure-triage .claude/skills/sdaf-failure-triage && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sdaf-failure-triage
GitHub stars
146
Token cost
~1.7k tokens
SKILL.md length
688 words
Files
2 (incl. references)
Skills in repo
19
Repo updated
First seen
Licence
MIT

At a glance

Triage a failed or suspicious SDAF run: distinguish a real failure from a green no-op or a clean plan reported as failure, map the observed symptom to a documented cause in…

  • Works in 4 steps: Establish exit-code intent → Map the symptom → Route to the stage-owning skill → …
  • A user says SDAF run failed
  • SKILL.md covers When to invoke, Preconditions, Recipe and Special cases (documented), plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Sdaf Failure Triage is an agent skill from Azure/sap-automation, published by the product's own GitHub organization. Triage a failed or suspicious SDAF run: distinguish a real failure from a green no-op or a clean plan reported as failure, map the observed symptom to a documented cause in docs/local/troubleshooting.md, and hand off to the stage-owning skill for retry. Use when a user says "SDAF run failed", "my deploy exited non-zero", "the plan was clean but exit 1", "the run said success but nothing was deployed", "state lock error", "unexpected replacement", "control plane stopped partway", "generated hosts.yaml missing"…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/symptom-map.md`).

It sits in DevOps & Cloud. The repository describes itself as: This is the repository supporting the SAP deployment automation framework on Azure. The licence is MIT.

When your agent uses it

  • A user says SDAF run failed
  • My deploy exited non-zero
  • The plan was clean but exit 1
  • The run said success but nothing was deployed

Example prompts

  • “SDAF run failed”
  • “my deploy exited non-zero”
  • “the plan was clean but exit 1”
  • “/sdaf-failure-triage”

Requirements

  • Pre-approved tools (allowed-tools): shell

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Establish exit-code intent
  2. Map the symptom
  3. Route to the stage-owning skill
  4. Confirm before retry

What it can do on your machine

Read from SKILL.md and the folder at commit 78835f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • shell

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sdaf Failure Triage loads about 1.7k tokens when it runs, and up to ~3k if it reads all its reference files. Until then it costs about 176 tokens; SKILL.md has 688 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~176
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Azure/sap-automation at commit 78835f0, republished under its MIT licence (© Azure). 688 words, ~1,734 tokens.

Download SKILL.mdSave it as .claude/skills/sdaf-failure-triage/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
sdaf-failure-triage
description
Triage a failed or suspicious SDAF run: distinguish a real failure from a green no-op or a clean plan reported as failure, map the observed symptom to a documented cause in `docs/local/troubleshooting.md`, and hand off to the stage-owning skill for retry. Use when a user says "SDAF run failed", "my deploy exited non-zero", "the plan was clean but exit 1", "the run said success but nothing was deployed", "state lock error", "unexpected replacement", "control plane stopped partway", "generated hosts.yaml missing", "workload-zone private endpoint failure", "workload-zone subnet policy failure", or "SDAF exit 2". Do NOT use to actually deploy or to redesign the workspace layout.
allowed-tools
shell
license
MIT

SDAF Failure Triage

Action-loop skill. The entry point for "it broke". Maps observed symptoms to the documented cause in docs/local/troubleshooting.md (and the stage docs' own § Validate / § Configuration preparation sections), then hands off to the stage-owning skill. This skill is a router, not the owner of any per-stage fix.

When to invoke

Trigger on: "SDAF run failed", "the plan was clean but the run exited 1", "success but nothing deployed", "state lock", "unexpected replacement", "control plane stopped partway", "generated hosts.yaml missing", "BOM files not found", "removal is incomplete", "Ansible playbook failed", "workload-zone private endpoint failure", "workload-zone subnet policy failure", "SDAF exit 2".

Do NOT trigger on: pre-flight readiness (sdaf-readiness-check), authoring tfvars (sdaf-workspace-and-tfvars), fresh deploy of a stage.

Preconditions

  • Collect the exit code and the last 200 lines of the run log before invoking. If you have neither, ask for them; do not guess.

Recipe

Step 1 — Establish exit-code intent
  • 0 — Terraform / script signalled success. Still worth verifying the expected artefacts exist (some symptoms below are "success but nothing deployed").
  • non-zero — real failure OR a validation-gate refusal (exit 2 from validate.sh or a stage script). Treat exit 2 as a validation gate that was intentionally tripped, not a bug; do not silence it.
Step 2 — Map the symptom

Walk references/symptom-map.md, which lists every named failure class from docs/local/troubleshooting.md plus the documented workload-zone private-endpoint / subnet-policy case, with the owning skill for removal/state/install/HA/media/sovereign routes. Pick the first row whose symptom matches the observed signature.

If nothing matches, walk the "canonical 'clean plan reported as failure' note" in references/symptom-map.md, which walks the stage § Validate sections. If nothing there matches either, say docs are silent on this symptom and stop — do not invent a cause.

Step 3 — Route to the stage-owning skill

Route to the stage-owning skill for the actual retry:

Stage / symptomOwning skill
Control planesdaf-control-plane-bootstrap
Workload zone (including the private-endpoint / subnet-policy workaround)sdaf-workload-zone
SAP systemsdaf-sap-system
SAP installation / numbered playbookssdaf-sap-installation
Media acquisition preconditions / clean downloader pathsdaf-media-acquisition
Media archive / checksum / extractor / BOM-processing failuressdaf-media-diagnostics
HA cluster evidence / crm / pcs / fencingsdaf-ha-diagnostics
Explicit Terraform state lock / import / remove / drift repairsdaf-state-management
Safe teardown / incomplete removal / control-plane step trapsdaf-safe-removal
Azure Government / sovereign-cloud deltassdaf-sovereign-cloud
WORKSPACES/tfvars authoring issuesdaf-workspace-and-tfvars
BOM file location / selectionsdaf-bom-selection
Step 4 — Confirm before retry

The retry safety rules — no --force / no --auto-approve / no state edits / no concurrent execution / do not delete .progress markers — are enforced by the shipped repo instructions (.github/copilot-instructions.md) and by each stage skill's own recipe. This triage skill defers to the stage-owning skill's retry section rather than restating those rules.

Show full SKILL.md (264 more words)Show less

Special cases (documented)

Interrupted control-plane removal reports success

docs/local/troubleshooting.md § Removal is incomplete and docs/local/07-00-operations.md § Remove resources: an interrupted control-plane removal can exit "successfully" after step=1 without deleting the deployer. Diagnostic path: inspect .sap_deployment_automation, persisted step, library destroy result, and remaining deployer state / resources. Do not edit step to bypass the guard. sdaf-safe-removal owns the documented diagnostic path and any approved retry.

success but nothing was deployed

Cross-check the expected artefacts for the stage:

  • Control plane: state in library storage account, metadata under .sap_deployment_automation, summary written (docs/local/03-00-control-plane.md § Validate, § Outcome).
  • Workload zone: .tfvars and backend metadata in the state account's tfvars container, zone Key Vault deployed (docs/local/04-00-workload-zone.md § Validate, § Outcome).
  • SAP system: infra deployed and <SID>_hosts.yaml / sap-parameters.yaml present (docs/local/05-00-sap-system.md § Validate, § Outcome).

If artefacts are missing, treat the "success" as false and route to the stage-owning skill.

Hard rules

  • Documented behaviour only (D19). If the symptom does not match a docs/local/troubleshooting.md section, a documented § Configuration preparation / § Validate cross-check, or a named owner row in references/symptom-map.md, say docs are silent and stop.
  • Do not silently pass a non-zero exit code.
  • Do not narrate benign log noise as a failure without a documented anchor.

What this skill does NOT do

  • Does not deploy or retry directly.
  • Does not repair Terraform state; sdaf-state-management owns that.
  • Does not restate stage-specific safety rules — the stage skills and .github/copilot-instructions.md own those.
  • Does not replace the documented media/install/HA/removal/sovereign owners.
  • Does not reason about undocumented failure modes (e.g. end-to-end Government beyond sdaf-sovereign-cloud, air-gapped, or undocumented ARM_ENVIRONMENT flows).

See also

  • sdaf-control-plane-bootstrap, sdaf-workload-zone, sdaf-sap-system, sdaf-sap-installation, sdaf-media-acquisition, sdaf-media-diagnostics, sdaf-ha-diagnostics, sdaf-state-management, sdaf-safe-removal, sdaf-sovereign-cloud, sdaf-workspace-and-tfvars, sdaf-bom-selection.
  • docs/local/troubleshooting.md, docs/local/07-00-operations.md, docs/local/04-00-workload-zone.md § Configuration preparation.

© Azure, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/sdaf-failure-triage of Azure/sap-automation.

  • SKILL.md
  • references/symptom-map.md

Open the folder on GitHubat commit 78835f0

Compare with similar skills

Sdaf Failure Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sdaf Failure Triage compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sdaf Failure Triage this skillAzure/sap-automation146—~1.7kAutomated safety check: PassMIT
Monitor CInrwl/nx29k6 repos~4.7kAutomated safety check: PassMIT
Terraform and OpenTofu Guideagentscope-ai/QwenPaw36k6 repos~4.2kAutomated safety check: PassApache-2.0
Vercel Optimize Auditvercel-labs/agent-skills32k8 repos~4.3kAutomated safety check: PassNone
Analyze GitHub Action Logswithastro/astro63k1 repos~1.3kAutomated safety check: PassCustom licence
Openclaw Live Updateropenclaw/openclaw392k—~3.7kAutomated safety check: PassMIT

Similar skills

  • Monitor CI

    nrwl/nx

    Monitor Nx Cloud CI pipeline and handle self-healing fixes. An agent skill from nrwl/nx.

    29k GitHub starsUsed in 6 repos~4.7k tokens
    DevOps & CloudAuto-check passed
  • Terraform and OpenTofu Guide

    agentscope-ai/QwenPaw

    Guidance for writing and testing Terraform and OpenTofu code: module structure, naming, test approaches, CI/CD workflows, state handling and security scanning.

    36k GitHub starsUsed in 6 repos~4.2k tokens
    DevOps & CloudAuto-check passed
  • Vercel Optimize Audit

    vercel-labs/agent-skills

    Official

    Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.

    32k GitHub starsUsed in 8 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Official

    Analyze recent GitHub Actions workflow runs to identify patterns, mistakes, and improvements.

    63k GitHub starsUsed in 1 repo~1.3k tokens
    DevOps & CloudAuto-check passed
  • Openclaw Live Updater

    openclaw/openclaw

    Maintain the canonical live OpenClaw main checkout, macOS LaunchAgent-managed Gateway, local macOS app, exact-head main CI, and recurring full release validation.

    392k GitHub stars~3.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Docs Learn PR Preview

    netdata/netdata

    Use only when the user explicitly asks to build, run, preview, inspect, or validate learn.netdata.cloud locally using the contents of a PR or documentation branch before merge.

    81k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed

More from Azure/sap-automation

All 19 skills in this repo
  • Sdaf Bom Selection

    Azure/sap-automation

    Official

    Pick the right SDAF BOM for a target SAP product / release / DB platform / version / kernel / topology.

    146 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Sdaf Orientation And Surface

    Azure/sap-automation

    Official

    Orient a newcomer to the SAP Deployment Automation Framework (SDAF): explain the spine (control plane → workload zone → SAP system → software → install → operate/remove), summarise the three…

    146 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Sdaf Quality Assurance

    Azure/sap-automation

    Official

    Validate a deployed SDAF SAP system through the SDAF-owned QA entry points: the local quality-assurance menu and the documented Azure DevOps pipeline 13 path.

    146 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Sdaf Sap Installation

    Azure/sap-automation

    Official

    Guide SDAF operating-system, database, and SAP installation after the SAP-system workspace and reviewed media are ready.

    146 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Sdaf Sovereign Cloud

    Azure/sap-automation

    Official

    Explain the current SDAF sovereign-cloud deltas without inventing a generic "all sovereigns" runbook.

    146 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Sdaf State Management

    Azure/sap-automation

    Official

    Inspect and repair SDAF Terraform state safely before any reviewed import/remove.

    146 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Sdaf Failure Triage

What does Sdaf Failure Triage do?

Triage a failed or suspicious SDAF run: distinguish a real failure from a green no-op or a clean plan reported as failure, map the observed symptom to a documented cause in…. Sdaf Failure Triage is an agent skill from Azure/sap-automation, published by the product's own GitHub organization.md, and hand off to the stage-owning skill for retry.

When should I use Sdaf Failure Triage?

Sdaf Failure Triage fits situations like: A user says SDAF run failed; my deploy exited non-zero; the plan was clean but exit 1; the run said success but nothing was deployed.

How do I install Sdaf Failure Triage in Claude Code?

Run `npx skills add Azure/sap-automation --skill sdaf-failure-triage -a claude-code`. Or copy the skill folder (skills/sdaf-failure-triage in Azure/sap-automation) into .claude/skills/sdaf-failure-triage in your project. Claude Code loads it when a task matches its description.

How do I install Sdaf Failure Triage in Codex?

Run `npx skills add Azure/sap-automation --skill sdaf-failure-triage -a codex`. Or copy the skill folder (skills/sdaf-failure-triage in Azure/sap-automation) into .agents/skills/sdaf-failure-triage in your project. Codex loads it when a task matches its description.

Can I use Sdaf Failure Triage in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Azure/sap-automation --skill sdaf-failure-triage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sdaf-failure-triage, .gemini/skills/sdaf-failure-triage, .github/skills/sdaf-failure-triage and .opencode/skills/sdaf-failure-triage in your project.

What does Sdaf Failure Triage need to run?

SKILL.md names no scripts, command-line tools or credentials: Sdaf Failure Triage is instructions for the agent only. Its frontmatter pre-approves these tools: shell.

Does Sdaf Failure Triage access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sdaf Failure Triage safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sdaf Failure Triage use?

Sdaf Failure Triage is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sdaf Failure Triage use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Sdaf Failure Triage?

Skills that share tags, products or a category with Sdaf Failure Triage: Monitor CI (nrwl/nx, 29k stars), Terraform and OpenTofu Guide (agentscope-ai/QwenPaw, 36k stars), Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars) and Analyze GitHub Action Logs (withastro/astro, 63k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sdaf Failure Triage?

Azure (a GitHub organization, an official publisher) maintains it in Azure/sap-automation, which has 146 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 8, 2026.

Source: Azure/sap-automation on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.