Agent skill

Analyze Disruption

by openshift-eng in openshift-eng/ai-helpers

A skill your agent uses when analyzing disruption from a DisruptionRegression alert, a Grafana disruption dashboard URL, or Prow CI job runs by examining interval data, audit logs, pod logs, and CPU…

Apache-2.0Auto-check passedDevOps & Cloud

Install Analyze Disruption

skills CLI
$ npx skills add openshift-eng/ai-helpers --skill analyze-disruption -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install openshift-eng/ai-helpers analyze-disruption --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ci/skills/analyze-disruption .claude/skills/analyze-disruption && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
analyze-disruption
GitHub stars
120
Token cost
~2.8k tokens
SKILL.md length
1,466 words
Files
13 (incl. references)
Skills in repo
118
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when analyzing disruption from a DisruptionRegression alert, a Grafana disruption dashboard URL, or Prow CI job runs by examining interval data, audit logs, pod logs, and CPU…

  • Works in 9 steps: Parse and Validate Input → 5: Resolve a Grafana URL or Alert to Job… → Download Artifacts for All Runs → …
  • Analyzing disruption from a DisruptionRegression alert
  • SKILL.md covers Prerequisites, Input Format, Bundled Resources and Implementation Steps, plus 2 more sections
  • Runs Python scripts from its folder; calls gcloud; reaches prow.ci.openshift.org and grafana-loki.ci.openshift.org; needs CLOUDSDK_AUTH_DISABLE_CREDENTIALS

What it does

Analyze Disruption is an agent skill from openshift-eng/ai-helpers. Use when analyzing disruption from a DisruptionRegression alert, a Grafana disruption dashboard URL, or Prow CI job runs by examining interval data, audit logs, pod logs, and CPU metrics

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including reference files (for example `download_timelines.py`, `find_disruption_runs.py` and `parse_disruption.py`).

It sits in DevOps & Cloud, covering Monitoring and alerting. It works with Grafana and Google Cloud. The repository describes itself as: Developer productivity tools for Claude Code & other AI assistants. The licence is Apache-2.0.

When your agent uses it

  • Analyzing disruption from a DisruptionRegression alert
  • A Grafana disruption dashboard URL
  • Prow CI job runs by examining interval data

Example prompts

  • “/analyze-disruption”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Parse and Validate Input
  2. 5: Resolve a Grafana URL or Alert to Job Runs
  3. Download Artifacts for All Runs
  4. Analyze Interval/Timeline Data
  5. Deep-Dive Artifact Download (Optional)
  6. Additional Diagnostic Checks
  7. Cross-Run Comparison (Multiple Runs Only)
  8. Generate Report
  9. Known Disruption Issue Lookup

What it can do on your machine

Read from SKILL.md and the folder at commit a627176. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • gcloud

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • prow.ci.openshift.org
    • grafana-loki.ci.openshift.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • CLOUDSDK_AUTH_DISABLE_CREDENTIALS

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Analyze Disruption loads about 2.8k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 51 tokens; SKILL.md has 1,466 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~51
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from openshift-eng/ai-helpers at commit a627176, republished under its Apache-2.0 licence (© openshift-eng). 1,466 words, ~2,805 tokens.

Download SKILL.mdSave it as .claude/skills/analyze-disruption/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
analyze-disruption
description
Use when analyzing disruption from a DisruptionRegression alert, a Grafana disruption dashboard URL, or Prow CI job runs by examining interval data, audit logs, pod logs, and CPU metrics

Analyze Disruption

This skill analyzes disruption events recorded in Prow CI job runs. It downloads interval/timeline data, audit logs, and pod logs, then correlates disruption across backends and job runs to identify root causes.

Prerequisites

  1. gcloud CLI Installation

    • Check if installed: which gcloud
    • The test-platform-results-public bucket is public. Do not run gcloud auth login for these artifacts, and do not change the user's gcloud config.
    • download_timelines.py sets CLOUDSDK_AUTH_DISABLE_CREDENTIALS=true on each gcloud subprocess (listing and copying, including parallel workers). That avoids refreshing expired local credentials.
    • Any direct gcloud storage command in this skill uses the same variable. Anonymous access still needs network execution when the command runs in a sandbox.
  2. Python 3 (3.7 or later)

Input Format

The user will provide one of the following as input:

Option A — Prow job URLs (direct analysis):

  1. One or more Prow job URLs (at least 1)
    • Example: https://prow.ci.openshift.org/view/gs/test-platform-results-public/logs/periodic-ci-openshift-release-master-ci-4.21-e2e-aws-ovn/1983307151598161920

Option B — Grafana disruption dashboard URL (run discovery + analysis):

  1. A Grafana disruption dashboard URL — the skill extracts filter parameters, finds matching job runs via Sippy, and presents candidates for the user to select before analysis
    • Example: https://grafana-loki.ci.openshift.org/d/gEdw_aLvk/disruption-for-5-0-os-agnostic?var-platform=gcp&var-backend=host-to-host-new-connections&var-upgrade_type=micro&var-architectures=amd64&var-topologies=ha&var-networks=ovn&var-releases=5.0
    • Recognized by hostname grafana-loki.ci.openshift.org and path starting with /d/
    • Every var-backend value is an exact backend name. Pass that full list as --backends unless the user overrides it. Do not reduce it to one API backend or to a stripped base name

Option C — disruption regression alert text (run discovery + analysis):

  1. A pasted DisruptionRegression alert — DisruptionRegressionP50, DisruptionRegressionP50Azure, DisruptionRegressionP75, or DisruptionRegressionP95, or a label block that contains alertname and backend
    • Example labels: backend, platform, upgrade_type, master_nodes_updated, architecture, topology, network, release, delta, feature_set, os, compare_release
    • A link: annotation, when present, is the dashboard URL. Pass the alert text through unchanged. Do not rebuild the URL, and do not drop master_nodes_updated, feature_set, or os
    • With no link, the finder uses the alert labels and a 3-day lookback, which is the regression window these alerts measure

Optional flags (all input options):

  1. --backends flag (optional) — comma-separated exact backend names to focus on

    • Example: --backends kube-api-new-connections,oauth-api-new-connections,openshift-api-new-connections
    • If omitted, analyze all backends that show disruption (Option A) or use every Grafana var-backend / alert backend value (Option B or C)
  2. --skip-jira flag (optional) — skip the Jira search for known disruption cards

    • By default, the skill searches TRT and OCPBUGS for existing disruption cards after analysis

Bundled Resources

Load these only at the step that needs them — not up front:

  • references/run-discovery.md — Grafana URL and alert resolution into Prow job runs (Step 1.5)
  • references/artifacts.md — timeline download and optional audit/etcd/PromQL deep dive (Steps 2 and 4)
  • references/timeline-analysis.md — parser modes, signal interpretation, and extra diagnostic checks (Steps 3 and 5)
  • references/cross-run-comparison.md — multi-run pattern detection and same-job clean comparison (Step 6)
  • references/report-guide.md — deep-link shapes, inline linking rules, and report structure (Step 7)
  • references/known-issues.md — Jira search and bug filing (Step 8)

Implementation Steps

Step 1: Parse and Validate Input
  1. Extract URLs and flags

    • Parse --backends flag if present, split on comma to get backend filter list
    • Parse --skip-jira flag as a boolean option (default: false)
    • Collect all positional URL arguments
  2. Detect input type:

    • Alert text: alertname is DisruptionRegressionP50, DisruptionRegressionP50Azure, DisruptionRegressionP75, or DisruptionRegressionP95, or the text is a label block with alertname and backend → proceed directly to Step 1.5 and pass the alert text as --alert-text. Skip item 3 (it runs after Step 1.5 resolves Prow URLs).
    • Grafana URL: hostname is grafana-loki.ci.openshift.org and path starts with /d/ → proceed directly to Step 1.5 to parse URL parameters and find job runs via Sippy. Do NOT fetch the URL, do NOT open or access the dashboard — it is behind SSO. Skip item 3 (it runs after Step 1.5 resolves Prow URLs).
    • Prow URL: any other URL (e.g., prow.ci.openshift.org, gcsweb-ci) → continue to step 3 below
    • Validate at least one URL or one alert is provided
  3. Parse each Prow URL to extract bucket path, job name, and build ID

    • Use the same URL parsing logic as the "prow-job-analysis" skill
    • Accept both prow.ci.openshift.org and gcsweb-ci URL formats
    • Extract build_id and job_name from each URL

    Deep-link shapes for the report are in references/report-guide.md. Construct them when writing the report (Step 7), not as a separate links table.

Step 1.5: Resolve a Grafana URL or Alert to Job Runs

Skip this step if the input is Prow job URL(s).

Read references/run-discovery.md (in this skill's directory) and follow it. Pass the Grafana URL or the alert text through unchanged, including every exact backend name. Do not fetch or open the Grafana URL, and do not query Sippy by hand. Do not strip names to a base such as kube-api or substitute derived probes.

The resolved Prow URLs proceed to Step 1, item 3, and then Step 2.

Step 2: Download Artifacts for All Runs

Read references/artifacts.md and follow Timeline download. Write files under .work/disruption-analysis/{date}/.

The downloader uses anonymous access to the public test-platform-results-public bucket. Do not run gcloud auth login. On exit 2, continue with the runs that downloaded. On exit 1, stop and report each run's error.

Step 3: Analyze Interval/Timeline Data

Read references/timeline-analysis.md and follow sections 3.1 through 3.4. Pass exact backend names as --backends. A shortened name is not the selected backend.

Step 4: Deep-Dive Artifact Download (Optional)

Only if the parser output from Step 3 is insufficient for root cause determination, read the Deep-dive artifact download section of references/artifacts.md and follow it. Prefix any direct gcloud storage command with CLOUDSDK_AUTH_DISABLE_CREDENTIALS=true. Do not run gcloud auth login.

Show full SKILL.md (568 more words)Show less
Step 5: Additional Diagnostic Checks

references/timeline-analysis.md includes the node-shutdown and endpoint-slice checks. Do them when disruption coincides with node events.

Step 6: Cross-Run Comparison (Multiple Runs Only)

When multiple job runs were provided, read references/cross-run-comparison.md and follow it.

Runs where ci-cluster-network-liveness is disrupted have unreliable disruption data. Keep them in the analysis and note the caveat. Do not draw conclusions solely from an unreliable run's disruption counts.

Step 7: Generate Report

Read references/report-guide.md and follow it. Save the report at .work/disruption-analysis/{date}/{backend_names}-analysis.md. The filename rule is in that guide.

Links go inline where the evidence is discussed. Do not add an Artifacts or Links table. Do not truncate job names.

Step 8: Known Disruption Issue Lookup

Skip this step if --skip-jira was passed. Otherwise read references/known-issues.md and follow it.

Error Handling

  1. No disruption found — If interval files show no disruption events, report that the run is clean and no disruption was detected. This is a valid result, not an error.

  2. Audit logs not available — Some jobs may not have audit logs. Note this in the report and continue analysis with available data.

  3. etcd logs not available — If etcd pod logs are not present in gather-extra, note this and skip etcd analysis.

  4. Interval files not found — If no interval/timeline files are found for a job run, this is a critical error for that run. Report it and skip that run if analyzing multiple runs.

  5. gcloud errors — Public-bucket commands use CLOUDSDK_AUTH_DISABLE_CREDENTIALS=true and do not require gcloud auth login. When a download fails, report the run's error text (it includes the gcloud diagnostic, with credential-like values removed). Exit 2 means some timeline files were saved: continue with those runs. Exit 1 means nothing was downloaded. A sandbox that blocks the network is separate from an auth failure; request network execution and run the downloader again.

  6. Jira MCP unavailable — If the Jira MCP tools are not available or authentication fails, skip Step 8 and note "Jira search skipped (MCP unavailable)" in the Known Disruption Issues section. Do not block the disruption analysis on Jira availability.

  7. Grafana URL or alert missing required parameters — If var-releases / release or var-backend / backend are missing, prompt the user for the missing values rather than failing.

7a. master_nodes_updated requested but absent from the disruption API — The finder keeps no runs and prints that the field was missing. Report that warning. Do not analyze the unfiltered job list.

  1. No Sippy results for Grafana filters — If no runs match the variant filters from the Grafana URL, suggest widening the time window or relaxing filters. Report the exact query parameters that were attempted so the user can diagnose the mismatch.

  2. No disruption test failures in matching runs — If matching runs exist but none have disruption test failures for the target backend, note this (disruption may be within threshold but elevated compared to baseline). Offer to analyze the most recent runs anyway.

  3. Sippy API unavailable — If the Sippy API is unreachable during Grafana URL resolution, report the error and suggest providing Prow job URLs directly as a fallback.

Performance Considerations

  • Download artifacts for multiple runs in parallel when analyzing more than one run
  • When analyzing multiple runs, process each run independently first, then perform cross-run comparison
  • Use --max-bytes limits when fetching large log files to avoid excessive downloads
  • Filter audit logs by timestamp range rather than downloading and scanning entire files when possible

© openshift-eng, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (references) in plugins/ci/skills/analyze-disruption of openshift-eng/ai-helpers.

  • SKILL.md
  • download_timelines.py
  • find_disruption_runs.py
  • parse_disruption.py
  • references/artifacts.md
  • references/cross-run-comparison.md
  • references/known-issues.md
  • references/report-guide.md
  • references/run-discovery.md
  • references/timeline-analysis.md
  • test_download_timelines.py
  • test_find_disruption_runs.py
  • test_parse_disruption.py

Open the folder on GitHubat commit a627176

Compare with similar skills

Analyze Disruption next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Analyze Disruption compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Analyze Disruption this skillopenshift-eng/ai-helpers120—~2.8kAutomated safety check: PassApache-2.0
Infrastructuregrafana/skills281—~1.2kAutomated safety check: PassApache-2.0
Private Connectivitygrafana/skills281—~1.1kAutomated safety check: PassApache-2.0
Mz Release SignoffMaterializeInc/materialize6.4k—~7.2kAutomated safety check: PassCustom licence
Axiom Dashboard Builderopenclaw/clawhub9.5k—~4.9kAutomated safety check: PassMIT
Happy Infra Metrics and Grafanaslopus/happy24k—~2kAutomated safety check: NotesMIT

Similar skills

  • Infrastructure

    grafana/skills

    Official

    Ship Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud — k8s-monitoring Helm chart for K8s clusters (metrics + logs + traces + events + cost), Alloy…

    281 GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Private Connectivity

    grafana/skills

    Official

    Set up private network connectivity to Grafana Cloud — AWS PrivateLink, Azure Private Link, GCP Private Service Connect, and Private Data Source Connect (PDC).

    281 GitHub stars~1.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Mz Release Signoff

    MaterializeInc/materialize

    Verify a release candidate on the Grafana dashboards and sign off in release.

    6.4k GitHub stars~7.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Axiom Dashboard Builder

    openclaw/clawhub

    Designs and deploys Axiom dashboards through the API, choosing chart types and writing APL or metrics queries, with templates and migration notes for Splunk and Grafana.

    9.5k GitHub stars~4.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.

    24k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Syncmeta

    pawurb/hotpath-rs

    Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta).

    1.9k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check: notes

More from openshift-eng/ai-helpers

All 118 skills in this repo
  • Investigate CI Reliability

    openshift-eng/ai-helpers

    Find and independently validate actionable reliability defects across OpenShift release jobs and presubmits, then export portable issue handoffs.

    120 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Address Review PR

    openshift-eng/ai-helpers

    Fetch and address all PR review comments — categorize by priority, make code changes, post replies, and push.

    120 GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check passed
  • Categorize Activity Types

    openshift-eng/ai-helpers

    Categorize Jira issues into Red Hat Sankey Activity Type categories using MCP Jira tools.

    120 GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • Has Review Work

    openshift-eng/ai-helpers

    Decide whether a GitHub PR has unanswered authorized review comments or new required CI failures worth a follow-up agent.

    120 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Must Gather Analyzer

    openshift-eng/ai-helpers

    Analyze OpenShift must-gather diagnostic data including cluster operators, pods, nodes, and network components.

    120 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Payload Autodl JSON

    openshift-eng/ai-helpers

    Schema for the autodl JSON data file produced by payload-analysis for database ingestion — you must use this skill whenever generating the autodl JSON file

    120 GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Analyze Disruption

What does Analyze Disruption do?

A skill your agent uses when analyzing disruption from a DisruptionRegression alert, a Grafana disruption dashboard URL, or Prow CI job runs by examining interval data, audit logs, pod logs, and CPU…. Analyze Disruption is an agent skill from openshift-eng/ai-helpers.

When should I use Analyze Disruption?

Analyze Disruption fits situations like: analyzing disruption from a DisruptionRegression alert; A Grafana disruption dashboard URL; prow CI job runs by examining interval data.

How do I install Analyze Disruption in Claude Code?

Run `npx skills add openshift-eng/ai-helpers --skill analyze-disruption -a claude-code`. Or copy the skill folder (plugins/ci/skills/analyze-disruption in openshift-eng/ai-helpers) into .claude/skills/analyze-disruption in your project. Claude Code loads it when a task matches its description.

How do I install Analyze Disruption in Codex?

Run `npx skills add openshift-eng/ai-helpers --skill analyze-disruption -a codex`. Or copy the skill folder (plugins/ci/skills/analyze-disruption in openshift-eng/ai-helpers) into .agents/skills/analyze-disruption in your project. Codex loads it when a task matches its description.

Can I use Analyze Disruption in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openshift-eng/ai-helpers --skill analyze-disruption -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-disruption, .gemini/skills/analyze-disruption, .github/skills/analyze-disruption and .opencode/skills/analyze-disruption in your project.

What does Analyze Disruption need to run?

Going by SKILL.md and its folder, Analyze Disruption needs Python for the scripts in its folder, the command-line tools its instructions call (gcloud) and credentials named CLOUDSDK_AUTH_DISABLE_CREDENTIALS. Our summary lists: Python 3.

Does Analyze Disruption access the network?

SKILL.md names 2 domains. In commands or code: prow.ci.openshift.org and grafana-loki.ci.openshift.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Analyze Disruption safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Analyze Disruption use?

Analyze Disruption is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Analyze Disruption use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 10k tokens, read only when the agent opens those files.

What are the alternatives to Analyze Disruption?

Skills that share tags, products or a category with Analyze Disruption: Infrastructure (grafana/skills, 281 stars), Private Connectivity (grafana/skills, 281 stars), Mz Release Signoff (MaterializeInc/materialize, 6.4k stars) and Axiom Dashboard Builder (openclaw/clawhub, 9.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Analyze Disruption?

openshift-eng (a GitHub organization) maintains it in openshift-eng/ai-helpers, which has 120 GitHub stars. The repository holds 118 skills in this directory. The repository was last updated on October 6, 2026.

Source: openshift-eng/ai-helpers on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.