Infrastructure
grafana/skills
Ship Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud — k8s-monitoring Helm chart for K8s clusters (metrics + logs + traces + events + cost), Alloy…
A skill your agent uses when analyzing disruption from a DisruptionRegression alert, a Grafana disruption dashboard URL, or Prow CI job runs by examining interval data, audit logs, pod logs, and CPU…
$ npx skills add openshift-eng/ai-helpers --skill analyze-disruption -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install openshift-eng/ai-helpers analyze-disruption --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ci/skills/analyze-disruption .claude/skills/analyze-disruption && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "analyze-disruption" agent skill from https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/analyze-disruption into .claude/skills/analyze-disruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-disruption", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/analyze-disruptionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add openshift-eng/ai-helpers --skill analyze-disruption -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install openshift-eng/ai-helpers analyze-disruption --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/ci/skills/analyze-disruption .agents/skills/analyze-disruption && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "analyze-disruption" agent skill from https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/analyze-disruption into .agents/skills/analyze-disruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-disruption", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add openshift-eng/ai-helpers --skill analyze-disruption -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install openshift-eng/ai-helpers analyze-disruption --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/ci/skills/analyze-disruption .cursor/skills/analyze-disruption && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "analyze-disruption" agent skill from https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/analyze-disruption into .cursor/skills/analyze-disruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-disruption", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/openshift-eng/ai-helpers.git --path plugins/ci/skills/analyze-disruption--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add openshift-eng/ai-helpers --skill analyze-disruption -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install openshift-eng/ai-helpers analyze-disruption --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/ci/skills/analyze-disruption .gemini/skills/analyze-disruption && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "analyze-disruption" agent skill from https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/analyze-disruption into .gemini/skills/analyze-disruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-disruption", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install openshift-eng/ai-helpers analyze-disruptionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add openshift-eng/ai-helpers --skill analyze-disruption -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/ci/skills/analyze-disruption .github/skills/analyze-disruption && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "analyze-disruption" agent skill from https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/analyze-disruption into .github/skills/analyze-disruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-disruption", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add openshift-eng/ai-helpers --skill analyze-disruption -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install openshift-eng/ai-helpers analyze-disruption --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/ci/skills/analyze-disruption .opencode/skills/analyze-disruption && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "analyze-disruption" agent skill from https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/analyze-disruption into .opencode/skills/analyze-disruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-disruption", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
analyze-disruptionA skill your agent uses when analyzing disruption from a DisruptionRegression alert, a Grafana disruption dashboard URL, or Prow CI job runs by examining interval data, audit logs, pod logs, and CPU…
Analyze Disruption is an agent skill from openshift-eng/ai-helpers. Use when analyzing disruption from a DisruptionRegression alert, a Grafana disruption dashboard URL, or Prow CI job runs by examining interval data, audit logs, pod logs, and CPU metrics
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including reference files (for example `download_timelines.py`, `find_disruption_runs.py` and `parse_disruption.py`).
It sits in DevOps & Cloud, covering Monitoring and alerting. It works with Grafana and Google Cloud. The repository describes itself as: Developer productivity tools for Claude Code & other AI assistants. The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a627176. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
gcloudFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
prow.ci.openshift.orggrafana-loki.ci.openshift.orgFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
CLOUDSDK_AUTH_DISABLE_CREDENTIALSFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Analyze Disruption loads about 2.8k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 51 tokens; SKILL.md has 1,466 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from openshift-eng/ai-helpers at commit a627176, republished under its Apache-2.0 licence (© openshift-eng). 1,466 words, ~2,805 tokens.
.claude/skills/analyze-disruption/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.This skill analyzes disruption events recorded in Prow CI job runs. It downloads interval/timeline data, audit logs, and pod logs, then correlates disruption across backends and job runs to identify root causes.
gcloud CLI Installation
which gcloudtest-platform-results-public bucket is public. Do not run gcloud auth login for these artifacts, and do not change the user's gcloud config.download_timelines.py sets CLOUDSDK_AUTH_DISABLE_CREDENTIALS=true on each gcloud subprocess (listing and copying, including parallel workers). That avoids refreshing expired local credentials.gcloud storage command in this skill uses the same variable. Anonymous access still needs network execution when the command runs in a sandbox.Python 3 (3.7 or later)
The user will provide one of the following as input:
Option A — Prow job URLs (direct analysis):
https://prow.ci.openshift.org/view/gs/test-platform-results-public/logs/periodic-ci-openshift-release-master-ci-4.21-e2e-aws-ovn/1983307151598161920Option B — Grafana disruption dashboard URL (run discovery + analysis):
https://grafana-loki.ci.openshift.org/d/gEdw_aLvk/disruption-for-5-0-os-agnostic?var-platform=gcp&var-backend=host-to-host-new-connections&var-upgrade_type=micro&var-architectures=amd64&var-topologies=ha&var-networks=ovn&var-releases=5.0grafana-loki.ci.openshift.org and path starting with /d/var-backend value is an exact backend name. Pass that full list as --backends unless the user overrides it. Do not reduce it to one API backend or to a stripped base nameOption C — disruption regression alert text (run discovery + analysis):
DisruptionRegression alert — DisruptionRegressionP50, DisruptionRegressionP50Azure, DisruptionRegressionP75, or DisruptionRegressionP95, or a label block that contains alertname and backendbackend, platform, upgrade_type, master_nodes_updated, architecture, topology, network, release, delta, feature_set, os, compare_releaselink: annotation, when present, is the dashboard URL. Pass the alert text through unchanged. Do not rebuild the URL, and do not drop master_nodes_updated, feature_set, or osOptional flags (all input options):
--backends flag (optional) — comma-separated exact backend names to focus on
--backends kube-api-new-connections,oauth-api-new-connections,openshift-api-new-connectionsvar-backend / alert backend value (Option B or C)--skip-jira flag (optional) — skip the Jira search for known disruption cards
Load these only at the step that needs them — not up front:
references/run-discovery.md — Grafana URL and alert resolution into Prow job runs (Step 1.5)references/artifacts.md — timeline download and optional audit/etcd/PromQL deep dive (Steps 2 and 4)references/timeline-analysis.md — parser modes, signal interpretation, and extra diagnostic checks (Steps 3 and 5)references/cross-run-comparison.md — multi-run pattern detection and same-job clean comparison (Step 6)references/report-guide.md — deep-link shapes, inline linking rules, and report structure (Step 7)references/known-issues.md — Jira search and bug filing (Step 8)Extract URLs and flags
--backends flag if present, split on comma to get backend filter list--skip-jira flag as a boolean option (default: false)Detect input type:
alertname is DisruptionRegressionP50, DisruptionRegressionP50Azure, DisruptionRegressionP75, or DisruptionRegressionP95, or the text is a label block with alertname and backend
→ proceed directly to Step 1.5 and pass the alert text as --alert-text.
Skip item 3 (it runs after Step 1.5 resolves Prow URLs).grafana-loki.ci.openshift.org and path starts with /d/
→ proceed directly to Step 1.5 to parse URL parameters and find job runs via Sippy.
Do NOT fetch the URL, do NOT open or access the dashboard — it is behind SSO.
Skip item 3 (it runs after Step 1.5 resolves Prow URLs).prow.ci.openshift.org, gcsweb-ci)
→ continue to step 3 belowParse each Prow URL to extract bucket path, job name, and build ID
prow.ci.openshift.org and gcsweb-ci URL formatsbuild_id and job_name from each URLDeep-link shapes for the report are in references/report-guide.md. Construct them when writing the report (Step 7), not as a separate links table.
Skip this step if the input is Prow job URL(s).
Read references/run-discovery.md (in this skill's directory) and follow it. Pass the Grafana URL or the alert text through unchanged, including every exact backend name. Do not fetch or open the Grafana URL, and do not query Sippy by hand. Do not strip names to a base such as kube-api or substitute derived probes.
The resolved Prow URLs proceed to Step 1, item 3, and then Step 2.
Read references/artifacts.md and follow Timeline download. Write files under .work/disruption-analysis/{date}/.
The downloader uses anonymous access to the public test-platform-results-public bucket. Do not run gcloud auth login. On exit 2, continue with the runs that downloaded. On exit 1, stop and report each run's error.
Read references/timeline-analysis.md and follow sections 3.1 through 3.4. Pass exact backend names as --backends. A shortened name is not the selected backend.
Only if the parser output from Step 3 is insufficient for root cause determination, read the Deep-dive artifact download section of references/artifacts.md and follow it. Prefix any direct gcloud storage command with CLOUDSDK_AUTH_DISABLE_CREDENTIALS=true. Do not run gcloud auth login.
references/timeline-analysis.md includes the node-shutdown and endpoint-slice checks. Do them when disruption coincides with node events.
When multiple job runs were provided, read references/cross-run-comparison.md and follow it.
Runs where ci-cluster-network-liveness is disrupted have unreliable disruption data. Keep them in the analysis and note the caveat. Do not draw conclusions solely from an unreliable run's disruption counts.
Read references/report-guide.md and follow it. Save the report at .work/disruption-analysis/{date}/{backend_names}-analysis.md. The filename rule is in that guide.
Links go inline where the evidence is discussed. Do not add an Artifacts or Links table. Do not truncate job names.
Skip this step if --skip-jira was passed. Otherwise read references/known-issues.md and follow it.
No disruption found — If interval files show no disruption events, report that the run is clean and no disruption was detected. This is a valid result, not an error.
Audit logs not available — Some jobs may not have audit logs. Note this in the report and continue analysis with available data.
etcd logs not available — If etcd pod logs are not present in gather-extra, note this and skip etcd analysis.
Interval files not found — If no interval/timeline files are found for a job run, this is a critical error for that run. Report it and skip that run if analyzing multiple runs.
gcloud errors — Public-bucket commands use CLOUDSDK_AUTH_DISABLE_CREDENTIALS=true and do not require gcloud auth login. When a download fails, report the run's error text (it includes the gcloud diagnostic, with credential-like values removed). Exit 2 means some timeline files were saved: continue with those runs. Exit 1 means nothing was downloaded. A sandbox that blocks the network is separate from an auth failure; request network execution and run the downloader again.
Jira MCP unavailable — If the Jira MCP tools are not available or authentication fails, skip Step 8 and note "Jira search skipped (MCP unavailable)" in the Known Disruption Issues section. Do not block the disruption analysis on Jira availability.
Grafana URL or alert missing required parameters — If var-releases / release or var-backend / backend are missing, prompt the user for the missing values rather than failing.
7a. master_nodes_updated requested but absent from the disruption API — The finder keeps no runs and prints that the field was missing. Report that warning. Do not analyze the unfiltered job list.
No Sippy results for Grafana filters — If no runs match the variant filters from the Grafana URL, suggest widening the time window or relaxing filters. Report the exact query parameters that were attempted so the user can diagnose the mismatch.
No disruption test failures in matching runs — If matching runs exist but none have disruption test failures for the target backend, note this (disruption may be within threshold but elevated compared to baseline). Offer to analyze the most recent runs anyway.
Sippy API unavailable — If the Sippy API is unreachable during Grafana URL resolution, report the error and suggest providing Prow job URLs directly as a fallback.
--max-bytes limits when fetching large log files to avoid excessive downloads© openshift-eng, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 12 other files (references) in plugins/ci/skills/analyze-disruption of openshift-eng/ai-helpers.
Open the folder on GitHubat commit a627176
Analyze Disruption next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Analyze Disruption this skillopenshift-eng/ai-helpers | 120 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Infrastructuregrafana/skills | 281 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Private Connectivitygrafana/skills | 281 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Mz Release SignoffMaterializeInc/materialize | 6.4k | — | ~7.2k | Automated safety check: Pass | Custom licence | |
| Axiom Dashboard Builderopenclaw/clawhub | 9.5k | — | ~4.9k | Automated safety check: Pass | MIT | |
| Happy Infra Metrics and Grafanaslopus/happy | 24k | — | ~2k | Automated safety check: Notes | MIT |
grafana/skills
Ship Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud — k8s-monitoring Helm chart for K8s clusters (metrics + logs + traces + events + cost), Alloy…
grafana/skills
Set up private network connectivity to Grafana Cloud — AWS PrivateLink, Azure Private Link, GCP Private Service Connect, and Private Data Source Connect (PDC).
MaterializeInc/materialize
Verify a release candidate on the Grafana dashboards and sign off in release.
openclaw/clawhub
Designs and deploys Axiom dashboards through the API, choosing chart types and writing APL or metrics queries, with templates and migration notes for Splunk and Grafana.
slopus/happy
Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.
pawurb/hotpath-rs
Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta).
openshift-eng/ai-helpers
Find and independently validate actionable reliability defects across OpenShift release jobs and presubmits, then export portable issue handoffs.
openshift-eng/ai-helpers
Fetch and address all PR review comments — categorize by priority, make code changes, post replies, and push.
openshift-eng/ai-helpers
Categorize Jira issues into Red Hat Sankey Activity Type categories using MCP Jira tools.
openshift-eng/ai-helpers
Decide whether a GitHub PR has unanswered authorized review comments or new required CI failures worth a follow-up agent.
openshift-eng/ai-helpers
Analyze OpenShift must-gather diagnostic data including cluster operators, pods, nodes, and network components.
openshift-eng/ai-helpers
Schema for the autodl JSON data file produced by payload-analysis for database ingestion — you must use this skill whenever generating the autodl JSON file
Works with
Categories
A skill your agent uses when analyzing disruption from a DisruptionRegression alert, a Grafana disruption dashboard URL, or Prow CI job runs by examining interval data, audit logs, pod logs, and CPU…. Analyze Disruption is an agent skill from openshift-eng/ai-helpers.
Analyze Disruption fits situations like: analyzing disruption from a DisruptionRegression alert; A Grafana disruption dashboard URL; prow CI job runs by examining interval data.
Run `npx skills add openshift-eng/ai-helpers --skill analyze-disruption -a claude-code`. Or copy the skill folder (plugins/ci/skills/analyze-disruption in openshift-eng/ai-helpers) into .claude/skills/analyze-disruption in your project. Claude Code loads it when a task matches its description.
Run `npx skills add openshift-eng/ai-helpers --skill analyze-disruption -a codex`. Or copy the skill folder (plugins/ci/skills/analyze-disruption in openshift-eng/ai-helpers) into .agents/skills/analyze-disruption in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openshift-eng/ai-helpers --skill analyze-disruption -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-disruption, .gemini/skills/analyze-disruption, .github/skills/analyze-disruption and .opencode/skills/analyze-disruption in your project.
Going by SKILL.md and its folder, Analyze Disruption needs Python for the scripts in its folder, the command-line tools its instructions call (gcloud) and credentials named CLOUDSDK_AUTH_DISABLE_CREDENTIALS. Our summary lists: Python 3.
SKILL.md names 2 domains. In commands or code: prow.ci.openshift.org and grafana-loki.ci.openshift.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Analyze Disruption is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 10k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Analyze Disruption: Infrastructure (grafana/skills, 281 stars), Private Connectivity (grafana/skills, 281 stars), Mz Release Signoff (MaterializeInc/materialize, 6.4k stars) and Axiom Dashboard Builder (openclaw/clawhub, 9.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
openshift-eng (a GitHub organization) maintains it in openshift-eng/ai-helpers, which has 120 GitHub stars. The repository holds 118 skills in this directory. The repository was last updated on October 6, 2026.
Source: openshift-eng/ai-helpers on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.