OpenLogi macOS Permissions Triage
AprilNEA/OpenLogi
Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.
A skill your agent uses for Component Readiness triage duty - holistically analyze and triage all untriaged regressions for a set of components in a view, clustering them into root-cause buckets
$ npx skills add openshift-eng/ai-helpers --skill bulk-triage-regressions -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install openshift-eng/ai-helpers bulk-triage-regressions --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ci/skills/bulk-triage-regressions .claude/skills/bulk-triage-regressions && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bulk-triage-regressions" agent skill from https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/bulk-triage-regressions into .claude/skills/bulk-triage-regressions/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bulk-triage-regressions", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/bulk-triage-regressionsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add openshift-eng/ai-helpers --skill bulk-triage-regressions -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install openshift-eng/ai-helpers bulk-triage-regressions --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/ci/skills/bulk-triage-regressions .agents/skills/bulk-triage-regressions && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bulk-triage-regressions" agent skill from https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/bulk-triage-regressions into .agents/skills/bulk-triage-regressions/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bulk-triage-regressions", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add openshift-eng/ai-helpers --skill bulk-triage-regressions -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install openshift-eng/ai-helpers bulk-triage-regressions --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/ci/skills/bulk-triage-regressions .cursor/skills/bulk-triage-regressions && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bulk-triage-regressions" agent skill from https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/bulk-triage-regressions into .cursor/skills/bulk-triage-regressions/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bulk-triage-regressions", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/openshift-eng/ai-helpers.git --path plugins/ci/skills/bulk-triage-regressions--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add openshift-eng/ai-helpers --skill bulk-triage-regressions -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install openshift-eng/ai-helpers bulk-triage-regressions --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/ci/skills/bulk-triage-regressions .gemini/skills/bulk-triage-regressions && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bulk-triage-regressions" agent skill from https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/bulk-triage-regressions into .gemini/skills/bulk-triage-regressions/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bulk-triage-regressions", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install openshift-eng/ai-helpers bulk-triage-regressionsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add openshift-eng/ai-helpers --skill bulk-triage-regressions -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/ci/skills/bulk-triage-regressions .github/skills/bulk-triage-regressions && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bulk-triage-regressions" agent skill from https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/bulk-triage-regressions into .github/skills/bulk-triage-regressions/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bulk-triage-regressions", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add openshift-eng/ai-helpers --skill bulk-triage-regressions -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install openshift-eng/ai-helpers bulk-triage-regressions --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/ci/skills/bulk-triage-regressions .opencode/skills/bulk-triage-regressions && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bulk-triage-regressions" agent skill from https://github.com/openshift-eng/ai-helpers/tree/main/plugins/ci/skills/bulk-triage-regressions into .opencode/skills/bulk-triage-regressions/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bulk-triage-regressions", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bulk-triage-regressionsA skill your agent uses for Component Readiness triage duty - holistically analyze and triage all untriaged regressions for a set of components in a view, clustering them into root-cause buckets
Bulk Triage Regressions is an agent skill from openshift-eng/ai-helpers. Use this skill for Component Readiness triage duty - holistically analyze and triage all untriaged regressions for a set of components in a view, clustering them into root-cause buckets
Its SKILL.md is about 17k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/case-notes.md`).
It sits in Development, covering Root cause analysis. The repository describes itself as: Developer productivity tools for Claude Code & other AI assistants. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a627176. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ghpython3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
sippy-auth.dptools.openshift.orgsippy.dptools.openshift.orgredhat.atlassian.netAlso links to:
id.atlassian.comredhat.enterprise.slack.comdocs.ci.openshift.orgFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
JIRA_API_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bulk Triage Regressions loads about 17k tokens when it runs, and up to ~20k if it reads all its reference files. Until then it costs about 52 tokens; SKILL.md has 9,525 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from openshift-eng/ai-helpers at commit a627176, republished under its Apache-2.0 licence (© openshift-eng). 9,525 words, ~17,164 tokens.
.claude/skills/bulk-triage-regressions/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.bulk-triage-regressions <view> [--components comp1 comp2 ...] [--auto-triage] [--existing-only]Example: bulk-triage-regressions 5.0-main --components Installer Unknown
This skill implements the Component Readiness triage duty workflow: it fetches all untriaged regressions for a set of components in a view (e.g., 5.0-main, components Installer and Unknown), analyzes them as a batch, clusters them into root-cause buckets, and then triages each bucket to a single JIRA bug (existing or new).
This differs from /ci:analyze-regression (which analyzes a single regression in depth). Triage duty requires a holistic view, because:
Many regressions, few root causes. One product bug commonly opens 5–30 regressions across variants (different platforms, arches, featuresets, upgrade modes) and across "wrapper" tests (install should succeed: overall, : cluster bootstrap, : cluster creation, verify the cluster readiness and stability, mass-failure tests, etc.). Analyzing regressions one-by-one wastes effort and risks filing duplicate bugs. Cluster first, deep-dive once per cluster.
Component attribution is often wrong. Regressions in Installer and Unknown are catch-all attributions. A failed installation or bootstrap is frequently caused by a specific component — e.g., a monitoring operator failing to go available blocks cluster creation, an etcd slowness issue breaks bootstrap, an MCO bug degrades nodes during install. The Sippy component label tells you which test failed, not whose bug it is. The real owner must be determined from artifacts (cluster operator status, log bundle, operator logs), and the JIRA bug must be filed against the actual owning component, not Installer.
Use this skill when doing triage duty for a view, or whenever a user asks to "look at all untriaged regressions from <components>" rather than a single regression ID.
Script invocation rules: Run Python skill scripts directly and analyze their JSON output with your own reasoning (pass --format json where the script offers the flag; scripts without it, such as list_regressions.py, emit JSON by default). Do not pipe script output through inline Python one-liners. Do not suppress stderr: if a script exits non-zero or returns invalid/empty JSON, stop and surface the error — an authentication or API failure must never be mistaken for an empty inventory ("nothing to triage").
Authentication: Read steps (listing, fetching details, test runs, GCS artifacts) require no auth — Phases 1–3 and a read-only report must work without any credentials. Two Sippy hosts exist and the split matters: https://sippy.dptools.openshift.org is public and read-only (regressions, test details, and the full symptom/label catalog — see the list-symptoms skill), while everything on https://sippy-auth.dptools.openshift.org requires a DPCR Bearer token, including the reevaluate probe, which has no public equivalent (the route returns 404 on the public host). The only Phase 3 step that needs a token is therefore the active symptom probe. When it is unavailable (no token, or a 401/403), treat the result as an unknown — never as "no symptoms matched", the same error as reading a failed listing script as an empty inventory, and work down this fallback: (1) passive label_summary / job_labels from fetch-regression-details; (2) list the catalog from the public API with the list-symptoms skill and grep each plausible symptom's match_string against its file_pattern in the GCS artifacts you are already fetching — this reproduces the probe locally for the handful of symptoms related to the bucket's stage and platform; (3) only then the full evidence ladder, recording that the active probe was unavailable so the next shift knows the catalog was checked by hand. Write credentials are validated only when writes are going to happen: with --auto-triage, validate both up front (so an expired token surfaces before hours of analysis); otherwise validate at the start of Phase 4, before the first write. Sippy writes (creating/updating triage records) require a Bearer token from the DPCR cluster (api.cr.j7t7.p1.openshiftapps.com:6443) — see the oc-auth skill and the token-extraction snippet in /ci:analyze-regression; check with an authenticated GET against https://sippy-auth.dptools.openshift.org/api/component_readiness/triages (200 vs 401/403).
JIRA writes (filing bugs, set-release-blocker, add-jira-triage-link) additionally require the JIRA_USERNAME and JIRA_API_TOKEN environment variables (API token from https://id.atlassian.com/manage-profile/security/api-tokens); verify with an authenticated GET against https://redhat.atlassian.net/rest/api/3/myself (Basic auth, 200 vs 401/403). If a required write credential is missing or invalid, pause before Phase 4 and ask the user to fix it — the analysis and report so far remain valid and must still be presented.
Load CI context: Read the files in plugins/ci/references/ (jobs.md, tests.md, sippy-apis.md) for conventions on tests, jobs, and Sippy APIs.
Parse arguments:
view: required, e.g. 5.0-main--components: component filter list, e.g. Installer Unknown. Matching is case-insensitive and hierarchy-aware: a filter matches the full component name or any /-separated segment of it, so Installer also covers Installer / openshift-installer, and Networking covers Networking / ovn-kubernetes, Networking / router, and every other Networking / * component. If omitted, ask the user which components the duty covers.--auto-triage: if present, triage buckets without per-bucket confirmation when confidence is high (see Phase 4). Default is to present findings and confirm before writing.--existing-only: never file new JIRA issues; write only for matches to existing issues (see "Existing-only mode" in Phase 4).List regressions with the list-regressions skill:
python3 plugins/teams/skills/list-regressions/list_regressions.py \
--view <view> --components <components...>Keep only open, untriaged regressions (empty triages array), but note recently-triaged ones — they are prime candidates for absorbing untriaged siblings.
Closed regressions are out of scope — even when untriaged. A regression whose closed field is set has already resolved itself; do not inventory it, cluster it, deep-dive it, or recommend retroactive triage for it. The duty batch consists solely of open untriaged regressions. Closed regressions may be consulted as evidence (e.g., a closed sibling that shares a root cause with an open bucket, or a closed sibling whose existing triage/JIRA an open bucket should reuse — see Pitfalls), but they must never appear as bucket members, action items, or "leftovers" in the report. The only exception is the explicit closed-set audit mode (--audit-closed, see below), which the user must request by name — it is never part of a normal duty run.
Build a batch inventory table: For every open untriaged regression record: regression ID, test name, component/capability, variants (Platform/Arch/Network/Topology/FeatureSet/Upgrade), opened date, failure/run counts. Present this table to the user up front so the scope of the duty run is visible. Do not include closed regressions in the inventory.
Stale-triage sweep (mandatory) — a 100% triaged board can still hide live defects. For every open, already-triaged regression, compare its last_failure against the state of its triage's JIRA: a regression whose bug is Closed/Verified/resolved but which has failed after the resolution date is an alarm, not a statistic. Its fresh failure window is either a failed fix or — more often — a different cause hiding behind the old triage record. For each such regression, verify the recent runs' signature (armed Sippy symptoms are a cheap first oracle: dry-run reevaluate the newest runs — label hits map windows to known causes instantly; unlabeled recent runs mean a new, uninvestigated cause; a hit label whose bugs differ from the triage's JIRA suggests a new cause, one listing the same closed bug a failed fix; a label linking several bugs is inconclusive — compare the runs' signature and timing with each) and add the correct triage(s) for the new window rather than trusting the stale one. (Case 4 in case-notes.md: three "triaged" regressions pointing at a Closed bug while failing daily from two new causes.)
Long-lived wrapper regressions are cause timelines, not single buckets. A wrapper regression that stays open for weeks accumulates causes, each window separately verified and separately triaged. When a previously-analyzed regression shows new last_failure dates, re-verify the new window from scratch — never assume the existing triage covers it. (Case 13: one record carried four causes.)
Before any deep log analysis, group regressions using signals already in hand:
install should succeed: overall + : cluster creation + verify the cluster readiness and stability all failing on azure/amd64/techpreview starting the same day is one bucket. Wrapper tests fail together.fetch-regression-details skill) and compare job_runs prowjob_run_ids. Regressions observed in the same failed runs are strong candidates for one bucket — but this is a clustering signal, not proof: the same run can contain independent defects or be a mass failure, so Phase 3 validation is still required. Also run the fetch-related-triages skill per regression; same_last_failure and similarly_named_test matches feed the clustering, and triaged_matches with confidence ≥5 immediately suggest an existing triage/bug for the whole bucket.fetch-regression-details returns label_summary (per job) and job_labels (per run). A label shared by most failed runs of several regressions is a strong bucket signal, and labels are precise (human-written matchers over artifacts). label_bugs maps each label to its linked Jira keys: labels sharing a bug are a bucket signal, and the bug is a candidate triage target. An empty label list means nothing was detected — not that nothing is wrong, and often just that the runs were never swept: Phase 3 opens with an active dry-run reevaluation that closes this gap.test_failures counts in job_runs mean the regression is likely collateral of a bigger event, not an independent issue.Output of this phase: a draft bucket list, each bucket with member regression IDs, the shared fingerprint (test/variant/date/job-run overlap), and any candidate existing triage or JIRA bug.
Treat buckets as hypotheses — Phase 3 must confirm or split them. Do not merge buckets merely because both are "install failures"; installs fail at different stages for unrelated reasons.
For each bucket, pick 2–5 representative failed job runs (spread across jobs/variants; include the newest) and analyze. 2–5 runs are sufficient only when they yield a consistent result (same error signature / failure stage across all of them). If the sample is mixed or unclear — different errors, different stages, or an inconclusive owner — extend the sample to 10–20 runs before concluding; a small ambiguous sample must never be the basis for splitting/merging a bucket or attributing an owner.
Cross-cutting analysis principles — most specific rules below (and most Pitfalls) are instances of these three; when a situation matches none of the specific rules, fall back to the principle:
job_runs by date (never a 3–4 run sample; Case 1), only runs from the regression's own job_runs list (never a broad job-filter query; Case 10), and source-code gating for any "X-only" scope claim.reevaluate), existing triage records, the owning component's recent JIRA bugs regardless of keywords, the owning repo's merge history at onset/cessation boundaries, and closed/dropped sibling regressions. (Case 3: one dry-run call resolved a board that cost a parallel run 264 shell commands.)Symptom-catalog oracle first (cheap, decisive — run before any artifact dive or subagent dispatch). The passive label_summary from Phase 2 only shows labels already applied by past sweeps; recent runs are typically unswept and show nothing. Actively probe the catalog instead: POST https://sippy-auth.dptools.openshift.org/api/jobs/runs/reevaluate with {"prow_job_build_ids": [...], "dry_run": true} (DPCR Bearer token) for 3–5 representative runs per bucket, spread across platforms. Reevaluation scans the runs' GCS artifacts server-side against every armed symptom matcher and returns symptoms_matched per run — one call can attribute an entire bucket to a known incident in seconds (Case 3 in case-notes.md). A matched label's bugs (see list-symptoms --labels) name the candidate triage target. Matched symptoms are a hypothesis with a named cause — still verify per the steps below (the label names the incident; confirm the causal chain to the bucket's specific test) — but they dictate where to look first and usually collapse the deep-dive to a single confirmatory artifact read. No matches means the cause is not yet cataloged — proceed with the full evidence ladder.
No token? The catalog is still public — degrade, never skip. A 401/403 from the probe is an unknown, never "no symptoms matched": use the credential-free fallback in the Authentication section above, and say in the report which rung was used.
Deep-dive at least one CI sample per bucket — always, even at high confidence. A Sippy triaged_matches confidence of 10 or a same-day triaged sibling is a hypothesis, not a verdict: Sippy matches on test names and shared job runs, which produces conf=10 for the same test name across unrelated platforms and root causes (see Pitfalls). Before accepting any disposition — including "extend existing triage" — read the actual failure evidence for at least one representative run of the bucket (failure output at minimum; installer logs / artifacts for install wrappers) and confirm it matches the target triage's root cause. You have deeper analysis capabilities than Sippy's heuristics — use your own judgement on the raw evidence, and if it contradicts the Sippy categorization, trust the evidence and re-bucket.
Failure outputs: fetch-test-runs skill with the bucket's test IDs and job run IDs — check whether error messages are consistent within the bucket. >90% same error is strong evidence of a single cause, not confirmation: wrapper tests and mass failures print identical error text for unrelated defects (e.g., "nodes not ready" covers disk, memory, and network deaths alike). Confirm with items 2–3 below before treating the bucket as one cause; inconsistent errors ⇒ split the bucket.
Job run context: fetch-job-run-summary skill per representative run — is the regressed test isolated, part of a consistent co-failure set, or one of hundreds of random failures? For Unknown-component and mass-failure regressions this is where the real component reveals itself: read the names of the co-failing tests.
Install/bootstrap failures — mandatory artifact dig: For any bucket whose tests include install should succeed (any stage) or bootstrap/cluster-readiness wrappers, invoke the prow-job-analysis skill per representative run. Do not stop at Sippy's generic "install failed" wrapper. From the GCS artifacts determine:
clusteroperators.json (or the installer log's "Cluster operator X is not available" lines) and the failing operator's pod logs from the log bundle. For bootstrap failures, read the bootstrap log bundle (etcd, bootkube, release-image pulls).clusteroperators.json, not pod existence: check etcd's StaticPodsAvailable/revision status and openshift-apiserver's APIServicesAvailable directly — a bootstrap-timed-out run whose gather still shows etcd Available=False: 0 nodes are active; 3 nodes are at revision 0 was blocked by etcd no matter what else looks broken. Do not infer "etcd/apiserver were fine" from workload pods running: during bootstrap the bootstrap-node etcd/apiserver serve the control plane, so pods run happily while cluster etcd never deploys. A crash-looping container's *_previous.log is primary evidence for that container's failure — but before assigning ownership, check whether its error names a missing upstream dependency: such a component is a victim, not an owner.LastTransitionTime/DurationSinceTransition: hundreds of transitions with DurationSinceTransition=1s across the stability window means the operator is flapping continuously (a sync-loop product bug that will never pass the stability check — permafail), whereas a single long Progressing=True window that eventually clears is a slow rollout (flaky race against the timeout). These two have opposite classifications and dispositions. (Case 1: a ~1 Hz flap, 1600+ transitions in 30 minutes, mislabeled "self-resolved slow rollout" from a 3-run sample.)creationTimestamps, Established/Available conditions, phase logs, first/last occurrence of the error). If X was ready before the errors began, that refutes this ordering theory — but not every timing theory: readiness at the source does not mean visibility at the consumer, so the next candidates are cache/informer warmup, an in-progress rollout, or a stale client. Name which one the artifacts support rather than dropping the timing angle entirely. Also measure the failure window's duration: it is part of the mechanism claim and belongs in the bug.2^32 − size, a type maximum, a page/alignment boundary, an epoch. Do the arithmetic; never pattern-match "big hex-ish number ⇒ overflow bug". (Case 12: 4294901760 + 65536 = 2^32 — an allocator boundary, handing you the whole mechanism.)etcdserver: request timed out on Azure ⇒ etcd bug; nodes degraded during install ⇒ MCO bug; quota/DNS/cloud-API errors ⇒ ci-infra, not a product bug at all.Recover the underlying condition message when the junit output is vague. Outputs like clusterversion not available: False (empty reason) carry no cause. Run this literal check against the e2e step's build-log.txt (and, if present, the monitor-intervals JSON under the step's artifacts/junit/):
curl -s <gcs-url-of-e2e-step>/build-log.txt | grep -oE '(Failing|Available|Degraded)=(True|False)[:,][^"\\]{0,160}' | sort | uniq -c | sort -rn | headThe condition messages recovered this way (e.g., a controller error string) name the real culprit. This step is complete only when the report quotes the recovered message for every vague-output run. Never attribute such a run from its co-failing tests: co-failures on techpreview jobs are usually unrelated background noise, and correlation with them has produced wrong owners in past duty runs.
Ask why the operator went unavailable, not just which operator. If everything is healthy at gather time, the wrapper caught a transient flap: read the operator's own pod log around the transition timestamps and find the trigger. Distinguish (a) the operator genuinely failing (⇒ product bug for that operator) from (b) a routine reconciliation/rollout triggered by a cluster mutation (config/secret change, node roll) — a rollout of a single-replica deployment flaps Available by design. A rollout-flap disposition must name the mutator, not stop at "the operator detected a configuration change": quote the operator-log line identifying which object changed (grep the operator pod log for object changed / secret and config names and read the surrounding lines), then identify who wrote it — audit logs (verb update/patch on that object: user + userAgent) or the job's test-harness step logs (search them for the object name or the value being set). If a test step caused the mutation mid-run, the bucket is test owned by that suite, not a product bug against the flapping operator; filing "spurious rollouts" against the operator without naming the mutator is an incomplete deep-dive. (Case 2 in case-notes.md: an image-registry flap drafted as a product bug until the operator log showed the harness had replaced the pull secret mid-run.)
Test-interference buckets — read the offending test's source before describing its scope. When the root cause is another test's or tool's behavior (a test creates pods/namespaces that break an invariant check, leaks resources, reboots nodes, etc.), do not infer why the failure is confined to certain variants from the regression's variant labels — correlation with a variant (techpreview, platform, upgrade mode) is not evidence of how the offending code is gated. Instead:
gh search code in openshift/origin or the relevant repo for the test name, namespace prefix, or error string) and read its skip/gate conditions (skipIf..., e2eskipper.Skipf, feature-gate checks, platform checks, suite membership).skipIfNotBaremetal"), and if that gating is broader than the regressed variants, say so — the regression's variant slice then reflects job-scheduling or sample-size effects, not the blast radius, and other variants are also at risk.gh api repos/<org>/<repo>/commits?path=...) — a merge date matching the regression onset both confirms the attribution and identifies the owning team.Onset and suspect PRs (when the bucket has a crisp start date): follow the "Determine Regression Start Date" and "Identify Suspect PRs in Payload" procedures from /ci:analyze-regression (first failing run → payload tag via fetch-prowjob-json → fetch-new-prs-in-payload → up to 5 candidate PRs vetted with gh). A LIKELY PR both strengthens the bucket and tells you the owning component/repo.
Cross-check globally: fetch-test-report skill (with --no-collapse) for the bucket's main test — confirms whether the issue is variant-specific or global, and surfaces open_bugs that may already cover the bucket.
Check Slack context (optional — only when Slack access is available): The TRT/release-oversight team discusses ongoing payload and CI issues in #forum-ocp-release-oversight (https://redhat.enterprise.slack.com/archives/C01CQA76KMX). Search/read the last 14 days of messages there for the bucket's signature (test name, error message, operator, platform, payload tag) — known payload-wide events, infra outages, and in-flight fixes are usually discussed there before triages/bugs exist, and a thread often names the owning team or an existing OCPBUGS ticket. If the agent has no Slack access (no Slack tooling/credentials), omit this step entirely — do not block or ask for access.
Mechanism self-review (mandatory before finalizing any bucket). Every rule above is an instance of one behavior: challenge your own conclusion before filing it. After drafting a bucket's root cause, run this adversarial pass on the mechanism claim itself:
LastTransitionTime, or the query that returned the result. A line number is one acceptable form, not a requirement — much of the evidence here (run metadata, API fields, object IDs, timestamps) has no stable line to cite, and inventing one is worse than quoting the field. What matters is that the next reader can go to the same place and see the same thing. An assumption with no such locator is downgraded in the filed text to an explicit unverified hypothesis ("unverified: possibly …") — never stated as fact.A bucket whose mechanism claim fails any of these checks is not finalizable — keep digging or downgrade honestly.
After deep-dive, finalize buckets. Depth is mandatory, not optional: no bucket may be finalized at LOW or MEDIUM confidence, and the duty run must not end with open "action items" like "needs artifact deep-dive" or "spot-check installer logs first". When one regression's failed runs split into multiple distinct failure signatures, every signature must be root-caused independently — each may get its own triage record (a regression can carry several), and no run may be left labeled "ambiguous"/"unclear" in a bucket claimed at HIGH confidence: an unexplained run either gets dug into until it joins a signature, or the bucket's confidence is honestly downgraded and the digging continues. Never bundle an unexplained sub-pattern into another pattern's triage "pragmatically" or "since the regression is open anyway" — triage covers only the runs whose root cause it actually explains; attaching unexplained runs to it hides them from the next duty shift and mis-scopes the bug. If a sub-pattern remains unexplained, the whole bucket is not finalizable — keep digging. If confidence is not HIGH after the steps above, keep digging until it is — escalate through the evidence ladder yourself: raw failure outputs → job-run summaries → GCS artifacts (prow-job-artifact-search, install/test-failure analysis skills) → audit logs (grep for the failing object/namespace to identify the creating user and userAgent — this reliably resolves "who created this pod/namespace" questions) → junit timing correlation (what else ran in the same window) → suspect-PR vetting. Triage taking longer is acceptable; leaving an unexplained bucket is not. The only permitted low-confidence outcome is when the evidence is genuinely exhausted (artifacts expired, logs missing), and then the report must say exactly what was checked and what was missing.
Additional finalization rules (each has caused a wrong disposition in a real duty run):
job_runs by date and confirm the newest runs' signature. A signature absent from a 3–4 run sample of a 20-run regression is not evidence it stopped. And when a sub-test starts "passing" recently, verify recent runs actually reach that stage: a bootstrap failure cannot "self-resolve" while newer runs of the same job die earlier at infrastructure provisioning — the earlier failure masks the later stage, it does not fix it.--existing-only, a new-issue bucket instead gets a "Proposed new bugs (not filed)" entry) — "collateral of noisy runs" is only a valid leftover justification when the failure has no independent mechanism (pure co-occurrence). A failure that is deterministically produced by another test's behavior has an independent mechanism and must be triaged as test.--triage-id); do not create a second triage record pointing at the same JIRA.*_previous.log reads for crash-looping containers). A missing mandatory quote means the sub-analysis is incomplete and must be redone, regardless of how confident its prose sounds. Batch size is never a reason to skip mandatory steps.Each bucket must have:
product / test / ci-infra / product-infraFor each bucket, before filing anything new:
Check bugs linked to the bucket's fired labels, then triaged_matches from fetch-related-triages (confidence ≥5 with an open JIRA is the default target) — but never act on a match, even conf=10, without the Phase 3 per-bucket CI-sample verification confirming the root cause actually matches.
Check open_bugs from the test report.
Search Jira for the root-cause signature (error message, operator name, component-regression label) in OCPBUGS against the owning component — the right bug may exist under Monitoring/etcd/MCO even though the regression sits under Installer.
Component-scoped JIRA listing (mandatory — keyword search is not enough). Owning teams describe defects in developer vocabulary that shares no keywords with the CI-side symptom. After determining the owning component, list its recent bugs regardless of keywords and read the summaries:
project = OCPBUGS AND component = "<owning component>" AND created >= -21d ORDER BY created DESCA bug whose creation date falls inside the bucket's failure window, on the owning component, is a duplicate candidate even with zero keyword overlap — open it and compare mechanisms.
Owning-repo merge-history check (mandatory when onset or cessation is dated). Query the owning repo for PRs merged around the bucket's onset and cessation dates:
gh pr list --repo <org>/<repo> --state merged --search "merged:<window>" --json number,title,mergedAtA merge at the cessation boundary is likely the fix — its OCPBUGS-* title prefix names the existing bug: triage to that bug instead of filing a new one. A merge at the onset boundary is a suspect trigger. The phrase "resolved by unidentified payload change" is banned from reports and bugs unless this check was run and came back empty. For currently-live breakage, also scan the owning repo's newest merges for revert PRs: a fresh Revert "..." title citing a TRT/OCPBUGS key hands you the trigger PR, the tracking ticket, and the expected recovery time in one query.
Then act (this is where --auto-triage applies; without it, confirm each bucket with the user):
triage-regression skill with --triage-id (additive merge is automatic; pass only the new IDs).triage-regression skill with --url, --type, and a one-sentence --description (<120 chars).--existing-only: do not file or triage; add a Phase 5 proposal instead): file with /jira:create bug (the create skill from the jira plugin) against the owning component, label component-regression, description per the bug-filing template in /ci:analyze-regression ("Prepare Bug Filing Recommendations" section: full test names in {code} blocks, test IDs, regression IDs, variants, error signature, Sippy test-details UI links for every member regression, suspect PRs). Every JIRA issue or comment created by this workflow must end with an AI-attribution footer as a separate, visually marked block — not a sentence buried in the text: place it after a divider, as its own paragraph or note panel, e.g. a rule followed by a panel (type note) in ADF containing "AI-generated content: This bug was filed by AI as part of Component Readiness triage duty. Please verify before acting on it." Release blocker is conditional on the triage type and impact, not automatic: mark the bug a release blocker (set-release-blocker skill) only for product bugs whose failures block or materially degrade blocking/informing payload jobs; test bugs (races, invariant-scan interference) and ci-infra issues (cloud capacity, registry outages) are not release blockers — state the blocker decision and its one-line justification in the report. Then create the triage record.add-jira-triage-link skill to put the triage URL into the JIRA description. The skill appends to the description — if the description ends with the AI-attribution footer, the appended link would land after it. After running the skill, verify the footer is still the final block; if not, move the footer back to the end of the description (or insert the triage link before the footer in the same update).manage-labels update --bugs <existing>,<new> (replaces the list; needs user confirmation, not covered by --auto-triage).With --auto-triage, only act autonomously when confidence is high: consistent error signature across the bucket, and either a confidence ≥5 triaged match or an unambiguous existing open bug. Buckets requiring a new bug, or with mixed signals, are always presented for confirmation.
--existing-only)This mode records matches to existing bugs but never files new JIRA issues. Analysis is unchanged. These rules apply to every write in any phase, including the Phase 1 stale-triage sweep:
Allowed writes: only for a verified match to an existing issue:
add-jira-triage-link;Everything else, including creating, reopening, or transitioning issues, goes in the report as a recommendation.
Otherwise, propose. New-bug buckets, failed fixes, and uncertain matches get a draft under "Proposed new bugs (not filed)" in the Phase 5 report. Their regressions stay untriaged. Never lower the match bar to avoid a proposal: a wrong triage hides the real defect.
Subagents do analysis only; the orchestrator does every write.
Present a final report:
bugs, matcher, file pattern, validation pair, retro-apply run set. Proposing costs nothing (creation still requires explicit user confirmation); a duty shift that root-caused a bucket and did not propose a symptom for an obviously grep-able signature has left cheap future-triage value on the table. If no bucket qualifies, say so and why (e.g., signature only visible via timing correlation, artifact expires, no unique string).--existing-only only, mandatory even if empty): one draft per proposal, following the Phase 4 new-bug template, with member regression IDs and the release-blocker recommendation. A human must be able to file it as written.When a bucket has a crisp, grep-able artifact signature (an error string or log line unique to the root cause), define a Sippy symptom so future runs are labeled automatically and the signature can be applied retroactively. Symptoms have repeatedly paid for themselves in real duty runs — a validated matcher has adjudicated a disputed root-cause attribution and machine-verified a 19-vs-1 bucket split (Case 6 in case-notes.md). Good candidates: recurring infra events that spawn regressions across many job families, single-occurrence product signatures worth a tripwire (n=1 bugs), and any signature that a "known recurring family" keeps regenerating. These are writes: always present the proposed label/symptom definitions and the target run list to the user and get explicit confirmation first — --auto-triage does not authorize them.
Check for existing coverage before creating anything. List the current catalog — this read is public and needs no token (GET https://sippy.dptools.openshift.org/api/jobs/labels, GET .../api/jobs/symptoms, or the list-symptoms skill) — and search it for the bucket's signature: the label/symptom may already exist from the shift that first triaged the bug (Case 6: a proposed console-crash symptom fully duplicated an existing one; the duplicate-key 400 from the label POST was what revealed it). Also search by the bug key (list-symptoms --labels --search OCPBUGS-NNN); a label linked to the bug counts as coverage only if its matcher targets this bucket's signature. A dry-run reevaluate of one bucket run is an equivalent cheap probe when a DPCR token is available: if an existing symptom targeting this signature fires on the run, the bucket is already covered and nothing should be created.
Create a label and symptom via the authenticated API (DPCR Bearer token, same as triage writes): POST https://sippy-auth.dptools.openshift.org/api/jobs/labels ({"label_title": ..., "explanation": ..., "bugs": ["<JIRA key>"]} — the title becomes the ID), then POST .../api/jobs/symptoms ({"summary": "<descriptive name>", "matcher_type": "string|regex|file", "file_pattern": "<glob under the job-run GCS root>", "match_string": ..., "label_ids": [...]}). The symptom ID is derived from the summary; gzipped artifacts are decompressed transparently. Keep Jira keys in the label's bugs array rather than embedding them in label explanations or symptom summaries. To change a symptom, use its full-replacement PUT endpoint.
Beware benign-transient false positives — presence is not pathology. Before finalizing a matcher, grep the target file in a healthy run: many error lines appear benignly during normal startup — sometimes once, sometimes in bulk (Case 6: a proposed match string appeared 120× in a healthy install's etcd-operator log as a normal warm-up transient; the failed run differed only in that the line persisted to log-end). String/regex matchers cannot count occurrences or check persistence, so if the line can appear benignly at all, re-anchor to an artifact that only exists — or only contains the string — in the failure mode: a container's *_previous.log (only present after a crash/restart) instead of its current log, an events.json string that healthy runs never emit, or best of all the gather-time end state in clusteroperators.json (e.g., StaticPodsAvailable: 0 nodes are active appears there only when etcd truly never deployed — zero hits in healthy runs by construction, because gather runs after the failure window).
Validate with a dry run against one known-positive and one known-negative run — the best negative is a near miss: a failed run of the same job family with a different root cause (it catches over-broad matchers that a green run would not). POST .../api/jobs/runs/reevaluate with {"prow_job_build_ids": [...], "dry_run": true}. Proceed only if both outcomes are correct; otherwise refine the matcher and re-validate. Note: reevaluation runs all symptoms, so unrelated pre-existing labels may legitimately appear on your negative — check only that your symptom's matching is correct.
Then apply with "dry_run": false to the bucket's prowjob_run_ids. The API caps requests at 50 IDs, but reevaluation scans GCS artifacts server-side and large batches time out at the gateway (504) — use batches of 5–10 with a retry, check each result's status, and stop and report on the first failed batch or any non-success result rather than continuing.
Match the underlying error, not a transient side effect: a matcher keyed on a crash/panic goes silently false-negative the moment a partial fix removes the crash while the defect persists. If the primary evidence lives in pod logs that sometimes fail to gather (dying clusters), add a fallback symptom against an artifact that survives (e.g., pods.json state/reason strings) mapped to the same label. When the same signature can surface under different step names (e.g., devscripts-driven vs plain IPI installs), create one symptom per file pattern, all mapped to the same label.
--audit-closed) — justify why every closed regression closedbulk-triage-regressions <view> --components ... --audit-closedAn explicitly-requested, read-only mode that answers a different question than triage duty: not "who owns this failure?" but "why did this regression close, and can we prove it?" Every closed untriaged regression must end the audit in exactly one of these justification classes:
| Class | Meaning | Proof required |
|---|---|---|
fixed | A fix merged; closure follows it | Fix PR/bug with merge/resolution date at or before the close boundary |
event-ended | An infra/payload event ended | Event window bracketing the failures (bad payload, repo outage, cloud capacity, credentials) |
intermittent-cause-open | Pass rate recovered but the defect is still open | Signature matched to an open bug; state explicitly that closure ≠ resolution |
collateral | Closed with the window of a tracked sibling event | Signature or run overlap with the tracked event |
evidence-expired | Run history and artifacts are gone | List exactly what was checked (regression details, GCS) and the statistical argument (age, sibling consistency) — flagged, never silently absorbed |
No regression may be left as "unknown". evidence-expired is the only permitted terminal state without a mechanism, and it must be flagged in its own report section.
artifacts/*/ipi-install-install*/build-log.txt, devscripts equivalent on metal) with an ordered signature catalog (most-specific first). Seed the catalog from the view's known events and extend it with whatever the duty history has verified (ign-push 403s, cipherSuites/rev-0, toolchain download, quay pulls, registry.ci 5xx, quota, capacity, provisioning, per-operator stabilization). Guard against substring traps (Case 7: lease matching inside release mis-binned 31 regressions).artifacts/*/<e2e-step>/artifacts/junit/junit_e2e*.xml) with a real XML parser (regex over XML mis-matched ~95% of testcases in practice) and take the <failure> of the exact testcase./api/component_readiness/triages) and index their regressions by test name — a closed untriaged record is very often the untriaged sibling of a triaged one, which supplies the bug and the closure mechanism for free.fixed requires the fix date ≤ close boundary (JIRA resolutiondate, gh pr list --search "merged:<window>" on the owning repo); event-ended requires the last failing run to fall inside the event window. The Phase 4 component-scoped JIRA listing and merge-history checks apply here unchanged.intermittent-cause-open the explicit warning that siblings will reopen.evidence-expired member.intermittent-cause-open) is a prediction of future regressions — surface these in the summary so the next duty shift expects them.Principles distilled from real mis-dispositions; full narratives in case-notes.md.
pass_sequence strings. Misreading direction inverts "regressed" vs "resolved".install should succeed regressions in duty batches were owned by other components or were infra noise. The installer is the messenger.Unknown component regressions (e.g., verify the cluster readiness and stability, verify all machines should be in Running state) are wrappers; the co-failing tests and operator states identify the owner.test_details_url to the sippy-ng UI form before putting it in bugs or reports, by replacing the base https://sippy.dptools.openshift.org/api/component_readiness/test_details with https://sippy-auth.dptools.openshift.org/sippy-ng/component_readiness/test_details (query parameters are identical).fetch-related-triages and a --test-name query without an open-only filter find these; reuse the existing triage/JIRA instead of filing a new bug.prowjob_run_id sets are the strongest clustering signal — but still not proof. When two regressions were opened from literally the same job runs, treat them as one draft bucket, then confirm in Phase 3 that the failure outputs actually point at one cause: a single run can carry independent defects, and in mass-failure runs co-occurrence is largely coincidental. Only merge into one triage after the error signatures/artifacts agree.etcdserver: request timed out during test setup — an etcd bug.)test. (Case 2.)clusterversion not available: False drafted as "UDN breakage" when the interval message named an unrelated CVO controller error.)test triage and a bug, regardless of how noisy the surrounding runs are.reevaluate a representative run and treat every label hit as a known-cause hypothesis to check. (Case 3: a 5,471-line log flood flagged as new was a downstream consequence of a cataloged defect.)ImagePullNeverCompletes on metal jobs (ci-infra), transient cloud quota or DNS provisioning errors. Search existing triages for the signature before opening anything.test even if it looks like a product failure; an etcd timeout that also hits customers is product; a registry outage is ci-infra even when it kills installs.job_runs list. Broad Sippy job-filter URLs ("all jobs failing test X") sweep in unrelated jobs and post-fix eras. Every run you cite must be a member of the regression (and predate any candidate fix). (Case 10: "3 random jobs" from a broad filter produced three confidently wrong root causes in one shift.)Deleting pod <step> that failed to start or pod pending for more than 1h: pod has not been scheduled (with 0/N nodes are available: ... untolerated taints / unschedulable), the OpenShift installer never ran — no product artifact exists to analyze. Classify as ci-infra, and name the build cluster from prowjob.json (.spec.cluster, e.g. build10) so the infra team knows where the capacity problem is. Trivial setup steps (rbac, hosted-loki) failing this way across a job family for days is a capacity outage worth its own tracking ticket.Unknown conditions are not resource pressure. When a kubelet stops posting status, all node conditions flip to Unknown, and the MCO controller logs Reporting unready: ... OutOfDisk=Unknown — this is a stale-status marker, not disk exhaustion. Only DiskPressure/MemoryPressure=True (or eviction/OOM events, df output, sosreport sar) is evidence of actual resource pressure.<view>: Component Readiness view name (e.g., 5.0-main). Required.--components: Space-separated component name filters, case-insensitive and hierarchy-aware (e.g., Installer also matches Installer / openshift-installer; Networking matches Networking / ovn-kubernetes, Networking / router, ...). Required in practice for duty scoping.--auto-triage: Allow high-confidence buckets to be triaged without per-bucket confirmation. New bug filing always requires confirmation.--existing-only: Record matches to existing triages/JIRA issues but never file new JIRA issues; new-bug buckets become report proposals (see "Existing-only mode" in Phase 4).--audit-closed: Read-only closed-set audit mode (see "Closed-set audit mode") — justifies why every closed untriaged regression closed. Mutually exclusive with normal duty triage and with --auto-triage; performs no writes./ci:analyze-regression — single-regression deep dive; this skill orchestrates its techniques across a batchcreate (jira plugin) — file new JIRA bugs via /jira:create bug (plugins/jira/skills/create/SKILL.md)list-regressions (teams plugin) — batch listing (plugins/teams/skills/list-regressions/SKILL.md)fetch-regression-details (plugins/ci/skills/fetch-regression-details/SKILL.md)fetch-related-triages (plugins/ci/skills/fetch-related-triages/SKILL.md)fetch-test-runs (plugins/ci/skills/fetch-test-runs/SKILL.md)fetch-job-run-summary (plugins/ci/skills/fetch-job-run-summary/SKILL.md)prow-job-analysis — GCS artifact analysis for install/bootstrap failures (plugins/ci/skills/prow-job-analysis/SKILL.md)fetch-test-report (plugins/ci/skills/fetch-test-report/SKILL.md)triage-regression (plugins/ci/skills/triage-regression/SKILL.md)add-jira-triage-link (plugins/ci/skills/add-jira-triage-link/SKILL.md)set-release-blocker (plugins/ci/skills/set-release-blocker/SKILL.md)oc-auth (plugins/ci/skills/oc-auth/SKILL.md)Evals before narratives. When a new duty-run mis-disposition surfaces, do not reach for a new case narrative or case-specific rule first. Instead: (1) capture the incident as an eval case under plugins/ci/evals/cases/bulk-triage-regressions/ (frozen regression snapshot + owning-team-verified ground truth + the wrong conclusion as a banned outcome); (2) run the eval against the unchanged skill — if the existing principles and the mechanism self-review already catch it, nothing needs adding; (3) only if the eval fails repeatedly does the incident earn a new generic rule (preferred) or narrative, and the eval then proves the addition works and guards it against regression. This is how the two source incidents of this rule set (the VAP/actuator and GID-mapping mis-diagnoses) were handled: a baseline-vs-variant eval matrix showed the generic rules catch both, so their narratives were dropped and their ground truth lives in eval cases 001/002. Every prompt token here costs money on every duty run — additions must pay for themselves under eval-bulk-triage-regressions.yaml.
© openshift-eng, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in plugins/ci/skills/bulk-triage-regressions of openshift-eng/ai-helpers.
Open the folder on GitHubat commit a627176
Bulk Triage Regressions next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bulk Triage Regressions this skillopenshift-eng/ai-helpers | 120 | — | ~17k | Automated safety check: Pass | Apache-2.0 | |
| OpenLogi macOS Permissions TriageAprilNEA/OpenLogi | 23k | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| Bug Finder for daisyUIsaadeghi/daisyui | 43k | — | ~2.3k | Automated safety check: Pass | MIT | |
| Root Cause Debugginggarrytan/gstack | 136k | — | ~1.4k | Automated safety check: Pass | MIT | |
| Review PRapache/shardingsphere | 21k | — | ~6.5k | Automated safety check: Pass | Apache-2.0 | |
| Graph-Based Bug Tracingtirth8205/code-review-graph | 32k | 1 repos | ~287 | Automated safety check: Pass | MIT |
AprilNEA/OpenLogi
Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.
saadeghi/daisyui
Investigates suspected bugs in the daisyUI monorepo through read-only analysis, then writes a decision-ready fix plan in tmp/bugs without changing any product code.
garrytan/gstack
Investigates bugs, errors and stack traces in phases and requires a root-cause hypothesis to be confirmed before any fix is written.
apache/shardingsphere
Review Apache ShardingSphere or user-authorized downstream pull requests and PR discussions from public or authorized repository evidence.
tirth8205/code-review-graph
Traces a bug through a code knowledge graph, following callers, callees and execution flow before opening source files, within a small token budget.
go-musicfox/go-musicfox
Fix or implement a tracker issue end to end from a single command — takes an issue id or a plain problem description (filed first via om-prepare-issue), classifies, then drives the bug autofix chain…
openshift-eng/ai-helpers
Find and independently validate actionable reliability defects across OpenShift release jobs and presubmits, then export portable issue handoffs.
openshift-eng/ai-helpers
Fetch and address all PR review comments — categorize by priority, make code changes, post replies, and push.
openshift-eng/ai-helpers
Categorize Jira issues into Red Hat Sankey Activity Type categories using MCP Jira tools.
openshift-eng/ai-helpers
Decide whether a GitHub PR has unanswered authorized review comments or new required CI failures worth a follow-up agent.
openshift-eng/ai-helpers
Analyze OpenShift must-gather diagnostic data including cluster operators, pods, nodes, and network components.
openshift-eng/ai-helpers
Schema for the autodl JSON data file produced by payload-analysis for database ingestion — you must use this skill whenever generating the autodl JSON file
Categories
A skill your agent uses for Component Readiness triage duty - holistically analyze and triage all untriaged regressions for a set of components in a view, clustering them into root-cause buckets. Bulk Triage Regressions is an agent skill from openshift-eng/ai-helpers.
Bulk Triage Regressions fits situations like: component Readiness triage duty - holistically analyze and triage all untriaged regressions for a set of components in a view; clustering them into root-cause buckets.
Run `npx skills add openshift-eng/ai-helpers --skill bulk-triage-regressions -a claude-code`. Or copy the skill folder (plugins/ci/skills/bulk-triage-regressions in openshift-eng/ai-helpers) into .claude/skills/bulk-triage-regressions in your project. Claude Code loads it when a task matches its description.
Run `npx skills add openshift-eng/ai-helpers --skill bulk-triage-regressions -a codex`. Or copy the skill folder (plugins/ci/skills/bulk-triage-regressions in openshift-eng/ai-helpers) into .agents/skills/bulk-triage-regressions in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openshift-eng/ai-helpers --skill bulk-triage-regressions -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bulk-triage-regressions, .gemini/skills/bulk-triage-regressions, .github/skills/bulk-triage-regressions and .opencode/skills/bulk-triage-regressions in your project.
Going by SKILL.md and its folder, Bulk Triage Regressions needs the command-line tools its instructions call (gh and python3) and credentials named JIRA_API_TOKEN. Our summary lists: Python 3; A credential in JIRA_API_TOKEN.
SKILL.md names 6 domains. In commands or code: sippy-auth.dptools.openshift.org, sippy.dptools.openshift.org and redhat.atlassian.net; the agent is likely to contact these when it follows the instructions. As links in the text: id.atlassian.com, redhat.enterprise.slack.com and docs.ci.openshift.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bulk Triage Regressions is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 17k tokens (SKILL.md is roughly 69k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Bulk Triage Regressions: OpenLogi macOS Permissions Triage (AprilNEA/OpenLogi, 23k stars), Bug Finder for daisyUI (saadeghi/daisyui, 43k stars), Root Cause Debugging (garrytan/gstack, 136k stars) and Review PR (apache/shardingsphere, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
openshift-eng (a GitHub organization) maintains it in openshift-eng/ai-helpers, which has 120 GitHub stars. The repository holds 118 skills in this directory. The repository was last updated on October 6, 2026.
Source: openshift-eng/ai-helpers on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.