Agent skill

Bulk Triage Regressions

by openshift-eng in openshift-eng/ai-helpers

A skill your agent uses for Component Readiness triage duty - holistically analyze and triage all untriaged regressions for a set of components in a view, clustering them into root-cause buckets

Apache-2.0Auto-check passedDevelopment

Install Bulk Triage Regressions

skills CLI
$ npx skills add openshift-eng/ai-helpers --skill bulk-triage-regressions -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install openshift-eng/ai-helpers bulk-triage-regressions --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ci/skills/bulk-triage-regressions .claude/skills/bulk-triage-regressions && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bulk-triage-regressions
GitHub stars
120
Token cost
~17k tokens
SKILL.md length
9,525 words
Files
2 (incl. references)
Skills in repo
118
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses for Component Readiness triage duty - holistically analyze and triage all untriaged regressions for a set of components in a view, clustering them into root-cause buckets

  • Works in 5 steps: Collect the full batch → Cluster into candidate buckets (cheap… → Deep-dive each bucket (confirm root… → …
  • Component Readiness triage duty - holistically analyze and triage all untriaged regressions for a set of components in a view
  • SKILL.md covers Input, Description, Implementation and Closed-set audit mode…, plus 4 more sections
  • Calls gh and python3; reaches sippy-auth.dptools.openshift.org and sippy.dptools.openshift.org; needs JIRA_API_TOKEN

What it does

Bulk Triage Regressions is an agent skill from openshift-eng/ai-helpers. Use this skill for Component Readiness triage duty - holistically analyze and triage all untriaged regressions for a set of components in a view, clustering them into root-cause buckets

Its SKILL.md is about 17k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/case-notes.md`).

It sits in Development, covering Root cause analysis. The repository describes itself as: Developer productivity tools for Claude Code & other AI assistants. The licence is Apache-2.0.

When your agent uses it

  • Component Readiness triage duty - holistically analyze and triage all untriaged regressions for a set of components in a view
  • Clustering them into root-cause buckets

Example prompts

  • “/bulk-triage-regressions”

Requirements

  • Python 3
  • A credential in JIRA_API_TOKEN

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Collect the full batch
  2. Cluster into candidate buckets (cheap signals first)
  3. Deep-dive each bucket (confirm root cause and real owner)
  4. Search for existing bugs, then triage each bucket
  5. Duty report

What it can do on your machine

Read from SKILL.md and the folder at commit a627176. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • sippy-auth.dptools.openshift.org
    • sippy.dptools.openshift.org
    • redhat.atlassian.net

    Also links to:

    • id.atlassian.com
    • redhat.enterprise.slack.com
    • docs.ci.openshift.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • JIRA_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bulk Triage Regressions loads about 17k tokens when it runs, and up to ~20k if it reads all its reference files. Until then it costs about 52 tokens; SKILL.md has 9,525 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~17k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~20k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from openshift-eng/ai-helpers at commit a627176, republished under its Apache-2.0 licence (© openshift-eng). 9,525 words, ~17,164 tokens.

Download SKILL.mdSave it as .claude/skills/bulk-triage-regressions/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
bulk-triage-regressions
description
Use this skill for Component Readiness triage duty - holistically analyze and triage all untriaged regressions for a set of components in a view, clustering them into root-cause buckets

Bulk Triage Regressions

Input

bulk-triage-regressions <view> [--components comp1 comp2 ...] [--auto-triage] [--existing-only]

Example: bulk-triage-regressions 5.0-main --components Installer Unknown

Description

This skill implements the Component Readiness triage duty workflow: it fetches all untriaged regressions for a set of components in a view (e.g., 5.0-main, components Installer and Unknown), analyzes them as a batch, clusters them into root-cause buckets, and then triages each bucket to a single JIRA bug (existing or new).

This differs from /ci:analyze-regression (which analyzes a single regression in depth). Triage duty requires a holistic view, because:

  1. Many regressions, few root causes. One product bug commonly opens 5–30 regressions across variants (different platforms, arches, featuresets, upgrade modes) and across "wrapper" tests (install should succeed: overall, : cluster bootstrap, : cluster creation, verify the cluster readiness and stability, mass-failure tests, etc.). Analyzing regressions one-by-one wastes effort and risks filing duplicate bugs. Cluster first, deep-dive once per cluster.

  2. Component attribution is often wrong. Regressions in Installer and Unknown are catch-all attributions. A failed installation or bootstrap is frequently caused by a specific component — e.g., a monitoring operator failing to go available blocks cluster creation, an etcd slowness issue breaks bootstrap, an MCO bug degrades nodes during install. The Sippy component label tells you which test failed, not whose bug it is. The real owner must be determined from artifacts (cluster operator status, log bundle, operator logs), and the JIRA bug must be filed against the actual owning component, not Installer.

Use this skill when doing triage duty for a view, or whenever a user asks to "look at all untriaged regressions from <components>" rather than a single regression ID.

Implementation

Script invocation rules: Run Python skill scripts directly and analyze their JSON output with your own reasoning (pass --format json where the script offers the flag; scripts without it, such as list_regressions.py, emit JSON by default). Do not pipe script output through inline Python one-liners. Do not suppress stderr: if a script exits non-zero or returns invalid/empty JSON, stop and surface the error — an authentication or API failure must never be mistaken for an empty inventory ("nothing to triage").

Authentication: Read steps (listing, fetching details, test runs, GCS artifacts) require no auth — Phases 1–3 and a read-only report must work without any credentials. Two Sippy hosts exist and the split matters: https://sippy.dptools.openshift.org is public and read-only (regressions, test details, and the full symptom/label catalog — see the list-symptoms skill), while everything on https://sippy-auth.dptools.openshift.org requires a DPCR Bearer token, including the reevaluate probe, which has no public equivalent (the route returns 404 on the public host). The only Phase 3 step that needs a token is therefore the active symptom probe. When it is unavailable (no token, or a 401/403), treat the result as an unknown — never as "no symptoms matched", the same error as reading a failed listing script as an empty inventory, and work down this fallback: (1) passive label_summary / job_labels from fetch-regression-details; (2) list the catalog from the public API with the list-symptoms skill and grep each plausible symptom's match_string against its file_pattern in the GCS artifacts you are already fetching — this reproduces the probe locally for the handful of symptoms related to the bucket's stage and platform; (3) only then the full evidence ladder, recording that the active probe was unavailable so the next shift knows the catalog was checked by hand. Write credentials are validated only when writes are going to happen: with --auto-triage, validate both up front (so an expired token surfaces before hours of analysis); otherwise validate at the start of Phase 4, before the first write. Sippy writes (creating/updating triage records) require a Bearer token from the DPCR cluster (api.cr.j7t7.p1.openshiftapps.com:6443) — see the oc-auth skill and the token-extraction snippet in /ci:analyze-regression; check with an authenticated GET against https://sippy-auth.dptools.openshift.org/api/component_readiness/triages (200 vs 401/403).

JIRA writes (filing bugs, set-release-blocker, add-jira-triage-link) additionally require the JIRA_USERNAME and JIRA_API_TOKEN environment variables (API token from https://id.atlassian.com/manage-profile/security/api-tokens); verify with an authenticated GET against https://redhat.atlassian.net/rest/api/3/myself (Basic auth, 200 vs 401/403). If a required write credential is missing or invalid, pause before Phase 4 and ask the user to fix it — the analysis and report so far remain valid and must still be presented.

Phase 1: Collect the full batch
  1. Load CI context: Read the files in plugins/ci/references/ (jobs.md, tests.md, sippy-apis.md) for conventions on tests, jobs, and Sippy APIs.

  2. Parse arguments:

    • view: required, e.g. 5.0-main
    • --components: component filter list, e.g. Installer Unknown. Matching is case-insensitive and hierarchy-aware: a filter matches the full component name or any /-separated segment of it, so Installer also covers Installer / openshift-installer, and Networking covers Networking / ovn-kubernetes, Networking / router, and every other Networking / * component. If omitted, ask the user which components the duty covers.
    • --auto-triage: if present, triage buckets without per-bucket confirmation when confidence is high (see Phase 4). Default is to present findings and confirm before writing.
    • --existing-only: never file new JIRA issues; write only for matches to existing issues (see "Existing-only mode" in Phase 4).
  3. List regressions with the list-regressions skill:

    bash
    python3 plugins/teams/skills/list-regressions/list_regressions.py \
      --view <view> --components <components...>

    Keep only open, untriaged regressions (empty triages array), but note recently-triaged ones — they are prime candidates for absorbing untriaged siblings.

    Closed regressions are out of scope — even when untriaged. A regression whose closed field is set has already resolved itself; do not inventory it, cluster it, deep-dive it, or recommend retroactive triage for it. The duty batch consists solely of open untriaged regressions. Closed regressions may be consulted as evidence (e.g., a closed sibling that shares a root cause with an open bucket, or a closed sibling whose existing triage/JIRA an open bucket should reuse — see Pitfalls), but they must never appear as bucket members, action items, or "leftovers" in the report. The only exception is the explicit closed-set audit mode (--audit-closed, see below), which the user must request by name — it is never part of a normal duty run.

  4. Build a batch inventory table: For every open untriaged regression record: regression ID, test name, component/capability, variants (Platform/Arch/Network/Topology/FeatureSet/Upgrade), opened date, failure/run counts. Present this table to the user up front so the scope of the duty run is visible. Do not include closed regressions in the inventory.

  5. Stale-triage sweep (mandatory) — a 100% triaged board can still hide live defects. For every open, already-triaged regression, compare its last_failure against the state of its triage's JIRA: a regression whose bug is Closed/Verified/resolved but which has failed after the resolution date is an alarm, not a statistic. Its fresh failure window is either a failed fix or — more often — a different cause hiding behind the old triage record. For each such regression, verify the recent runs' signature (armed Sippy symptoms are a cheap first oracle: dry-run reevaluate the newest runs — label hits map windows to known causes instantly; unlabeled recent runs mean a new, uninvestigated cause; a hit label whose bugs differ from the triage's JIRA suggests a new cause, one listing the same closed bug a failed fix; a label linking several bugs is inconclusive — compare the runs' signature and timing with each) and add the correct triage(s) for the new window rather than trusting the stale one. (Case 4 in case-notes.md: three "triaged" regressions pointing at a Closed bug while failing daily from two new causes.)

  6. Long-lived wrapper regressions are cause timelines, not single buckets. A wrapper regression that stays open for weeks accumulates causes, each window separately verified and separately triaged. When a previously-analyzed regression shows new last_failure dates, re-verify the new window from scratch — never assume the existing triage covers it. (Case 13: one record carried four causes.)

Phase 2: Cluster into candidate buckets (cheap signals first)

Before any deep log analysis, group regressions using signals already in hand:

  • Same test, different variants — almost always one bucket.
  • Same variant fingerprint, different tests — e.g., install should succeed: overall + : cluster creation + verify the cluster readiness and stability all failing on azure/amd64/techpreview starting the same day is one bucket. Wrapper tests fail together.
  • Same opened date — regressions opened the same day across components often share one payload-level cause.
  • Shared job runs — fetch details for each regression (fetch-regression-details skill) and compare job_runs prowjob_run_ids. Regressions observed in the same failed runs are strong candidates for one bucket — but this is a clustering signal, not proof: the same run can contain independent defects or be a mass failure, so Phase 3 validation is still required. Also run the fetch-related-triages skill per regression; same_last_failure and similarly_named_test matches feed the clustering, and triaged_matches with confidence ≥5 immediately suggest an existing triage/bug for the whole bucket.
  • Symptom labels — fetch-regression-details returns label_summary (per job) and job_labels (per run). A label shared by most failed runs of several regressions is a strong bucket signal, and labels are precise (human-written matchers over artifacts). label_bugs maps each label to its linked Jira keys: labels sharing a bug are a bucket signal, and the bug is a candidate triage target. An empty label list means nothing was detected — not that nothing is wrong, and often just that the runs were never swept: Phase 3 opens with an active dry-run reevaluation that closes this gap.
  • Mass-failure marker: high test_failures counts in job_runs mean the regression is likely collateral of a bigger event, not an independent issue.

Output of this phase: a draft bucket list, each bucket with member regression IDs, the shared fingerprint (test/variant/date/job-run overlap), and any candidate existing triage or JIRA bug.

Treat buckets as hypotheses — Phase 3 must confirm or split them. Do not merge buckets merely because both are "install failures"; installs fail at different stages for unrelated reasons.

Phase 3: Deep-dive each bucket (confirm root cause and real owner)

For each bucket, pick 2–5 representative failed job runs (spread across jobs/variants; include the newest) and analyze. 2–5 runs are sufficient only when they yield a consistent result (same error signature / failure stage across all of them). If the sample is mixed or unclear — different errors, different stages, or an inconclusive owner — extend the sample to 10–20 runs before concluding; a small ambiguous sample must never be the basis for splitting/merging a bucket or attributing an owner.

Cross-cutting analysis principles — most specific rules below (and most Pitfalls) are instances of these three; when a situation matches none of the specific rules, fall back to the principle:

  • Messenger vs. owner. The component that reports a failure is presumed the messenger until artifacts prove it is the producer. This covers: the installer for install failures, wrapper tests for whatever blocked them, a test whose setup failed on another component's error (Case 9 in case-notes.md), a crash-looper whose error names a missing upstream dependency, and any component that names a bad value someone else computed (Case 12). Ask "who produced the condition this error describes?" and walk upstream until the answer is "this component itself".
  • Population claims require population evidence. Any claim quantifying over runs — "resolved", "stopped", "all runs show X", "only variant Y" — must enumerate the population it quantifies over: all of the regression's job_runs by date (never a 3–4 run sample; Case 1), only runs from the regression's own job_runs list (never a broad job-filter query; Case 10), and source-code gating for any "X-only" scope claim.
  • Consult known-cause catalogs before deriving from scratch. Cheap indices exist for almost every question and must be queried before artifact dives or new-bug filings: the armed symptom catalog (dry-run reevaluate), existing triage records, the owning component's recent JIRA bugs regardless of keywords, the owning repo's merge history at onset/cessation boundaries, and closed/dropped sibling regressions. (Case 3: one dry-run call resolved a board that cost a parallel run 264 shell commands.)

Symptom-catalog oracle first (cheap, decisive — run before any artifact dive or subagent dispatch). The passive label_summary from Phase 2 only shows labels already applied by past sweeps; recent runs are typically unswept and show nothing. Actively probe the catalog instead: POST https://sippy-auth.dptools.openshift.org/api/jobs/runs/reevaluate with {"prow_job_build_ids": [...], "dry_run": true} (DPCR Bearer token) for 3–5 representative runs per bucket, spread across platforms. Reevaluation scans the runs' GCS artifacts server-side against every armed symptom matcher and returns symptoms_matched per run — one call can attribute an entire bucket to a known incident in seconds (Case 3 in case-notes.md). A matched label's bugs (see list-symptoms --labels) name the candidate triage target. Matched symptoms are a hypothesis with a named cause — still verify per the steps below (the label names the incident; confirm the causal chain to the bucket's specific test) — but they dictate where to look first and usually collapse the deep-dive to a single confirmatory artifact read. No matches means the cause is not yet cataloged — proceed with the full evidence ladder.

No token? The catalog is still public — degrade, never skip. A 401/403 from the probe is an unknown, never "no symptoms matched": use the credential-free fallback in the Authentication section above, and say in the report which rung was used.

Deep-dive at least one CI sample per bucket — always, even at high confidence. A Sippy triaged_matches confidence of 10 or a same-day triaged sibling is a hypothesis, not a verdict: Sippy matches on test names and shared job runs, which produces conf=10 for the same test name across unrelated platforms and root causes (see Pitfalls). Before accepting any disposition — including "extend existing triage" — read the actual failure evidence for at least one representative run of the bucket (failure output at minimum; installer logs / artifacts for install wrappers) and confirm it matches the target triage's root cause. You have deeper analysis capabilities than Sippy's heuristics — use your own judgement on the raw evidence, and if it contradicts the Sippy categorization, trust the evidence and re-bucket.

  1. Failure outputs: fetch-test-runs skill with the bucket's test IDs and job run IDs — check whether error messages are consistent within the bucket. >90% same error is strong evidence of a single cause, not confirmation: wrapper tests and mass failures print identical error text for unrelated defects (e.g., "nodes not ready" covers disk, memory, and network deaths alike). Confirm with items 2–3 below before treating the bucket as one cause; inconsistent errors ⇒ split the bucket.

  2. Job run context: fetch-job-run-summary skill per representative run — is the regressed test isolated, part of a consistent co-failure set, or one of hundreds of random failures? For Unknown-component and mass-failure regressions this is where the real component reveals itself: read the names of the co-failing tests.

  3. Install/bootstrap failures — mandatory artifact dig: For any bucket whose tests include install should succeed (any stage) or bootstrap/cluster-readiness wrappers, invoke the prow-job-analysis skill per representative run. Do not stop at Sippy's generic "install failed" wrapper. From the GCS artifacts determine:

    • Failure stage: infrastructure provisioning / bootstrap / cluster creation (operators rolling out) / stability window.
    • The blocking condition: for cluster-creation failures, read clusteroperators.json (or the installer log's "Cluster operator X is not available" lines) and the failing operator's pod logs from the log bundle. For bootstrap failures, read the bootstrap log bundle (etcd, bootkube, release-image pulls).
    • Bootstrap-era logs are full of normal transients — validate every causal theory against the gathered end-state. During any bootstrap, operator logs contain scary-looking messages that resolve on their own ("the server could not find the requested resource (post routes.route.openshift.io)" before openshift-apiserver is up, static pods "0 nodes at revision 0" early on). Quoting one of these as the root cause is only valid if the final cluster state agrees. The end-state oracle for bootstrap-stage failures is clusteroperators.json, not pod existence: check etcd's StaticPodsAvailable/revision status and openshift-apiserver's APIServicesAvailable directly — a bootstrap-timed-out run whose gather still shows etcd Available=False: 0 nodes are active; 3 nodes are at revision 0 was blocked by etcd no matter what else looks broken. Do not infer "etcd/apiserver were fine" from workload pods running: during bootstrap the bootstrap-node etcd/apiserver serve the control plane, so pods run happily while cluster etcd never deploys. A crash-looping container's *_previous.log is primary evidence for that container's failure — but before assigning ownership, check whether its error names a missing upstream dependency: such a component is a victim, not an owner.
    • "Operators were not stable" — count condition transitions before classifying. "Healthy at gather time" does NOT imply a harmless transient. Grep the install log for the operator's condition lines and look at LastTransitionTime/DurationSinceTransition: hundreds of transitions with DurationSinceTransition=1s across the stability window means the operator is flapping continuously (a sync-loop product bug that will never pass the stability check — permafail), whereas a single long Progressing=True window that eventually clears is a slow rollout (flaky race against the timeout). These two have opposite classifications and dispositions. (Case 1: a ~1 Hz flap, 1600+ transitions in 30 minutes, mislabeled "self-resolved slow rollout" from a 3-run sample.)
    • Race/ordering theories require timestamp proof — from artifacts, not from the error string. Any "X was not ready when Y ran" mechanism must be proven with timestamps the artifacts already contain (creationTimestamps, Established/Available conditions, phase logs, first/last occurrence of the error). If X was ready before the errors began, that refutes this ordering theory — but not every timing theory: readiness at the source does not mean visibility at the consumer, so the next candidates are cache/informer warmup, an in-progress rollout, or a stale client. Name which one the artifacts support rather than dropping the timing angle entirely. Also measure the failure window's duration: it is part of the mechanism claim and belongs in the bug.
    • Bad-value failures: trace the value to its producer — the error-emitter is usually just the messenger. When a component rejects or fails writing a value (a mapping, an ID range, a path, a quota number), walk the data flow upstream until you reach whoever computed it, however many hops away: the component that names the bad value in its error is the victim's messenger. The producer is then the prime suspect, not the automatic owner — before filing, confirm the value was already wrong as produced: check that no intermediate hop transformed a valid value into an invalid one, and that the value does not violate a documented input contract the consumer was itself obliged to enforce. Whichever hop first turned a valid value into an invalid one owns the bug; name the hops you cleared. (Case 12 in case-notes.md: a GID-mapping error filed against CRI-O when the producer was kubelet's ID allocator two hops upstream — pinns and CRI-O were verified pass-throughs.)
    • Decode suspicious constants before writing "malformed" or "overflow". A weird-looking number is a mechanism clue, not evidence of corruption: check whether it is semantically meaningful — 2^32 − size, a type maximum, a page/alignment boundary, an epoch. Do the arithmetic; never pattern-match "big hex-ish number ⇒ overflow bug". (Case 12: 4294901760 + 65536 = 2^32 — an allocator boundary, handing you the whole mechanism.)
    • The claimed mechanism must predict the observed frequency. A deterministic-bug theory cannot explain n=1 across many runs exercising the same path — a low rate demands a stochastic or boundary mechanism, and the expected probability should be stated in the bug. A mechanism/frequency mismatch means the theory is wrong; resolve it before filing, and mark any remaining mechanism prose as an explicit unverified hypothesis ("unverified: possibly …"), never as "suggesting X bug in component Y". (Case 12: the filed theory predicted every-pod failure; the real mechanism predicted ~1/65535 per pod — matching n=1.)
    • Differential recovery check — when the same error hits many objects but only some stay broken, explain the asymmetry before finalizing. Survivors show the failure was not uniformly permanent; before concluding the trigger was transient, rule out the alternatives (survivors had different exposure, a retry path, or different state). Either way the lasting failure lives in whatever failed to recover, and that component owns the bug. Pattern: transient trigger vs. persistent damage — ask which candidate defect explains the permanence of the failure, and lead with that one.
    • Read the owning repo's source before claiming a missing safeguard. A claim that "component X lacks ordering/gating/protection Y" is only valid after locating (or proving absent) that safeguard in the repo — the same rule as reading a test's skip/gates before describing its scope, applied to product code.
    • The real owner: the component whose operator/pods are actually failing. Examples from past duty: "cluster creation failed" ⇒ monitoring operator degraded ⇒ Monitoring bug; "bootstrap failed" ⇒ etcdserver: request timed out on Azure ⇒ etcd bug; nodes degraded during install ⇒ MCO bug; quota/DNS/cloud-API errors ⇒ ci-infra, not a product bug at all.
    • Stability-window failures — "operator X not available" is a symptom, not a root cause. For cluster-readiness/stability wrappers that fail on a ClusterVersion or ClusterOperator condition, two extra steps are mandatory before assigning an owner:
      1. Recover the underlying condition message when the junit output is vague. Outputs like clusterversion not available: False (empty reason) carry no cause. Run this literal check against the e2e step's build-log.txt (and, if present, the monitor-intervals JSON under the step's artifacts/junit/):

        bash
        curl -s <gcs-url-of-e2e-step>/build-log.txt | grep -oE '(Failing|Available|Degraded)=(True|False)[:,][^"\\]{0,160}' | sort | uniq -c | sort -rn | head

        The condition messages recovered this way (e.g., a controller error string) name the real culprit. This step is complete only when the report quotes the recovered message for every vague-output run. Never attribute such a run from its co-failing tests: co-failures on techpreview jobs are usually unrelated background noise, and correlation with them has produced wrong owners in past duty runs.

      2. Ask why the operator went unavailable, not just which operator. If everything is healthy at gather time, the wrapper caught a transient flap: read the operator's own pod log around the transition timestamps and find the trigger. Distinguish (a) the operator genuinely failing (⇒ product bug for that operator) from (b) a routine reconciliation/rollout triggered by a cluster mutation (config/secret change, node roll) — a rollout of a single-replica deployment flaps Available by design. A rollout-flap disposition must name the mutator, not stop at "the operator detected a configuration change": quote the operator-log line identifying which object changed (grep the operator pod log for object changed / secret and config names and read the surrounding lines), then identify who wrote it — audit logs (verb update/patch on that object: user + userAgent) or the job's test-harness step logs (search them for the object name or the value being set). If a test step caused the mutation mid-run, the bucket is test owned by that suite, not a product bug against the flapping operator; filing "spurious rollouts" against the operator without naming the mutator is an incomplete deep-dive. (Case 2 in case-notes.md: an image-registry flap drafted as a product bug until the operator log showed the harness had replaced the pull secret mid-run.)

  4. Test-interference buckets — read the offending test's source before describing its scope. When the root cause is another test's or tool's behavior (a test creates pods/namespaces that break an invariant check, leaks resources, reboots nodes, etc.), do not infer why the failure is confined to certain variants from the regression's variant labels — correlation with a variant (techpreview, platform, upgrade mode) is not evidence of how the offending code is gated. Instead:

    • Locate the offending code (gh search code in openshift/origin or the relevant repo for the test name, namespace prefix, or error string) and read its skip/gate conditions (skipIf..., e2eskipper.Skipf, feature-gate checks, platform checks, suite membership).
    • State the actual gating in the report (e.g., "gated on baremetal platform via skipIfNotBaremetal"), and if that gating is broader than the regressed variants, say so — the regression's variant slice then reflects job-scheduling or sample-size effects, not the blast radius, and other variants are also at risk.
    • Check the file's merge history (gh api repos/<org>/<repo>/commits?path=...) — a merge date matching the regression onset both confirms the attribution and identifies the owning team.
    • Never write "X-only" (techpreview-only, platform-only, arch-only) about a test or tool in the report or a bug unless the source code gating has been read and confirms it.
  5. Onset and suspect PRs (when the bucket has a crisp start date): follow the "Determine Regression Start Date" and "Identify Suspect PRs in Payload" procedures from /ci:analyze-regression (first failing run → payload tag via fetch-prowjob-json → fetch-new-prs-in-payload → up to 5 candidate PRs vetted with gh). A LIKELY PR both strengthens the bucket and tells you the owning component/repo.

  6. Cross-check globally: fetch-test-report skill (with --no-collapse) for the bucket's main test — confirms whether the issue is variant-specific or global, and surfaces open_bugs that may already cover the bucket.

  7. Check Slack context (optional — only when Slack access is available): The TRT/release-oversight team discusses ongoing payload and CI issues in #forum-ocp-release-oversight (https://redhat.enterprise.slack.com/archives/C01CQA76KMX). Search/read the last 14 days of messages there for the bucket's signature (test name, error message, operator, platform, payload tag) — known payload-wide events, infra outages, and in-flight fixes are usually discussed there before triages/bugs exist, and a thread often names the owning team or an existing OCPBUGS ticket. If the agent has no Slack access (no Slack tooling/credentials), omit this step entirely — do not block or ask for access.

Mechanism self-review (mandatory before finalizing any bucket). Every rule above is an instance of one behavior: challenge your own conclusion before filing it. After drafting a bucket's root cause, run this adversarial pass on the mechanism claim itself:

  1. Enumerate the load-bearing assumptions in the claim — every "because", "not yet", "race", "missing", "computes wrong", "X-only". For each, cite a specific, checkable evidence locator: an artifact path plus the quoted line, a source-code location, a job-run ID, a structured API/JSON field, a condition's LastTransitionTime, or the query that returned the result. A line number is one acceptable form, not a requirement — much of the evidence here (run metadata, API fields, object IDs, timestamps) has no stable line to cite, and inventing one is worse than quoting the field. What matters is that the next reader can go to the same place and see the same thing. An assumption with no such locator is downgraded in the filed text to an explicit unverified hypothesis ("unverified: possibly …") — never stated as fact.
  2. Check the theory's predictions against everything observed, not just the failing line: does it predict the observed frequency (deterministic theory vs. n=1), the measured window duration, the recovery asymmetry (who healed vs. who stayed broken), and the variant/platform spread? Any mismatch means the theory is wrong or incomplete — resolve it before filing.
  3. Argue the strongest alternative. Spend one honest paragraph on the best competing mechanism (warmup instead of ordering, upstream producer instead of the error-emitter, boundary condition instead of corruption, unrecovered state instead of the transient trigger) and state the artifact evidence that discriminates between them. If nothing discriminates, the confidence is not HIGH.

A bucket whose mechanism claim fails any of these checks is not finalizable — keep digging or downgrade honestly.

After deep-dive, finalize buckets. Depth is mandatory, not optional: no bucket may be finalized at LOW or MEDIUM confidence, and the duty run must not end with open "action items" like "needs artifact deep-dive" or "spot-check installer logs first". When one regression's failed runs split into multiple distinct failure signatures, every signature must be root-caused independently — each may get its own triage record (a regression can carry several), and no run may be left labeled "ambiguous"/"unclear" in a bucket claimed at HIGH confidence: an unexplained run either gets dug into until it joins a signature, or the bucket's confidence is honestly downgraded and the digging continues. Never bundle an unexplained sub-pattern into another pattern's triage "pragmatically" or "since the regression is open anyway" — triage covers only the runs whose root cause it actually explains; attaching unexplained runs to it hides them from the next duty shift and mis-scopes the bug. If a sub-pattern remains unexplained, the whole bucket is not finalizable — keep digging. If confidence is not HIGH after the steps above, keep digging until it is — escalate through the evidence ladder yourself: raw failure outputs → job-run summaries → GCS artifacts (prow-job-artifact-search, install/test-failure analysis skills) → audit logs (grep for the failing object/namespace to identify the creating user and userAgent — this reliably resolves "who created this pod/namespace" questions) → junit timing correlation (what else ran in the same window) → suspect-PR vetting. Triage taking longer is acceptable; leaving an unexplained bucket is not. The only permitted low-confidence outcome is when the evidence is genuinely exhausted (artifacts expired, logs missing), and then the report must say exactly what was checked and what was missing.

Additional finalization rules (each has caused a wrong disposition in a real duty run):

  • "Resolved"/"stopped"/"no recurrence" claims require the full run list, not a sample. Before classifying a signature as resolved or transient, enumerate all of the regression's job_runs by date and confirm the newest runs' signature. A signature absent from a 3–4 run sample of a 20-run regression is not evidence it stopped. And when a sub-test starts "passing" recently, verify recent runs actually reach that stage: a bootstrap failure cannot "self-resolve" while newer runs of the same job die earlier at infrastructure provisioning — the earlier failure masks the later stage, it does not fix it.
  • "Leave untriaged" is not a permitted disposition for a bucket with an identified root cause and owner. If the deep-dive named the mechanism and the responsible component/suite, the bucket gets a triage (to an existing or new issue; under --existing-only, a new-issue bucket instead gets a "Proposed new bugs (not filed)" entry) — "collateral of noisy runs" is only a valid leftover justification when the failure has no independent mechanism (pure co-occurrence). A failure that is deterministically produced by another test's behavior has an independent mechanism and must be triaged as test.
  • Extend, don't duplicate. When an existing triage record already covers the bucket's bug, extend that triage with the new regression IDs (--triage-id); do not create a second triage record pointing at the same JIRA.
  • Subagent outputs must be verified, not trusted. If bucket deep-dives are delegated (subagents, parallel tasks), the orchestrator must check each returned bucket against this section's requirements before accepting it — in particular that the mandatory quotes are present (recovered condition messages for vague wrappers, mutator identification for rollout flaps, transition counts for stability failures, *_previous.log reads for crash-looping containers). A missing mandatory quote means the sub-analysis is incomplete and must be redone, regardless of how confident its prose sounds. Batch size is never a reason to skip mandatory steps.
  • Never end your turn to "wait" for dispatched subagents — the session ends the moment you stop. There is no background execution across turns: an assistant message that ends with a status narration ("waiting for X analysis to complete") and no tool call terminates the run, and in CI the harness will tear the process down at that point. Collect every subagent's result within the turn that needs it, and treat the duty report as a hard checkpoint: if the run were killed right after your current message, the report file must already exist on disk — write intermediate versions early and update them, rather than deferring all writing to a final step that may never come. (Case 5 in case-notes.md: $12 of analysis, no report.)

Each bucket must have:

  • Member regression IDs (re-check the untriaged list — new siblings may have opened during analysis)
  • Root cause summary (one paragraph) and failure classification (permafail / flaky / resolved / recent)
  • Owning component (may differ from the Sippy component — state both)
  • Triage type: product / test / ci-infra / product-infra
  • Disposition: existing triage to extend / existing JIRA to create a triage for / new JIRA needed / no action (resolved or pure infra noise — say so explicitly and leave untriaged only with justification)
Phase 4: Search for existing bugs, then triage each bucket

For each bucket, before filing anything new:

  1. Check bugs linked to the bucket's fired labels, then triaged_matches from fetch-related-triages (confidence ≥5 with an open JIRA is the default target) — but never act on a match, even conf=10, without the Phase 3 per-bucket CI-sample verification confirming the root cause actually matches.

  2. Check open_bugs from the test report.

  3. Search Jira for the root-cause signature (error message, operator name, component-regression label) in OCPBUGS against the owning component — the right bug may exist under Monitoring/etcd/MCO even though the regression sits under Installer.

  4. Component-scoped JIRA listing (mandatory — keyword search is not enough). Owning teams describe defects in developer vocabulary that shares no keywords with the CI-side symptom. After determining the owning component, list its recent bugs regardless of keywords and read the summaries:

    project = OCPBUGS AND component = "<owning component>" AND created >= -21d ORDER BY created DESC

    A bug whose creation date falls inside the bucket's failure window, on the owning component, is a duplicate candidate even with zero keyword overlap — open it and compare mechanisms.

  5. Owning-repo merge-history check (mandatory when onset or cessation is dated). Query the owning repo for PRs merged around the bucket's onset and cessation dates:

    bash
    gh pr list --repo <org>/<repo> --state merged --search "merged:<window>" --json number,title,mergedAt

    A merge at the cessation boundary is likely the fix — its OCPBUGS-* title prefix names the existing bug: triage to that bug instead of filing a new one. A merge at the onset boundary is a suspect trigger. The phrase "resolved by unidentified payload change" is banned from reports and bugs unless this check was run and came back empty. For currently-live breakage, also scan the owning repo's newest merges for revert PRs: a fresh Revert "..." title citing a TRT/OCPBUGS key hands you the trigger PR, the tracking ticket, and the expected recovery time in one query.

Then act (this is where --auto-triage applies; without it, confirm each bucket with the user):

  • Extend existing triage: triage-regression skill with --triage-id (additive merge is automatic; pass only the new IDs).
  • New triage to existing bug: triage-regression skill with --url, --type, and a one-sentence --description (<120 chars).
  • New bug (under --existing-only: do not file or triage; add a Phase 5 proposal instead): file with /jira:create bug (the create skill from the jira plugin) against the owning component, label component-regression, description per the bug-filing template in /ci:analyze-regression ("Prepare Bug Filing Recommendations" section: full test names in {code} blocks, test IDs, regression IDs, variants, error signature, Sippy test-details UI links for every member regression, suspect PRs). Every JIRA issue or comment created by this workflow must end with an AI-attribution footer as a separate, visually marked block — not a sentence buried in the text: place it after a divider, as its own paragraph or note panel, e.g. a rule followed by a panel (type note) in ADF containing "AI-generated content: This bug was filed by AI as part of Component Readiness triage duty. Please verify before acting on it." Release blocker is conditional on the triage type and impact, not automatic: mark the bug a release blocker (set-release-blocker skill) only for product bugs whose failures block or materially degrade blocking/informing payload jobs; test bugs (races, invariant-scan interference) and ci-infra issues (cloud capacity, registry outages) are not release blockers — state the blocker decision and its one-line justification in the report. Then create the triage record.
  • Suggested fixes carry a higher evidentiary bar than attribution — and triage duty never opens fix PRs. A "suggested fix" in a filed bug is optional; when included it requires (a) the mechanism timeline-proven per Phase 3 (timestamps, window duration, alternatives refuted), and (b) a target that matches what the measured window actually proves. A seconds-long trigger window does not by itself justify changing steady-state behaviour for everyone — but it very often does justify a permanent resilience change (retry, re-discovery, gating, idempotent recovery) in whatever failed to survive that window. Aim the fix at the thing that stayed broken, not at the thing that blinked. Never propose weakening a safety mechanism (failurePolicy, validation, admission gating, health checks) as the primary fix; if failing open seems attractive, flag it as a question for the owning team instead. If the analysis surfaces a secondary resilience defect (e.g., a controller that cannot recover from one failed patch), do not demote it to a "defense-in-depth" footnote — re-run the differential-recovery check: it is often the actual primary bug. Opening code-change PRs is out of scope for this workflow entirely.
  • Always finish a triage by running the add-jira-triage-link skill to put the triage URL into the JIRA description. The skill appends to the description — if the description ends with the AI-attribution footer, the appended link would land after it. After running the skill, verify the footer is still the final block; if not, move the footer back to the end of the description (or insert the triage link before the footer in the same update).
  • If a label that fired on the bucket's runs does not list the triaged bug, propose manage-labels update --bugs <existing>,<new> (replaces the list; needs user confirmation, not covered by --auto-triage).
  • While triaging each bucket, note whether its signature is symptom-worthy (crisp grep-able line in a durable artifact) — the Phase 5 report must carry a symptom proposal for every bucket where it is (see "Label the bucket's signature in Sippy").

With --auto-triage, only act autonomously when confidence is high: consistent error signature across the bucket, and either a confidence ≥5 triaged match or an unambiguous existing open bug. Buckets requiring a new bug, or with mixed signals, are always presented for confirmation.

Show full SKILL.md (3,458 more words)Show less
Existing-only mode (--existing-only)

This mode records matches to existing bugs but never files new JIRA issues. Analysis is unchanged. These rules apply to every write in any phase, including the Phase 1 stale-triage sweep:

  • Allowed writes: only for a verified match to an existing issue:

    • extend a triage, or create one for the existing issue;
    • add-jira-triage-link;
    • comment on that issue.

    Everything else, including creating, reopening, or transitioning issues, goes in the report as a recommendation.

  • Otherwise, propose. New-bug buckets, failed fixes, and uncertain matches get a draft under "Proposed new bugs (not filed)" in the Phase 5 report. Their regressions stay untriaged. Never lower the match bar to avoid a proposal: a wrong triage hides the real defect.

  • Subagents do analysis only; the orchestrator does every write.

Phase 5: Duty report

Present a final report:

  1. Inventory: N untriaged regressions found → M buckets.
  2. Per bucket: member regression IDs, root cause, owning component (vs. Sippy component), classification, evidence highlights (error signature, stage, representative run links), action taken (triage ID + JIRA link) or recommendation awaiting confirmation.
  3. Leftovers: regressions deliberately left untriaged (resolved / one-off flake / inconclusive) with justification and what evidence would change the call.
  4. Proposed Sippy symptoms (mandatory section, even if empty): for every bucket whose root cause has a crisp, grep-able single-line signature in a durable artifact, include a concrete symptom proposal per the "Label the bucket's signature in Sippy" section below — label name, bugs, matcher, file pattern, validation pair, retro-apply run set. Proposing costs nothing (creation still requires explicit user confirmation); a duty shift that root-caused a bucket and did not propose a symptom for an obviously grep-able signature has left cheap future-triage value on the table. If no bucket qualifies, say so and why (e.g., signature only visible via timing correlation, artifact expires, no unique string).
  5. Cross-cutting observations: payload-wide events, infra instability windows, techpreview-only patterns — useful context for the next duty shift.
  6. Session usage (when available): if the run is orchestrated by a harness that captures usage telemetry (e.g. the CI job appends a "Session usage" section with model, turns, token counts, and cost after the session ends), do not fabricate these numbers yourself — the model cannot observe its own final token totals mid-session. In interactive runs simply omit the section.
  7. Proposed new bugs (not filed) (--existing-only only, mandatory even if empty): one draft per proposal, following the Phase 4 new-bug template, with member regression IDs and the release-blocker recommendation. A human must be able to file it as written.
Label the bucket's signature in Sippy (propose always, create on confirmation)

When a bucket has a crisp, grep-able artifact signature (an error string or log line unique to the root cause), define a Sippy symptom so future runs are labeled automatically and the signature can be applied retroactively. Symptoms have repeatedly paid for themselves in real duty runs — a validated matcher has adjudicated a disputed root-cause attribution and machine-verified a 19-vs-1 bucket split (Case 6 in case-notes.md). Good candidates: recurring infra events that spawn regressions across many job families, single-occurrence product signatures worth a tripwire (n=1 bugs), and any signature that a "known recurring family" keeps regenerating. These are writes: always present the proposed label/symptom definitions and the target run list to the user and get explicit confirmation first — --auto-triage does not authorize them.

  1. Check for existing coverage before creating anything. List the current catalog — this read is public and needs no token (GET https://sippy.dptools.openshift.org/api/jobs/labels, GET .../api/jobs/symptoms, or the list-symptoms skill) — and search it for the bucket's signature: the label/symptom may already exist from the shift that first triaged the bug (Case 6: a proposed console-crash symptom fully duplicated an existing one; the duplicate-key 400 from the label POST was what revealed it). Also search by the bug key (list-symptoms --labels --search OCPBUGS-NNN); a label linked to the bug counts as coverage only if its matcher targets this bucket's signature. A dry-run reevaluate of one bucket run is an equivalent cheap probe when a DPCR token is available: if an existing symptom targeting this signature fires on the run, the bucket is already covered and nothing should be created.

  2. Create a label and symptom via the authenticated API (DPCR Bearer token, same as triage writes): POST https://sippy-auth.dptools.openshift.org/api/jobs/labels ({"label_title": ..., "explanation": ..., "bugs": ["<JIRA key>"]} — the title becomes the ID), then POST .../api/jobs/symptoms ({"summary": "<descriptive name>", "matcher_type": "string|regex|file", "file_pattern": "<glob under the job-run GCS root>", "match_string": ..., "label_ids": [...]}). The symptom ID is derived from the summary; gzipped artifacts are decompressed transparently. Keep Jira keys in the label's bugs array rather than embedding them in label explanations or symptom summaries. To change a symptom, use its full-replacement PUT endpoint.

  3. Beware benign-transient false positives — presence is not pathology. Before finalizing a matcher, grep the target file in a healthy run: many error lines appear benignly during normal startup — sometimes once, sometimes in bulk (Case 6: a proposed match string appeared 120× in a healthy install's etcd-operator log as a normal warm-up transient; the failed run differed only in that the line persisted to log-end). String/regex matchers cannot count occurrences or check persistence, so if the line can appear benignly at all, re-anchor to an artifact that only exists — or only contains the string — in the failure mode: a container's *_previous.log (only present after a crash/restart) instead of its current log, an events.json string that healthy runs never emit, or best of all the gather-time end state in clusteroperators.json (e.g., StaticPodsAvailable: 0 nodes are active appears there only when etcd truly never deployed — zero hits in healthy runs by construction, because gather runs after the failure window).

  4. Validate with a dry run against one known-positive and one known-negative run — the best negative is a near miss: a failed run of the same job family with a different root cause (it catches over-broad matchers that a green run would not). POST .../api/jobs/runs/reevaluate with {"prow_job_build_ids": [...], "dry_run": true}. Proceed only if both outcomes are correct; otherwise refine the matcher and re-validate. Note: reevaluation runs all symptoms, so unrelated pre-existing labels may legitimately appear on your negative — check only that your symptom's matching is correct.

  5. Then apply with "dry_run": false to the bucket's prowjob_run_ids. The API caps requests at 50 IDs, but reevaluation scans GCS artifacts server-side and large batches time out at the gateway (504) — use batches of 5–10 with a retry, check each result's status, and stop and report on the first failed batch or any non-success result rather than continuing.

  6. Match the underlying error, not a transient side effect: a matcher keyed on a crash/panic goes silently false-negative the moment a partial fix removes the crash while the defect persists. If the primary evidence lives in pod logs that sometimes fail to gather (dying clusters), add a fallback symptom against an artifact that survives (e.g., pods.json state/reason strings) mapped to the same label. When the same signature can surface under different step names (e.g., devscripts-driven vs plain IPI installs), create one symptom per file pattern, all mapped to the same label.

Closed-set audit mode (--audit-closed) — justify why every closed regression closed

bulk-triage-regressions <view> --components ... --audit-closed

An explicitly-requested, read-only mode that answers a different question than triage duty: not "who owns this failure?" but "why did this regression close, and can we prove it?" Every closed untriaged regression must end the audit in exactly one of these justification classes:

ClassMeaningProof required
fixedA fix merged; closure follows itFix PR/bug with merge/resolution date at or before the close boundary
event-endedAn infra/payload event endedEvent window bracketing the failures (bad payload, repo outage, cloud capacity, credentials)
intermittent-cause-openPass rate recovered but the defect is still openSignature matched to an open bug; state explicitly that closure ≠ resolution
collateralClosed with the window of a tracked sibling eventSignature or run overlap with the tracked event
evidence-expiredRun history and artifacts are goneList exactly what was checked (regression details, GCS) and the statistical argument (age, sibling consistency) — flagged, never silently absorbed

No regression may be left as "unknown". evidence-expired is the only permitted terminal state without a mechanism, and it must be flagged in its own report section.

Scale technique (a closed set is 5–30x a duty batch — per-regression deep-dives do not scale)
  1. Inventory and cluster first: fetch details for all members (parallel, ~8 workers is safe for the details endpoint), then cluster by (test-family, platform, featureset, close-window). Expect heavy super-clusters — always look for same-day open/close waves across platforms before analyzing anything individually (Case 7 in case-notes.md: one bad nightly payload explained 21% of an entire closed set).
  2. Bypass the Sippy runs API for bulk signature extraction — it rate-limits (HTTP 429) far below audit volume, and backoff does not help at this scale. Go to GCS directly (unthrottled, parallel-safe):
    • Install-family tests: classify from the installer log (artifacts/*/ipi-install-install*/build-log.txt, devscripts equivalent on metal) with an ordered signature catalog (most-specific first). Seed the catalog from the view's known events and extend it with whatever the duty history has verified (ign-push 403s, cipherSuites/rev-0, toolchain download, quay pulls, registry.ci 5xx, quota, capacity, provisioning, per-operator stabilization). Guard against substring traps (Case 7: lease matching inside release mis-binned 31 regressions).
    • Everything else: parse the junit XML under the e2e step (artifacts/*/<e2e-step>/artifacts/junit/junit_e2e*.xml) with a real XML parser (regex over XML mis-matched ~95% of testcases in practice) and take the <failure> of the exact testcase.
    • Sample the 2 newest failed runs per regression; extend only where signatures disagree.
  3. Match against the full triage catalog: fetch all existing triage records once (/api/component_readiness/triages) and index their regressions by test name — a closed untriaged record is very often the untriaged sibling of a triaged one, which supplies the bug and the closure mechanism for free.
  4. Date-anchor every closure claim: fixed requires the fix date ≤ close boundary (JIRA resolutiondate, gh pr list --search "merged:<window>" on the owning repo); event-ended requires the last failing run to fall inside the event window. The Phase 4 component-scoped JIRA listing and merge-history checks apply here unchanged.
  5. Checkpoint long-running collection to disk and run it detached (nohup + periodic JSON dumps): bulk GCS extraction takes tens of minutes, harness timeouts will kill foreground loops, and two concurrent writers on one results file destroy each other — one writer, atomic-ish checkpoints, verify counts after every stage.
Audit report structure
  1. Executive summary: category table (name, member count, one-line verdict).
  2. Method: exactly which artifact paths / APIs were used and why.
  3. Per-category narrative: mechanism, proof, closure explanation, and for intermittent-cause-open the explicit warning that siblings will reopen.
  4. Per-regression appendix — one row per member, no exceptions (ID, test, platform/featureset, opened→closed, category, evidence sample).
  5. Honest-limitations section listing every evidence-expired member.
Audit-specific pitfalls
  • Sippy retention creates a hard evidence horizon: regressions from the view's first tracking week may have no run history at all, and GCS artifacts expire (~3 months). Do not let the horizon silently shrink the audit — count and flag such members.
  • The view-baseline start date masquerades as a mass event: many unrelated regressions "open" on the view's first tracking day. Check whether a suspicious same-day open wave is simply the earliest date in the dataset before hunting for a common cause.
  • "Closed" is not a verdict: a closed regression whose signature maps to a still-open bug (intermittent-cause-open) is a prediction of future regressions — surface these in the summary so the next duty shift expects them.

Pitfalls (learned from real duty runs)

Principles distilled from real mis-dispositions; full narratives in case-notes.md.

  • Left = newest in pass_sequence strings. Misreading direction inverts "regressed" vs "resolved".
  • Do not file bugs against Installer by default. In practice a majority of install should succeed regressions in duty batches were owned by other components or were infra noise. The installer is the messenger.
  • Unknown component regressions (e.g., verify the cluster readiness and stability, verify all machines should be in Running state) are wrappers; the co-failing tests and operator states identify the owner.
  • One bucket can span components: Monitoring + Test Framework + Unknown + Installer regressions have all belonged to a single MCO bug. Don't let the component column fragment a bucket.
  • Techpreview variants often fail for techpreview-only reasons (new feature gates); check whether the same job without techpreview passes before assuming a general regression.
  • A regression's variant slice is not the offending code's gating. "Techpreview-only"/"platform-only" claims about a test or tool must come from its source-code skip/gate conditions (Phase 3 item 4), never from the variants Sippy happened to flag.
  • API vs UI URLs: convert test_details_url to the sippy-ng UI form before putting it in bugs or reports, by replacing the base https://sippy.dptools.openshift.org/api/component_readiness/test_details with https://sippy-auth.dptools.openshift.org/sippy-ng/component_readiness/test_details (query parameters are identical).
  • Re-list before writing: new regressions open continuously; refresh the untriaged list right before creating/updating triages so siblings opened mid-analysis are included.
  • Check closed/dropped siblings too. A regression that looks novel is often a new open instance of a root cause whose earlier sibling was already triaged and has since dropped out of the active view (its regression closed). fetch-related-triages and a --test-name query without an open-only filter find these; reuse the existing triage/JIRA instead of filing a new bug.
  • Identical prowjob_run_id sets are the strongest clustering signal — but still not proof. When two regressions were opened from literally the same job runs, treat them as one draft bucket, then confirm in Phase 3 that the failure outputs actually point at one cause: a single run can carry independent defects, and in mass-failure runs co-occurrence is largely coincidental. Only merge into one triage after the error signatures/artifacts agree.
  • The failing monitor/test is often just the messenger (see the messenger-vs-owner principle in Phase 3). Always read the actual error text, including setup/preparation errors, before trusting the test's subject area. (Case 9: "Networking" connectivity tests failing on etcdserver: request timed out during test setup — an etcd bug.)
  • "Operator not available" during a stability window is not automatically that operator's bug. Before filing against a flapping operator, find the mutation that triggered the reconcile — if the job's harness caused it, the bucket is test. (Case 2.)
  • A vague wrapper message hides a specific controller error — extract it, don't guess. When the junit output has an empty reason, grep the build log/monitor intervals for the condition-change events and their messages before attributing; never attribute from co-failing tests. (Case 8: bare clusterversion not available: False drafted as "UDN breakage" when the interval message named an unrelated CVO controller error.)
  • Bootstrap-transient theories that contradict the end-state. Pods running proves nothing about cluster etcd during bootstrap, and a crash-looping component whose error names a missing upstream dependency is a victim, not an owner.
  • A validated Sippy symptom detects a signature, not a cause. When a symptom's signature can be produced downstream of a different defect, its label text must say so, and a second symptom keyed on the upstream discriminator should exist to adjudicate.
  • Flapping ≠ slow rollout, and a small sample ≠ cessation. Count condition transitions and enumerate all run dates before choosing between flaky-transient and permafail. (Case 1.)
  • Identifying a root cause and then leaving the regression untriaged is a contradiction. Deterministic test interference is an independent mechanism: it gets a test triage and a bug, regardless of how noisy the surrounding runs are.
  • No triage record ≠ no bug, and keyword search ≠ component search. Run the component-scoped JIRA listing and the owning-repo merge-history check (Phase 4 items 4–5) before every new-bug disposition, and treat an unexplained cessation as a strong hint that a fix already landed.
  • A negative grep for one error string does not make a signature "new" — run the symptom catalog before declaring a distinct mechanism. Known-cause bugs surface through several messenger strings; the armed catalog encodes the reliable ones. Dry-run reevaluate a representative run and treat every label hit as a known-cause hypothesis to check. (Case 3: a 5,471-line log flood flagged as new was a downstream consequence of a cataloged defect.)
  • Known recurring families: some failure signatures come back shift after shift and usually already have a triage — e.g., etcd slowness / slow fdatasync on Azure masters, quay.io 502s / ImagePullNeverCompletes on metal jobs (ci-infra), transient cloud quota or DNS provisioning errors. Search existing triages for the signature before opening anything.
  • Similar symptom ≠ same bucket. Two image-pull triages can coexist for different causes (e.g., a quay ci-infra outage vs. an MCO bootimage product bug). Match on the full signature — error text, platform, job family, timing — not just the headline symptom, and pick the triage whose root cause matches, not the first one found.
  • Triage type follows the root cause, not the component: a flaky test with an external dependency is test even if it looks like a product failure; an etcd timeout that also hits customers is product; a registry outage is ci-infra even when it kills installs.
  • A closed/MODIFIED bug can still be the right triage target when the failures predate the fix landing; check the fix-merge date against the newest failed run before dismissing it — but if failures continue after the fix, that's a failed fix (analysis_status -1000) and needs the bug reopened or a new one.
  • Sample runs only from the regression's own job_runs list. Broad Sippy job-filter URLs ("all jobs failing test X") sweep in unrelated jobs and post-fix eras. Every run you cite must be a member of the regression (and predate any candidate fix). (Case 10: "3 random jobs" from a broad filter produced three confidently wrong root causes in one shift.)
  • A CI step pod that never started is a build-cluster problem, not a product one. If the prow build log shows Deleting pod <step> that failed to start or pod pending for more than 1h: pod has not been scheduled (with 0/N nodes are available: ... untolerated taints / unschedulable), the OpenShift installer never ran — no product artifact exists to analyze. Classify as ci-infra, and name the build cluster from prowjob.json (.spec.cluster, e.g. build10) so the infra team knows where the capacity problem is. Trivial setup steps (rbac, hosted-loki) failing this way across a job family for days is a capacity outage worth its own tracking ticket.
  • Unknown conditions are not resource pressure. When a kubelet stops posting status, all node conditions flip to Unknown, and the MCO controller logs Reporting unready: ... OutOfDisk=Unknown — this is a stale-status marker, not disk exhaustion. Only DiskPressure/MemoryPressure=True (or eviction/OOM events, df output, sosreport sar) is evidence of actual resource pressure.
  • Superficially different per-run mechanisms can share one upstream cause — and two defects can chain. When run-level root causes look "completely different", look one level up (what leaked, what was stuck, what accumulated) before declaring them unrelated; and when a symptom vanishes without its component's fix merging, suspect an interacting defect that got fixed instead. (Case 11: maxPods saturation, memory livelock, and an OVS flow storm — all downstream of one namespace-deletion blockage.)

Arguments

  • <view>: Component Readiness view name (e.g., 5.0-main). Required.
  • --components: Space-separated component name filters, case-insensitive and hierarchy-aware (e.g., Installer also matches Installer / openshift-installer; Networking matches Networking / ovn-kubernetes, Networking / router, ...). Required in practice for duty scoping.
  • --auto-triage: Allow high-confidence buckets to be triaged without per-bucket confirmation. New bug filing always requires confirmation.
  • --existing-only: Record matches to existing triages/JIRA issues but never file new JIRA issues; new-bug buckets become report proposals (see "Existing-only mode" in Phase 4).
  • --audit-closed: Read-only closed-set audit mode (see "Closed-set audit mode") — justifies why every closed untriaged regression closed. Mutually exclusive with normal duty triage and with --auto-triage; performs no writes.

See Also

  • Related Command: /ci:analyze-regression — single-regression deep dive; this skill orchestrates its techniques across a batch
  • Related Skill: create (jira plugin) — file new JIRA bugs via /jira:create bug (plugins/jira/skills/create/SKILL.md)
  • Related Skill: list-regressions (teams plugin) — batch listing (plugins/teams/skills/list-regressions/SKILL.md)
  • Related Skill: fetch-regression-details (plugins/ci/skills/fetch-regression-details/SKILL.md)
  • Related Skill: fetch-related-triages (plugins/ci/skills/fetch-related-triages/SKILL.md)
  • Related Skill: fetch-test-runs (plugins/ci/skills/fetch-test-runs/SKILL.md)
  • Related Skill: fetch-job-run-summary (plugins/ci/skills/fetch-job-run-summary/SKILL.md)
  • Related Skill: prow-job-analysis — GCS artifact analysis for install/bootstrap failures (plugins/ci/skills/prow-job-analysis/SKILL.md)
  • Related Skill: fetch-test-report (plugins/ci/skills/fetch-test-report/SKILL.md)
  • Related Skill: triage-regression (plugins/ci/skills/triage-regression/SKILL.md)
  • Related Skill: add-jira-triage-link (plugins/ci/skills/add-jira-triage-link/SKILL.md)
  • Related Skill: set-release-blocker (plugins/ci/skills/set-release-blocker/SKILL.md)
  • Related Skill: oc-auth (plugins/ci/skills/oc-auth/SKILL.md)
  • TRT Documentation: https://docs.ci.openshift.org/docs/release-oversight/troubleshooting-failures/

Maintaining this skill (for authors, not duty runs)

Evals before narratives. When a new duty-run mis-disposition surfaces, do not reach for a new case narrative or case-specific rule first. Instead: (1) capture the incident as an eval case under plugins/ci/evals/cases/bulk-triage-regressions/ (frozen regression snapshot + owning-team-verified ground truth + the wrong conclusion as a banned outcome); (2) run the eval against the unchanged skill — if the existing principles and the mechanism self-review already catch it, nothing needs adding; (3) only if the eval fails repeatedly does the incident earn a new generic rule (preferred) or narrative, and the eval then proves the addition works and guards it against regression. This is how the two source incidents of this rule set (the VAP/actuator and GID-mapping mis-diagnoses) were handled: a baseline-vs-variant eval matrix showed the generic rules catch both, so their narratives were dropped and their ground truth lives in eval cases 001/002. Every prompt token here costs money on every duty run — additions must pay for themselves under eval-bulk-triage-regressions.yaml.

© openshift-eng, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/ci/skills/bulk-triage-regressions of openshift-eng/ai-helpers.

  • SKILL.md
  • references/case-notes.md

Open the folder on GitHubat commit a627176

Compare with similar skills

Bulk Triage Regressions next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bulk Triage Regressions compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bulk Triage Regressions this skillopenshift-eng/ai-helpers120—~17kAutomated safety check: PassApache-2.0
OpenLogi macOS Permissions TriageAprilNEA/OpenLogi23k—~2.5kAutomated safety check: NotesApache-2.0
Bug Finder for daisyUIsaadeghi/daisyui43k—~2.3kAutomated safety check: PassMIT
Root Cause Debugginggarrytan/gstack136k—~1.4kAutomated safety check: PassMIT
Review PRapache/shardingsphere21k—~6.5kAutomated safety check: PassApache-2.0
Graph-Based Bug Tracingtirth8205/code-review-graph32k1 repos~287Automated safety check: PassMIT

Similar skills

  • Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.

    23k GitHub stars~2.5k tokensUpdated today
    DevelopmentAuto-check: notes
  • Bug Finder for daisyUI

    saadeghi/daisyui

    Investigates suspected bugs in the daisyUI monorepo through read-only analysis, then writes a decision-ready fix plan in tmp/bugs without changing any product code.

    43k GitHub stars~2.3k tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Root Cause Debugging

    garrytan/gstack

    Investigates bugs, errors and stack traces in phases and requires a root-cause hypothesis to be confirmed before any fix is written.

    136k GitHub stars~1.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Review PR

    apache/shardingsphere

    Review Apache ShardingSphere or user-authorized downstream pull requests and PR discussions from public or authorized repository evidence.

    21k GitHub stars~6.5k tokensUpdated today
    DevelopmentAuto-check passed
  • Graph-Based Bug Tracing

    tirth8205/code-review-graph

    Traces a bug through a code knowledge graph, following callers, callees and execution flow before opening source files, within a small token budget.

    32k GitHub starsUsed in 1 repo~287 tokens
    DevelopmentAuto-check passed
  • Om Auto Fix Issue

    go-musicfox/go-musicfox

    Fix or implement a tracker issue end to end from a single command — takes an issue id or a plain problem description (filed first via om-prepare-issue), classifies, then drives the bug autofix chain…

    2.6k GitHub starsUsed in 1 repo~5k tokens
    DevelopmentAuto-check: notes

More from openshift-eng/ai-helpers

All 118 skills in this repo
  • Investigate CI Reliability

    openshift-eng/ai-helpers

    Find and independently validate actionable reliability defects across OpenShift release jobs and presubmits, then export portable issue handoffs.

    120 GitHub stars~1.9k tokensUpdated 4 days ago
    Auto-check passed
  • Address Review PR

    openshift-eng/ai-helpers

    Fetch and address all PR review comments — categorize by priority, make code changes, post replies, and push.

    120 GitHub stars~2.9k tokensUpdated 4 days ago
    Auto-check passed
  • Categorize Activity Types

    openshift-eng/ai-helpers

    Categorize Jira issues into Red Hat Sankey Activity Type categories using MCP Jira tools.

    120 GitHub stars~2.4k tokensUpdated 4 days ago
    Auto-check passed
  • Has Review Work

    openshift-eng/ai-helpers

    Decide whether a GitHub PR has unanswered authorized review comments or new required CI failures worth a follow-up agent.

    120 GitHub stars~1.9k tokensUpdated 4 days ago
    Auto-check passed
  • Must Gather Analyzer

    openshift-eng/ai-helpers

    Analyze OpenShift must-gather diagnostic data including cluster operators, pods, nodes, and network components.

    120 GitHub stars~2.3k tokensUpdated 4 days ago
    Auto-check passed
  • Payload Autodl JSON

    openshift-eng/ai-helpers

    Schema for the autodl JSON data file produced by payload-analysis for database ingestion — you must use this skill whenever generating the autodl JSON file

    120 GitHub stars~2.6k tokensUpdated 4 days ago
    Auto-check passed

Categories

Questions about Bulk Triage Regressions

What does Bulk Triage Regressions do?

A skill your agent uses for Component Readiness triage duty - holistically analyze and triage all untriaged regressions for a set of components in a view, clustering them into root-cause buckets. Bulk Triage Regressions is an agent skill from openshift-eng/ai-helpers.

When should I use Bulk Triage Regressions?

Bulk Triage Regressions fits situations like: component Readiness triage duty - holistically analyze and triage all untriaged regressions for a set of components in a view; clustering them into root-cause buckets.

How do I install Bulk Triage Regressions in Claude Code?

Run `npx skills add openshift-eng/ai-helpers --skill bulk-triage-regressions -a claude-code`. Or copy the skill folder (plugins/ci/skills/bulk-triage-regressions in openshift-eng/ai-helpers) into .claude/skills/bulk-triage-regressions in your project. Claude Code loads it when a task matches its description.

How do I install Bulk Triage Regressions in Codex?

Run `npx skills add openshift-eng/ai-helpers --skill bulk-triage-regressions -a codex`. Or copy the skill folder (plugins/ci/skills/bulk-triage-regressions in openshift-eng/ai-helpers) into .agents/skills/bulk-triage-regressions in your project. Codex loads it when a task matches its description.

Can I use Bulk Triage Regressions in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openshift-eng/ai-helpers --skill bulk-triage-regressions -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bulk-triage-regressions, .gemini/skills/bulk-triage-regressions, .github/skills/bulk-triage-regressions and .opencode/skills/bulk-triage-regressions in your project.

What does Bulk Triage Regressions need to run?

Going by SKILL.md and its folder, Bulk Triage Regressions needs the command-line tools its instructions call (gh and python3) and credentials named JIRA_API_TOKEN. Our summary lists: Python 3; A credential in JIRA_API_TOKEN.

Does Bulk Triage Regressions access the network?

SKILL.md names 6 domains. In commands or code: sippy-auth.dptools.openshift.org, sippy.dptools.openshift.org and redhat.atlassian.net; the agent is likely to contact these when it follows the instructions. As links in the text: id.atlassian.com, redhat.enterprise.slack.com and docs.ci.openshift.org. This is read from the text; nothing was executed.

Is Bulk Triage Regressions safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bulk Triage Regressions use?

Bulk Triage Regressions is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bulk Triage Regressions use?

About 17k tokens (SKILL.md is roughly 69k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.8k tokens, read only when the agent opens those files.

What are the alternatives to Bulk Triage Regressions?

Skills that share tags, products or a category with Bulk Triage Regressions: OpenLogi macOS Permissions Triage (AprilNEA/OpenLogi, 23k stars), Bug Finder for daisyUI (saadeghi/daisyui, 43k stars), Root Cause Debugging (garrytan/gstack, 136k stars) and Review PR (apache/shardingsphere, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bulk Triage Regressions?

openshift-eng (a GitHub organization) maintains it in openshift-eng/ai-helpers, which has 120 GitHub stars. The repository holds 118 skills in this directory. The repository was last updated on October 6, 2026.

Source: openshift-eng/ai-helpers on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.