Agent skill

Prow Job Analysis

by openshift-eng in openshift-eng/ai-helpers

A skill your agent uses when debugging a failed Prow CI job.

Apache-2.0Auto-check passedDevelopment

Install Prow Job Analysis

skills CLI
$ npx skills add openshift-eng/ai-helpers --skill prow-job-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install openshift-eng/ai-helpers prow-job-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ci/skills/prow-job-analysis .claude/skills/prow-job-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
prow-job-analysis
GitHub stars
120
Token cost
~2.8k tokens
SKILL.md length
1,050 words
Files
18 (incl. references)
Skills in repo
118
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when debugging a failed Prow CI job.

  • Works in 5 steps: Parse URL and Extract Metadata → Fetch prowjob.json → Classify Job Type from Name → …
  • Debugging a failed Prow CI job
  • SKILL.md covers Input Format, Prerequisites, Investigation Workflow and Failure Routing Table, plus 3 more sections
  • Runs Python scripts from its folder; calls gcloud; reaches prow.ci.openshift.org and gcsweb-ci.apps.ci.l2s4.p1.openshiftapps.com

What it does

Prow Job Analysis is an agent skill from openshift-eng/ai-helpers. Use this skill when debugging a failed Prow CI job.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including reference files (for example `prow_job_artifact_search.py`, `references/aggregated.md` and `references/artifacts.md`).

It sits in Development. The repository describes itself as: Developer productivity tools for Claude Code & other AI assistants. The licence is Apache-2.0.

When your agent uses it

  • Debugging a failed Prow CI job

Example prompts

  • “/prow-job-analysis”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Parse URL and Extract Metadata
  2. Fetch prowjob.json
  3. Classify Job Type from Name
  4. Download Key Artifacts
  5. Classify Failure and Route to Reference

What it can do on your machine

Read from SKILL.md and the folder at commit a627176. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • gcloud

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • prow.ci.openshift.org
    • gcsweb-ci.apps.ci.l2s4.p1.openshiftapps.com
    • storage.googleapis.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prow Job Analysis loads about 2.8k tokens when it runs, and up to ~114k if it reads all its reference files. Until then it costs about 17 tokens; SKILL.md has 1,050 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~17
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~114k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from openshift-eng/ai-helpers at commit a627176, republished under its Apache-2.0 licence (© openshift-eng). 1,050 words, ~2,826 tokens.

Download SKILL.mdSave it as .claude/skills/prow-job-analysis/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
prow-job-analysis
description
Use this skill when debugging a failed Prow CI job.

Prow Job Analysis

Analyze failures in OpenShift Prow CI jobs. Identify the job type, inspect artifacts, classify the failure, and route to the specialized reference for deep analysis.

Input Format

The user will provide:

  1. Prow job URL (required) — Prow UI or gcsweb URL

    • https://prow.ci.openshift.org/view/gs/test-platform-results-public/logs/<job>/<build_id>
    • https://gcsweb-ci.apps.ci.l2s4.p1.openshiftapps.com/gcs/test-platform-results-public/...
  2. Test name (optional) — specific failed test to focus on

  3. Flags (optional):

    • --backends <list> — focus disruption analysis on specific backends

Prerequisites

  • Python 3.7+: which python3
  • jq: which jq
  • gcloud CLI (recommended, not required): which gcloud — fastest access to the public bucket (no auth required). Without it, every artifact operation works over plain HTTPS: prow_job_artifact_search.py (stdlib-only list/search/fetch) or curl against https://storage.googleapis.com/test-platform-results-public/....

Investigation Workflow

Step 1: Parse URL and Extract Metadata
  1. Find /view/gs/<bucket>/ (Prow UI) or /gcs/<bucket>/ (gcsweb) in the URL. Any non-empty bucket segment is accepted.
  2. Extract the object path after the bucket name, then build_id — pattern (\d{10,}) in the path.
  3. Construct GCS base: gs://{bucket}/{bucket-path}/ using the URL bucket as-is, including prow-artifact-archive. A private bucket will 403.
Step 2: Fetch prowjob.json

Use the fetch-prowjob-json skill to get job metadata. Extract:

  • Job name from .spec.job
  • Target from --target= in ci-operator args
  • Job state from .status.state
  • Refs (org, repo, PR number) from .spec.refs
Step 3: Classify Job Type from Name

Parse the job name to determine the environment and expected failure modes:

Pattern in NameJob TypeKey Implications
upgradeUpgrade jobInstalls first, then upgrades — see upgrade reference
metal, baremetalBare metalUses dev-scripts + Metal3/Ironic — see metal install reference
hypershiftHyperShiftHosted control planes — see hypershift reference
fipsFIPS-enabledWatch for crypto/TLS errors
ipv6, dualstackIPv6/dualstackOften disconnected, uses mirror registry
single-node, snoSingle-nodeResource exhaustion more likely
aggregated- prefixAggregatedStatistical analysis of multiple runs — see aggregated reference
aws, gcp, azureCloud platformPlatform-specific errors — see cloud provider reference
techpreviewTech previewFeature gates enabled, features may be unstable
rhcos9, rhcos10, rhcos9_10, rtRHCOS variant / RT kernelOS variant pinned or heterogeneous; OS-level differences (kernel/systemd/SELinux) — see operating system changes reference
Step 4: Download Key Artifacts
bash
mkdir -p .work/prow-job-analysis/{build_id}/logs

# Build log (always)
gcloud storage cp gs://{bucket}/{bucket-path}/build-log.txt \
  .work/prow-job-analysis/{build_id}/logs/ --no-user-output-enabled

# JUnit XML (always — identifies failed tests/steps)
gcloud storage ls "gs://{bucket}/{bucket-path}/artifacts/**/junit*.xml" 2>/dev/null

# Node journals (always, when the job created a cluster) — required input for the
# Step 5 OS-layer check. Gzip-compressed WITHOUT a .gz extension: zcat/zgrep only.
gcloud storage cp -r \
  "gs://{bucket}/{bucket-path}/artifacts/{target}/gather-extra/artifacts/nodes" \
  .work/prow-job-analysis/{build_id}/ --no-user-output-enabled 2>/dev/null || true
Step 5: Classify Failure and Route to Reference

Examine the build log and JUnit results to classify the failure, then consult the appropriate reference file for detailed analysis procedures.

OS-layer evidence check (mandatory for every job, before routing)

Operating-system (RHCOS) layer breakage frequently masquerades as an unrelated product failure: a single RHCOS bump swaps the kernel, cri-o, systemd, NetworkManager, and SELinux policy across the whole cluster at once, so the real cause surfaces as a symptom in some other domain. Before selecting a row from the routing table, complete BOTH steps:

1. Compare runtime versions across boots in the node journals (downloaded in Step 4; gzip-compressed without a .gz extension — plain grep silently matches nothing, use zcat/zgrep):

bash
# Runtime versions per boot. End-of-run snapshots (oc_cmds/nodes, nodes.json)
# show only the final version; changes within the run are visible only here.
zgrep -hE "Starting CRI-O, version|Container runtime initialized" \
  .work/prow-job-analysis/{build_id}/nodes/*/journal | sort | uniq -c

2. Scan the build log, JUnit, oc_cmds (node / clusteroperator status), MachineConfig data, and the journals for these signals:

  • NetworkPluginNotReady, or a missing CNI config (/etc/cni/net.d empty / no CNI plugin)
  • A ContainerRuntimeVersion change on nodes (cri-o version bump between runs)
  • A MachineConfigDaemon (MCD) rendered-config diff touching passwd, files, or units
  • Multiple nodes going NotReady after a reboot
  • CreateContainerError, RunContainerError, or OCI runtime errors (crun / runc)
  • Kernel panic, BUG, Oops, or soft lockup in node journals or the serial console
  • avc: denied / SELinux denials
  • The same failure spanning multiple unrelated jobs at a payload boundary

If step 1 shows more than one runtime version on any node, or any step-2 signal is present, the RHCOS layer is implicated: still route via the table below using whichever reference matches the surface symptom, but also read operating-system-changes.md alongside it. Never clear the OS layer from end-of-run snapshots alone.

Show full SKILL.md (455 more words)Show less

Failure Routing Table

Failure SignalReferenceWhen to Use
install should succeed fails in JUnitInstall — GeneralInstall failed at config/infra/bootstrap/cluster-creation/operator-stability stage
Metal/baremetal job + install failureInstall — MetalBare-metal install (dev-scripts, Metal3/Ironic, libvirt) — use alongside Install — General
A test failed (start here)Flaky Test IdentificationTriage entry for any failing test: classify infra vs product regression vs flake, then route onward
Confirmed regression in a plain e2e testTest Failure Root-CauseRoot-cause a real product regression in a plain (non-extension/install/upgrade) e2e test — e.g. [sig-network] ... should serve endpoints: test source, cluster state, originating error
*-tests-ext extension binary errorTest Extension BinariesOTE extension-binary extraction/discovery/version-skew failures — not core openshift-tests
Disruption events in intervalsDisruptionAPI backends stopped responding; interpret interval/timeline data (cause vs symptom vs noise)
Upgrade-phase failure or regressionUpgradeCVO stuck, operators degraded, MCO drain/reboot stalls, or version skew during upgrade
HyperShift / HCP job failureHyperShiftHosted control planes — correlate management and hosted clusters
aggregated- job failureAggregated JobsStatistical regression analysis across parallel child runs
Cloud API errors, quota, throttlingCloud Provider ErrorsAWS/GCP/Azure API/quota/provisioning failures before/during cluster creation
Node NotReady, OOM, disk pressureResource ExhaustionCPU/memory/disk/PID/etcd exhaustion, eviction, unschedulable pods
DNS, OVN, registry/pull, ingress errorsNetworkingOVN-Kubernetes/SDN, DNS, image pull/registry, load balancer/ingress, network policy
Container-start (cri-o), kernel panic, NetworkManager, RHCOS variant-isolated failureOperating System ChangesNode OS (RHCOS) layer — cri-o/crun, kernel, systemd, NetworkManager, SELinux, or an RHCOS bump in the payload
Lease/quota, ci-operator, Prow infraCI InfrastructureDistinguish "product broke" from "CI config changed"; ci-operator, step registry, leases
Need a specific artifact fileArtifactsArtifact directory structure, paths, and gcloud fetch commands

Job-name routing (Step 3) picks which reference to read. Failure classification (install | test | upgrade | infra) follows the root cause, not the job name.

Common Artifact Paths

These are the most frequently needed artifacts. See artifacts reference for the complete directory structure.

PathDescription
build-log.txtTop-level ci-operator log
artifacts/{target}/openshift-e2e-test/build-log.txtE2E test console log
artifacts/{target}/openshift-e2e-test/artifacts/junit/JUnit XML results
artifacts/{target}/openshift-e2e-test/artifacts/junit/e2e-timelines_spyglass_*.jsonDisruption timeline data
artifacts/{target}/gather-extra/artifacts/oc_cmds/Cluster state snapshots
artifacts/{target}/gather-extra/artifacts/pods/Pod logs by namespace
artifacts/{target}/gather-extra/artifacts/audit_logs/API server audit logs
artifacts/{target}/gather-must-gather/artifacts/must-gather.tarMust-gather archive
prowjob.jsonJob metadata and timing

URL Formats

Both formats are accepted and interchangeable:

text
# Prow UI
https://prow.ci.openshift.org/view/gs/test-platform-results-public/logs/{job}/{build_id}

# gcsweb (direct GCS browser)
https://gcsweb-ci.apps.ci.l2s4.p1.openshiftapps.com/gcs/test-platform-results-public/logs/{job}/{build_id}

Use the bucket from Step 1 as named in the URL (test-platform-results-public or prow-artifact-archive when the URL names it).

Tips

  • Start with build-log.txt — it shows the ci-operator orchestration and which steps failed
  • JUnit XML is the source of truth for test pass/fail status
  • Job name encodes environment — always parse it before diving into logs
  • Check prowjob.json for timing, payload tag, and whether the job timed out
  • Upgrade jobs install first — an "upgrade" job failing at install is an install failure, not an upgrade failure
  • Aggregated jobs need statistical analysis, not individual test debugging
  • Use .work/prow-job-analysis/{build_id}/ as the working directory for downloads

© openshift-eng, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (references) in plugins/ci/skills/prow-job-analysis of openshift-eng/ai-helpers.

  • SKILL.md
  • prow_job_artifact_search.py
  • references/aggregated.md
  • references/artifacts.md
  • references/ci-infrastructure-changes.md
  • references/cloud-provider-errors.md
  • references/disruption.md
  • references/flaky-test-identification.md
  • references/hypershift.md
  • references/install/general.md
  • references/install/metal.md
  • references/networking.md
  • references/operating-system-changes.md
  • references/resource-exhaustion.md
  • references/test-extension-binaries.md
  • references/test-failure.md
  • references/upgrade.md
  • test_prow_job_artifact_search.py

Open the folder on GitHubat commit a627176

Compare with similar skills

Prow Job Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prow Job Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prow Job Analysis this skillopenshift-eng/ai-helpers120—~2.8kAutomated safety check: PassApache-2.0
Vercel Composition Patternssupabase/supabase111k58 repos~726Automated safety check: PassMIT
Finishing a Development Branchobra/superpowers297k5 repos~1.9kAutomated safety check: PassMIT
Typescript Advanced Typesrolling-scopes/rsschool-app10k25 repos~4.2kAutomated safety check: PassMPL-2.0
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Official

    React composition patterns that scale. An agent skill from supabase/supabase.

    111k GitHub starsUsed in 58 repos~726 tokens
    DevelopmentAuto-check passed
  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    297k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Typescript Advanced Types

    rolling-scopes/rsschool-app

    Master TypeScript's advanced type system including generics, conditional types, mapped types, template literals, and utility types for building type-safe applications.

    10k GitHub starsUsed in 25 repos~4.2k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Greploop

    onyx-dot-app/onyx

    Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.

    32k GitHub starsUsed in 4 repos~3.3k tokens
    DevelopmentAuto-check passed

More from openshift-eng/ai-helpers

All 118 skills in this repo
  • Investigate CI Reliability

    openshift-eng/ai-helpers

    Find and independently validate actionable reliability defects across OpenShift release jobs and presubmits, then export portable issue handoffs.

    120 GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check passed
  • Address Review PR

    openshift-eng/ai-helpers

    Fetch and address all PR review comments — categorize by priority, make code changes, post replies, and push.

    120 GitHub stars~2.9k tokensUpdated 3 days ago
    Auto-check passed
  • Categorize Activity Types

    openshift-eng/ai-helpers

    Categorize Jira issues into Red Hat Sankey Activity Type categories using MCP Jira tools.

    120 GitHub stars~2.4k tokensUpdated 3 days ago
    Auto-check passed
  • Has Review Work

    openshift-eng/ai-helpers

    Decide whether a GitHub PR has unanswered authorized review comments or new required CI failures worth a follow-up agent.

    120 GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check passed
  • Must Gather Analyzer

    openshift-eng/ai-helpers

    Analyze OpenShift must-gather diagnostic data including cluster operators, pods, nodes, and network components.

    120 GitHub stars~2.3k tokensUpdated 3 days ago
    Auto-check passed
  • Payload Autodl JSON

    openshift-eng/ai-helpers

    Schema for the autodl JSON data file produced by payload-analysis for database ingestion — you must use this skill whenever generating the autodl JSON file

    120 GitHub stars~2.6k tokensUpdated 3 days ago
    Auto-check passed

Categories

Questions about Prow Job Analysis

What does Prow Job Analysis do?

A skill your agent uses when debugging a failed Prow CI job. Prow Job Analysis is an agent skill from openshift-eng/ai-helpers. Use this skill when debugging a failed Prow CI job.

When should I use Prow Job Analysis?

Prow Job Analysis fits situations like: debugging a failed Prow CI job.

How do I install Prow Job Analysis in Claude Code?

Run `npx skills add openshift-eng/ai-helpers --skill prow-job-analysis -a claude-code`. Or copy the skill folder (plugins/ci/skills/prow-job-analysis in openshift-eng/ai-helpers) into .claude/skills/prow-job-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Prow Job Analysis in Codex?

Run `npx skills add openshift-eng/ai-helpers --skill prow-job-analysis -a codex`. Or copy the skill folder (plugins/ci/skills/prow-job-analysis in openshift-eng/ai-helpers) into .agents/skills/prow-job-analysis in your project. Codex loads it when a task matches its description.

Can I use Prow Job Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openshift-eng/ai-helpers --skill prow-job-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prow-job-analysis, .gemini/skills/prow-job-analysis, .github/skills/prow-job-analysis and .opencode/skills/prow-job-analysis in your project.

What does Prow Job Analysis need to run?

Going by SKILL.md and its folder, Prow Job Analysis needs Python for the scripts in its folder and the command-line tools its instructions call (gcloud). Our summary lists: Python 3.

Does Prow Job Analysis access the network?

SKILL.md names 3 domains. In commands or code: prow.ci.openshift.org, gcsweb-ci.apps.ci.l2s4.p1.openshiftapps.com and storage.googleapis.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Prow Job Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prow Job Analysis use?

Prow Job Analysis is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prow Job Analysis use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 111k tokens, read only when the agent opens those files.

What are the alternatives to Prow Job Analysis?

Skills that share tags, products or a category with Prow Job Analysis: Vercel Composition Patterns (supabase/supabase, 111k stars), Finishing a Development Branch (obra/superpowers, 297k stars), Typescript Advanced Types (rolling-scopes/rsschool-app, 10k stars) and PR Babysitter (openinterpreter/openinterpreter, 69k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prow Job Analysis?

openshift-eng (a GitHub organization) maintains it in openshift-eng/ai-helpers, which has 120 GitHub stars. The repository holds 118 skills in this directory. The repository was last updated on October 6, 2026.

Source: openshift-eng/ai-helpers on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.