Agent skill

Analyzing CI Flakiness

by opsmill in opsmill/infrahub

Analyzes recent CI failures on pull requests to identify flaky tests, using retry outcomes (failed attempt → green re-run) and cross-PR recurrence as evidence, and maintains a local longitudinal…

Apache-2.0Auto-check passedTesting & QA

Install Analyzing CI Flakiness

skills CLI
$ npx skills add opsmill/infrahub --skill analyzing-ci-flakiness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install opsmill/infrahub analyzing-ci-flakiness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/opsmill/infrahub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/analyzing-ci-flakiness .claude/skills/analyzing-ci-flakiness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
analyzing-ci-flakiness
GitHub stars
533
Token cost
~2k tokens
SKILL.md length
914 words
Files
2 (incl. scripts)
Skills in repo
32
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyzes recent CI failures on pull requests to identify flaky tests, using retry outcomes (failed attempt → green re-run) and cross-PR recurrence as evidence, and maintains a local longitudinal…

  • Works in 5 steps: Parse arguments → Collect → Investigate what the script could not name → …
  • : the user wants to find flaky tests
  • SKILL.md covers Introduction, Step 1 — Parse arguments, Step 2 — Collect and Step 3 — Investigate what the…, plus 2 more sections
  • Runs Python scripts from its folder; calls python3 and docker

What it does

Analyzing CI Flakiness is an agent skill from opsmill/infrahub. Analyzes recent CI failures on pull requests to identify flaky tests, using retry outcomes (failed attempt → green re-run) and cross-PR recurrence as evidence, and maintains a local longitudinal ledger so flakiness can be tracked over time. TRIGGER when: the user wants to find flaky tests, correlate recent CI failures, check which tests fail across PRs or recover on retry, or refresh the flakiness trend report. DO NOT TRIGGER when: babysitting a single PR's CI until green → monitoring-pull-requests; diagnosing or…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/collect.py`). Compatibility notes: Requires the gh CLI authenticated against the repo. Python 3 (stdlib only). Writes a cache under ~/ci-cache.

It sits in Testing & QA, covering Failing and flaky tests. The repository describes itself as: Infrahub is a graph-based data management platform with built-in version control, CI workflows, peer review, and API access. It’s purpose-built to power reliable infrastructure… The licence is Apache-2.0.

When your agent uses it

  • : the user wants to find flaky tests
  • Correlate recent CI failures
  • Check which tests fail across PRs
  • Recover on retry

Example prompts

  • “Use the analyzing-ci-flakiness skill to analyz recent CI failures on pull requests to identify flaky tests, using retry outcomes (failed attempt →…”
  • “/analyzing-ci-flakiness”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): Requires the gh CLI authenticated against the repo. Python 3 (stdlib only). Writes a cache under ~/ci-cache.
  • Pre-approved tools (allowed-tools): Bash(python3 .agents/skills/analyzing-ci-flakiness/scripts/collect.py:*)

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Parse arguments
  2. Collect
  3. Investigate what the script could not name
  4. Judge: flake vs regression
  5. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 2e1f1eb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(python3 .agents/skills/analyzing-ci-flakiness/scripts/collect.py:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the gh CLI authenticated against the repo. Python 3 (stdlib only). Writes a cache under ~/ci-cache.

    From compatibility in the SKILL.md frontmatter.

Context cost

Analyzing CI Flakiness loads about 2k tokens when it runs. Until then it costs about 150 tokens; SKILL.md has 914 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~150
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from opsmill/infrahub at commit 2e1f1eb, republished under its Apache-2.0 licence (© opsmill). 914 words, ~2,004 tokens.

Download SKILL.mdSave it as .claude/skills/analyzing-ci-flakiness/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
analyzing-ci-flakiness
description
Analyzes recent CI failures on pull requests to identify flaky tests, using retry outcomes (failed attempt → green re-run) and cross-PR recurrence as evidence, and maintains a local longitudinal ledger so flakiness can be tracked over time. TRIGGER when: the user wants to find flaky tests, correlate recent CI failures, check which tests fail across PRs or recover on retry, or refresh the flakiness trend report. DO NOT TRIGGER when: babysitting a single PR's CI until green → monitoring-pull-requests; diagnosing or fixing one specific failing test → the bug-analysis skills.
allowed-tools
Bash(python3 .agents/skills/analyzing-ci-flakiness/scripts/collect.py:*)
compatibility
Requires the gh CLI authenticated against the repo. Python 3 (stdlib only). Writes a cache under ~/ci-cache.
argument-hint
Optional base-branch glob(s) and window, e.g. `release-1.11 14` (default: all bases, last 7 days)
metadata.version
0.1.0
metadata.author
OpsMill

CI Flakiness Analyzer

Introduction

A test is flaky when its failure does not reproduce on the same code: the run was retried and went green, or the same test fails on unrelated PRs. This skill mines both signals from GitHub Actions history, downloads the failed job logs once into a local cache, and appends every observation to a ledger (~/ci-cache/<owner>-<name>/ledger.jsonl) so repeated invocations — weekly, or ad hoc — accumulate trend data instead of starting from scratch.

The mechanical part (fetching, caching, test-name extraction, known-signature classification) is done by the bundled script. Your job is the judgment part: separating flakes from real regressions, spotting new systemic signatures, and writing the report.

Step 1 — Parse arguments

  • Base-branch filter: any arguments that look like branch names or globs (release-1.11, release-*, stable). Default: no filter (all PR bases), which is usually what "how flaky is CI" means. Filter when the user names a branch.
  • Window: a bare integer is a number of days (default 7). An ISO date means "since that date".

Step 2 — Collect

Run the bundled collector (repo-root relative):

bash
python3 .agents/skills/analyzing-ci-flakiness/scripts/collect.py \
  [--base <glob> ...] [--days N | --since YYYY-MM-DD] [--repo owner/name]

It prints a JSON report to stdout and writes everything under ~/ci-cache/<owner>-<name>/windows/<since>_<until>/:

  • runs.jsonl — every pull_request workflow run created in the window
  • failed_jobs_with_tests.json — failed jobs of the interesting run-attempts, with extracted failing tests, systemic-bucket tags, and a recovered_same_run flag
  • report-data.json — headline numbers, ranked per-test table, per-bucket incident counts (bucket_incidents: distinct jobs/runs/PRs per systemic bucket), and the ledger's weekly history
  • joblogs/<job_id>.log — raw logs (ANSI intact; strip with sed 's/\x1b\[[0-9;]*m//g')

Notes the script already accounts for — don't re-derive them:

  • The runs API's pull_requests field is empty for many runs; the script joins runs to PRs through every PR head commit SHA as well. Don't trust the field alone.
  • "Interesting attempts" = every earlier attempt of a retried run (that's what the retry fixed) plus final attempts that failed. Runs cancelled on attempt 1 are concurrency noise and skipped.
  • Logs already on disk are never re-downloaded; the ledger is deduplicated by (job, test). Old logs expire on GitHub's side (~90 days) — an empty joblogs/*.log means expired, not passing.

Step 3 — Investigate what the script could not name

For failed jobs with an empty tests list and no bucket tag, read the log yourself (grep for ##[error], FAILED, Error:, Timeout). Two outcomes:

  • It matches a new systemic signature (infra failure that cascades over many tests). Add a regex for it to BUCKETS in collect.py and to the table below, so future runs classify it.
  • It's a genuine test failure the extraction regexes missed — note the test manually and consider extending extract_tests.
Known systemic signatures (as of 2026-08 — keep in sync with BUCKETS in collect.py)
BucketSignatureMeaning
stack-readinessServerNotResponsiveError … /api/schema/loadSeeded testcontainers stack not ready; the whole pytest-playwright shard errors. One incident, not N flaky tests.
vitest-mock-corruptionTypeError: vi.mocked(...).mockX is not a functionvitest browser-mode module-mocking race; hits a different test file each time.
prefect-setup-triggers-timeoutSetup triggers task ReadTimeoutPrefect hang at session setup; downstream tests hit their own timeouts.
neo4j-deadlockNeo.TransientError.Transaction.DeadlockDetectedConcurrent-write deadlock, usually integration suites under xdist.
compose-boot-failuredocker compose … up --wait non-zero exitStack never booted; job-level infra failure.
sqlite-locked(sqlite3.OperationalError) database is locked (also matches the raw sqlite3.OperationalError: form)Prefect's sqlite under contention.
runner-oomProcess completed with exit code 137Runner OOM/SIGKILL; the mass test failures in the same job are casualties, not flakes.
docker-network-pool-exhaustedall predefined address pools have been fully subnettedLeaked compose networks exhausted the docker address pools on a self-hosted runner.
actions-download-429Failed to download action … 429GitHub rate-limited its own action download; pure platform flake.
prefect-task-manager-wedgedRuntimeError: Prefect task manager setup already failed for http…The memoized task-manager setup (backend/tests/helpers/task_manager.py) timed out once against a Prefect test server; every later class fail-fasts on the remembered failure. One incident, hundreds of cascaded ERRORs.
pytest-green-exit-1green pytest summary (no failed, no errors) directly followed by exit 1Session-teardown/plugin abort after all tests passed (e.g. testcontainers result reporting).
Show full SKILL.md (270 more words)Show less

Step 4 — Judge: flake vs regression

For each test in the ranked table, classify:

  • Flaky (strong) — fails on ≥2 unrelated PRs, or recovered_on_retry > 0. The more distinct PRs, the stronger.
  • Flaky (weak) — single occurrence with an infra-flavored error (locator timeout, transient branch not found) and the PR later went green. List, but rank low.
  • Suspect regression, not a flake — the same test fails on every attempt of the same commit and the PR is still red, or the failures started only after a specific merge. Say so explicitly; do not bury it in the flake list. Cross-check: does the test fail on any PR that does not contain the suspect change?
  • Systemic bucket — tests whose only failures carry a bucket tag are casualties, not causes. Report the bucket (with the incident count from bucket_incidents in report-data.json), not the individual tests.

Different tests failing on successive attempts of the same run = two independent flakes, not a regression.

Step 5 — Report

Write ANALYSIS.md into the window directory, then give the user a summary. Lead with the ranked flake candidates. Include:

  1. Headline numbers: PRs in scope, runs matched, retried runs, retried-and-recovered runs (pure-flake evidence), hard failures.
  2. Ranked flake candidates — test id, distinct PRs/runs, recovered-on-retry count, one-line error cause. Group systemic buckets as single entries.
  3. Suspected real regressions, clearly separated.
  4. Trend — from weekly_history in report-data.json: which offenders are new this window, which recur week over week, which disappeared (likely fixed). This section is the reason the ledger exists; don't skip it once ≥2 windows of data exist.

Do not propose fixes unless asked; the deliverable is the evidence-ranked candidate list.

© opsmill, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in .agents/skills/analyzing-ci-flakiness of opsmill/infrahub.

  • SKILL.md
  • scripts/collect.py

Open the folder on GitHubat commit 2e1f1eb

Compare with similar skills

Analyzing CI Flakiness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Analyzing CI Flakiness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Analyzing CI Flakiness this skillopsmill/infrahub533—~2kAutomated safety check: PassApache-2.0
Swig Testswig/swig6.3k—~2.3kAutomated safety check: PassCustom licence
Triage CI FailureDataDog/datadog-agent3.8k—~2.3kAutomated safety check: PassApache-2.0
Dynamo Jira TicketDynamoDS/Dynamo2k—~1.1kAutomated safety check: PassApache-2.0
Fix Ready PRsfastrepl/anarlog9.5k—~1.4kAutomated safety check: PassMIT
Trx Analysismicrosoft/vstest969—~1.8kAutomated safety check: PassMIT

Similar skills

  • Swig Test

    swig/swig

    Run SWIG test suite for specific languages. An agent skill from swig/swig.

    6.3k GitHub stars~2.3k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Triage CI Failure

    DataDog/datadog-agent

    Official

    Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.

    3.8k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Dynamo Jira Ticket

    DynamoDS/Dynamo

    Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests.

    2k GitHub stars~1.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Fix Ready PRs

    fastrepl/anarlog

    Inspect every open non-draft PR for CI failures and unresolved Cursor Bugbot findings, then fix them on the existing PR branches.

    9.5k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Trx Analysis

    microsoft/vstest

    Official

    Parse and analyze Visual Studio TRX test result files. An agent skill from microsoft/vstest.

    969 GitHub stars~1.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Wio

    workersio/skills

    Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing test-suite health.

    200 GitHub stars~5.8k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed

More from opsmill/infrahub

All 32 skills in this repo
  • Audit Docs

    opsmill/infrahub

    Audits internal (dev/) and external (docs/) documentation completeness for a feature, subject, or set of existing docs, maps changes indicated by the user, across Infrahub's documentation layers…

    533 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Commit

    opsmill/infrahub

    Stages and commits the current changes onto a safe working branch, enforcing branch discipline and optionally pushing upstream.

    533 GitHub stars~2.8k tokensUpdated today
    Auto-check: notes
  • A skill your agent uses when you've fixed a bug, added a feature, or made any user-facing change in a project that uses Towncrier and need to record it for the changelog — before committing or…

    533 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Creating Issues

    opsmill/infrahub

    Turns a single feature idea, improvement, or bug into ONE well-structured GitHub issue.

    533 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Creating Prd

    opsmill/infrahub

    Synthesises the current conversation context into a Product Requirements Document and publishes it to GitHub (as a comment on a referenced issue, or a new issue).

    533 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Grilling Ideas

    opsmill/infrahub

    Stress-tests a fuzzy or vague feature idea before any PRD, spec, or ticket is written.

    533 GitHub stars~3.8k tokensUpdated today
    Auto-check passed

Categories

Questions about Analyzing CI Flakiness

What does Analyzing CI Flakiness do?

Analyzes recent CI failures on pull requests to identify flaky tests, using retry outcomes (failed attempt → green re-run) and cross-PR recurrence as evidence, and maintains a local longitudinal…. Analyzing CI Flakiness is an agent skill from opsmill/infrahub. Analyzes recent CI failures on pull requests to identify flaky tests, using retry outcomes (failed attempt → green re-run) and cross-PR recurrence as evidence, and maintains a local longitudinal ledger so flakiness can be tracked over time.

When should I use Analyzing CI Flakiness?

Analyzing CI Flakiness fits situations like: : the user wants to find flaky tests; correlate recent CI failures; check which tests fail across PRs; recover on retry.

How do I install Analyzing CI Flakiness in Claude Code?

Run `npx skills add opsmill/infrahub --skill analyzing-ci-flakiness -a claude-code`. Or copy the skill folder (.agents/skills/analyzing-ci-flakiness in opsmill/infrahub) into .claude/skills/analyzing-ci-flakiness in your project. Claude Code loads it when a task matches its description.

How do I install Analyzing CI Flakiness in Codex?

Run `npx skills add opsmill/infrahub --skill analyzing-ci-flakiness -a codex`. Or copy the skill folder (.agents/skills/analyzing-ci-flakiness in opsmill/infrahub) into .agents/skills/analyzing-ci-flakiness in your project. Codex loads it when a task matches its description.

Can I use Analyzing CI Flakiness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add opsmill/infrahub --skill analyzing-ci-flakiness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyzing-ci-flakiness, .gemini/skills/analyzing-ci-flakiness, .github/skills/analyzing-ci-flakiness and .opencode/skills/analyzing-ci-flakiness in your project.

What does Analyzing CI Flakiness need to run?

Going by SKILL.md and its folder, Analyzing CI Flakiness needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and docker). Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Bash(python3 .agents/skills/analyzing-ci-flakiness/scripts/collect.py:*). Compatibility (from SKILL.md): Requires the gh CLI authenticated against the repo. Python 3 (stdlib only). Writes a cache under ~/ci-cache..

Does Analyzing CI Flakiness access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Analyzing CI Flakiness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Analyzing CI Flakiness use?

Analyzing CI Flakiness is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Analyzing CI Flakiness use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Analyzing CI Flakiness?

Skills that share tags, products or a category with Analyzing CI Flakiness: Swig Test (swig/swig, 6.3k stars), Triage CI Failure (DataDog/datadog-agent, 3.8k stars), Dynamo Jira Ticket (DynamoDS/Dynamo, 2k stars) and Fix Ready PRs (fastrepl/anarlog, 9.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Analyzing CI Flakiness?

opsmill (a GitHub organization) maintains it in opsmill/infrahub, which has 533 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 9, 2026.

Source: opsmill/infrahub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.