Official agent skill

Flaky Test Catcher

by elastic in elastic/terraform-provider-elasticstack

Detects broken and flaky acceptance tests from recent CI failures on main and opens structured GitHub issues for automated remediation.

OfficialApache-2.0Auto-check passedTesting & QA

Install Flaky Test Catcher

skills CLI
$ npx skills add elastic/terraform-provider-elasticstack --skill flaky-test-catcher -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install elastic/terraform-provider-elasticstack flaky-test-catcher --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/elastic/terraform-provider-elasticstack.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/flaky-test-catcher .claude/skills/flaky-test-catcher && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
flaky-test-catcher
GitHub stars
210
Token cost
~2.9k tokens
SKILL.md length
1,365 words
Files
1
Skills in repo
21
Repo updated
First seen
Licence
Apache-2.0

At a glance

Detects broken and flaky acceptance tests from recent CI failures on main and opens structured GitHub issues for automated remediation.

  • Works in 10 steps: Overview / Purpose → Inputs from pre-activation context → Fetching job logs → …
  • Tasks that involve Failing and flaky tests
  • SKILL.md covers 1. Overview / Purpose, 2. Inputs from pre-activation…, 3. Fetching job logs and 4. --- FAIL: extraction pattern, plus 7 more sections
  • Calls gh and git

What it does

Flaky Test Catcher is an agent skill from elastic/terraform-provider-elasticstack, published by the product's own GitHub organization. Detects broken and flaky acceptance tests from recent CI failures on main and opens structured GitHub issues for automated remediation. Follow this skill strictly when analyzing CI failures.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Failing and flaky tests and End-to-end testing. It works with GitHub. The repository describes itself as: Terraform provider for Elastic Stack. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Failing and flaky tests
  • Tasks that involve End-to-end testing

Example prompts

  • “Use the flaky-test-catcher skill to detect broken and flaky acceptance tests from recent CI failures on main and opens structured GitHub issues for…”
  • “/flaky-test-catcher”

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Overview / Purpose
  2. Inputs from pre-activation context
  3. Fetching job logs
  4. FAIL: extraction pattern
  5. Fail-rate formula and thresholds
  6. Base-test-name grouping rule
  7. Commit analysis steps
  8. Issue deduplication
  9. Required issue body sections
  10. Noop conditions

What it can do on your machine

Read from SKILL.md and the folder at commit b6bbc21. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Flaky Test Catcher loads about 2.9k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 1,365 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from elastic/terraform-provider-elasticstack at commit b6bbc21, republished under its Apache-2.0 licence (© elastic). 1,365 words, ~2,912 tokens.

Download SKILL.mdSave it as .claude/skills/flaky-test-catcher/SKILL.md (or your agent's skills folder).
name
flaky-test-catcher
description
Detects broken and flaky acceptance tests from recent CI failures on main and opens structured GitHub issues for automated remediation. Follow this skill strictly when analyzing CI failures.

Flaky Test Catcher — Analysis Protocol

1. Overview / Purpose

This skill defines the end-to-end protocol for detecting broken and flaky acceptance tests by analyzing recent CI failures on main, then opening structured GitHub issues so that automated remediation workflows can address the root causes.

You must follow this protocol strictly. Do not improvise or skip steps.

2. Inputs from pre-activation context

The workflow pre-activation step has already computed all run-level data. Do not re-query GitHub for run lists or issue counts. Use the values injected into your prompt:

VariableMeaning
failed_run_idsJSON array of run IDs of failed test.yml runs on main in the last 3 days
total_run_countCount of completed runs with meaningful conclusions (success, failure, timed_out, neutral, action_required) — cancelled runs excluded
open_issuesCurrent count of open flaky-test issues
issue_slots_availableHow many new issues you may create (max 3)

Parse failed_run_ids as a JSON array immediately. Example: ["12345678","87654321"].

3. Fetching job logs

For each run ID in failed_run_ids:

3.1 List jobs for the run
gh api /repos/{owner}/{repo}/actions/runs/{run_id}/jobs?per_page=100

Replace {owner} and {repo} with the values from the repository's GitHub remote (visible in git remote get-url origin).

3.2 Filter to relevant failing jobs

From the returned jobs array, keep only jobs where both conditions hold:

  • conclusion == "failure"
  • The job name contains Matrix Acceptance Test (this is the job group that runs acceptance tests)

Ignore infrastructure jobs (e.g. lint, build, generate) — they do not produce --- FAIL: lines.

3.3 Fetch the log for each failing job

The GitHub API log endpoint returns a ZIP archive. Use the gh CLI to stream plain-text logs for a job:

bash
gh run view --job {job_id} --log | grep '^--- FAIL:'

Log size warning: Job logs can be very large (10 MB+). Do not load the full log into context. Instead:

  • Stream the log output through grep so only matching lines are retained.
  • Scan only for lines matching the --- FAIL: pattern (see §4).
  • Use grep -B3 -A3 or similar when you also need surrounding context for the "Sample Failure Output" issue section.
  • Capture a small surrounding context (3–5 lines before/after each --- FAIL: line) for the "Sample Failure Output" issue section.
3.4 Pagination

If a job list response has 100 items and there may be more, check the Link header for a next page URL and repeat the request. In practice, a single run rarely has more than 100 jobs.

4. --- FAIL: extraction pattern

Scan each log for lines that match this exact pattern:

^--- FAIL: TestName (timing)

Examples of matching lines:

--- FAIL: TestAccResourceAgentConfiguration_alternateEnvironment (12.34s)
--- FAIL: TestAccSomeResource_basic (0.45s)

Rules:

  • The line must start with --- FAIL: (three hyphens, a space, FAIL:, a space).
  • The test name follows Go test naming: TestAcc... with optional underscore-separated sub-name.
  • Extract only the test function name — strip the timing suffix (12.34s).
  • Ignore bare FAIL lines without the --- prefix; those are package-level failure markers, not individual test failures.

Collect all extracted test names across all runs and all jobs. A test may appear multiple times (once per run where it failed) — track counts.

Same-run deduplication: If the same test name appears in multiple failing jobs within a single run (e.g. multiple shards both failing the same test), count it only once for that run. Deduplication is by run ID, not job ID — use a set per run when accumulating test names.

5. Fail-rate formula and thresholds

For each unique test name:

fail_rate = fail_count / total_run_count

Where:

  • fail_count = number of distinct run IDs in which this test appeared as --- FAIL:
  • total_run_count = the value from pre-activation context (already excludes cancelled runs)
ClassificationConditionAction
Brokenfail_rate == 1.0 (fails in 100% of runs)Create issue
Flakyfail_rate >= 0.20 and < 1.0Create issue
Noisefail_rate < 0.20Ignore — do not create issues

6. Base-test-name grouping rule

Extract the base test name from each test function name using this rule:

Take the substring from the beginning up to (but not including) the first underscore _.

Pattern: TestAcc[^_]+

Examples:

Full test nameBase test name
TestAccResourceAgentConfiguration_alternateEnvironmentTestAccResourceAgentConfiguration
TestAccResourceAgentConfiguration_minimalTestAccResourceAgentConfiguration
TestAccResourceAgentConfigurationTestAccResourceAgentConfiguration
TestAccSomeResource_basicTestAccSomeResource

One issue per base test name. All scenario variants (subtests/suffixes) belonging to the same base name are consolidated into a single issue. List each specific variant inside the issue body.

Fallback for non-TestAcc tests: All acceptance tests in this project follow the TestAcc prefix convention. If a non-TestAcc test name appears in the logs, treat everything up to the first _ (or the full name if no _) as the base name.

7. Commit analysis steps

For each base test name that will receive an issue, investigate whether any recent commit may already address the failure:

7.1 Find the oldest failing run timestamp

From the failed_run_ids list, identify the oldest run's created_at timestamp. You can get metadata for a single run:

gh api /repos/{owner}/{repo}/actions/runs/{run_id}

Take the minimum created_at across all failed runs.

7.2 Fetch commits on main since that timestamp
gh api "/repos/{owner}/{repo}/commits?sha=main&since={timestamp}&per_page=50"

Replace {timestamp} with the ISO 8601 value from step 7.1 (e.g. 2024-01-15T12:00:00Z).

7.3 For each commit, check relevance

For each commit returned:

a. Commit message relevance — does it reference any of:

  • The base test name (e.g. TestAccResourceAgentConfiguration)
  • The resource name (derive from the test name, e.g. agent_configuration → AgentConfiguration)
  • Keywords: fix, flaky, test, revert

b. Changed file relevance — fetch the full commit detail to get changed file paths:

gh api /repos/{owner}/{repo}/commits/{sha}

Check if any file in files[].filename matches patterns like:

  • *_test.go files whose name contains a token from the resource name
  • Files in the same Go package directory as the test
Show full SKILL.md (519 more words)Show less
7.4 Include findings in issue body
  • If a relevant commit is found:
    ⚠️ may already be addressed in `{short_sha}` — {one-line message summary}
  • If no relevant commits found:
    No recent commits appear to address this failure.

Frame the analysis as "has this been fixed yet?" — not as blame attribution.

Do not suppress issue creation: Even if a fix commit is found, always proceed with creating the issue and include the fix-detection note in the Commit Analysis section. The issue serves as the remediation trigger regardless.

8. Issue deduplication

Before creating an issue for a base test name, check whether one already exists:

gh api "/repos/{owner}/{repo}/issues?labels=flaky-test&state=open&per_page=100"
  • Inspect each returned issue's title.
  • The full rendered title for flaky-test issues is: [flaky-test] {BaseTestName} (the [flaky-test] prefix is applied automatically by the create-issue safe output). When checking for duplicates, compare against this full title as it appears in GitHub.
  • If an exact title match exists, skip creating a new issue for that base test name.
  • Do not re-query or recalculate issue_slots_available; use only the value from pre-activation.
  • If the response contains exactly 100 results, check for a next link in the Link response header and repeat the request for subsequent pages until all open issues are fetched.

9. Required issue body sections

Issue title: Pass only {BaseTestName} to the create-issue safe output — the [flaky-test] prefix is added automatically.

Every issue you create must contain exactly these 5 sections in this order:

markdown
## Broken Tests

List each test function name (including scenario suffix) that failed in 100% of runs:
- ❌ `TestAccResourceFoo_basic` — failed in 5/5 runs

## Flaky Tests

List each test function name that failed in ≥ 20% but < 100% of runs, with the observed rate:
- ⚠️ `TestAccResourceFoo_update` — failed in 3/5 runs (60%)
- ⚠️ `TestAccResourceFoo_import` — failed in 1/5 runs (20%)

## Commit Analysis

{Output from §7. Either a ⚠️ note about a possible fix commit, or the "No recent commits" message. Include commit SHA, message, and affected file paths when relevant.}

## Sample Failure Output

{Short excerpt (5–15 lines) of the actual log output surrounding a `--- FAIL:` line. Include any immediately preceding error messages for context.}

## Affected Stack Versions

{List the Elastic Stack versions / matrix dimension values (e.g. Elasticsearch version, Kibana version) from the failing job names or log metadata. If not determinable, write "Unknown — not present in log output".}

Formatting rules:

  • Use ❌ for broken tests (100% fail rate).
  • Use ⚠️ for flaky tests (20%–99% fail rate), and always include the fraction and percentage.
  • Keep "Sample Failure Output" to the most informative excerpt; do not paste hundreds of lines.

Cap enforcement: Before creating each issue, verify that the number of issues created so far in this run has not reached issue_slots_available. Stop creating issues once the cap is reached, even if additional base test names remain.

10. Noop conditions

Call noop with a descriptive explanation (do not create any issues) when any of these conditions holds:

  1. All failures already have open issues — every qualifying base test name matched an existing open flaky-test issue during deduplication; nothing new to open.

  2. All failures are below the 20% threshold — every observed --- FAIL: test has fail_rate < 0.20; there are no broken or flaky tests to report.

  3. No --- FAIL: patterns found — none of the logs for the provided failed_run_ids contained a --- FAIL: line; the CI failures were likely infrastructure failures (network timeouts, setup errors, etc.) rather than test logic failures.

When calling noop, state which condition applied and include basic counts (e.g. "3 failures observed, all below 20% threshold").

Summary of execution order

  1. Parse failed_run_ids from pre-activation context.
  2. For each run ID: list jobs → filter to failing Matrix Acceptance Test jobs → fetch and scan logs for --- FAIL: lines.
  3. Aggregate per-test fail counts across all runs.
  4. Apply fail-rate thresholds; discard noise (< 20%).
  5. Group surviving tests by base test name.
  6. Deduplicate against existing open flaky-test issues.
  7. For each remaining base test name (up to issue_slots_available): run commit analysis, then create an issue with all 5 required sections.
  8. If no issues were created, call noop.

© elastic, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/flaky-test-catcher of elastic/terraform-provider-elasticstack.

Open the folder on GitHubat commit b6bbc21

Compare with similar skills

Flaky Test Catcher next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Flaky Test Catcher compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Flaky Test Catcher this skillelastic/terraform-provider-elasticstack210—~2.9kAutomated safety check: PassApache-2.0
Detect Flaky Testsagent-substrate/substrate4.5k—~3kAutomated safety check: PassApache-2.0
MAUI UI Test Writerdotnet/maui23k—~3kAutomated safety check: PassMIT
Test Fix WorkflowGoogleCloudPlatform/magic-modules974—~1.3kAutomated safety check: PassCustom licence
Cucumber and Playwright E2E Testslanggenius/dify158k—~682Automated safety check: PassCustom licence
Triage CI Flakepayloadcms/payload45k—~4.4kAutomated safety check: PassMIT

Similar skills

  • Detect Flaky Tests

    agent-substrate/substrate

    Detects flaky Go tests by analyzing GitHub Actions workflow runs across the last 7 days and all PRs — covering both the run-tests job (unit/integration) and the e2e-test job (gVisor and microVM…

    4.5k GitHub stars~3k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Writes UI tests that reproduce a GitHub issue in .NET MAUI and keeps iterating until the tests actually fail, proving they catch the bug.

    23k GitHub stars~3k tokensUpdated today
    Testing & QAAuto-check passed
  • Test Fix Workflow

    GoogleCloudPlatform/magic-modules

    Workflow for diagnosing, fixing, and verifying failing Terraform acceptance tests from GitHub issue URLs (detecting test-failure labels), direct prompts, or log files using Failure Scenario Decision…

    974 GitHub stars~1.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions.

    158k GitHub stars~682 tokensUpdated today
    Testing & QAAuto-check passed
  • Triage CI Flake

    payloadcms/payload

    A skill your agent uses when CI tests fail on main branch after PR merge, when investigating flaky test failures, or when user provides a PR URL/number to aggregate all failing tests

    45k GitHub stars~4.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.

    6.7k GitHub stars~4.4k tokensUpdated today
    Testing & QAAuto-check passed

More from elastic/terraform-provider-elasticstack

All 21 skills in this repo
  • Openspec Explore

    elastic/terraform-provider-elasticstack

    Official

    Enter explore mode - a thinking partner for exploring ideas, investigating problems, and clarifying requirements.

    210 GitHub starsUsed in 87 repos~4.6k tokens
    Auto-check passed
  • Openspec Apply Change

    elastic/terraform-provider-elasticstack

    Official

    Implement tasks from an OpenSpec change. An agent skill from elastic/terraform-provider-elasticstack.

    210 GitHub starsUsed in 92 repos~2.1k tokens
    Auto-check passed
  • Openspec Archive Change

    elastic/terraform-provider-elasticstack

    Official

    Archive a completed change in the experimental workflow. An agent skill from elastic/terraform-provider-elasticstack.

    210 GitHub starsUsed in 86 repos~2.7k tokens
    Auto-check passed
  • PR Monitoring Loop

    elastic/terraform-provider-elasticstack

    Official

    Monitor GitHub pull requests through a subagent-based loop that watches CI checks, review comments, PR comments, review state, merge conflicts, and branch freshness.

    210 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Openspec Plus Proposal

    elastic/terraform-provider-elasticstack

    Official

    MANDATORY skill that activates whenever the OpenSpec proposal phase begins.

    210 GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Openspec Continue Change

    elastic/terraform-provider-elasticstack

    Official

    Continue working on an OpenSpec change by creating the next artifact.

    210 GitHub starsUsed in 31 repos~1.7k tokens
    Auto-check passed

Works with

Categories

Questions about Flaky Test Catcher

What does Flaky Test Catcher do?

Detects broken and flaky acceptance tests from recent CI failures on main and opens structured GitHub issues for automated remediation. Flaky Test Catcher is an agent skill from elastic/terraform-provider-elasticstack, published by the product's own GitHub organization. Detects broken and flaky acceptance tests from recent CI failures on main and opens structured GitHub issues for automated remediation.

When should I use Flaky Test Catcher?

Flaky Test Catcher fits situations like: tasks that involve Failing and flaky tests; tasks that involve End-to-end testing.

How do I install Flaky Test Catcher in Claude Code?

Run `npx skills add elastic/terraform-provider-elasticstack --skill flaky-test-catcher -a claude-code`. Or copy the skill folder (.agents/skills/flaky-test-catcher in elastic/terraform-provider-elasticstack) into .claude/skills/flaky-test-catcher in your project. Claude Code loads it when a task matches its description.

How do I install Flaky Test Catcher in Codex?

Run `npx skills add elastic/terraform-provider-elasticstack --skill flaky-test-catcher -a codex`. Or copy the skill folder (.agents/skills/flaky-test-catcher in elastic/terraform-provider-elasticstack) into .agents/skills/flaky-test-catcher in your project. Codex loads it when a task matches its description.

Can I use Flaky Test Catcher in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add elastic/terraform-provider-elasticstack --skill flaky-test-catcher -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/flaky-test-catcher, .gemini/skills/flaky-test-catcher, .github/skills/flaky-test-catcher and .opencode/skills/flaky-test-catcher in your project.

What does Flaky Test Catcher need to run?

Going by SKILL.md and its folder, Flaky Test Catcher needs the command-line tools its instructions call (gh and git).

Does Flaky Test Catcher access the network?

SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Flaky Test Catcher safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Flaky Test Catcher use?

Flaky Test Catcher is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Flaky Test Catcher use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Flaky Test Catcher?

Skills that share tags, products or a category with Flaky Test Catcher: Detect Flaky Tests (agent-substrate/substrate, 4.5k stars), MAUI UI Test Writer (dotnet/maui, 23k stars), Test Fix Workflow (GoogleCloudPlatform/magic-modules, 974 stars) and Cucumber and Playwright E2E Tests (langgenius/dify, 158k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Flaky Test Catcher?

elastic (a GitHub organization, an official publisher) maintains it in elastic/terraform-provider-elasticstack, which has 210 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 8, 2026.

Source: elastic/terraform-provider-elasticstack on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.