Agent skill

Test Failure Investigator

by proffesor-for-testing in proffesor-for-testing/agentic-qe

A skill your agent uses when a test is failing and you need to determine root cause: is it flaky, an environment issue, or a real regression?

MITAuto-check passedTesting & QA

Install Test Failure Investigator

skills CLI
$ npx skills add proffesor-for-testing/agentic-qe --skill test-failure-investigator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install proffesor-for-testing/agentic-qe test-failure-investigator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/proffesor-for-testing/agentic-qe.git skills-src && mkdir -p .claude/skills && cp -r skills-src/assets/skills/test-failure-investigator .claude/skills/test-failure-investigator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-failure-investigator
GitHub stars
495
Token cost
~811 tokens
SKILL.md length
216 words
Files
1
Skills in repo
93
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when a test is failing and you need to determine root cause: is it flaky, an environment issue, or a real regression?

  • Works in 5 steps: Classify the Failure → Narrow the Scope → Root Cause Analysis → …
  • A test is failing and you need to determine root cause: is it flaky
  • SKILL.md covers Activation, Investigation Flow, Composition and Gotchas
  • Calls git, npx and npm

What it does

Test Failure Investigator is an agent skill from proffesor-for-testing/agentic-qe. Use when a test is failing and you need to determine root cause: is it flaky, an environment issue, or a real regression? Traces failure from symptom to fix.

Its SKILL.md is about 810 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Failing and flaky tests and Root cause analysis. The repository describes itself as: Agentic QE Fleet is an open-source AI-powered QA/QE platform designed for use with Coding Agents (works best with Claude Code) featuring specialized agents and skills to support… The licence is MIT.

When your agent uses it

  • A test is failing and you need to determine root cause: is it flaky
  • An environment issue
  • A real regression?

Example prompts

  • “/test-failure-investigator”

Requirements

  • Node.js
  • Docker

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Classify the Failure
  2. Narrow the Scope
  3. Root Cause Analysis
  4. Flaky Test Investigation
  5. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 1363bc7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • npx
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, npx and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Failure Investigator loads about 811 tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 216 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~811

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from proffesor-for-testing/agentic-qe at commit 1363bc7, republished under its MIT licence (© proffesor-for-testing). 216 words, ~811 tokens.

Download SKILL.mdSave it as .claude/skills/test-failure-investigator/SKILL.md (or your agent's skills folder).
name
test-failure-investigator
description
Use when a test is failing and you need to determine root cause: is it flaky, an environment issue, or a real regression? Traces failure from symptom to fix.
user-invocable
true

Test Failure Investigator

Runbook-style skill for systematic test failure investigation. Given a failing test, determines root cause and recommends action.

Activation

/test-failure-investigator [test-name-or-file]

Investigation Flow

Step 1: Classify the Failure

Run the test 3 times and classify:

Result PatternClassificationAction
Fails consistentlyRegression or EnvironmentContinue to Step 2
Fails intermittentlyFlakySkip to Step 4
Passes nowTransientCheck CI logs, environment diff
bash
# Run test 3 times
for i in 1 2 3; do npx jest {{test_file}} 2>&1 | tail -5; echo "--- Run $i ---"; done
Step 2: Narrow the Scope
bash
# When did it start failing?
git log --oneline -20 -- {{related_source_files}}

# What changed recently?
git diff HEAD~5 -- {{related_source_files}}

# Does it fail in isolation?
npx jest {{test_file}} --testNamePattern="{{test_name}}"

# Does it fail with other tests?
npx jest --runInBand  # sequential execution
Step 3: Root Cause Analysis
SymptomLikely CauseInvestigation
TimeoutNetwork/DB dependencyCheck external service availability
Assertion mismatchLogic changeCompare expected vs actual, check git blame
Import errorDependency changeCheck package.json changes, run npm ci
Permission deniedEnvironmentCheck file permissions, Docker volumes
Out of memoryResource leakProfile with --detectOpenHandles
Step 4: Flaky Test Investigation
bash
# Run 10 times to confirm flakiness
for i in $(seq 1 10); do npx jest {{test_file}} --forceExit 2>&1 | grep -E 'PASS|FAIL'; done

# Common flaky causes:
# - Shared state between tests (missing cleanup)
# - Time-dependent assertions (use fake timers)
# - Race conditions (missing await)
# - Port conflicts (use random ports)
# - Order dependency (run with --randomize)
Step 5: Report
markdown
## Test Failure Report
- **Test**: {{test_name}}
- **File**: {{test_file}}
- **Classification**: Regression / Flaky / Environment / Transient
- **Root Cause**: {{description}}
- **First Failed**: {{commit_hash}} ({{date}})
- **Fix**: {{recommended_action}}
- **Verified**: [ ] Fix applied and test passes 3x consecutively

Composition

After investigation, compose with:

  • /bug-reporting-excellence — if regression found, file a bug report
  • /regression-testing — if regression, add to regression suite
  • /qe-test-execution — for re-running tests after fix

Gotchas

  • Agent may guess at root cause without running the test — always reproduce first
  • "Works on my machine" is not a diagnosis — compare environments (node version, OS, deps)
  • Flaky tests that pass 9/10 times will still be reported as "passing" by CI — run 10+ times
  • Test isolation failures are the #1 cause of flaky tests — check for shared state in beforeAll/afterAll

© proffesor-for-testing, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in assets/skills/test-failure-investigator of proffesor-for-testing/agentic-qe.

Open the folder on GitHubat commit 1363bc7

Compare with similar skills

Test Failure Investigator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Failure Investigator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Failure Investigator this skillproffesor-for-testing/agentic-qe495—~811Automated safety check: PassMIT
Trx Analysismicrosoft/vstest969—~1.8kAutomated safety check: PassMIT
Evanflow Debugevanklem/evanflow418—~760Automated safety check: PassCustom licence
Issue TracerZaxbyHub/opencode-swarm494—~383Automated safety check: PassMIT
Issue TracerZaxbyHub/opencode-swarm494—~397Automated safety check: PassMIT
Pester Failure AnalysisPowerShell/PowerShell56k—~5.1kAutomated safety check: PassMIT

Similar skills

  • Trx Analysis

    microsoft/vstest

    Official

    Parse and analyze Visual Studio TRX test result files. An agent skill from microsoft/vstest.

    969 GitHub stars~1.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Evanflow Debug

    evanklem/evanflow

    Root-cause discipline for bugs, test failures, and unexpected behavior.

    418 GitHub stars~760 tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Issue Tracer

    ZaxbyHub/opencode-swarm

    A skill your agent uses when asked to trace, investigate, root-cause, reproduce, plan, fix, resolve, close, or prepare a PR for an issue, bug report, defect, regression, failing test, or confusing…

    494 GitHub stars~383 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Issue Tracer

    ZaxbyHub/opencode-swarm

    A skill your agent uses when asked to trace, investigate, root-cause, reproduce, plan, fix, resolve, close, or prepare a PR for an issue, bug report, defect, regression, failing test, or confusing…

    494 GitHub stars~397 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Pester Failure Analysis

    PowerShell/PowerShell

    Investigates failing Pester tests in PowerShell CI jobs by following a six-step workflow from pull request status to documented fix recommendations.

    56k GitHub stars~5.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report.

    41k GitHub stars~1.5k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from proffesor-for-testing/agentic-qe

All 93 skills in this repo
  • Contract Testing

    proffesor-for-testing/agentic-qe

    Consumer-driven contract testing for microservices using Pact, schema validation, API versioning, and backward compatibility testing.

    495 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Mutation Testing

    proffesor-for-testing/agentic-qe

    Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate.

    495 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Performance Testing

    proffesor-for-testing/agentic-qe

    Profiles application performance under load using k6, Artillery, or JMeter to measure latency, throughput, and error rates.

    495 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Code Review Quality

    proffesor-for-testing/agentic-qe

    Conduct context-driven code reviews focusing on quality, testability, and maintainability.

    495 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Security Testing

    proffesor-for-testing/agentic-qe

    Scans for security vulnerabilities including XSS, SQL injection, CSRF, and auth flaws using OWASP Top 10 methodology.

    495 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes
  • Database Testing

    proffesor-for-testing/agentic-qe

    Database schema validation, data integrity testing, migration testing, transaction isolation, and query performance.

    495 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed

Categories

Questions about Test Failure Investigator

What does Test Failure Investigator do?

A skill your agent uses when a test is failing and you need to determine root cause: is it flaky, an environment issue, or a real regression? Test Failure Investigator is an agent skill from proffesor-for-testing/agentic-qe. Use when a test is failing and you need to determine root cause: is it flaky, an environment issue, or a real regression?

When should I use Test Failure Investigator?

Test Failure Investigator fits situations like: A test is failing and you need to determine root cause: is it flaky; an environment issue; A real regression?.

How do I install Test Failure Investigator in Claude Code?

Run `npx skills add proffesor-for-testing/agentic-qe --skill test-failure-investigator -a claude-code`. Or copy the skill folder (assets/skills/test-failure-investigator in proffesor-for-testing/agentic-qe) into .claude/skills/test-failure-investigator in your project. Claude Code loads it when a task matches its description.

How do I install Test Failure Investigator in Codex?

Run `npx skills add proffesor-for-testing/agentic-qe --skill test-failure-investigator -a codex`. Or copy the skill folder (assets/skills/test-failure-investigator in proffesor-for-testing/agentic-qe) into .agents/skills/test-failure-investigator in your project. Codex loads it when a task matches its description.

Can I use Test Failure Investigator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add proffesor-for-testing/agentic-qe --skill test-failure-investigator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-failure-investigator, .gemini/skills/test-failure-investigator, .github/skills/test-failure-investigator and .opencode/skills/test-failure-investigator in your project.

What does Test Failure Investigator need to run?

Going by SKILL.md and its folder, Test Failure Investigator needs the command-line tools its instructions call (git, npx and npm). Our summary lists: Node.js; Docker.

Does Test Failure Investigator access the network?

SKILL.md contains no URLs. Its commands use git, npx and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Failure Investigator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Failure Investigator use?

Test Failure Investigator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Failure Investigator use?

About 811 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Failure Investigator?

Skills that share tags, products or a category with Test Failure Investigator: Trx Analysis (microsoft/vstest, 969 stars), Evanflow Debug (evanklem/evanflow, 418 stars), Issue Tracer (ZaxbyHub/opencode-swarm, 494 stars) and Issue Tracer (ZaxbyHub/opencode-swarm, 494 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Failure Investigator?

proffesor-for-testing (a GitHub user) maintains it in proffesor-for-testing/agentic-qe, which has 495 GitHub stars. The repository holds 93 skills in this directory. The repository was last updated on October 9, 2026.

Source: proffesor-for-testing/agentic-qe on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.