Agent skill

Must Gather Investigation

by scylladb in scylladb/scylla-operator

Investigate failed e2e tests from Ginkgo JSON reports and must-gather artifacts, systematically analyzing logs, events, and resource states to identify root causes.

Apache-2.0Auto-check passedTesting & QA

Install Must Gather Investigation

skills CLI
$ npx skills add scylladb/scylla-operator --skill must-gather-investigation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scylladb/scylla-operator must-gather-investigation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scylladb/scylla-operator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/must-gather-investigation .claude/skills/must-gather-investigation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
must-gather-investigation
GitHub stars
401
Token cost
~2.4k tokens
SKILL.md length
780 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Investigate failed e2e tests from Ginkgo JSON reports and must-gather artifacts, systematically analyzing logs, events, and resource states to identify root causes.

  • Works in 7 steps: Extract Failed Tests from e2e.json → Locate the Test Source Code → Examine Test Namespace Artifacts → …
  • Tasks that involve End-to-end testing
  • SKILL.md covers Inputs, Phase 1: Extract Failed Tests…, Phase 2: Locate the Test… and Phase 3: Examine Test…, plus 6 more sections
  • Calls jq

What it does

Must Gather Investigation is an agent skill from scylladb/scylla-operator. Investigate failed e2e tests from Ginkgo JSON reports and must-gather artifacts, systematically analyzing logs, events, and resource states to identify root causes.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering End-to-end testing, Root cause analysis and Container orchestration. It works with Kubernetes. The repository describes itself as: The Kubernetes Operator for ScyllaDB. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve End-to-end testing
  • Tasks that involve Root cause analysis
  • Tasks that involve Container orchestration

Example prompts

  • “/must-gather-investigation”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Extract Failed Tests from e2e.json
  2. Locate the Test Source Code
  3. Examine Test Namespace Artifacts
  4. Examine Operator Logs
  5. Examine Infrastructure Logs
  6. Timeline Reconstruction
  7. Root Cause Analysis

What it can do on your machine

Read from SKILL.md and the folder at commit 57e552d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Must Gather Investigation loads about 2.4k tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 780 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scylladb/scylla-operator at commit 57e552d, republished under its Apache-2.0 licence (© scylladb). 780 words, ~2,388 tokens.

Download SKILL.mdSave it as .claude/skills/must-gather-investigation/SKILL.md (or your agent's skills folder).
name
must-gather-investigation
description
Investigate failed e2e tests from Ginkgo JSON reports and must-gather artifacts, systematically analyzing logs, events, and resource states to identify root causes.
metadata.audience
maintainers

Must-Gather Investigation

You are a Kubernetes operator e2e test failure investigator. Your goal is to analyze test artifacts from a failed test run, reconstruct the sequence of events, and identify the root cause.

Inputs

The user provides:

  • A path to the root directory containing e2e.json (a Ginkgo JSON report) and associated artifacts. All artifact paths below are relative to this root directory. You should assume you are already in this directory.
  • Optionally, a specific test name to investigate

Phase 1: Extract Failed Tests from e2e.json

The e2e.json file is a Ginkgo JSON report. Use jq to navigate it.

Key jq paths

List all failed tests:

bash
jq -r '.[] | .SpecReports[] | select(.State == "failed") | .ContainerHierarchyTexts + [.LeafNodeText] | join(" > ")' e2e.json

Extract failure details for a specific failed test:

bash
jq '.[] | .SpecReports[] | select(.State == "failed") | {
  name: (.ContainerHierarchyTexts + [.LeafNodeText] | join(" > ")),
  state: .State,
  startTime: .StartTime,
  endTime: .EndTime,
  runTime: .RunTime,
  failureMessage: .Failure.Message,
  failureLocation: (.Failure.Location.FileName + ":" + (.Failure.Location.LineNumber | tostring)),
  ginkgoOutput: .CapturedGinkgoWriterOutput
}' e2e.json

Extract SpecEvents timeline for a failed test:

bash
jq '.[] | .SpecReports[] | select(.State == "failed") | .SpecEvents[] | {type: .SpecEventType, message: .Message, duration: .Duration, codeLocation: (.CodeLocation.FileName + ":" + (.CodeLocation.LineNumber | tostring))}' e2e.json
Key fields in a SpecReport
FieldDescription
.State"passed", "failed", "skipped", "pending"
.ContainerHierarchyTextsArray of Describe/Context block names (outermost first)
.LeafNodeTextThe It block name
.StartTime / .EndTimeISO 8601 timestamps
.RunTimeDuration in nanoseconds
.Failure.MessageThe assertion error message
.Failure.Location{FileName, LineNumber} of the failing assertion
.CapturedGinkgoWriterOutputAll GinkgoWriter output during the test (contains namespace names, resource names, log lines)
.SpecEvents[]Timeline of By() steps, DeferCleanup calls, etc. Each has .SpecEventType, .Message, .Duration, .CodeLocation
Extracting the test namespace

The test namespace is typically logged in CapturedGinkgoWriterOutput. Search for patterns like e2e-test-* or look for lines containing "namespace" or "ns".

bash
jq -r '.[] | .SpecReports[] | select(.State == "failed") | .CapturedGinkgoWriterOutput' e2e.json | grep -oE 'e2e-test-[a-z0-9-]+'

Phase 2: Locate the Test Source Code

Search test/e2e/ in the repository for the test name (the LeafNodeText or a unique substring from it):

bash
grep -rn "LEAF_NODE_TEXT_SUBSTRING" test/e2e/

Read the test to understand:

  • What resources it creates (ScyllaCluster, ScyllaDBDatacenter, etc.)
  • What it waits for (rollout, conditions, specific states)
  • What it asserts (cleanup jobs completed, connections succeed, etc.)
  • Timeout values and polling intervals
  • Any Eventually/Consistently blocks — these are where timeouts cause failures

Phase 3: Examine Test Namespace Artifacts

Navigate to e2e/cluster/namespaces/<test-namespace>/. This contains the state of all resources in the test namespace at the time the must-gather was collected (typically after the test failed, during namespace teardown).

Priority order for examination
  1. Events (events.events.k8s.io/*.yaml): Chronological record of what happened. Look for warnings, errors, and unusual sequences.

  2. ScyllaCluster / ScyllaDBDatacenter status (scyllaclusters.scylla.scylladb.com/*.yaml or scylladbdatacenters.scylla.scylladb.com/*.yaml): Check .status.conditions — especially Available, Progressing, Degraded. The reason and message fields explain why a condition is set.

  3. Pod status and logs (pods/<pod-name>/):

    • <container>.current — current container logs
    • <container>.terminated — logs from a previous container instance (if it restarted)
    • <pod-name>.yaml — full pod spec and status, including conditions, container states, restart counts
    • df.log — disk usage (for Scylla data pods)
    • nodetool-status.log — Scylla cluster membership
    • nodetool-gossipinfo.log — Scylla gossip state
  4. Jobs (jobs/*.yaml): Check .status for completionTime, conditions, ready, active, failed counts. Compare job UIDs with pod controller-uid labels to verify ownership.

  5. Services (services/*.yaml): Check annotations — CurrentTokenRingHash, LastCleanedUpTokenRingHash, HostID, etc. Compare across nodes.

  6. StatefulSets (statefulsets.apps/*.yaml): Check .status.readyReplicas, .status.currentRevision, .status.updateRevision.

  7. Other resources: ConfigMaps, Secrets, PVCs, Ingresses, EndpointSlices — as relevant to the test.

Show full SKILL.md (315 more words)Show less

Phase 4: Examine Operator Logs

Operator logs are at must-gather/cluster/namespaces/scylla-operator/pods/<operator-pod>/scylla-operator.current.

These are structured JSON logs (one JSON object per line). Key fields:

  • "ts" — timestamp
  • "msg" — log message
  • "controller" — which controller emitted the log
  • "namespace" / "name" — the resource being reconciled
  • "err" — error details

Filter by the test namespace to find relevant reconciliation activity:

bash
grep '<test-namespace>' scylla-operator.current

Look for:

  • Reconciliation start/end and duration
  • Error messages or warnings
  • Resource creation, update, deletion events
  • Status condition changes
  • Queuing and re-queuing patterns

Phase 5: Examine Infrastructure Logs

Depending on the test, check logs from infrastructure components:

  • HAProxy ingress (must-gather/cluster/namespaces/haproxy-ingress/): Backend configuration, reload events, connection logs
  • Scylla Manager (must-gather/cluster/namespaces/scylla-manager/): Task scheduling, repair/backup operations
  • cert-manager: Certificate issuance and renewal
  • Scylla Manager Agent (sidecar in Scylla pods, scylla-manager-agent.current): API calls, health checks

Phase 6: Timeline Reconstruction

Build a chronological timeline from all log sources, correlating timestamps. Include:

  • Pod lifecycle events (created, scheduled, started, ready)
  • Controller reconciliation actions
  • Resource state changes
  • The test's own actions (from CapturedGinkgoWriterOutput and SpecEvents)
  • Infrastructure events (reloads, connection attempts)

This timeline is the core artifact for identifying the root cause. It should make the causal chain visible.

Phase 7: Root Cause Analysis

Trace the causal chain from the failure backward:

  1. What assertion failed, and what was the actual vs expected state?
  2. Why was the resource/condition in that state?
  3. What controller/component was responsible for getting it to the expected state?
  4. What prevented it from doing so?
  5. Was it a timing issue, a logic bug, an infrastructure failure, or a test design issue?

Artifact Structure Reference

./
├── e2e.json                          # Ginkgo JSON test report
├── junit.e2e.xml                     # JUnit XML test report
├── deploy/                           # Deployment manifests used
│   ├── operator/
│   ├── manager/
│   ├── prometheus-operator/
│   └── haproxy-ingress/
├── e2e/cluster/                      # Resources collected during test execution
│   ├── cluster-scoped/               # Cluster-wide resources
│   │   ├── nodes/
│   │   ├── persistentvolumes/
│   │   └── ...
│   └── namespaces/
│       └── <test-namespace>/         # Test-specific namespace
│           ├── pods/
│           │   └── <pod-name>/
│           │       ├── <container>.current          # Container logs
│           │       ├── <container>.terminated       # Previous container logs
│           │       ├── df.log                       # Disk usage
│           │       ├── nodetool-gossipinfo.log      # Scylla gossip info
│           │       └── nodetool-status.log          # Scylla cluster status
│           ├── events.events.k8s.io/
│           ├── statefulsets.apps/
│           ├── jobs/
│           ├── services/
│           ├── configmaps/
│           ├── secrets/
│           ├── scyllaclusters.scylla.scylladb.com/
│           └── scylladbdatacenters.scylla.scylladb.com/
└── must-gather/cluster/              # Must-gather output
    ├── cluster-scoped/
    └── namespaces/
        ├── scylla-operator/
        │   ├── pods/
        │   │   └── <operator-pod>/
        │   │       └── scylla-operator.current      # Operator logs
        │   ├── events.events.k8s.io/
        │   └── ...
        ├── scylla-manager/
        ├── haproxy-ingress/
        └── ...

Important Notes

  • Always reference specific file paths, line numbers, and timestamps when citing evidence.
  • Focus on the causal chain — what led to what.
  • Consider whether the issue is in the test, the operator, or the infrastructure.
  • If you need more information from a file, say what you need and why.
  • This skill provides investigation methodology only. The calling agent determines the output format.

© scylladb, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/must-gather-investigation of scylladb/scylla-operator.

Open the folder on GitHubat commit 57e552d

Compare with similar skills

Must Gather Investigation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Must Gather Investigation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Must Gather Investigation this skillscylladb/scylla-operator401—~2.4kAutomated safety check: PassApache-2.0
Kaniop Developmentpando85/kaniop132—~3.1kAutomated safety check: PassAGPL-3.0
Unblock Dependabot PRkubernetes-sigs/cloud-provider-azure294—~1.5kAutomated safety check: PassApache-2.0
Run E2E Testkubernetes-sigs/cloud-provider-azure294—~3.8kAutomated safety check: PassApache-2.0
Debug E2E Pipelinekubernetes-sigs/cloud-provider-azure294—~3.4kAutomated safety check: PassApache-2.0
Azure Diagnosticsmicrosoft/GitHub-Copilot-for-Azure255—~1.9kAutomated safety check: PassMIT

Similar skills

  • Kaniop Development

    pando85/kaniop

    Kaniop architecture, Rust controller conventions, commands, testing, and repository workflows.

    132 GitHub stars~3.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Unblock Dependabot PR

    kubernetes-sigs/cloud-provider-azure

    Official

    Diagnose and unblock failed Dependabot pull requests in cloud-provider-azure by closing Kubernetes minor-version dependency bumps, classifying CI failures, syncing Go modules, retesting quota-flaked…

    294 GitHub stars~1.5k tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Run E2E Test

    kubernetes-sigs/cloud-provider-azure

    Official

    Parse a Go e2e test from tests/e2e/, translate each step to kubectl and az CLI commands, and interactively replay the test against a live cluster.

    294 GitHub stars~3.8k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Debug E2E Pipeline

    kubernetes-sigs/cloud-provider-azure

    Official

    Fetch and analyze Prow e2e pipeline failures for cloud-provider-azure.

    294 GitHub stars~3.4k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Azure Diagnostics

    microsoft/GitHub-Copilot-for-Azure

    Official

    Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage.

    255 GitHub stars~1.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Multi Cluster API Data Mismatch

    divinevideo/divine-mobile

    Debug "API returns data that doesn't exist in database" when multiple Kubernetes clusters exist (production, staging, POC).

    266 GitHub stars~1.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from scylladb/scylla-operator

  • Release Notes Generator

    scylladb/scylla-operator

    An experienced Kubernetes operator developer agent that generates structured, concise release notes for the ScyllaDB Operator by analyzing git history and PR context.

    401 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Must Gather Investigation

What does Must Gather Investigation do?

Investigate failed e2e tests from Ginkgo JSON reports and must-gather artifacts, systematically analyzing logs, events, and resource states to identify root causes. Must Gather Investigation is an agent skill from scylladb/scylla-operator. Investigate failed e2e tests from Ginkgo JSON reports and must-gather artifacts, systematically analyzing logs, events, and resource states to identify root causes.

When should I use Must Gather Investigation?

Must Gather Investigation fits situations like: tasks that involve End-to-end testing; tasks that involve Root cause analysis; tasks that involve Container orchestration.

How do I install Must Gather Investigation in Claude Code?

Run `npx skills add scylladb/scylla-operator --skill must-gather-investigation -a claude-code`. Or copy the skill folder (.claude/skills/must-gather-investigation in scylladb/scylla-operator) into .claude/skills/must-gather-investigation in your project. Claude Code loads it when a task matches its description.

How do I install Must Gather Investigation in Codex?

Run `npx skills add scylladb/scylla-operator --skill must-gather-investigation -a codex`. Or copy the skill folder (.claude/skills/must-gather-investigation in scylladb/scylla-operator) into .agents/skills/must-gather-investigation in your project. Codex loads it when a task matches its description.

Can I use Must Gather Investigation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scylladb/scylla-operator --skill must-gather-investigation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/must-gather-investigation, .gemini/skills/must-gather-investigation, .github/skills/must-gather-investigation and .opencode/skills/must-gather-investigation in your project.

What does Must Gather Investigation need to run?

Going by SKILL.md and its folder, Must Gather Investigation needs the command-line tools its instructions call (jq).

Does Must Gather Investigation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Must Gather Investigation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Must Gather Investigation use?

Must Gather Investigation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Must Gather Investigation use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Must Gather Investigation?

Skills that share tags, products or a category with Must Gather Investigation: Kaniop Development (pando85/kaniop, 132 stars), Unblock Dependabot PR (kubernetes-sigs/cloud-provider-azure, 294 stars), Run E2E Test (kubernetes-sigs/cloud-provider-azure, 294 stars) and Debug E2E Pipeline (kubernetes-sigs/cloud-provider-azure, 294 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Must Gather Investigation?

scylladb (a GitHub organization) maintains it in scylladb/scylla-operator, which has 401 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 10, 2026.

Source: scylladb/scylla-operator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.