Agent skill

Troubleshoot Cassandra

by Kilo-Org in Kilo-Org/kilo-marketplace

A skill your agent uses when diagnosing issues with Apache Cassandra: gc death spiral, compaction death spiral, tombstone storm, disk space exhaustion, or hint overflow.

Apache-2.0Auto-check passedDatabases

Install Troubleshoot Cassandra

skills CLI
$ npx skills add Kilo-Org/kilo-marketplace --skill troubleshoot-cassandra -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Kilo-Org/kilo-marketplace troubleshoot-cassandra --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/troubleshoot-cassandra .claude/skills/troubleshoot-cassandra && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
troubleshoot-cassandra
GitHub stars
190
Token cost
~2.5k tokens
SKILL.md length
1,024 words
Files
12
Skills in repo
86
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when diagnosing issues with Apache Cassandra: gc death spiral, compaction death spiral, tombstone storm, disk space exhaustion, or hint overflow.

  • Works in 9 steps: Confirm the Apache Cassandra service is… → Pull the last 15 minutes of signals for… → Check for GC Death Spiral. Heap pressure… → …
  • Diagnosing issues with Apache Cassandra: gc death spiral
  • SKILL.md covers When to use this skill, Key facts, Step-by-step and Common mistakes, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Troubleshoot Cassandra is an agent skill from Kilo-Org/kilo-marketplace. Use when diagnosing issues with Apache Cassandra: gc death spiral, compaction death spiral, tombstone storm, disk space exhaustion, or hint overflow. Queries Netdata via MCP for node liveness (failure detector), native transport active, client request rate (read/write), client request latency (coordinator), applies the diagnostic tree from the Netdata operator playbook, and recommends remediation.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files (for example `README.md`, `rules/availability.md` and `rules/cache-performance.md`).

It sits in Databases, covering NoSQL databases. It works with Model Context Protocol. The repository describes itself as: Kilo Marketplace - A curated collection of Skills, MCP Servers, and Modes for enhancing AI agent capabilities across the Kilo ecosystem—including Kilo Code (VS Code extension)… The licence is Apache-2.0.

When your agent uses it

  • Diagnosing issues with Apache Cassandra: gc death spiral
  • Compaction death spiral
  • Tombstone storm
  • Disk space exhaustion

Example prompts

  • “/troubleshoot-cassandra”

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Confirm the Apache Cassandra service is up. Query Netdata via MCP with list_nodes and filter by
  2. Pull the last 15 minutes of signals for the target. Use query_metrics against the contexts
  3. Check for GC Death Spiral. Heap pressure then long GC pauses then gossip failures then node
  4. Check for Compaction Death Spiral. Write rate exceeds compaction throughput then SSTables
  5. Check for Tombstone Storm. Accumulated tombstones (deletes/expired TTLs) force reads to scan
  6. Check for Disk Space Exhaustion. Compaction backlog + snapshots + hints consume space then
  7. Check for Hint Overflow. Long node outage then hints accumulate on coordinators then hints
  8. Correlate with host-level signals (system.cpu.utilization, system.memory.usage,
  9. Apply the remediation hinted at in the matching rule file or the operator playbook. Re-run the

What it can do on your machine

Read from SKILL.md and the folder at commit ff51758. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Troubleshoot Cassandra loads about 2.5k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 1,024 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Kilo-Org/kilo-marketplace at commit ff51758, republished under its Apache-2.0 licence (© Kilo-Org). 1,024 words, ~2,511 tokens.

Download SKILL.mdSave it as .claude/skills/troubleshoot-cassandra/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
troubleshoot-cassandra
description
Use when diagnosing issues with Apache Cassandra: gc death spiral, compaction death spiral, tombstone storm, disk space exhaustion, or hint overflow. Queries Netdata via MCP for node liveness (failure detector), native transport active, client request rate (read/write), client request latency (coordinator), applies the diagnostic tree from the Netdata operator playbook, and recommends remediation.
metadata.category
observability

Troubleshoot Apache Cassandra

When to use this skill

  • GC Death Spiral: Heap pressure then long GC pauses then gossip failures then node marked DOWN then client retries flood then more heap pressure. Self-reinforcing.
  • Compaction Death Spiral: Write rate exceeds compaction throughput then SSTables accumulate then read amplification increases then latency spikes then more compaction needed then disk I/O saturated.
  • Tombstone Storm: Accumulated tombstones (deletes/expired TTLs) force reads to scan massive amounts of dead data then read latency spikes, possible query abortion at 100K tombstones.
  • Disk Space Exhaustion: Compaction backlog + snapshots + hints consume space then compaction cannot run (needs temporary space) then writes blocked.
  • Hint Overflow: Long node outage then hints accumulate on coordinators then hints expire (max_hint_window, default 3h) then data permanently inconsistent unless repaired.
  • Any time the user reports a Apache Cassandra service behaving outside its expected envelope (elevated errors, latency, saturation, resource exhaustion, or unexpected restarts).
  • An on-call engineer is paging on a Netdata alert tied to a Apache Cassandra instance and wants a structured triage path.

Key facts

  • This skill wraps the Netdata operator playbook for Apache Cassandra. It does not replace the playbook; it routes a coding agent through MCP queries against the same signals the playbook relies on.
  • The playbook decomposes Apache Cassandra health into 8 signal domains: Availability, Throughput, Latency, Errors, Saturation / Internal State, Replication / Consistency. Each domain maps to one rule file in this skill.
  • Dominant failure archetypes the playbook calls out: GC Death Spiral; Compaction Death Spiral; Tombstone Storm; Disk Space Exhaustion; Hint Overflow.
  • Netdata observes the signals listed in the rule files via its native collectors, plus any OpenTelemetry-shipped metrics that your Apache Cassandra instrumentation adds. Both paths end at the same MCP query surface.
  • Netdata's cassandra collector emits 28 context(s) under cassandra.*. The rule files enumerate which contexts surface which domain; the Verification section below names the load-bearing ones explicitly.

Step-by-step

  1. Confirm the Apache Cassandra service is up. Query Netdata via MCP with list_nodes and filter by the host running the target. A missing node means the symptom is at the network or orchestrator layer, not inside the service.
  2. Pull the last 15 minutes of signals for the target. Use query_metrics against the contexts listed in the domain rule files. Run find_anomalous_metrics in parallel over the same window; anomalies frame which rule file to read first.
  3. Check for GC Death Spiral. Heap pressure then long GC pauses then gossip failures then node marked DOWN then client retries flood then more heap pressure. Self-reinforcing. Inspect the rule file whose signals move first for this mode.
  4. Check for Compaction Death Spiral. Write rate exceeds compaction throughput then SSTables accumulate then read amplification increases then latency spikes then more compaction needed then disk I/O saturated. Inspect the rule file whose signals move first for this mode.
  5. Check for Tombstone Storm. Accumulated tombstones (deletes/expired TTLs) force reads to scan massive amounts of dead data then read latency spikes, possible query abortion at 100K tombstones. Inspect the rule file whose signals move first for this mode.
  6. Check for Disk Space Exhaustion. Compaction backlog + snapshots + hints consume space then compaction cannot run (needs temporary space) then writes blocked. Inspect the rule file whose signals move first for this mode.
  7. Check for Hint Overflow. Long node outage then hints accumulate on coordinators then hints expire (max_hint_window, default 3h) then data permanently inconsistent unless repaired. Inspect the rule file whose signals move first for this mode.
  8. Correlate with host-level signals (system.cpu.utilization, system.memory.usage, system.disk.io_time). Many service-level failures have a host-resource precursor.
  9. Apply the remediation hinted at in the matching rule file or the operator playbook. Re-run the MCP queries from the Verification section to confirm the signals returned to expected ranges. A fix that does not move the signal back is not a fix.
Show full SKILL.md (390 more words)Show less
Handy MCP call templates
text
# Discover metrics from Apache Cassandra
list_metrics with q="cassandra"

# Pull a specific context over the last window
query_metrics with context="cassandra.dropped_messages_rate", relative_window=-15m

# Rank anomalies for the service or host
find_anomalous_metrics with node=<host> and context_pattern="cassandra.*"

# Correlate a known problem context with others
find_correlated_metrics around the incident window

# Show current alert state
list_raised_alerts scoped to the node

Common mistakes

  • Treating Apache Cassandra as a generic HTTP or process health check. Apache Cassandra has specific failure archetypes (see Key facts) that generic checks miss.
  • Stopping at the first anomalous metric. Several archetypes produce correlated spikes; use find_correlated_metrics to widen the search before concluding a root cause.
  • Quoting percentile latency without the sample count. Low traffic plus a single slow request moves p99 by seconds.
  • Reading dashboards for a window shorter than the failure's fingerprint. Slow-brew failures (queue growth, bloat, memory fragmentation) need 30+ minutes of data to see the trend.
  • Skipping the host-level correlation. A process-level fix for a noisy-neighbour problem does not hold.
  • Assuming alert thresholds are tuned for your workload. Tune against observed Apache Cassandra traffic before escalating an alert configuration issue.

Verification

Run these MCP queries against the Netdata instance that sees the Apache Cassandra service. Every context listed below is a real Netdata chart name; the agent does not need to guess.

text
1. list_metrics filtered by q="cassandra" (returns every cassandra.* context Netdata sees)
2. query_metrics with contexts=[cassandra.dropped_messages_rate, cassandra.client_requests_timeouts_rate, cassandra.client_requests_failures_rate, cassandra.client_requests_rate, cassandra.client_requests_latency, cassandra.row_cache_hit_rate] and relative_window=-30m
3. find_anomalous_metrics filtered by node=<host> and context_pattern="cassandra.*"

Load-bearing contexts for this service:

  • cassandra.dropped_messages_rate: Dropped messages rate (messages/s). Dimensions: dropped.
  • cassandra.client_requests_timeouts_rate: Client requests timeouts rate (timeout/s). Dimensions: read, write.
  • cassandra.client_requests_failures_rate: Client requests failures rate (failures/s). Dimensions: read, write.
  • cassandra.client_requests_rate: Client requests rate (requests/s). Dimensions: read, write.
  • cassandra.client_requests_latency: Client requests total latency (seconds). Dimensions: read, write.
  • cassandra.row_cache_hit_rate: Key cache hit rate (events/s). Dimensions: hits, misses.

A clean result means every context is within its expected band and the find_anomalous_metrics list is empty or contains only already-acknowledged items. If the fix was real, re-running the same queries 10 minutes after applying it will show a clean result. If it does not, revert and look deeper.

When the fix does not hold

If signals drift back into the anomalous range within 30 minutes of a remediation, the cause was deeper than the applied change. Typical misdiagnoses for Apache Cassandra:

  • Host-resource pressure masquerading as application bug.
  • Dependent service (DB, cache, upstream) causing a secondary symptom in the instrumented service.
  • Configuration change that was never reloaded (some subsystems only pick up config on full restart).

Escalate by widening the query window: 2-6 hours instead of 15 minutes. Slow-moving causes are invisible at triage window sizes.

References

© Kilo-Org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files in skills/troubleshoot-cassandra of Kilo-Org/kilo-marketplace.

  • SKILL.md
  • LICENSE
  • README.md
  • rules/availability.md
  • rules/cache-performance.md
  • rules/errors.md
  • rules/latency.md
  • rules/other-contexts.md
  • rules/replication-consistency.md
  • rules/saturation-internal-state.md
  • rules/security-integrity.md
  • rules/throughput.md

Open the folder on GitHubat commit ff51758

Compare with similar skills

Troubleshoot Cassandra next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Troubleshoot Cassandra compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Troubleshoot Cassandra this skillKilo-Org/kilo-marketplace190—~2.5kAutomated safety check: PassApache-2.0
Azure Storagemicrosoft/GitHub-Copilot-for-Azure2552 repos~1.3kAutomated safety check: PassMIT
Mongodb Search And AImongodb/agent-skills1891 repos~1.7kAutomated safety check: PassApache-2.0
Cloudbase CLITencentCloudBase/CloudBase-AI-Toolkit1.1k1 repos~1.7kAutomated safety check: PassMIT
Ak Add Capabilitiesyaalalabs/agent-kernel192—~13kAutomated safety check: PassApache-2.0
AWS Storageaws/agent-toolkit-for-aws2.8k—~5.8kAutomated safety check: PassApache-2.0

Similar skills

  • Azure Storage

    microsoft/GitHub-Copilot-for-Azure

    Official

    Azure Storage Services including Blob Storage, File Shares, Queue Storage, Table Storage, and Data Lake.

    255 GitHub starsUsed in 2 repos~1.3k tokens
    DatabasesAuto-check passed
  • Mongodb Search And AI

    mongodb/agent-skills

    Official

    Guides MongoDB users through implementing and optimizing Atlas Search (full-text), Vector Search (semantic), and Hybrid Search solutions.

    189 GitHub starsUsed in 1 repo~1.7k tokens
    DatabasesAuto-check passed
  • Cloudbase CLI

    TencentCloudBase/CloudBase-AI-Toolkit

    CloudBase CLI (tcb, 云开发CLI, Tencent CloudBase命令行) resource management skill.

    1.1k GitHub starsUsed in 1 repo~1.7k tokens
    DatabasesAuto-check passed
  • Ak Add Capabilities

    yaalalabs/agent-kernel

    Add capabilities to an existing Agent Kernel project. An agent skill from yaalalabs/agent-kernel.

    192 GitHub stars~13k tokensUpdated yesterday
    DatabasesAuto-check passed
  • AWS Storage

    aws/agent-toolkit-for-aws

    Official

    Selects, investigates, and compares AWS object, file, and block storage services, and answers cost, performance, configuration, security, and troubleshooting questions about storage services.

    2.8k GitHub stars~5.8k tokensUpdated today
    DatabasesAuto-check passed
  • Mongodb Query Optimizer

    mongodb/agent-skills

    Official

    Help with MongoDB query optimization and indexing. An agent skill from mongodb/agent-skills.

    189 GitHub starsUsed in 2 repos~2.6k tokens
    DatabasesAuto-check passed

More from Kilo-Org/kilo-marketplace

All 86 skills in this repo
  • AzureML Project Scaffolding

    Kilo-Org/kilo-marketplace

    Sets up and maintains AzureML-ready Python projects as uv workspaces with devcontainers, a Makefile and job YAML, so local runs match cloud jobs and experiments stay reproducible.

    190 GitHub stars~3.1k tokensUpdated 11 days ago
    Auto-check: notes
  • Jupyter Notebook Builder

    Kilo-Org/kilo-marketplace

    Creates, inspects, edits and runs Jupyter notebooks, scaffolding experiment or tutorial notebooks from templates and preferring a Jupyter MCP server over raw JSON edits.

    190 GitHub stars~1.3k tokensUpdated 11 days ago
    Auto-check passed
  • Tableau Dashboard Creator

    Kilo-Org/kilo-marketplace

    Takes a plain-language dashboard request through brand setup, data exploration, planning, an interactive HTML mock and a Tableau implementation spec.

    190 GitHub stars~3.8k tokensUpdated 11 days ago
    Auto-check: notes
  • Elasticsearch File Ingest

    Kilo-Org/kilo-marketplace

    Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.

    190 GitHub stars~2.8k tokensUpdated 11 days ago
    Auto-check passed
  • Nifi Flow Layout

    Kilo-Org/kilo-marketplace

    A skill your agent uses when arranging Apache NiFi processors, process groups, ports, comments, numbering, crossing connections, dense fan-in/fan-out, or reusable readable canvas layouts.

    190 GitHub stars~1.5k tokensUpdated 11 days ago
    Auto-check passed
  • Splunk Ingest Processor Setup

    Kilo-Org/kilo-marketplace

    Render Cisco Data Fabric ingest-time routing workflows and Splunk Cloud Platform Ingest Processor setup plans with SPL2 pipelines, source types, destinations, lifecycle handoffs, queue and…

    190 GitHub stars~1.2k tokensUpdated 11 days ago
    Auto-check passed

Categories

Questions about Troubleshoot Cassandra

What does Troubleshoot Cassandra do?

A skill your agent uses when diagnosing issues with Apache Cassandra: gc death spiral, compaction death spiral, tombstone storm, disk space exhaustion, or hint overflow. Troubleshoot Cassandra is an agent skill from Kilo-Org/kilo-marketplace. Use when diagnosing issues with Apache Cassandra: gc death spiral, compaction death spiral, tombstone storm, disk space exhaustion, or hint overflow.

When should I use Troubleshoot Cassandra?

Troubleshoot Cassandra fits situations like: diagnosing issues with Apache Cassandra: gc death spiral; compaction death spiral; tombstone storm; disk space exhaustion.

How do I install Troubleshoot Cassandra in Claude Code?

Run `npx skills add Kilo-Org/kilo-marketplace --skill troubleshoot-cassandra -a claude-code`. Or copy the skill folder (skills/troubleshoot-cassandra in Kilo-Org/kilo-marketplace) into .claude/skills/troubleshoot-cassandra in your project. Claude Code loads it when a task matches its description.

How do I install Troubleshoot Cassandra in Codex?

Run `npx skills add Kilo-Org/kilo-marketplace --skill troubleshoot-cassandra -a codex`. Or copy the skill folder (skills/troubleshoot-cassandra in Kilo-Org/kilo-marketplace) into .agents/skills/troubleshoot-cassandra in your project. Codex loads it when a task matches its description.

Can I use Troubleshoot Cassandra in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Kilo-Org/kilo-marketplace --skill troubleshoot-cassandra -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/troubleshoot-cassandra, .gemini/skills/troubleshoot-cassandra, .github/skills/troubleshoot-cassandra and .opencode/skills/troubleshoot-cassandra in your project.

What does Troubleshoot Cassandra need to run?

SKILL.md names no scripts, command-line tools or credentials: Troubleshoot Cassandra is instructions for the agent only.

Does Troubleshoot Cassandra access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Troubleshoot Cassandra safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Troubleshoot Cassandra use?

Troubleshoot Cassandra is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Troubleshoot Cassandra use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Troubleshoot Cassandra?

Skills that share tags, products or a category with Troubleshoot Cassandra: Azure Storage (microsoft/GitHub-Copilot-for-Azure, 255 stars), Mongodb Search And AI (mongodb/agent-skills, 189 stars), Cloudbase CLI (TencentCloudBase/CloudBase-AI-Toolkit, 1.1k stars) and Ak Add Capabilities (yaalalabs/agent-kernel, 192 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Troubleshoot Cassandra?

Kilo-Org (a GitHub organization) maintains it in Kilo-Org/kilo-marketplace, which has 190 GitHub stars. The repository holds 86 skills in this directory. The repository was last updated on September 28, 2026.

Source: Kilo-Org/kilo-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.