Agent skill

Clickhouse Incident Runbook

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

ClickHouse incident response — triage, diagnose, and remediate server issues using system tables, kill stuck queries, and execute recovery procedures.

MITAuto-check passedDevOps & Cloud

Install Clickhouse Incident Runbook

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill clickhouse-incident-runbook -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace clickhouse-incident-runbook --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/clickhouse-incident-runbook .claude/skills/clickhouse-incident-runbook && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
clickhouse-incident-runbook
GitHub stars
2.8k
Token cost
~1.7k tokens
SKILL.md length
469 words
Files
3 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

ClickHouse incident response — triage, diagnose, and remediate server issues using system tables, kill stuck queries, and execute recovery procedures.

  • Works in 4 steps: Quick triage (run first) → Decision tree — classify the failure → Apply the matching remediation → …
  • ClickHouse is slow
  • SKILL.md covers Overview, Prerequisites, Severity Levels and Instructions, plus 4 more sections
  • Calls curl; reaches status.clickhouse.cloud

What it does

Clickhouse Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. ClickHouse incident response — triage, diagnose, and remediate server issues using system tables, kill stuck queries, and execute recovery procedures. Use when ClickHouse is slow, unresponsive, OOM-killed, out of disk, backing up merges, or producing errors in production and you need an on-call playbook. Trigger with "clickhouse incident", "clickhouse outage", "clickhouse down", "clickhouse emergency", "clickhouse on-call", "clickhouse broken".

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/evidence-and-comms.md` and `references/remediation-procedures.md`). Compatibility notes: Designed for Claude Code

It sits in DevOps & Cloud, covering Data warehousing, Incident response and Runbooks and postmortems. It works with ClickHouse. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • ClickHouse is slow
  • Backing up merges
  • Producing errors in production and you need an on-call playbook
  • With clickhouse incident

Example prompts

  • “clickhouse incident”
  • “clickhouse outage”
  • “clickhouse down”
  • “/clickhouse-incident-runbook”

Requirements

  • Docker
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Bash(kubectl:*), Bash(curl:*)

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Quick triage (run first)
  2. Decision tree — classify the failure
  3. Apply the matching remediation
  4. Collect evidence and communicate

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash(kubectl:*)
    • Bash(curl:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • status.clickhouse.cloud

    Also links to:

    • clickhouse.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Clickhouse Incident Runbook loads about 1.7k tokens when it runs, and up to ~2.7k if it reads all its reference files. Until then it costs about 119 tokens; SKILL.md has 469 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 469 words, ~1,651 tokens.

Download SKILL.mdSave it as .claude/skills/clickhouse-incident-runbook/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
clickhouse-incident-runbook
description
ClickHouse incident response — triage, diagnose, and remediate server issues using system tables, kill stuck queries, and execute recovery procedures. Use when ClickHouse is slow, unresponsive, OOM-killed, out of disk, backing up merges, or producing errors in production and you need an on-call playbook. Trigger with "clickhouse incident", "clickhouse outage", "clickhouse down", "clickhouse emergency", "clickhouse on-call", "clickhouse broken".
allowed-tools
Read, Bash(kubectl:*), Bash(curl:*)
compatibility
Designed for Claude Code
version
1.7.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, database, analytics, clickhouse, olap

ClickHouse Incident Runbook

Overview

Step-by-step procedures for triaging and resolving ClickHouse incidents using built-in system tables and SQL commands. Start here: assess severity, run quick triage, walk the decision tree, then jump to the matching remediation procedure.

Prerequisites

  • Network access to the ClickHouse HTTP interface (default port 8123) or a working clickhouse-client.
  • A user with rights to read system.* tables and issue KILL QUERY / ALTER.
  • Shell access to the host or container for P1 restarts (systemctl, docker, or kubectl).

Severity Levels

LevelDefinitionResponseExamples
P1ClickHouse unreachable / all queries failing< 15 minServer down, OOM, disk full
P2Degraded performance / partial failures< 1 hourSlow queries, merge backlog
P3Minor impact / non-critical errors< 4 hoursSingle table issue, warnings
P4No user impactNext business dayMonitoring gaps, optimization

Instructions

Work the incident top to bottom: triage, classify with the decision tree, then apply the matching procedure.

1. Quick triage (run first)
bash
# 1. Is ClickHouse alive? (8123 is the default ClickHouse HTTP interface port)
curl -sf 'http://localhost:8123/ping' && echo "UP" || echo "DOWN"

# 2. Can it answer a query?
curl -sf 'http://localhost:8123/?query=SELECT+1' && echo "OK" || echo "QUERY FAILED"

# 3. Check ClickHouse Cloud status
curl -sf 'https://status.clickhouse.cloud' | head -5
sql
-- 4. Server health snapshot (run if server responds)
SELECT
    version()                          AS version,
    formatReadableTimeDelta(uptime())  AS uptime,
    (SELECT count() FROM system.processes) AS running_queries,
    (SELECT value FROM system.metrics WHERE metric = 'MemoryTracking')
        AS memory_bytes,
    (SELECT count() FROM system.merges) AS active_merges;

-- 5. Recent errors
SELECT event_time, exception_code, exception, substring(query, 1, 200) AS q
FROM system.query_log
WHERE type = 'ExceptionWhileProcessing'
  AND event_time >= now() - INTERVAL 10 MINUTE
ORDER BY event_time DESC
LIMIT 10;
2. Decision tree — classify the failure
Server responds to ping?
├─ NO → Check process/container status, disk space, OOM killer logs
│       └─ Container/process dead → Restart, check logs
│       └─ Disk full → Emergency: drop old partitions, expand disk
│       └─ OOM killed → Reduce max_memory_usage, add RAM
└─ YES → Queries succeeding?
    ├─ NO → Check error codes below
    │   └─ Auth errors (516) → Verify credentials, check user exists
    │   └─ Too many queries (202) → Kill stuck queries, reduce concurrency
    │   └─ Memory exceeded (241) → Kill large queries, reduce max_threads
    └─ YES but slow → Performance triage below
3. Apply the matching remediation

Each branch maps to a full procedure (SQL + shell, copy-paste ready) in references/remediation-procedures.md:

  • P1: Server down / OOM — inspect dmesg/journalctl, restart, verify.
  • P1: Disk full — find largest tables, drop old partitions, check system.disks.
  • P2: Stuck queries — inspect system.processes, KILL QUERY by id/user/elapsed.
  • P2: Too many parts — check part counts, raise parts_to_throw_insert, batch inserts.
  • P2: Memory pressure — rank by memory_usage, kill the largest, cap max_memory_usage.
  • P3: Replication lag — inspect system.replicas for queue and replica gaps.
4. Collect evidence and communicate

Once mitigated, export the error window and post status updates. Templates and INTO OUTFILE exports are in references/evidence-and-comms.md.

Show full SKILL.md (209 more words)Show less

Output

Working through this runbook produces:

  • A severity classification (P1–P4) and the identified failure class.
  • Remediation actions applied — killed query ids, dropped partitions, restarted service, or adjusted settings.
  • A recovery confirmation (SELECT version() / SELECT 1 succeeds again).
  • Forensic artifacts for the postmortem: /tmp/incident-queries.json and /tmp/incident-metrics.tsv, plus a filled-in postmortem document.

Error Handling

SymptomLikely CauseFirst Action
All queries failServer downCheck process, restart
Inserts failToo many partsKILL QUERY long merges, raise limit
Selects slowMemory pressureKill large queries, add filters
Disk alertsNo TTL / no cleanupDrop old partitions
Replication lagNetwork / merge backlogCheck system.replicas

If the server does not respond to ping at all, do not keep issuing SQL — move straight to the P1 host-level checks (process, disk, OOM logs) before anything else.

Examples

Kill a runaway query (P2). Triage shows one query pinning memory; classify as "queries succeeding but slow", then from the stuck-query procedure:

sql
KILL QUERY WHERE query_id = 'abc-123-def';

Emergency disk reclaim (P1). Ping fails and the host is out of disk; the disk-full procedure drops the oldest partition to restore writes:

sql
ALTER TABLE analytics.events DROP PARTITION '202301';

Full multi-step walkthroughs for every severity live in references/remediation-procedures.md; the post-incident export and comms templates live in references/evidence-and-comms.md.

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/.curated/clickhouse-incident-runbook of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/evidence-and-comms.md
  • references/remediation-procedures.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Clickhouse Incident Runbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Clickhouse Incident Runbook compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Clickhouse Incident Runbook this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.7kAutomated safety check: PassMIT
Incident Responsechmonitor/chmonitor300—~3.4kAutomated safety check: PassGPL-3.0
Monitoring Ingestion PipelinePostHog/posthog40k—~9.1kAutomated safety check: PassCustom licence
Verify Productionchmonitor/chmonitor300—~2.2kAutomated safety check: NotesGPL-3.0
Cluster Operationschmonitor/chmonitor300—~634Automated safety check: PassGPL-3.0
Funnelcake Deployment Workflowdivinevideo/divine-mobile266—~3.6kAutomated safety check: PassMPL-2.0

Similar skills

  • Incident Response

    chmonitor/chmonitor

    Structured triage recipes for common ClickHouse incidents: disk, errors, replication, mutations, cluster health, and slow queries.

    300 GitHub stars~3.4k tokensUpdated 5 days ago
    DatabasesAuto-check passed
  • Official

    Guide for using the Grafana MCP to monitor and diagnose the Node.js ingestion pipeline workers in production.

    40k GitHub stars~9.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Verify Production

    chmonitor/chmonitor

    Prove the deployed chmonitor product actually works, not just that the Worker answers.

    300 GitHub stars~2.2k tokensUpdated 5 days ago
    DevOps & CloudAuto-check: notes
  • Cluster Operations

    chmonitor/chmonitor

    Cluster management: distributed tables, ON CLUSTER DDL, node lifecycle, resharding, load balancing, and Keeper migration.

    300 GitHub stars~634 tokensUpdated 5 days ago
    DevOps & CloudAuto-check passed
  • Funnelcake Deployment Workflow

    divinevideo/divine-mobile

    Deploy funnelcake (api + relay) to ANY environment (production, staging, poc) on GKE via ArgoCD.

    266 GitHub stars~3.6k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Multi Cluster API Data Mismatch

    divinevideo/divine-mobile

    Debug "API returns data that doesn't exist in database" when multiple Kubernetes clusters exist (production, staging, POC).

    266 GitHub stars~1.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Clickhouse Incident Runbook

What does Clickhouse Incident Runbook do?

ClickHouse incident response — triage, diagnose, and remediate server issues using system tables, kill stuck queries, and execute recovery procedures. Clickhouse Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. ClickHouse incident response — triage, diagnose, and remediate server issues using system tables, kill stuck queries, and execute recovery procedures.

When should I use Clickhouse Incident Runbook?

Clickhouse Incident Runbook fits situations like: clickHouse is slow; backing up merges; producing errors in production and you need an on-call playbook; with clickhouse incident.

How do I install Clickhouse Incident Runbook in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill clickhouse-incident-runbook -a claude-code`. Or copy the skill folder (skills/.curated/clickhouse-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/clickhouse-incident-runbook in your project. Claude Code loads it when a task matches its description.

How do I install Clickhouse Incident Runbook in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill clickhouse-incident-runbook -a codex`. Or copy the skill folder (skills/.curated/clickhouse-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/clickhouse-incident-runbook in your project. Codex loads it when a task matches its description.

Can I use Clickhouse Incident Runbook in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill clickhouse-incident-runbook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/clickhouse-incident-runbook, .gemini/skills/clickhouse-incident-runbook, .github/skills/clickhouse-incident-runbook and .opencode/skills/clickhouse-incident-runbook in your project.

What does Clickhouse Incident Runbook need to run?

Going by SKILL.md and its folder, Clickhouse Incident Runbook needs the command-line tools its instructions call (curl). Our summary lists: Docker. Its frontmatter pre-approves these tools: Read, Bash(kubectl:*), Bash(curl:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Clickhouse Incident Runbook access the network?

SKILL.md names 2 domains. In commands or code: status.clickhouse.cloud; the agent is likely to contact it when it follows the instructions. As links in the text: clickhouse.com. This is read from the text; nothing was executed.

Is Clickhouse Incident Runbook safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Clickhouse Incident Runbook use?

Clickhouse Incident Runbook is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Clickhouse Incident Runbook use?

About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.1k tokens, read only when the agent opens those files.

What are the alternatives to Clickhouse Incident Runbook?

Skills that share tags, products or a category with Clickhouse Incident Runbook: Incident Response (chmonitor/chmonitor, 300 stars), Monitoring Ingestion Pipeline (PostHog/posthog, 40k stars), Verify Production (chmonitor/chmonitor, 300 stars) and Cluster Operations (chmonitor/chmonitor, 300 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Clickhouse Incident Runbook?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.