DB Ops Sop
OpenDCAI/DataMind
Database operations runbook — backup, recovery, performance tuning, troubleshooting.
Structured triage recipes for common ClickHouse incidents: disk, errors, replication, mutations, cluster health, and slow queries.
$ npx skills add chmonitor/chmonitor --skill incident-response -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install chmonitor/chmonitor incident-response --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/chmonitor/chmonitor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/incident-response .claude/skills/incident-response && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "incident-response" agent skill from https://github.com/chmonitor/chmonitor/tree/main/.agents/skills/incident-response into .claude/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/chmonitor/chmonitor/tree/main/.agents/skills/incident-responseType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add chmonitor/chmonitor --skill incident-response -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install chmonitor/chmonitor incident-response --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/chmonitor/chmonitor.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/incident-response .agents/skills/incident-response && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "incident-response" agent skill from https://github.com/chmonitor/chmonitor/tree/main/.agents/skills/incident-response into .agents/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add chmonitor/chmonitor --skill incident-response -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install chmonitor/chmonitor incident-response --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/chmonitor/chmonitor.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/incident-response .cursor/skills/incident-response && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "incident-response" agent skill from https://github.com/chmonitor/chmonitor/tree/main/.agents/skills/incident-response into .cursor/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/chmonitor/chmonitor.git --path .agents/skills/incident-response--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add chmonitor/chmonitor --skill incident-response -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install chmonitor/chmonitor incident-response --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/chmonitor/chmonitor.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/incident-response .gemini/skills/incident-response && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "incident-response" agent skill from https://github.com/chmonitor/chmonitor/tree/main/.agents/skills/incident-response into .gemini/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install chmonitor/chmonitor incident-responseInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add chmonitor/chmonitor --skill incident-response -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/chmonitor/chmonitor.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/incident-response .github/skills/incident-response && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "incident-response" agent skill from https://github.com/chmonitor/chmonitor/tree/main/.agents/skills/incident-response into .github/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add chmonitor/chmonitor --skill incident-response -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install chmonitor/chmonitor incident-response --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/chmonitor/chmonitor.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/incident-response .opencode/skills/incident-response && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "incident-response" agent skill from https://github.com/chmonitor/chmonitor/tree/main/.agents/skills/incident-response into .opencode/skills/incident-response/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "incident-response", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
incident-responseStructured triage recipes for common ClickHouse incidents: disk, errors, replication, mutations, cluster health, and slow queries.
Incident Response is an agent skill from chmonitor/chmonitor. Structured triage recipes for common ClickHouse incidents: disk, errors, replication, mutations, cluster health, and slow queries.
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Databases, covering Incident response, Data warehousing and Query optimization. It works with ClickHouse. The repository describes itself as: Open-source operational advisor for ClickHouse — real-time monitoring plus AI-driven index/partition/materialized-view recommendations. The licence is GPL-3.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit fc39ef0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are sql).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Incident Response loads about 3.4k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 1,196 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from chmonitor/chmonitor at commit fc39ef0, republished under its GPL-3.0 licence (© chmonitor). 1,196 words, ~3,374 tokens.
.claude/skills/incident-response/SKILL.md (or your agent's skills folder).Structured, step-by-step recipes for the most common ClickHouse incidents. Each recipe names the exact tool or system table query, what to look for, and the standard remediation. For multi-step incidents, pair with update_plan + the plan-and-verify skill to track progress and verify each fix before moving to the next step.
Triggers: disk usage > 80 %, insert errors mentioning DB::Exception: Not enough space, alerts on free bytes.
Step 1 — Check free space per disk
Use the get_disk_usage tool or query directly:
SELECT name, path,
formatReadableSize(free_space) AS free,
formatReadableSize(total_space) AS total,
round(100 - free_space * 100.0 / total_space, 1) AS used_pct
FROM system.disks
ORDER BY used_pct DESCAlert threshold: used_pct > 85. Note which disk is filling and which storage policy it belongs to.
Step 2 — Find the biggest tables and parts
Use list_tables (or query system.parts) or:
SELECT database, table,
formatReadableSize(sum(bytes_on_disk)) AS disk,
sum(rows) AS rows,
count() AS parts
FROM system.parts
WHERE active
GROUP BY database, table
ORDER BY sum(bytes_on_disk) DESC
LIMIT 20Also look at system.parts with high modification_time age — old unmerged parts waste space.
Step 3 — Estimate insert rate and time-to-full
Use the forecast_disk_capacity tool. Manual estimate:
SELECT toStartOfHour(event_time) AS hour,
sum(rows) AS rows_inserted,
formatReadableSize(sum(bytes_written_to_disk)) AS written
FROM system.part_log
WHERE event_type = 'NewPart'
AND event_time >= now() - INTERVAL 24 HOUR
GROUP BY hour
ORDER BY hourDivide current free bytes by hourly write rate → hours until full.
Step 4 — Remediation
TTL on high-volume tables (ALTER TABLE t MODIFY TTL date + INTERVAL 30 DAY).ALTER TABLE t DROP PARTITION '2024-01' for old data (irreversible — confirm first).ALTER TABLE t MOVE PARTITION '...' TO DISK 'cold'.Cross-references: storage-optimization (TTL syntax, tiered storage), plan-and-verify (track multi-partition drops safely).
Triggers: spike in failed queries, user-facing 500s, alert on system.errors.
Step 1 — Recent errors from system.errors
SELECT name, code, value, remote,
last_error_time, last_error_message
FROM system.errors
WHERE last_error_time >= now() - INTERVAL 1 HOUR
ORDER BY value DESC
LIMIT 20value is the cumulative error count since startup; look for codes with recent last_error_time and high value.
Step 2 — Failed queries with exceptions
Use get_failed_queries or:
SELECT exception_code,
count() AS cnt,
topK(5)(exception) AS samples,
topK(5)(query) AS queries
FROM system.query_log
WHERE type IN ('ExceptionWhileProcessing', 'ExceptionBeforeStart')
AND event_time >= now() - INTERVAL 1 HOUR
GROUP BY exception_code
ORDER BY cnt DESCStep 3 — Interpret error codes
| Code | Meaning | Typical fix |
|---|---|---|
| 60 | Table not found | Verify table/database name; check if DROP happened |
| 47 | Unknown column | Use get_table_schema; column may have been dropped |
| 241 | Memory limit exceeded | Reduce scope; add LIMIT; raise max_memory_usage |
| 159 | Timeout | Add time filter; check for table lock (mutations) |
| 252 | Too many parts | Wait for merges; run recipe 4 |
| 285 | Quorum write failed | Check replica availability (recipe 3) |
| 999 | ZooKeeper/Keeper error | Check Keeper health (recipe 3) |
Step 4 — Check resource pressure
Use get_metrics or:
SELECT metric, value
FROM system.metrics
WHERE metric IN (
'Query', 'BackgroundMergesAndMutationsPoolTask',
'ZooKeeperRequest', 'MemoryTracking'
)Correlate a memory or Keeper spike with the error burst.
Step 5 — Remediation
Fix the specific error code (table, column, limit, Keeper). If a single bad query is causing the spike, kill it:
KILL QUERY WHERE query_id = '...' ASYNCCross-references: troubleshooting (error code details, OOM), anomaly-detection (automated spike detection).
Triggers: replica is behind, reads from replica return stale data, absolute_delay alert.
Step 1 — Check per-table replication lag
Use get_replication_status or query system.replication_queue after get_table_schema:
SELECT database, table, replica_name, replica_path,
is_leader, is_readonly,
absolute_delay,
queue_size, inserts_in_queue, merges_in_queue,
last_queue_update
FROM system.replicas
WHERE absolute_delay > 0 OR queue_size > 0
ORDER BY absolute_delay DESCabsolute_delay > 300 (5 min) is a concern. is_readonly = 1 means the replica cannot write — usually a Keeper connectivity problem.
Step 2 — Inspect the replication queue
SELECT database, table, type, source_replica,
parts_to_merge, create_time,
last_attempt_time, last_exception,
num_tries
FROM system.replication_queue
WHERE last_exception != ''
ORDER BY num_tries DESC
LIMIT 20Repeated failures with the same last_exception point to a stuck task. High num_tries means the replica has been retrying for a while.
Step 3 — Keeper health
SELECT *
FROM system.zookeeper
WHERE path = '/clickhouse'Or query system.zookeeper with a path filter after get_table_schema. Look for high zookeeper_sessions in system.metrics and watch for ZooKeeperRequest latency spikes in system.asynchronous_metrics.
Also check the distributed DDL queue for stuck operations:
SELECT *
FROM system.distributed_ddl_queue
WHERE status != 'Finished'
ORDER BY entry_timeStep 4 — Remediation
SYSTEM RESTART REPLICA db.table.SYSTEM DROP REPLICA 'bad_host' FROM TABLE db.table to remove a dead peer, then let the replica re-sync.DETACH TABLE, restore from another replica's data, ATTACH TABLE.Cross-references: replication-guide (full recovery playbook, detach/reattach steps, Keeper tuning).
Triggers: ALTER TABLE ... UPDATE/DELETE never completes, parts_to_do stays non-zero, merge backlog growing.
Step 1 — Find stuck mutations
Query system.mutations (columns in system-tables-reference) or:
SELECT database, table, mutation_id,
command, create_time,
parts_to_do, is_done,
latest_fail_reason, latest_fail_time
FROM system.mutations
WHERE is_done = 0
ORDER BY create_timelatest_fail_reason != '' tells you exactly why it is stuck (disk space, missing column, Keeper timeout).
Step 2 — Find long-running merges
SELECT database, table, elapsed,
formatReadableSize(total_size_bytes_compressed) AS size,
progress, merge_type,
partition_id
FROM system.merges
ORDER BY elapsed DESC
LIMIT 10elapsed > 3600 (1 hour) for a merge is unusual. progress not advancing over several minutes means the merge is likely stuck.
Step 3 — Check the part count
SELECT database, table, partition_id,
count() AS parts
FROM system.parts
WHERE active
GROUP BY database, table, partition_id
HAVING parts > 100
ORDER BY parts DESCExcessive parts (>300 in a partition) cause mutations and merges to slow or hang.
Step 4 — Remediation
KILL MUTATION WHERE mutation_id = 'mutation_8.txt'background_pool_size load by pausing inserts temporarily.latest_fail_reason mentions disk space, run recipe 1 first.max_bytes_to_merge_at_max_space_in_pool or temporarily OPTIMIZE TABLE t PARTITION 'part' during off-peak.ReplacingMergeTree + CollapsingMergeTree over heavy UPDATE/DELETE mutations.Cross-references: troubleshooting (mutation + merge details), storage-optimization (part count management).
Triggers: routine check-in, "how is the cluster", pre-maintenance review, post-deploy verification.
Run these in order. Each surfaces a different failure class.
Step 1 — Server uptime and version
Use get_metrics or:
SELECT version(), uptime(), now()Unexpected recent uptime means a crash restart occurred.
Step 2 — Disk headroom
Run step 1 of recipe 1. Flag any disk above 80 %.
Step 3 — Recent slow queries
SELECT query_id,
round(query_duration_ms / 1000, 1) AS secs,
read_rows, formatReadableSize(read_bytes) AS read,
memory_usage, query
FROM system.query_log
WHERE type = 'QueryFinish'
AND event_time >= now() - INTERVAL 15 MINUTE
AND query_duration_ms > 5000
ORDER BY query_duration_ms DESC
LIMIT 10Step 4 — Recent errors
Use get_failed_queries or run step 1 of recipe 2.
Step 5 — Merge backlog
SELECT count() AS active_merges,
sum(elapsed) AS total_elapsed_s,
max(elapsed) AS longest_s
FROM system.mergesactive_merges > 50 or longest_s > 600 warrants investigation (recipe 4).
Step 6 — Replication status
SELECT count() AS lagging_tables,
max(absolute_delay) AS max_delay_s,
sum(queue_size) AS total_queue
FROM system.replicas
WHERE absolute_delay > 30 OR queue_size > 10Any lagging_tables > 0 with growing max_delay_s → recipe 3.
Step 7 — Summarize
Report findings in priority order: disk → replication → merges → errors → slow queries. Use update_plan to record findings and track follow-up actions.
Cross-references: anomaly-detection (automated sweep), plan-and-verify (track remediation steps).
Triggers: a specific query is slow, user reports latency, query_duration_ms alert.
Step 1 — Currently running queries
Use get_running_queries tool or:
SELECT query_id, user, elapsed,
read_rows, formatReadableSize(read_bytes) AS read,
memory_usage,
query
FROM system.processes
ORDER BY elapsed DESCA query running for minutes is usually either doing a full scan or waiting for a merge/mutation to release a lock.
Step 2 — Historical slow queries
SELECT query_id,
event_time,
round(query_duration_ms / 1000, 1) AS secs,
read_rows,
formatReadableSize(read_bytes) AS read,
ProfileEvents['SelectedMarks'] AS marks_selected,
ProfileEvents['MergeTreeDataSelectExecutorReadRows'] AS rows_from_disk,
exception,
query
FROM system.query_log
WHERE type = 'QueryFinish'
AND event_time >= now() - INTERVAL 1 HOUR
AND query_duration_ms > 3000
ORDER BY query_duration_ms DESC
LIMIT 20marks_selected high relative to actual result rows → full granule scan, missing primary key filtering.
Step 3 — Explain the query plan
Use the explain_query tool. Look for:
ReadFromMergeTree with no key condition → full table scanFilter after ReadFromMergeTree instead of before → PREWHERE opportunityPartialSortingTransform with huge row counts → missing ORDER BY key alignmentHashJoin build side → consider join_algorithm = 'partial_merge'Step 4 — Check table schema and sorting key
Use get_table_schema tool. Verify:
Step 5 — Remediation
PREWHERE for selective low-cardinality filters applied before reading full columns.INDEX idx col TYPE bloom_filter) for high-cardinality equality lookups.KILL QUERY WHERE query_id = '...' ASYNC.Cross-references: query-tuning-advisor (index selection, MV design), plan-and-verify (test the rewrite safely before production).
event_time filters prevent query_log scans from timing out.update_plan to record what you found, what you changed, and what to verify — especially for multi-step incidents.© chmonitor, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/incident-response of chmonitor/chmonitor.
Open the folder on GitHubat commit fc39ef0
Incident Response next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Incident Response this skillchmonitor/chmonitor | 299 | — | ~3.4k | Automated safety check: Pass | GPL-3.0 | |
| DB Ops SopOpenDCAI/DataMind | 423 | — | ~388 | Automated safety check: Pass | Apache-2.0 | |
| Clickhouse Ioaffaan-m/ECC | 275k | 1 repos | ~2.7k | Automated safety check: Pass | MIT | |
| Generating Clickhouse Query Performance ReportsPostHog/posthog | 40k | — | ~5.1k | Automated safety check: Pass | Custom licence | |
| Clickhouse Incident Runbookjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Clickhouse Performance Tuningjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~1.2k | Automated safety check: Pass | MIT |
OpenDCAI/DataMind
Database operations runbook — backup, recovery, performance tuning, troubleshooting.
affaan-m/ECC
ClickHouse database patterns, query optimization, analytics, and data engineering best practices for high-performance analytical workloads.
PostHog/posthog
Produce and structure slow-query performance reports for PostHog's production ClickHouse (US and EU).
jeremylongshore/tons-of-skills-marketplace
ClickHouse incident response — triage, diagnose, and remediate server issues using system tables, kill stuck queries, and execute recovery procedures.
jeremylongshore/tons-of-skills-marketplace
Optimize ClickHouse query performance with indexing, projections, settings tuning, and query analysis using system tables.
hellangleZ/burn-in-cceverywhere-ralph
ClickHouse database patterns, query optimization, analytics, and data engineering best practices for high-performance analytical workloads.
chmonitor/chmonitor
Non-animation creative direction for HyperFrames videos. An agent skill from chmonitor/chmonitor.
chmonitor/chmonitor
Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…
chmonitor/chmonitor
Port an existing Remotion (React) composition to HyperFrames HTML.
chmonitor/chmonitor
A skill your agent uses when the user has a music track (an audio file, or a video to pull audio from) and wants a beat-synced HyperFrames video, calm to hard-hitting.
chmonitor/chmonitor
All animation knowledge for HyperFrames — atomic motion rules, multi-phase scene blueprints, scene transitions, broader motion-design techniques, AND the seven runtime adapters (GSAP default, plus…
chmonitor/chmonitor
turn arbitrary text — an article, notes, a topic, a brief — into a faceless explainer video, up to ~3 min (sweet spot 30-90s), where every visual is invented (typography, abstract graphics…
Works with
Categories
Structured triage recipes for common ClickHouse incidents: disk, errors, replication, mutations, cluster health, and slow queries. Incident Response is an agent skill from chmonitor/chmonitor. Structured triage recipes for common ClickHouse incidents: disk, errors, replication, mutations, cluster health, and slow queries.
Incident Response fits situations like: tasks that involve Incident response; tasks that involve Data warehousing; tasks that involve Query optimization.
Run `npx skills add chmonitor/chmonitor --skill incident-response -a claude-code`. Or copy the skill folder (.agents/skills/incident-response in chmonitor/chmonitor) into .claude/skills/incident-response in your project. Claude Code loads it when a task matches its description.
Run `npx skills add chmonitor/chmonitor --skill incident-response -a codex`. Or copy the skill folder (.agents/skills/incident-response in chmonitor/chmonitor) into .agents/skills/incident-response in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add chmonitor/chmonitor --skill incident-response -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/incident-response, .gemini/skills/incident-response, .github/skills/incident-response and .opencode/skills/incident-response in your project.
SKILL.md names no scripts, command-line tools or credentials: Incident Response is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Incident Response is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Incident Response: DB Ops Sop (OpenDCAI/DataMind, 423 stars), Clickhouse Io (affaan-m/ECC, 275k stars), Generating Clickhouse Query Performance Reports (PostHog/posthog, 40k stars) and Clickhouse Incident Runbook (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
chmonitor (a GitHub organization) maintains it in chmonitor/chmonitor, which has 299 GitHub stars. The repository holds 53 skills in this directory. The repository was last updated on October 5, 2026.
Source: chmonitor/chmonitor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.