Agent skill

Cluster Operations

by chmonitor in chmonitor/chmonitor

Cluster management: distributed tables, ON CLUSTER DDL, node lifecycle, resharding, load balancing, and Keeper migration.

GPL-3.0Auto-check passedDevOps & Cloud

Install Cluster Operations

skills CLI
$ npx skills add chmonitor/chmonitor --skill cluster-operations -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install chmonitor/chmonitor cluster-operations --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/chmonitor/chmonitor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/cluster-operations .claude/skills/cluster-operations && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cluster-operations
GitHub stars
298
Token cost
~634 tokens
SKILL.md length
278 words
Files
1
Skills in repo
53
Repo updated
First seen
Licence
GPL-3.0

At a glance

Cluster management: distributed tables, ON CLUSTER DDL, node lifecycle, resharding, load balancing, and Keeper migration.

  • Works in 5 steps: Install ClickHouse on new node → Configure Keeper/ZooKeeper connection → Update cluster config (remote_servers)… → …
  • Tasks that involve Container orchestration
  • SKILL.md covers Distributed Tables, ON CLUSTER DDL, Load Balancing and Read Routing and Adding Nodes, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cluster Operations is an agent skill from chmonitor/chmonitor. Cluster management: distributed tables, ON CLUSTER DDL, node lifecycle, resharding, load balancing, and Keeper migration.

Its SKILL.md is about 630 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Container orchestration, Cloud networking and Data warehousing. It works with ClickHouse. The repository describes itself as: Open-source operational advisor for ClickHouse — real-time monitoring plus AI-driven index/partition/materialized-view recommendations. The licence is GPL-3.0.

When your agent uses it

  • Tasks that involve Container orchestration
  • Tasks that involve Cloud networking
  • Tasks that involve Data warehousing

Example prompts

  • “/cluster-operations”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Install ClickHouse on new node
  2. Configure Keeper/ZooKeeper connection
  3. Update cluster config (remote_servers) on all nodes
  4. Create local tables on new node
  5. ReplicatedMergeTree: data syncs automatically; non-replicated: copy or re-insert

What it can do on your machine

Read from SKILL.md and the folder at commit fc39ef0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cluster Operations loads about 634 tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 278 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~634

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from chmonitor/chmonitor at commit fc39ef0, republished under its GPL-3.0 licence (© chmonitor). 278 words, ~634 tokens.

Download SKILL.mdSave it as .claude/skills/cluster-operations/SKILL.md (or your agent's skills folder).
name
cluster-operations
description
Cluster management: distributed tables, ON CLUSTER DDL, node lifecycle, resharding, load balancing, and Keeper migration.

Cluster Operations

Distributed Tables

  • CREATE TABLE dist ENGINE = Distributed(cluster, db, local_table, sharding_key)
  • Sharding key: rand() for even distribution, cityHash64(user_id) for user affinity
  • Reads: query all shards in parallel; Writes: route to correct shard or write locally

ON CLUSTER DDL

  • ALTER TABLE t ON CLUSTER '{cluster}' ADD COLUMN col Type — propagate schema to all replicas
  • CREATE TABLE t ON CLUSTER '{cluster}' AS template_db.template_table — clone across shards
  • distributed_ddl_output_mode: throw (fail on error), null (ignore), none, active
  • Check status: SELECT * FROM system.distributed_ddl_queue

Load Balancing and Read Routing

  • load_balancing: random, in_order, first_or_random, nearest_hostname
  • max_replica_delay_for_distributed_queries — skip lagging replicas
  • fallback_to_stale_replicas_for_distributed_queries=1 — use stale when all delayed

Adding Nodes

  1. Install ClickHouse on new node
  2. Configure Keeper/ZooKeeper connection
  3. Update cluster config (remote_servers) on all nodes
  4. Create local tables on new node
  5. ReplicatedMergeTree: data syncs automatically; non-replicated: copy or re-insert

Removing Nodes

  1. Stop writes, wait for replication queue to drain
  2. SYSTEM DROP REPLICA for replicated tables
  3. Remove from cluster config, restart remaining nodes

Resharding

  • No native online resharding — create new distributed table with new sharding scheme
  • INSERT INTO new_dist SELECT * FROM old_dist or clickhouse-copier for large migrations

Cluster Recovery

  • SYSTEM RESTART REPLICA ON CLUSTER '{cluster}' — restart replication across all nodes
  • SYSTEM SYNC REPLICA ON CLUSTER '{cluster}' — force sync from ZooKeeper
  • SYSTEM FETCH PARTS ON CLUSTER '{cluster}' — pull missing parts from other replicas

Monitoring Clusters

  • system.clusters — topology; system.distributed_ddl_queue — DDL status; system.replicas — replication
  • Cross-shard queries: use Distributed table or remote() function

Keeper Migration (ZooKeeper to ClickHouse Keeper)

  1. Deploy ClickHouse Keeper alongside ZooKeeper
  2. Snapshot data: clickhouse-keeper-converter or zk-dump.sh
  3. Configure Keeper with converted snapshot
  4. Update zookeeper config, restart one node at a time
  5. Verify replication recovers, then remove ZooKeeper

© chmonitor, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/cluster-operations of chmonitor/chmonitor.

Open the folder on GitHubat commit fc39ef0

Compare with similar skills

Cluster Operations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cluster Operations compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cluster Operations this skillchmonitor/chmonitor298—~634Automated safety check: PassGPL-3.0
Funnelcake Deployment Workflowdivinevideo/divine-mobile265—~3.6kAutomated safety check: PassMPL-2.0
Multi Cluster API Data Mismatchdivinevideo/divine-mobile265—~1.3kAutomated safety check: PassMPL-2.0
Opensourcefaqdigoal/blog8.6k—~966Automated safety check: PassGPL-2.0
Gke Cost Analysisgoogle/skills21k—~1.5kAutomated safety check: PassApache-2.0
Clickhouse System Log Disk Exhaustiondivinevideo/divine-mobile265—~2kAutomated safety check: PassMPL-2.0

Similar skills

  • Funnelcake Deployment Workflow

    divinevideo/divine-mobile

    Deploy funnelcake (api + relay) to ANY environment (production, staging, poc) on GKE via ArgoCD.

    265 GitHub stars~3.6k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Multi Cluster API Data Mismatch

    divinevideo/divine-mobile

    Debug "API returns data that doesn't exist in database" when multiple Kubernetes clusters exist (production, staging, POC).

    265 GitHub stars~1.3k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Opensourcefaq

    digoal/blog

    解答与开源产品有关的深度技术问题,输出图文并茂的 Markdown 技术文章。触发条件:用户提出与开源项目(如 PostgreSQL、Redis、Kafka、Kubernetes、ClickHouse、Flink 等)相关的技术问题,并提供源码目录或 URL、deepwiki repo 名称。即使用户只说"帮我解答这个开源问题"或"分析一下这个项目的某个机制",也应使用本…

    8.6k GitHub stars~966 tokensUpdated 9 days ago
    DatabasesAuto-check passed
  • Gke Cost Analysis

    google/skills

    Official

    Answer natural language questions and perform analysis on GKE cluster and workload costs using BigQuery billing exports, cost allocation data, and live cluster monitoring metrics.

    21k GitHub stars~1.5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Clickhouse System Log Disk Exhaustion

    divinevideo/divine-mobile

    Fix ClickHouse "Code: 243 Cannot reserve 1.00 MiB, not enough space" errors caused by system log tables (textlog, tracelog, processorsprofilelog, querylog) filling the disk.

    265 GitHub stars~2k tokensUpdated today
    DatabasesAuto-check passed
  • Monitoring Ingestion Pipeline

    PostHog/posthog-foss

    Official

    Guide for using the Grafana MCP to monitor and diagnose the Node.js ingestion pipeline workers in production.

    721 GitHub stars~9.1k tokensUpdated today
    DevOps & CloudAuto-check passed

More from chmonitor/chmonitor

All 53 skills in this repo
  • Hyperframes Creative

    chmonitor/chmonitor

    Non-animation creative direction for HyperFrames videos. An agent skill from chmonitor/chmonitor.

    298 GitHub starsUsed in 5 repos~1.3k tokens
    Auto-check passed
  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    298 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check: notes
  • Remotion To Hyperframes

    chmonitor/chmonitor

    Port an existing Remotion (React) composition to HyperFrames HTML.

    298 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Music To Video

    chmonitor/chmonitor

    A skill your agent uses when the user has a music track (an audio file, or a video to pull audio from) and wants a beat-synced HyperFrames video, calm to hard-hitting.

    298 GitHub starsUsed in 1 repo~4k tokens
    Auto-check: notes
  • Hyperframes Animation

    chmonitor/chmonitor

    All animation knowledge for HyperFrames — atomic motion rules, multi-phase scene blueprints, scene transitions, broader motion-design techniques, AND the seven runtime adapters (GSAP default, plus…

    298 GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • Faceless Explainer

    chmonitor/chmonitor

    turn arbitrary text — an article, notes, a topic, a brief — into a faceless explainer video, up to ~3 min (sweet spot 30-90s), where every visual is invented (typography, abstract graphics…

    298 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check: notes

Works with

Questions about Cluster Operations

What does Cluster Operations do?

Cluster management: distributed tables, ON CLUSTER DDL, node lifecycle, resharding, load balancing, and Keeper migration. Cluster Operations is an agent skill from chmonitor/chmonitor. Cluster management: distributed tables, ON CLUSTER DDL, node lifecycle, resharding, load balancing, and Keeper migration.

When should I use Cluster Operations?

Cluster Operations fits situations like: tasks that involve Container orchestration; tasks that involve Cloud networking; tasks that involve Data warehousing.

How do I install Cluster Operations in Claude Code?

Run `npx skills add chmonitor/chmonitor --skill cluster-operations -a claude-code`. Or copy the skill folder (.agents/skills/cluster-operations in chmonitor/chmonitor) into .claude/skills/cluster-operations in your project. Claude Code loads it when a task matches its description.

How do I install Cluster Operations in Codex?

Run `npx skills add chmonitor/chmonitor --skill cluster-operations -a codex`. Or copy the skill folder (.agents/skills/cluster-operations in chmonitor/chmonitor) into .agents/skills/cluster-operations in your project. Codex loads it when a task matches its description.

Can I use Cluster Operations in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add chmonitor/chmonitor --skill cluster-operations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cluster-operations, .gemini/skills/cluster-operations, .github/skills/cluster-operations and .opencode/skills/cluster-operations in your project.

What does Cluster Operations need to run?

SKILL.md names no scripts, command-line tools or credentials: Cluster Operations is instructions for the agent only.

Does Cluster Operations access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cluster Operations safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cluster Operations use?

Cluster Operations is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cluster Operations use?

About 634 tokens (SKILL.md is roughly 2.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cluster Operations?

Skills that share tags, products or a category with Cluster Operations: Funnelcake Deployment Workflow (divinevideo/divine-mobile, 265 stars), Multi Cluster API Data Mismatch (divinevideo/divine-mobile, 265 stars), Opensourcefaq (digoal/blog, 8.6k stars) and Gke Cost Analysis (google/skills, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cluster Operations?

chmonitor (a GitHub organization) maintains it in chmonitor/chmonitor, which has 298 GitHub stars. The repository holds 53 skills in this directory. The repository was last updated on October 5, 2026.

Source: chmonitor/chmonitor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.