Agent skill

Data Analysis

by chmonitor in chmonitor/chmonitor

Analyze cluster data and metrics using raw SQL recipes against system.querylog and system.parts when no dedicated tool exists.

GPL-3.0Auto-check passedData & Analytics

Install Data Analysis

skills CLI
$ npx skills add chmonitor/chmonitor --skill data-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install chmonitor/chmonitor data-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/chmonitor/chmonitor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/data-analysis .claude/skills/data-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-analysis
GitHub stars
298
Token cost
~2k tokens
SKILL.md length
492 words
Files
1
Skills in repo
53
Repo updated
First seen
Licence
GPL-3.0

At a glance

Analyze cluster data and metrics using raw SQL recipes against system.querylog and system.parts when no dedicated tool exists.

  • Tasks that involve Data analysis
  • SKILL.md covers Largest Data Scan Ever, Most Expensive Queries, Query Fingerprint Patterns and Query Volume / Activity Over…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve SQL

What it does

Data Analysis is an agent skill from chmonitor/chmonitor. Analyze cluster data and metrics using raw SQL recipes against system.querylog and system.parts when no dedicated tool exists.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data analysis and SQL. It works with SQL. The repository describes itself as: Open-source operational advisor for ClickHouse — real-time monitoring plus AI-driven index/partition/materialized-view recommendations. The licence is GPL-3.0.

When your agent uses it

  • Tasks that involve Data analysis
  • Tasks that involve SQL

Example prompts

  • “/data-analysis”

What it can do on your machine

Read from SKILL.md and the folder at commit fc39ef0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are sql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Analysis loads about 2k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 492 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from chmonitor/chmonitor at commit fc39ef0, republished under its GPL-3.0 licence (© chmonitor). 492 words, ~2,042 tokens.

Download SKILL.mdSave it as .claude/skills/data-analysis/SKILL.md (or your agent's skills folder).
name
data-analysis
description
Analyze cluster data and metrics using raw SQL recipes against system.query_log and system.parts when no dedicated tool exists.

Data Analysis

Use this skill when dedicated tools (get_slow_queries, list_slow_query_patterns, etc.) don't cover the specific aggregation you need. All recipes below are read-only. Always verify column names against system-tables-reference before running — system.query_log columns vary across ClickHouse versions.

Core filter for finished queries:

sql
WHERE type = 'QueryFinish'
  AND event_time >= now() - INTERVAL 24 HOUR
  AND is_initial_query = 1

Omit is_initial_query = 1 only when you want internal sub-queries too.


Largest Data Scan Ever

Returns the single query that read the most bytes from disk. Useful for spotting runaway full-table scans in history.

sql
SELECT
    query_id,
    user,
    event_time,
    formatReadableSize(read_bytes)  AS read_size,
    formatReadableQuantity(read_rows) AS read_rows_fmt,
    query_duration_ms,
    substring(query, 1, 400)        AS query_text
FROM system.query_log
WHERE type = 'QueryFinish'
  AND is_initial_query = 1
ORDER BY read_bytes DESC
LIMIT 1

For the largest scan in a time window add AND event_time >= now() - INTERVAL 7 DAY. read_bytes is the compressed bytes read from storage; read_rows is the row count before filtering. Column availability: both present since v21.x.


Most Expensive Queries

Rank finished queries by memory, bytes read, or wall-clock duration. Run one variant or UNION all three for a combined view.

sql
-- By memory_usage
SELECT
    query_id,
    user,
    event_time,
    query_duration_ms,
    formatReadableSize(memory_usage)  AS memory,
    formatReadableSize(read_bytes)    AS read_size,
    substring(query, 1, 300)          AS query_text
FROM system.query_log
WHERE type = 'QueryFinish'
  AND is_initial_query = 1
  AND event_time >= now() - INTERVAL 24 HOUR
ORDER BY memory_usage DESC
LIMIT 20

Swap ORDER BY memory_usage DESC for read_bytes DESC or query_duration_ms DESC to rank by a different cost axis. The list_slow_query_patterns / get_slow_queries cover the common case; use raw SQL when you need a custom time window or additional columns like tables or ProfileEvents.


Query Fingerprint Patterns

Groups parameterized variants of the same logical query using normalized_query_hash. Shows which query shapes dominate load.

sql
SELECT
    normalized_query_hash,
    count()                                          AS calls,
    avg(query_duration_ms)                           AS avg_duration_ms,
    quantile(0.95)(query_duration_ms)                AS p95_duration_ms,
    sum(read_bytes)                                  AS total_read_bytes,
    formatReadableSize(sum(read_bytes))              AS total_read_size,
    formatReadableQuantity(sum(read_rows))           AS total_read_rows,
    any(substring(query, 1, 200))                    AS sample_query
FROM system.query_log
WHERE type = 'QueryFinish'
  AND is_initial_query = 1
  AND event_time >= now() - INTERVAL 24 HOUR
GROUP BY normalized_query_hash
ORDER BY total_read_bytes DESC
LIMIT 30

normalized_query_hash replaces literals with ? placeholders before hashing so SELECT 1 and SELECT 2 share a hash. Available since v20.6. quantile(0.95) requires no extra setup; use quantiles(0.5, 0.95, 0.99) for multiple percentiles in one pass.


Query Volume / Activity Over Time

Counts of finished queries bucketed by hour. Useful for spotting traffic spikes or quiet periods.

sql
SELECT
    toStartOfHour(event_time)                AS hour,
    count()                                  AS queries,
    countIf(exception_code != 0)             AS errors,
    avg(query_duration_ms)                   AS avg_duration_ms,
    formatReadableSize(sum(read_bytes))      AS total_read
FROM system.query_log
WHERE type = 'QueryFinish'
  AND is_initial_query = 1
  AND event_time >= now() - INTERVAL 48 HOUR
GROUP BY hour
ORDER BY hour

Swap toStartOfHour for toStartOfFifteenMinutes or toStartOfDay to change granularity. exception_code is 0 on success; a non-zero value with type = 'QueryFinish' means the query completed but reported an error.


Show full SKILL.md (214 more words)Show less

Top Tables by Size and Rows

Aggregates system.parts (active parts only) to rank user tables by compressed disk usage and row count.

sql
SELECT
    database,
    table,
    formatReadableSize(sum(bytes_on_disk))          AS disk_size,
    formatReadableSize(sum(data_compressed_bytes))  AS compressed,
    formatReadableSize(sum(data_uncompressed_bytes)) AS uncompressed,
    round(sum(data_uncompressed_bytes) /
          nullIf(sum(data_compressed_bytes), 0), 2) AS compression_ratio,
    formatReadableQuantity(sum(rows))               AS rows,
    count()                                         AS parts
FROM system.parts
WHERE active = 1
  AND database NOT IN ('system', 'information_schema', 'INFORMATION_SCHEMA')
GROUP BY database, table
ORDER BY sum(bytes_on_disk) DESC
LIMIT 30

bytes_on_disk includes index files and marks; data_compressed_bytes covers only column data. Filter active = 1 to exclude detached/obsolete parts. All columns present since v21.x. The get_disk_usage tool covers per-disk summaries; this recipe breaks down by table.


Period-over-Period Comparison

Compares query load between two equal-length windows using conditional aggregation — no subqueries, single scan.

sql
-- Compare last 24h vs previous 24h
SELECT
    normalized_query_hash,
    any(substring(query, 1, 200))                    AS sample_query,

    -- Current window
    countIf(event_time >= now() - INTERVAL 24 HOUR)             AS calls_now,
    avgIf(query_duration_ms, event_time >= now() - INTERVAL 24 HOUR) AS avg_ms_now,
    formatReadableSize(
        sumIf(read_bytes, event_time >= now() - INTERVAL 24 HOUR)
    )                                                            AS read_now,

    -- Previous window
    countIf(event_time < now() - INTERVAL 24 HOUR)              AS calls_prev,
    avgIf(query_duration_ms, event_time < now() - INTERVAL 24 HOUR) AS avg_ms_prev,
    formatReadableSize(
        sumIf(read_bytes, event_time < now() - INTERVAL 24 HOUR)
    )                                                            AS read_prev,

    -- Delta
    round(
        (countIf(event_time >= now() - INTERVAL 24 HOUR) -
         countIf(event_time  < now() - INTERVAL 24 HOUR)) * 100.0 /
        nullIf(countIf(event_time < now() - INTERVAL 24 HOUR), 0),
    1)                                                           AS call_delta_pct
FROM system.query_log
WHERE type = 'QueryFinish'
  AND is_initial_query = 1
  AND event_time >= now() - INTERVAL 48 HOUR
GROUP BY normalized_query_hash
HAVING calls_now > 5 OR calls_prev > 5
ORDER BY abs(call_delta_pct) DESC NULLS LAST
LIMIT 30

The single WHERE event_time >= now() - INTERVAL 48 HOUR covers both windows. *If aggregates split the data by condition. Adjust the INTERVAL literals consistently for wider windows (e.g. 7 DAY / 14 DAY).


Usage Notes

  • query_log is sampled on busy clusters; if log_queries_probability < 1 is set, counts will be proportional estimates, not exact totals.
  • Column availability varies across versions. If a column reference fails with code 47 ("Unknown column"), load system-tables-reference and check the real schema with get_table_schema('system.query_log').
  • memory_usage in query_log reflects peak memory during the query, not average. For CPU approximation use ProfileEvents['OSCPUVirtualTimeMicroseconds'].
  • system.query_log is flushed asynchronously; the most recent ~1–2 seconds of finished queries may not appear immediately.

Cross-references

  • Load system-tables-reference for exact column lists and common pitfalls before hand-writing SQL.
  • Load query-optimization for EXPLAIN-driven tuning once expensive patterns are identified.
  • Load troubleshooting for error-code-driven diagnosis of failed queries (type = 'ExceptionWhileProcessing').

© chmonitor, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/data-analysis of chmonitor/chmonitor.

Open the folder on GitHubat commit fc39ef0

Compare with similar skills

Data Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Analysis this skillchmonitor/chmonitor298—~2kAutomated safety check: PassGPL-3.0
Dinobase Business Data Querieskappa90/dinobase263—~1.5kAutomated safety check: PassCustom licence
Analytics Engineerborghei/Claude-Skills874—~3.4kAutomated safety check: PassMIT
Network Data Analysisautomateyournetwork/netclaw674—~1.3kAutomated safety check: NotesApache-2.0
Data Analytics Data Analystchendongqi/OPB-Skills125—~1.1kAutomated safety check: PassNone
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT

Similar skills

  • Sets up Dinobase, a local DuckDB database that syncs data from 100+ business sources, then answers questions across them with SQL joins and previewed write-backs.

    263 GitHub stars~1.5k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    874 GitHub stars~3.4k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Network Data Analysis

    automateyournetwork/netclaw

    Ad-hoc read-only SQL analysis over exported network data (Zeek logs, Suricata eve.json, generated reports) using DuckDB.

    674 GitHub stars~1.3k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check: notes
  • Data Analytics Data Analyst

    chendongqi/OPB-Skills

    数据分析助手 - 专业的数据分析与业务洞察专家。适用场景: (1) 数据可视化与报表设计 (2) 业务指标体系构建 (3) 趋势分析与预测建模 (4) 异常检测与根因分析 (5) A/B测试设计与分析 (6) 数据报告撰写与汇报 (7) SQL/数据提取指导 触发关键词:数据分析、数据可视化、指标体系、趋势分析、异常检测、AB测试、数据报告、SQL、数据洞察、Dashboard、数据驱动

    125 GitHub stars~1.1k tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Analyzing Data

    astronomer/agents

    Queries the data warehouse with SQL and answers business questions about data.

    450 GitHub stars~1.3k tokensUpdated 2 days ago
    DatabasesAuto-check passed

More from chmonitor/chmonitor

All 53 skills in this repo
  • Hyperframes Creative

    chmonitor/chmonitor

    Non-animation creative direction for HyperFrames videos. An agent skill from chmonitor/chmonitor.

    298 GitHub starsUsed in 5 repos~1.3k tokens
    Auto-check passed
  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    298 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check: notes
  • Remotion To Hyperframes

    chmonitor/chmonitor

    Port an existing Remotion (React) composition to HyperFrames HTML.

    298 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Music To Video

    chmonitor/chmonitor

    A skill your agent uses when the user has a music track (an audio file, or a video to pull audio from) and wants a beat-synced HyperFrames video, calm to hard-hitting.

    298 GitHub starsUsed in 1 repo~4k tokens
    Auto-check: notes
  • Hyperframes Animation

    chmonitor/chmonitor

    All animation knowledge for HyperFrames — atomic motion rules, multi-phase scene blueprints, scene transitions, broader motion-design techniques, AND the seven runtime adapters (GSAP default, plus…

    298 GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • Faceless Explainer

    chmonitor/chmonitor

    turn arbitrary text — an article, notes, a topic, a brief — into a faceless explainer video, up to ~3 min (sweet spot 30-90s), where every visual is invented (typography, abstract graphics…

    298 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check: notes

Works with

Questions about Data Analysis

What does Data Analysis do?

Analyze cluster data and metrics using raw SQL recipes against system.querylog and system.parts when no dedicated tool exists. Data Analysis is an agent skill from chmonitor/chmonitor.parts when no dedicated tool exists.

When should I use Data Analysis?

Data Analysis fits situations like: tasks that involve Data analysis; tasks that involve SQL.

How do I install Data Analysis in Claude Code?

Run `npx skills add chmonitor/chmonitor --skill data-analysis -a claude-code`. Or copy the skill folder (.agents/skills/data-analysis in chmonitor/chmonitor) into .claude/skills/data-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Data Analysis in Codex?

Run `npx skills add chmonitor/chmonitor --skill data-analysis -a codex`. Or copy the skill folder (.agents/skills/data-analysis in chmonitor/chmonitor) into .agents/skills/data-analysis in your project. Codex loads it when a task matches its description.

Can I use Data Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add chmonitor/chmonitor --skill data-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-analysis, .gemini/skills/data-analysis, .github/skills/data-analysis and .opencode/skills/data-analysis in your project.

What does Data Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Analysis is instructions for the agent only.

Does Data Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Analysis use?

Data Analysis is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Analysis use?

About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Analysis?

Skills that share tags, products or a category with Data Analysis: Dinobase Business Data Queries (kappa90/dinobase, 263 stars), Analytics Engineer (borghei/Claude-Skills, 874 stars), Network Data Analysis (automateyournetwork/netclaw, 674 stars) and Data Analytics Data Analyst (chendongqi/OPB-Skills, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Analysis?

chmonitor (a GitHub organization) maintains it in chmonitor/chmonitor, which has 298 GitHub stars. The repository holds 53 skills in this directory. The repository was last updated on October 5, 2026.

Source: chmonitor/chmonitor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.