Agent skill

Anomaly Detection

by chmonitor in chmonitor/chmonitor

Detect cluster abnormalities by comparing recent activity (last 1h) to a baseline (preceding 24h) using the raw query tool.

GPL-3.0Auto-check passedData & Analytics

Install Anomaly Detection

skills CLI
$ npx skills add chmonitor/chmonitor --skill anomaly-detection -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install chmonitor/chmonitor anomaly-detection --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/chmonitor/chmonitor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/anomaly-detection .claude/skills/anomaly-detection && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
anomaly-detection
GitHub stars
299
Token cost
~2.3k tokens
SKILL.md length
353 words
Files
1
Skills in repo
53
Repo updated
First seen
Licence
GPL-3.0

At a glance

Detect cluster abnormalities by comparing recent activity (last 1h) to a baseline (preceding 24h) using the raw query tool.

  • Tasks that involve Anomaly detection
  • SKILL.md covers Error-Rate Spike, Query Duration Regression, Query Volume Spike / Drop and Memory Anomalies, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Anomaly Detection is an agent skill from chmonitor/chmonitor. Detect cluster abnormalities by comparing recent activity (last 1h) to a baseline (preceding 24h) using the raw query tool.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Anomaly detection. The repository describes itself as: Open-source operational advisor for ClickHouse — real-time monitoring plus AI-driven index/partition/materialized-view recommendations. The licence is GPL-3.0.

When your agent uses it

  • Tasks that involve Anomaly detection

Example prompts

  • “/anomaly-detection”

What it can do on your machine

Read from SKILL.md and the folder at commit fc39ef0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are sql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Anomaly Detection loads about 2.3k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 353 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from chmonitor/chmonitor at commit fc39ef0, republished under its GPL-3.0 licence (© chmonitor). 353 words, ~2,277 tokens.

Download SKILL.mdSave it as .claude/skills/anomaly-detection/SKILL.md (or your agent's skills folder).
name
anomaly-detection
description
Detect cluster abnormalities by comparing recent activity (last 1h) to a baseline (preceding 24h) using the raw query tool.

Anomaly Detection

Use this skill when no dedicated tool covers anomaly detection. All recipes below use the query tool with raw SQL. Compare a recent window (last 1 hour) to a baseline window (preceding 24 hours) and surface ratios or deltas that exceed the thresholds in the interpretation section.

Column availability varies by ClickHouse version — load system-tables-reference before modifying any query.


Error-Rate Spike

Source: system.query_log (per-query granularity) + system.errors (server-wide counters).

sql
-- Ratio of failed queries: recent vs baseline
-- recent  = last 1 hour, baseline = prior 24 hours
SELECT
    countIf(type = 'ExceptionWhileProcessing' AND event_time >= now() - INTERVAL 1 HOUR)  AS recent_failures,
    countIf(type != 'QueryStart'              AND event_time >= now() - INTERVAL 1 HOUR)  AS recent_total,
    countIf(type = 'ExceptionWhileProcessing' AND event_time  < now() - INTERVAL 1 HOUR
                                              AND event_time >= now() - INTERVAL 25 HOUR) AS baseline_failures,
    countIf(type != 'QueryStart'              AND event_time  < now() - INTERVAL 1 HOUR
                                              AND event_time >= now() - INTERVAL 25 HOUR) AS baseline_total,
    round(recent_failures  / nullIf(recent_total,   0) * 100, 2) AS recent_error_pct,
    round(baseline_failures / nullIf(baseline_total, 0) * 100, 2) AS baseline_error_pct
FROM system.query_log
WHERE is_initial_query = 1
sql
-- Top error codes in the recent window
SELECT exception_code, count() AS n, any(exception) AS sample
FROM system.query_log
WHERE type = 'ExceptionWhileProcessing'
  AND is_initial_query = 1
  AND event_time >= now() - INTERVAL 1 HOUR
GROUP BY exception_code
ORDER BY n DESC
LIMIT 20
sql
-- system.errors: codes whose value jumped since last hour
-- (no time column — compare last_error_time as a proxy)
SELECT name, code, value, last_error_time, last_error_message
FROM system.errors
WHERE last_error_time >= now() - INTERVAL 1 HOUR
ORDER BY value DESC
LIMIT 20

Query Duration Regression

Source: system.query_log.

sql
-- p95 duration: recent 1h vs baseline 24h
SELECT
    quantileIf(0.95)(query_duration_ms,
        event_time >= now() - INTERVAL 1 HOUR)  AS p95_recent_ms,
    quantileIf(0.95)(query_duration_ms,
        event_time  < now() - INTERVAL 1 HOUR
        AND event_time >= now() - INTERVAL 25 HOUR) AS p95_baseline_ms,
    round(p95_recent_ms / nullIf(p95_baseline_ms, 0), 2) AS ratio
FROM system.query_log
WHERE type = 'QueryFinish'
  AND is_initial_query = 1
  AND event_time >= now() - INTERVAL 25 HOUR
sql
-- Break down by user to find who regressed
SELECT
    user,
    quantileIf(0.95)(query_duration_ms,
        event_time >= now() - INTERVAL 1 HOUR)  AS p95_recent_ms,
    quantileIf(0.95)(query_duration_ms,
        event_time  < now() - INTERVAL 1 HOUR
        AND event_time >= now() - INTERVAL 25 HOUR) AS p95_baseline_ms,
    round(p95_recent_ms / nullIf(p95_baseline_ms, 0), 2) AS ratio
FROM system.query_log
WHERE type = 'QueryFinish'
  AND is_initial_query = 1
  AND event_time >= now() - INTERVAL 25 HOUR
GROUP BY user
ORDER BY ratio DESC
LIMIT 20

Query Volume Spike / Drop

Source: system.query_log.

sql
-- Queries per minute: recent 1h vs baseline hourly average
SELECT
    countIf(event_time >= now() - INTERVAL 1 HOUR) AS recent_count,
    countIf(event_time  < now() - INTERVAL 1 HOUR
        AND event_time >= now() - INTERVAL 25 HOUR) / 24 AS baseline_hourly_avg,
    round(recent_count / nullIf(baseline_hourly_avg, 0), 2) AS ratio
FROM system.query_log
WHERE type != 'QueryStart'
  AND is_initial_query = 1
  AND event_time >= now() - INTERVAL 25 HOUR
sql
-- Per-minute breakdown of the last 2 hours (spot sudden spikes or drops)
SELECT
    toStartOfMinute(event_time) AS minute,
    count() AS queries
FROM system.query_log
WHERE type != 'QueryStart'
  AND is_initial_query = 1
  AND event_time >= now() - INTERVAL 2 HOUR
GROUP BY minute
ORDER BY minute

Memory Anomalies

Source: system.query_log (memory_usage is per-query peak).

sql
-- Peak memory_usage: recent p95 vs baseline p95
SELECT
    quantileIf(0.95)(memory_usage,
        event_time >= now() - INTERVAL 1 HOUR)  AS p95_recent_bytes,
    quantileIf(0.95)(memory_usage,
        event_time  < now() - INTERVAL 1 HOUR
        AND event_time >= now() - INTERVAL 25 HOUR) AS p95_baseline_bytes,
    formatReadableSize(p95_recent_bytes)   AS p95_recent,
    formatReadableSize(p95_baseline_bytes) AS p95_baseline,
    round(p95_recent_bytes / nullIf(p95_baseline_bytes, 0), 2) AS ratio
FROM system.query_log
WHERE type = 'QueryFinish'
  AND is_initial_query = 1
  AND event_time >= now() - INTERVAL 25 HOUR
sql
-- Top memory consumers in the last hour
SELECT
    user, substring(query, 1, 200) AS query_preview,
    formatReadableSize(memory_usage) AS mem,
    query_duration_ms
FROM system.query_log
WHERE type = 'QueryFinish'
  AND is_initial_query = 1
  AND event_time >= now() - INTERVAL 1 HOUR
ORDER BY memory_usage DESC
LIMIT 10
sql
-- Current server-wide memory tracking (instantaneous)
SELECT metric, value, description
FROM system.metrics
WHERE metric IN ('MemoryTracking', 'MemoryAllocated')

Part-Count Explosion

Source: system.parts. High active-part counts per table indicate merge backpressure or excessive insert batching.

sql
-- Tables with the most active parts right now
SELECT
    database,
    table,
    count() AS active_parts,
    sum(rows) AS total_rows,
    formatReadableSize(sum(bytes_on_disk)) AS disk_size
FROM system.parts
WHERE active = 1
GROUP BY database, table
ORDER BY active_parts DESC
LIMIT 30
sql
-- Flag tables that exceeded a threshold in the last day
-- (join today vs yesterday snapshot via modification_time proxy)
SELECT
    database,
    table,
    countIf(modification_time >= now() - INTERVAL 1 HOUR)  AS parts_added_1h,
    countIf(modification_time >= now() - INTERVAL 25 HOUR) AS parts_total_25h,
    count() AS active_parts
FROM system.parts
WHERE active = 1
GROUP BY database, table
HAVING active_parts > 300 OR parts_added_1h > 50
ORDER BY active_parts DESC
LIMIT 20

Use the get_merge_status tool to check if background merges are keeping up. See troubleshooting for error code 252 (too many parts).


Replication Lag

Source: system.replicas. absolute_delay is seconds behind the leader.

sql
-- Tables with non-zero replication lag
SELECT
    database,
    table,
    is_leader,
    is_readonly,
    absolute_delay,
    queue_size,
    inserts_in_queue,
    merges_in_queue,
    active_replicas,
    total_replicas
FROM system.replicas
WHERE absolute_delay > 0 OR is_readonly = 1 OR active_replicas < total_replicas
ORDER BY absolute_delay DESC
sql
-- Trend: max absolute_delay across all replicated tables
SELECT
    database,
    table,
    absolute_delay,
    formatReadableTimeDelta(absolute_delay) AS lag_human
FROM system.replicas
ORDER BY absolute_delay DESC
LIMIT 10

For diagnosis and recovery steps, load the replication-guide skill.


How to Interpret Results

SignalThreshold (guide, not absolute)Action
Error-rate ratio recent_error_pct / baseline_error_pct> 2×Investigate top error codes; load troubleshooting
p95 duration ratio> 1.5×Check new/changed queries, index usage; load query-optimization
Query volume ratio< 0.3× or > 3×Verify application health; check for batch jobs or traffic anomaly
Memory p95 ratio> 2×Find heavy queries; consider max_memory_usage; load troubleshooting
active_parts per table> 300 (MergeTree default warn zone)Check merge queue via get_merge_status; consider OPTIMIZE
absolute_delay> 300 s (5 min)Replication is lagging; load replication-guide
Show full SKILL.md (124 more words)Show less

Noise vs real anomaly: a ratio spike in a 1-hour window with fewer than ~50 total events is likely noise (small sample). Check the raw counts (recent_total, baseline_total) before escalating. A sustained anomaly across two consecutive 1-hour windows is more actionable.

Baseline window caveat: the 24-hour baseline includes the recent 1 hour in some queries above for simplicity. If you want a strictly prior baseline, replace INTERVAL 25 HOUR with INTERVAL 24 HOUR and add AND event_time < now() - INTERVAL 1 HOUR to the baseline predicate.


Cross-references

  • system-tables-reference — exact column names before modifying any SQL above
  • troubleshooting — error codes, OOM, stuck mutations
  • replication-guide — replication lag diagnosis and recovery
  • query-optimization — EXPLAIN, PREWHERE, JOIN tuning for duration regressions
  • storage-optimization — disk pressure if part-count explosion fills the disk

© chmonitor, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/anomaly-detection of chmonitor/chmonitor.

Open the folder on GitHubat commit fc39ef0

Compare with similar skills

Anomaly Detection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Anomaly Detection compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Anomaly Detection this skillchmonitor/chmonitor299—~2.3kAutomated safety check: PassGPL-3.0
TimesFM Forecastinggoogle-research/timesfm34k—~4.7kAutomated safety check: PassApache-2.0
Anomalib Adding A Modelopen-edge-platform/anomalib6.2k—~1.9kAutomated safety check: PassApache-2.0
Anomalib Tiled Ensembleopen-edge-platform/anomalib6.2k—~1.4kAutomated safety check: PassApache-2.0
Kqlmicrosoft/fabric-rti-mcp131—~6.2kAutomated safety check: PassMIT
Time Series Analytics Useropen-edge-platform/edge-ai-libraries169—~3.1kAutomated safety check: PassApache-2.0

Similar skills

  • TimesFM Forecasting

    google-research/timesfm

    Forecasts any univariate time series zero-shot with Google's TimesFM model, returning point forecasts and calibrated prediction intervals without training.

    34k GitHub stars~4.7k tokensUpdated 9 days ago
    Data & AnalyticsAuto-check passed
  • Anomalib Adding A Model

    open-edge-platform/anomalib

    Adds a new anomaly-detection model to anomalib under src/anomalib/models/.

    6.2k GitHub stars~1.9k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Anomalib Tiled Ensemble

    open-edge-platform/anomalib

    Runs and configures the anomalib tiled-ensemble pipeline, which trains/evaluates one model per image tile and merges results (with optional seam smoothing) for high-resolution anomaly detection.

    6.2k GitHub stars~1.4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Kql

    microsoft/fabric-rti-mcp

    Official

    KQL language expertise for writing correct, efficient Kusto queries using the Fabric RTI MCP tools.

    131 GitHub stars~6.2k tokensUpdated 7 days ago
    Data & AnalyticsAuto-check passed
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    169 GitHub stars~3.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Dt Obs Analytics

    Dynatrace/dynatrace-for-ai

    Analyze dashboards and notebooks using Davis analyzers — anomaly detection, novelty scoring, and correlation.

    161 GitHub stars~3.9k tokensUpdated 7 days ago
    Data & AnalyticsAuto-check passed

More from chmonitor/chmonitor

All 53 skills in this repo
  • Hyperframes Creative

    chmonitor/chmonitor

    Non-animation creative direction for HyperFrames videos. An agent skill from chmonitor/chmonitor.

    299 GitHub starsUsed in 5 repos~1.3k tokens
    Auto-check passed
  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    299 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check: notes
  • Remotion To Hyperframes

    chmonitor/chmonitor

    Port an existing Remotion (React) composition to HyperFrames HTML.

    299 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Music To Video

    chmonitor/chmonitor

    A skill your agent uses when the user has a music track (an audio file, or a video to pull audio from) and wants a beat-synced HyperFrames video, calm to hard-hitting.

    299 GitHub starsUsed in 1 repo~4k tokens
    Auto-check: notes
  • Hyperframes Animation

    chmonitor/chmonitor

    All animation knowledge for HyperFrames — atomic motion rules, multi-phase scene blueprints, scene transitions, broader motion-design techniques, AND the seven runtime adapters (GSAP default, plus…

    299 GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • Faceless Explainer

    chmonitor/chmonitor

    turn arbitrary text — an article, notes, a topic, a brief — into a faceless explainer video, up to ~3 min (sweet spot 30-90s), where every visual is invented (typography, abstract graphics…

    299 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check: notes

Questions about Anomaly Detection

What does Anomaly Detection do?

Detect cluster abnormalities by comparing recent activity (last 1h) to a baseline (preceding 24h) using the raw query tool. Anomaly Detection is an agent skill from chmonitor/chmonitor. Detect cluster abnormalities by comparing recent activity (last 1h) to a baseline (preceding 24h) using the raw query tool.

When should I use Anomaly Detection?

Anomaly Detection fits situations like: tasks that involve Anomaly detection.

How do I install Anomaly Detection in Claude Code?

Run `npx skills add chmonitor/chmonitor --skill anomaly-detection -a claude-code`. Or copy the skill folder (.agents/skills/anomaly-detection in chmonitor/chmonitor) into .claude/skills/anomaly-detection in your project. Claude Code loads it when a task matches its description.

How do I install Anomaly Detection in Codex?

Run `npx skills add chmonitor/chmonitor --skill anomaly-detection -a codex`. Or copy the skill folder (.agents/skills/anomaly-detection in chmonitor/chmonitor) into .agents/skills/anomaly-detection in your project. Codex loads it when a task matches its description.

Can I use Anomaly Detection in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add chmonitor/chmonitor --skill anomaly-detection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/anomaly-detection, .gemini/skills/anomaly-detection, .github/skills/anomaly-detection and .opencode/skills/anomaly-detection in your project.

What does Anomaly Detection need to run?

SKILL.md names no scripts, command-line tools or credentials: Anomaly Detection is instructions for the agent only.

Does Anomaly Detection access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Anomaly Detection safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Anomaly Detection use?

Anomaly Detection is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Anomaly Detection use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Anomaly Detection?

Skills that share tags, products or a category with Anomaly Detection: TimesFM Forecasting (google-research/timesfm, 34k stars), Anomalib Adding A Model (open-edge-platform/anomalib, 6.2k stars), Anomalib Tiled Ensemble (open-edge-platform/anomalib, 6.2k stars) and Kql (microsoft/fabric-rti-mcp, 131 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Anomaly Detection?

chmonitor (a GitHub organization) maintains it in chmonitor/chmonitor, which has 299 GitHub stars. The repository holds 53 skills in this directory. The repository was last updated on October 5, 2026.

Source: chmonitor/chmonitor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.