Agent skill

Plan And Verify

by chmonitor in chmonitor/chmonitor

Decompose multi-step tasks into an explicit updateplan checklist, then verify each result before stating it as fact.

GPL-3.0Auto-check passed

Install Plan And Verify

skills CLI
$ npx skills add chmonitor/chmonitor --skill plan-and-verify -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install chmonitor/chmonitor plan-and-verify --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/chmonitor/chmonitor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/plan-and-verify .claude/skills/plan-and-verify && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
plan-and-verify
GitHub stars
299
Token cost
~2k tokens
SKILL.md length
835 words
Files
1
Skills in repo
53
Repo updated
First seen
Licence
GPL-3.0

At a glance

Decompose multi-step tasks into an explicit updateplan checklist, then verify each result before stating it as fact.

  • Works in 2 steps: The expected effect (e.g., "reduces part… → How the user can measure it: the exact…
  • SKILL.md covers When to Plan, Using the update_plan Tool, The VERIFY Discipline and Reporting: Verified vs.…, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Plan And Verify is an agent skill from chmonitor/chmonitor. Decompose multi-step tasks into an explicit updateplan checklist, then verify each result before stating it as fact.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Open-source operational advisor for ClickHouse — real-time monitoring plus AI-driven index/partition/materialized-view recommendations. The licence is GPL-3.0.

Example prompts

  • “/plan-and-verify”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. The expected effect (e.g., "reduces part count in this partition from ~800 to ~200 over one merge cycle")
  2. How the user can measure it: the exact query or metric to check before and after

What it can do on your machine

Read from SKILL.md and the folder at commit fc39ef0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are sql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Plan And Verify loads about 2k tokens when it runs. Until then it costs about 33 tokens; SKILL.md has 835 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from chmonitor/chmonitor at commit fc39ef0, republished under its GPL-3.0 licence (© chmonitor). 835 words, ~2,044 tokens.

Download SKILL.mdSave it as .claude/skills/plan-and-verify/SKILL.md (or your agent's skills folder).
name
plan-and-verify
description
Decompose multi-step tasks into an explicit update_plan checklist, then verify each result before stating it as fact.

Plan-and-Verify

Use this discipline for any task that spans 3 or more distinct actions. The goal is to avoid the two most common agent mistakes: stating a finding before it is confirmed, and losing track of what has actually been done.

When to Plan

Use update_plan when the work genuinely has multiple phases:

  • Investigations: 3+ queries or tool calls needed to reach a conclusion
  • Changes / recommendations: any action where a wrong answer has real cost (e.g., index advice, setting change, schema alter)
  • Anomaly diagnosis: root cause requires cross-checking multiple signals
  • Multi-table analysis: correlating data across system tables or time windows

Skip it for single-shot answers — if one query call settles the question, call it and respond directly. update_plan exists to make complex work transparent, not to add ceremony to simple requests.

Using the update_plan Tool

Call once up front to lay out the full plan. Set the first step to in_progress and every other step to pending (omitting status defaults to pending):

update_plan(steps=[
  { title: "Scan query_log for slow queries (last 24 h)", status: "in_progress" },
  { title: "Check merge backlog on affected tables" },
  { title: "Verify finding against a narrower time window" },
  { title: "Summarize confirmed findings" },
])

Call again after each step completes to advance the checklist. Mark the finished step completed, the next one in_progress, and leave the rest pending:

update_plan(steps=[
  { title: "Scan query_log for slow queries (last 24 h)", status: "completed" },
  { title: "Check merge backlog on affected tables",       status: "in_progress" },
  { title: "Verify finding against a narrower window" },
  { title: "Summarize confirmed findings" },
])

Rules:

  • Exactly one step is in_progress at any moment.
  • Keep plans to ≤ 7 steps. If something needs more, it is two separate tasks.
  • Titles are action-oriented and short (≤ 140 chars): "Scan query_log …", "Run EXPLAIN on both versions", "Cross-check baseline".
  • Add a note field when the current status needs a one-line callout: note: "High merge backlog confirmed — checking root cause".

The VERIFY Discipline

Produce a result. Then confirm it before reporting it. "Looked right" is not verification.

Data findings — re-run a tighter query

A wide-window query identifies a candidate. Before calling it a finding, re-run against a narrower window or a second system table to confirm the signal is real and not an artifact of the aggregation window.

sql
-- Initial: top tables by read_bytes, last 7 days
SELECT tables[1] AS tbl, avg(read_bytes) FROM system.query_log
WHERE event_date >= today() - 7 AND type = 'QueryFinish'
GROUP BY tbl ORDER BY avg(read_bytes) DESC LIMIT 10

-- Verify: same table, last 1 day — does the pattern hold?
SELECT tables[1] AS tbl, count(), avg(read_bytes)
FROM system.query_log
WHERE event_date = today() AND type = 'QueryFinish' AND tables[1] = '<candidate>'
GROUP BY tbl

If the narrower window contradicts the wide one, report the discrepancy — do not average the two.

Query rewrites — compare with EXPLAIN

Never claim a rewrite is "faster" without evidence. Use explain_query on both the original and the rewrite. Compare rows_read estimates. Report the ratio, not just "better".

explain_query(query="SELECT ... -- original")
explain_query(query="SELECT ... -- rewrite")

If rows_read is identical, the rewrite does not improve scan cost — say so even if the SQL looks cleaner.

Settings / schema recommendations — state the measurement

A recommendation without a measurable effect is a hypothesis. Every recommendation must include:

  1. The expected effect (e.g., "reduces part count in this partition from ~800 to ~200 over one merge cycle")
  2. How the user can measure it: the exact query or metric to check before and after

Example: recommending a lower merge_max_block_size:

Expected: smaller memory peaks per merge. Measure: SELECT max(memory_usage) FROM system.merges before and ~30 min after applying the setting.

Anomaly claims — confirm baseline and signal-to-noise

Before flagging a metric as anomalous:

  1. Establish the baseline window (e.g., same hour yesterday, or last 7-day average).
  2. Confirm the current value exceeds the baseline by a meaningful margin (not just rounding noise).
  3. Check whether the anomaly is isolated to one host or cluster-wide.
sql
-- Baseline: avg query duration same hour yesterday
SELECT avg(query_duration_ms) FROM system.query_log
WHERE type = 'QueryFinish' AND event_time BETWEEN yesterday() + INTERVAL 14 HOUR AND yesterday() + INTERVAL 15 HOUR

-- Current: same hour today
SELECT avg(query_duration_ms) FROM system.query_log
WHERE type = 'QueryFinish' AND event_time >= now() - INTERVAL 1 HOUR

If the current value is 1.05× the baseline, it is noise. If it is 4×, it is a finding.

Show full SKILL.md (315 more words)Show less

Reporting: Verified vs. Hypotheses

Always separate what you confirmed from what you inferred.

Structure your final response as:

  • Confirmed — findings backed by at least two data points or a re-run verification query
  • Likely — single data point, plausible but not cross-checked
  • Hypothesis — pattern that warrants investigation but was not verified in this session

Never present a hypothesis as a confirmed finding. Surface uncertainty explicitly: "This looks like a merge backlog issue, but I was not able to verify the part count trend — check system.part_log to confirm."

Worked Example: 4-Step Investigation

Request: "Why are queries on events slow since yesterday afternoon?"

Step 1 — Plan

update_plan(steps=[
  { title: "Scan query_log for slow queries on events (last 48 h)", status: "in_progress" },
  { title: "Check part count and merge backlog on events" },
  { title: "Verify: narrow window + cross-check parts growth" },
  { title: "Report confirmed findings with measurement guidance" },
])

Step 2 — Scan query_log
Query returns: avg(query_duration_ms) jumped from 420 ms to 3 100 ms starting at 14:00 yesterday. First guess: memory pressure.

Step 3 — Check merge backlog
get_merge_status shows 0 active merges. system.parts WHERE table = 'events' AND active = 1 returns 1 840 parts — unusually high. This contradicts the memory-pressure guess.

Step 4 — Verify
Re-run with event_time >= yesterday() + INTERVAL 13 HOUR (one hour before the slowdown). Parts were 340 at 13:00, grew to 1 840 by 15:00. The part explosion matches a batch insert job. Memory pressure ruled out.

update_plan(steps=[
  { title: "Scan query_log for slow queries on events (last 48 h)", status: "completed" },
  { title: "Check part count and merge backlog on events",           status: "completed" },
  { title: "Verify: narrow window + cross-check parts growth",       status: "completed" },
  { title: "Report confirmed findings with measurement guidance",    status: "in_progress" },
])

Report:

  • Confirmed: Part count on events grew from ~340 to ~1 840 between 13:00 and 15:00 yesterday, correlating directly with the query slowdown (4× duration increase). Caused by a high-frequency insert job creating many small parts.
  • Recommendation: Merge parts with OPTIMIZE TABLE events PARTITION <partition> FINAL (off-peak). Then reduce insert frequency or batch larger inserts. Measure: SELECT count() FROM system.parts WHERE table = 'events' AND active = 1 — should drop below 400 after merge completes.
  • Hypothesis (not verified): the insert job may be running on a fixed schedule — check system.query_log WHERE query LIKE '%INSERT INTO events%' grouped by hour to confirm recurrence.

Cross-references

  • anomaly-detection — baseline calculation patterns and signal thresholds
  • query-tuning-advisor — EXPLAIN interpretation and rewrite patterns
  • troubleshooting — error-code diagnosis workflows that pair well with this loop

© chmonitor, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/plan-and-verify of chmonitor/chmonitor.

Open the folder on GitHubat commit fc39ef0

Compare with similar skills

Plan And Verify next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Plan And Verify compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Plan And Verify this skillchmonitor/chmonitor299—~2kAutomated safety check: PassGPL-3.0
DecomposeFrkAk/piyaz194—~7.5kAutomated safety check: PassAGPL-3.0
Decompose Gateshappier-dev/happier1.9k—~1.5kAutomated safety check: PassMIT
Time Series Decomposerjeremylongshore/tons-of-skills-marketplace2.8k—~570Automated safety check: PassMIT
Task DecomposerMathews-Tom/armory328—~2.8kAutomated safety check: PassMIT
First Principles Decomposersundial-org/awesome-openclaw-skills663—~706Automated safety check: PassNone

Similar skills

  • Decompose

    FrkAk/piyaz

    A skill your agent uses when a Piyaz project exists with a description but few or no tasks, and the user wants it broken into an implementable graph (project-level decomposition).

    194 GitHub stars~7.5k tokensUpdated 8 days ago
    Agent WorkflowsAuto-check passed
  • Decompose Gates

    happier-dev/happier

    Decompose a hard or multi-part task into independently checkable pieces with explicit verification gates and risk-weighted ordering.

    1.9k GitHub stars~1.5k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Time Series Decomposer

    jeremylongshore/tons-of-skills-marketplace

    Manage time series decomposer operations. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~570 tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Task Decomposer

    Mathews-Tom/armory

    Produces phased task boards from feature requests: dependency-mapped work items, parallelization flags, risk flags, edge cases, test matrices.

    328 GitHub stars~2.8k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • First Principles Decomposer

    sundial-org/awesome-openclaw-skills

    Break any problem down to fundamental truths, then rebuild solutions from atoms up.

    663 GitHub stars~706 tokensUpdated 7 mo ago
    Auto-check passed
  • A skill your agent uses when the user wants to add a new feature, capability, or cluster of work to an existing active Piyaz project.

    194 GitHub stars~4.7k tokensUpdated 8 days ago
    Agent WorkflowsAuto-check passed

More from chmonitor/chmonitor

All 53 skills in this repo
  • Hyperframes Creative

    chmonitor/chmonitor

    Non-animation creative direction for HyperFrames videos. An agent skill from chmonitor/chmonitor.

    299 GitHub starsUsed in 5 repos~1.3k tokens
    Auto-check passed
  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    299 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check: notes
  • Remotion To Hyperframes

    chmonitor/chmonitor

    Port an existing Remotion (React) composition to HyperFrames HTML.

    299 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Music To Video

    chmonitor/chmonitor

    A skill your agent uses when the user has a music track (an audio file, or a video to pull audio from) and wants a beat-synced HyperFrames video, calm to hard-hitting.

    299 GitHub starsUsed in 1 repo~4k tokens
    Auto-check: notes
  • Hyperframes Animation

    chmonitor/chmonitor

    All animation knowledge for HyperFrames — atomic motion rules, multi-phase scene blueprints, scene transitions, broader motion-design techniques, AND the seven runtime adapters (GSAP default, plus…

    299 GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • Faceless Explainer

    chmonitor/chmonitor

    turn arbitrary text — an article, notes, a topic, a brief — into a faceless explainer video, up to ~3 min (sweet spot 30-90s), where every visual is invented (typography, abstract graphics…

    299 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check: notes

Questions about Plan And Verify

What does Plan And Verify do?

Decompose multi-step tasks into an explicit updateplan checklist, then verify each result before stating it as fact. Plan And Verify is an agent skill from chmonitor/chmonitor. Decompose multi-step tasks into an explicit updateplan checklist, then verify each result before stating it as fact.

How do I install Plan And Verify in Claude Code?

Run `npx skills add chmonitor/chmonitor --skill plan-and-verify -a claude-code`. Or copy the skill folder (.agents/skills/plan-and-verify in chmonitor/chmonitor) into .claude/skills/plan-and-verify in your project. Claude Code loads it when a task matches its description.

How do I install Plan And Verify in Codex?

Run `npx skills add chmonitor/chmonitor --skill plan-and-verify -a codex`. Or copy the skill folder (.agents/skills/plan-and-verify in chmonitor/chmonitor) into .agents/skills/plan-and-verify in your project. Codex loads it when a task matches its description.

Can I use Plan And Verify in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add chmonitor/chmonitor --skill plan-and-verify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/plan-and-verify, .gemini/skills/plan-and-verify, .github/skills/plan-and-verify and .opencode/skills/plan-and-verify in your project.

What does Plan And Verify need to run?

SKILL.md names no scripts, command-line tools or credentials: Plan And Verify is instructions for the agent only.

Does Plan And Verify access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Plan And Verify safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Plan And Verify use?

Plan And Verify is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Plan And Verify use?

About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Plan And Verify?

Skills that share tags, products or a category with Plan And Verify: Decompose (FrkAk/piyaz, 194 stars), Decompose Gates (happier-dev/happier, 1.9k stars), Time Series Decomposer (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Task Decomposer (Mathews-Tom/armory, 328 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Plan And Verify?

chmonitor (a GitHub organization) maintains it in chmonitor/chmonitor, which has 299 GitHub stars. The repository holds 53 skills in this directory. The repository was last updated on October 5, 2026.

Source: chmonitor/chmonitor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.