Agent skill

Verify Production

by chmonitor in chmonitor/chmonitor

Prove the deployed chmonitor product actually works, not just that the Worker answers.

GPL-3.0Auto-check: notesDevOps & Cloud

Install Verify Production

skills CLI
$ npx skills add chmonitor/chmonitor --skill verify-production -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install chmonitor/chmonitor verify-production --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/chmonitor/chmonitor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/verify-production .claude/skills/verify-production && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verify-production
GitHub stars
299
Token cost
~2.2k tokens
SKILL.md length
1,009 words
Files
1
Skills in repo
53
Repo updated
First seen
Licence
GPL-3.0

At a glance

Prove the deployed chmonitor product actually works, not just that the Worker answers.

  • Works in 3 steps: Revert PR. git revert the deployed… → wrangler rollback only when service is… → Record the version ids and the action in…
  • Tasks that involve Data warehousing
  • SKILL.md covers What each signal is worth, --skip-auth passing does not…, Run it and Agent and guest-model probes, plus 3 more sections
  • Calls curl, bun and pnpm; reaches dash.chmonitor.dev and preview.dash.chmonitor.dev; needs CHM_API_KEY_SECRET and CLOUDFLARE_API_TOKEN

What it does

Verify Production is an agent skill from chmonitor/chmonitor. Prove the deployed chmonitor product actually works, not just that the Worker answers. Use after every deploy to dash.chmonitor.dev (or a preview), when triaging "no data" / "blank overview" / "agent does not answer", when a CI deploy was green but production looks broken, and when deciding whether to revert or roll back. Covers the verify-deploy contract, what the health endpoints do and do not prove, the agent/guest-model probes, usage and quota watch, and the restore-service order (revert PR first, wrangler…

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Data warehousing. It works with Cloudflare Workers and ClickHouse. The repository describes itself as: Open-source operational advisor for ClickHouse — real-time monitoring plus AI-driven index/partition/materialized-view recommendations. The licence is GPL-3.0.

When your agent uses it

  • Tasks that involve Data warehousing

Example prompts

  • “no data”
  • “blank overview”
  • “agent does not answer”
  • “/verify-production”

Requirements

  • A credential in CHM_API_KEY_SECRET
  • A credential in CLOUDFLARE_API_TOKEN

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Revert PR. git revert the deployed commit on a branch, then
  2. wrangler rollback only when service is down now, a revert PR cannot
  3. Record the version ids and the action in the run's changes.md.

What it can do on your machine

Read from SKILL.md and the folder at commit fc39ef0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • bun
    • pnpm
    • git
    • gh
    • wrangler
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • dash.chmonitor.dev
    • preview.dash.chmonitor.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • CHM_API_KEY_SECRET
    • CLOUDFLARE_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verify Production loads about 2.2k tokens when it runs. Until then it costs about 200 tokens; SKILL.md has 1,009 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~200
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:101
    ror. The secret lives in `apps/dashboard/.env.local` and the GitHub secret of

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from chmonitor/chmonitor at commit fc39ef0, republished under its GPL-3.0 licence (© chmonitor). 1,009 words, ~2,194 tokens.

Download SKILL.mdSave it as .claude/skills/verify-production/SKILL.md (or your agent's skills folder).
name
verify-production
description
Prove the deployed chmonitor product actually works, not just that the Worker answers. Use after every deploy to dash.chmonitor.dev (or a preview), when triaging "no data" / "blank overview" / "agent does not answer", when a CI deploy was green but production looks broken, and when deciding whether to revert or roll back. Covers the verify-deploy contract, what the health endpoints do and do not prove, the agent/guest-model probes, usage and quota watch, and the restore-service order (revert PR first, wrangler rollback only as an opt-in emergency). Triggers: "verify deploy", "production is down", "deploy looks green", "revert", "rollback", "wrangler rollback", "agent not responding", "guest AI", "model unavailable", "usage", "quota", "no data in production", "downtime".

Verify production

The dashboard deploys itself: .github/workflows/cloudflare.yml builds, ships, and runs bun scripts/verify-deploy.ts on every push to main. So CI proves the deploy path. It does not prove the product works, because the checks CI runs stop at "the endpoints answered".

This skill is the knowledge for the second pair of eyes: what to probe, what a green probe actually means, and what to do when it is red.

Script contract, gotchas, and the assertion list: scripts/verify-deploy.ts (apps/dashboard/). The desk job that runs all of this on a schedule is local:prod — see docs/herdr-desk/prod-watch.md.

What each signal is worth

ProbeProvesDoes NOT prove
GET /api/health (anon)the Worker is upwhich version is deployed
GET /api/health (authed)gitSha + buildTimestampthat the UI renders
GET /overview?host=0the shell serves and the entry bundle is referencedthat data loads
GET /api/v1/host-status?hostId=0 (anon)the edge can reach ClickHouse and the Worker can query it— (the best anonymous signal there is)
GET /api/v1/agents/config-check (anon)the agent has keys and a base URLthat any model id routes
GET /api/v1/agents/models (anon)the registry resolves to a listthat a specific id still routes upstream
POST /api/v1/auth/api-key → GET /api/v1/menu-counts?hostId=Hthe worker reaches ClickHouseanything about the agent
POST /api/v1/agenta real answer, end to end— (this is the only full path)

Anonymous /api/health returns {status, timestamp} only. Deployment metadata is deliberately withheld from anonymous callers (#1768), so you cannot compare gitSha to origin/main without a token. With CHM_API_KEY_SECRET in the env, verify-deploy.ts does the authenticated call for you.

--skip-auth passing does not mean ClickHouse is reachable

This is the trap, and it cost a full day of a broken guest experience before anyone noticed. The authenticated half of verify-deploy.ts is the only check that proves the Worker can query ClickHouse, and it needs CHM_API_KEY_SECRET. Without it, run the compensating check:

sh
curl -sS -o /dev/null -w '%{http_code}\n' \
  'https://dash.chmonitor.dev/api/v1/host-status?hostId=0'
curl -sS https://dash.chmonitor.dev/api/healthz | head -c 400
  • 500 with {"error":"error code: 1016"} — Cloudflare cannot resolve the origin domain. The signature of a ClickHouse host that only resolves on a tailnet: a *.ts.net MagicDNS name has no public DNS record, so it answers from a tailnet-connected box and nowhere else. Workers resolve through public DNS, so the edge gets 1016 while the host itself is alive and healthy.
  • 503 with hosts[0].status: "down" — the Worker is up; upstream is not. /api/healthz returns 503 for the whole deployment when one host is down, so read it as a host signal, not a Worker signal.

In cloud mode the env host list is the public demo (docs/knowledge/cloud-saas-mode.md) and an anonymous visitor is shown nothing else — so one unreachable demo host means every guest data read fails while /api/health, /overview, the Clerk-key check and the agent gateway all stay green. Check cloudMode.mismatch in the healthz payload to rule out a build/runtime split-brain.

If you are fixing the host: note that filterToDemoHosts (lib/cloud/demo-hosts.ts) fails open — an allowlist matching zero hosts shows all env hosts rather than emptying the demo — so renaming allowlist entries neither breaks nor fixes this. Fix the host, then reconcile the names.

Run it

sh
cd apps/dashboard

# anonymous only — no secret needed
bun scripts/verify-deploy.ts --skip-auth

# full, including ClickHouse connectivity
CHM_API_KEY_SECRET=… bun scripts/verify-deploy.ts --hosts 0

# a preview deploy
bun scripts/verify-deploy.ts \
  --base-url https://preview.dash.chmonitor.dev --json

Flags: --base-url <url> (default https://dash.chmonitor.dev), --hosts 0,1, --json, --skip-auth. Exit 0 = all pass, 1 = failures, 2 = harness error. The secret lives in apps/dashboard/.env.local and the GitHub secret of the same name — source it into the env, never print it.

When --skip-auth is all you can run (the secret is unset locally), add the host-status probe below. Four green unauthenticated checks and a broken guest product is a combination that actually happens.

Agent and guest-model probes

A dead model id is invisible in every health check, and that is exactly how it broke before: the guest default is the constant GUEST_DEFAULT_AGENT_MODEL in apps/dashboard/src/lib/billing/guest-ai.ts. That id can go stale while other surfaces still answer.

sh
curl -sS https://dash.chmonitor.dev/api/v1/agents/config-check | head -c 400
curl -sS https://dash.chmonitor.dev/api/v1/agents/models     | head -c 400

configured.apiKey: true with a model list that 404s upstream is a routing / account fact, not a code bug. Record it; do not file a repo issue against the code for it, and never change a model default or a quota as a drive-by. A default id is a product decision: it needs a human.

Show full SKILL.md (367 more words)Show less

Usage and quota

  • Guest AI: 3 requests/day per IP, 5/min (CHM_GUEST_AI_REQUESTS_PER_DAY, RATE_LIMIT_AGENT_GUEST_PER_MIN). Counts live in D1 ai_usage_daily (lib/billing/ai-usage-store.ts), keyed guest:<sha256-prefix> per IP.
  • GET /api/v1/billing/usage returns the owner's meters vs. plan caps (routes/api/v1/billing/usage.ts). Auth mirrors the other billing routes; anonymous cloud visitors get the slim guest payload.
  • Cap pressure that is not explained by real traffic means a retry loop, not users. That is an issue with the loop in it, not a quota change.
  • Never log a token, key, cookie, or user id. Aggregate counts only.

Restore service, in this order

  1. Revert PR. git revert the deployed commit on a branch, then gh pr merge --auto --squash. It takes the same required checks and the same deploy path as any other change, and it is reviewable. This is the default for every reason.

  2. wrangler rollback only when service is down now, a revert PR cannot land in time, and both CLOUDFLARE_API_TOKEN and CHM_ALLOW_INSTANT_ROLLBACK=1 are in the environment:

    sh
    cd apps/dashboard
    pnpm exec wrangler rollback          # to the previous version
    pnpm exec wrangler deployments list  # record the version ids either way

    Then open the revert PR anyway. An instant rollback with no follow-up leaves main and production diverged, which is a worse incident than the one you just stopped.

  3. Record the version ids and the action in the run's changes.md.

Gotchas learned

  • /api/v1/data rejects arbitrary SQL by design (permission_error). Only the registry endpoints (/api/v1/charts/[name], /api/v1/tables/[name], menu-counts, …) run queries — use those to prove connectivity.
  • The data API param is hostId, not host.
  • The ClickHouse hosts may be Tailscale Funnel URLs (*.ts.net): on the tailnet MagicDNS gives a private 100.x address, but public DNS resolves the funnel ingress, so https://<host>/ping → 200 and the CF edge can reach it.
  • A chm_ token is JWT-shaped (chm_<payload>.<sig>). Extract the full value including the .; a chm_[A-Za-z0-9_]+ regex truncates it and the failure looks like "malformed token".
  • Give the Worker a few seconds to propagate after a deploy before probing (CI sleeps 5).
  • The chart smoke test (scripts/smoke-test.ts) is continue-on-error: true in CI — informational, not a gate.
  • Claude Issues (red: missing Anthropic credentials) and promptfoo (red: gateway 404 model_unavailable) are known-informational on main. They are not code defects and must not gate anything.

© chmonitor, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/verify-production of chmonitor/chmonitor.

Open the folder on GitHubat commit fc39ef0

Compare with similar skills

Verify Production next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verify Production compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verify Production this skillchmonitor/chmonitor299—~2.2kAutomated safety check: NotesGPL-3.0
Monitoring Ingestion PipelinePostHog/posthog40k—~9.1kAutomated safety check: PassCustom licence
Clickhouse Observabilityjeremylongshore/tons-of-skills-marketplace2.8k—~1.4kAutomated safety check: PassMIT
Funnelcake Deployment Workflowdivinevideo/divine-mobile266—~3.6kAutomated safety check: PassMPL-2.0
Multi Cluster API Data Mismatchdivinevideo/divine-mobile266—~1.3kAutomated safety check: PassMPL-2.0
Releasehypequery/hypequery103—~623Automated safety check: PassCustom licence

Similar skills

  • Official

    Guide for using the Grafana MCP to monitor and diagnose the Node.js ingestion pipeline workers in production.

    40k GitHub stars~9.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Clickhouse Observability

    jeremylongshore/tons-of-skills-marketplace

    Monitor ClickHouse with Prometheus metrics, Grafana dashboards, system table queries, and alerting for query performance, merge health, and resource usage.

    2.8k GitHub stars~1.4k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Funnelcake Deployment Workflow

    divinevideo/divine-mobile

    Deploy funnelcake (api + relay) to ANY environment (production, staging, poc) on GKE via ArgoCD.

    266 GitHub stars~3.6k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Multi Cluster API Data Mismatch

    divinevideo/divine-mobile

    Debug "API returns data that doesn't exist in database" when multiple Kubernetes clusters exist (production, staging, POC).

    266 GitHub stars~1.3k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Release

    hypequery/hypequery

    Cut a stable hypequery release via Changesets, or explain/check the canary flow.

    103 GitHub stars~623 tokensUpdated today
    DatabasesAuto-check passed
  • Analyze Cloud Costs

    langfuse/langfuse

    Analyze Langfuse Cloud infrastructure cost structure using Metabase cost marts.

    36k GitHub stars~672 tokensUpdated today
    DevOps & CloudAuto-check passed

More from chmonitor/chmonitor

All 53 skills in this repo
  • Hyperframes Creative

    chmonitor/chmonitor

    Non-animation creative direction for HyperFrames videos. An agent skill from chmonitor/chmonitor.

    299 GitHub starsUsed in 5 repos~1.3k tokens
    Auto-check passed
  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    299 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check: notes
  • Remotion To Hyperframes

    chmonitor/chmonitor

    Port an existing Remotion (React) composition to HyperFrames HTML.

    299 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Music To Video

    chmonitor/chmonitor

    A skill your agent uses when the user has a music track (an audio file, or a video to pull audio from) and wants a beat-synced HyperFrames video, calm to hard-hitting.

    299 GitHub starsUsed in 1 repo~4k tokens
    Auto-check: notes
  • Hyperframes Animation

    chmonitor/chmonitor

    All animation knowledge for HyperFrames — atomic motion rules, multi-phase scene blueprints, scene transitions, broader motion-design techniques, AND the seven runtime adapters (GSAP default, plus…

    299 GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • Faceless Explainer

    chmonitor/chmonitor

    turn arbitrary text — an article, notes, a topic, a brief — into a faceless explainer video, up to ~3 min (sweet spot 30-90s), where every visual is invented (typography, abstract graphics…

    299 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check: notes

Questions about Verify Production

What does Verify Production do?

Prove the deployed chmonitor product actually works, not just that the Worker answers. Verify Production is an agent skill from chmonitor/chmonitor. Prove the deployed chmonitor product actually works, not just that the Worker answers.

When should I use Verify Production?

Verify Production fits situations like: tasks that involve Data warehousing.

How do I install Verify Production in Claude Code?

Run `npx skills add chmonitor/chmonitor --skill verify-production -a claude-code`. Or copy the skill folder (.claude/skills/verify-production in chmonitor/chmonitor) into .claude/skills/verify-production in your project. Claude Code loads it when a task matches its description.

How do I install Verify Production in Codex?

Run `npx skills add chmonitor/chmonitor --skill verify-production -a codex`. Or copy the skill folder (.claude/skills/verify-production in chmonitor/chmonitor) into .agents/skills/verify-production in your project. Codex loads it when a task matches its description.

Can I use Verify Production in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add chmonitor/chmonitor --skill verify-production -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verify-production, .gemini/skills/verify-production, .github/skills/verify-production and .opencode/skills/verify-production in your project.

What does Verify Production need to run?

Going by SKILL.md and its folder, Verify Production needs the command-line tools its instructions call (curl, bun, pnpm, git, gh and wrangler) and credentials named CHM_API_KEY_SECRET and CLOUDFLARE_API_TOKEN. Our summary lists: A credential in CHM_API_KEY_SECRET; A credential in CLOUDFLARE_API_TOKEN.

Does Verify Production access the network?

SKILL.md names 2 domains. In commands or code: dash.chmonitor.dev and preview.dash.chmonitor.dev; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Verify Production safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Verify Production use?

Verify Production is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verify Production use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Verify Production?

Skills that share tags, products or a category with Verify Production: Monitoring Ingestion Pipeline (PostHog/posthog, 40k stars), Clickhouse Observability (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Funnelcake Deployment Workflow (divinevideo/divine-mobile, 266 stars) and Multi Cluster API Data Mismatch (divinevideo/divine-mobile, 266 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verify Production?

chmonitor (a GitHub organization) maintains it in chmonitor/chmonitor, which has 299 GitHub stars. The repository holds 53 skills in this directory. The repository was last updated on October 5, 2026.

Source: chmonitor/chmonitor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.