Agent skill

Regression Watch

by hoangsonww in hoangsonww/Claude-Code-Agent-Monitor

Detect quality and efficiency regressions over time using Agent Monitor data — rising error rate (APIError events), falling cache hit rate, growing compaction frequency, and climbing cost-per-session.

MITAuto-check passed

Install Regression Watch

skills CLI
$ npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill regression-watch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hoangsonww/Claude-Code-Agent-Monitor regression-watch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hoangsonww/Claude-Code-Agent-Monitor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ccam-insights/skills/regression-watch .claude/skills/regression-watch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
regression-watch
GitHub stars
1.1k
Token cost
~1k tokens
SKILL.md length
427 words
Files
2
Skills in repo
78
Repo updated
First seen
Licence
MIT

At a glance

Detect quality and efficiency regressions over time using Agent Monitor data — rising error rate (APIError events), falling cache hit rate, growing compaction frequency, and climbing cost-per-session.

  • Works in 6 steps: Windowing → Error Rate Regression → Cache Hit Rate Regression → …
  • Checking whether things are degrading
  • SKILL.md covers Input, Data Sources, Report Sections and Output
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Regression Watch is an agent skill from hoangsonww/Claude-Code-Agent-Monitor. Detect quality and efficiency regressions over time using Agent Monitor data — rising error rate (APIError events), falling cache hit rate, growing compaction frequency, and climbing cost-per-session. Splits history into an earlier baseline window and a recent window and reports which metrics are getting worse, by how much, and where. Use when checking whether things are degrading or trending in the wrong direction.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

The repository describes itself as: 🚀 A real-time monitoring dashboard for Claude Code & Codex, built with SQLite3, Node.js, Express, React, Vite, TailwindCSS, & WebSockets. It tracks sessions, agent activity… The licence is MIT.

When your agent uses it

  • Checking whether things are degrading
  • Trending in the wrong direction

Example prompts

  • “/regression-watch”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Windowing
  2. Error Rate Regression
  3. Cache Hit Rate Regression
  4. Compaction Frequency Regression
  5. Cost-per-Session Regression
  6. Verdict

What it can do on your machine

Read from SKILL.md and the folder at commit a06db03. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Regression Watch loads about 1k tokens when it runs. Until then it costs about 109 tokens; SKILL.md has 427 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~109
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hoangsonww/Claude-Code-Agent-Monitor at commit a06db03, republished under its MIT licence (© hoangsonww). 427 words, ~1,018 tokens.

Download SKILL.mdSave it as .claude/skills/regression-watch/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
regression-watch
description
Detect quality and efficiency regressions over time using Agent Monitor data — rising error rate (APIError events), falling cache hit rate, growing compaction frequency, and climbing cost-per-session. Splits history into an earlier baseline window and a recent window and reports which metrics are getting worse, by how much, and where. Use when checking whether things are degrading or trending in the wrong direction.

Regression Watch

Detect whether Claude Code sessions are getting worse over time across quality and efficiency metrics, using Agent Monitor data.

Input

The user provides: $ARGUMENTS

This may be:

  • empty or "all" — check every regression metric (default)
  • "errors" — error-rate regression only
  • "cache" — cache hit-rate regression only
  • "compaction" — compaction-frequency regression only
  • "cost" — cost-per-session regression only
  • A window like "last 30d" or "30 vs 90" — set the recent vs baseline window sizes

Data Sources

EndpointReturns
GET /api/analyticsdaily_events (365d), daily_sessions (365d), event_types, tokens (total_input, total_output, total_cache_read, total_cache_write — baselines pre-summed), avg_events_per_session
GET /api/events?session_id=XEvent stream incl. APIError, Compaction, PreToolUse/PostToolUse — used to localize regressions to specific sessions
GET /api/pricing/cost{ total_cost, breakdown[...] } — total cost to derive cost-per-session
GET /api/pricing/cost/{sessionId}Per-session cost — used to compare recent vs baseline session cost
GET /api/workflows/{sessionId}compaction (impact), errorPropagation (by depth), effectiveness — per-session quality signals
GET /api/sessions?limit=NSessions with started_at, cost, metadata — to bucket sessions into time windows

Report Sections

1. Windowing

Split history into a baseline window (older) and a recent window (newer). Default: recent = last 30 days, baseline = the 30–90 day range before it. Use daily_events/daily_sessions for series metrics and GET /api/sessions?limit=N to assign sessions to each window by started_at.

2. Error Rate Regression
  • Recent error rate = APIError count / total events in the recent window (from event_types and daily_events, or per-session GET /api/events).
  • Compare to the baseline rate. Flag if recent is higher.
  • Report the absolute and relative change and which sessions contributed most APIError events.
Show full SKILL.md (188 more words)Show less
3. Cache Hit Rate Regression
  • Cache hit rate = total_cache_read / (total_cache_read + total_input).
  • Compute for each window (per-window input/cache_read from session metadata or the pricing breakdown). Flag a falling hit rate — that means more uncached input tokens and higher cost.
4. Compaction Frequency Regression
  • Compaction frequency = Compaction events / session per window (from event_types / daily_events, confirmed via per-session GET /api/workflows/{id} compaction). Flag a rising rate — context is overflowing more often.
5. Cost-per-Session Regression
  • Cost-per-session = window total cost / window session count, using GET /api/pricing/cost overall and GET /api/pricing/cost/{id} for the sessions in each window. Flag a climbing value.
6. Verdict

Roll up which metrics regressed, rank by relative worsening, and name the most likely driver (e.g., cache hit rate fell → cost per session climbed).

Output

  • A Markdown table: metric | baseline | recent | Δ | direction (▲ worse / ▼ better) | verdict.
  • Tag each regressed metric 🔴 (clear regression), 🟡 (mild/within noise), or 🟢 (improved).
  • Currency in USD to 4 decimals; rates as percentages to 2 decimals.
  • List the specific session IDs that contributed most to any regression.
  • End with the single highest-priority regression to address and a concrete next step.
  • Read-only: only report what the API returns; never fabricate baselines.

© hoangsonww, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in plugins/ccam-insights/skills/regression-watch of hoangsonww/Claude-Code-Agent-Monitor.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit a06db03

Compare with similar skills

Regression Watch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Regression Watch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Regression Watch this skillhoangsonww/Claude-Code-Agent-Monitor1.1k—~1kAutomated safety check: PassMIT
Detecting Performance Regressionsjeremylongshore/tons-of-skills-marketplace2.8k—~1.1kAutomated safety check: PassMIT
Detection And Monitoringcbrock84/headcount2k—~1.3kAutomated safety check: PassMIT
Canary Watchaffaan-m/ECC276k1 repos~770Automated safety check: PassMIT
Detecting Performance Regressionsforyourhealth111-pixel/Vibe-Skills3.6k—~355Automated safety check: PassMIT
Visual Regressionthedaviddias/Front-End-Checklist74k—~493Automated safety check: PassMIT

Similar skills

  • Detecting Performance Regressions

    jeremylongshore/tons-of-skills-marketplace

    Automatically detect performance regressions in CI/CD pipelines by comparing metrics against baselines.

    2.8k GitHub stars~1.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Detection And Monitoring

    cbrock84/headcount

    Builds the capability to notice an attack in progress — deciding what to log and retain, centralizing it somewhere tamper-resistant, writing detections that fire on attacker behavior rather than on…

    2k GitHub stars~1.3k tokensUpdated 22 days ago
    DevOps & CloudAuto-check passed
  • Canary Watch

    affaan-m/ECC

    A skill your agent uses to monitor and verify a deployed URL after releases — checks HTTP endpoints, SSE streams, static assets, console errors, and performance regressions after deploys, merges, or…

    276k GitHub starsUsed in 1 repo~770 tokens
    DevOps & CloudAuto-check passed
  • Detecting Performance Regressions

    foryourhealth111-pixel/Vibe-Skills

    Compare current benchmark results against historical baselines to spot performance regressions.

    3.6k GitHub stars~355 tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Visual Regression

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing CI coverage, automated checks, or test strategy related to Use visual regression testing.

    74k GitHub stars~493 tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Implementing File Integrity Monitoring With Aide

    mukul975/Anthropic-Cybersecurity-Skills

    Configures AIDE (Advanced Intrusion Detection Environment) for file integrity monitoring on Linux, covering baseline database creation, scheduled integrity checks via cron, change detection, and…

    34k GitHub stars~642 tokensUpdated 1 mo ago
    SecurityAuto-check: notes

More from hoangsonww/Claude-Code-Agent-Monitor

All 78 skills in this repo
  • Version Release

    hoangsonww/Claude-Code-Agent-Monitor

    Choose and apply the correct semantic version bump for this repository.

    1.1k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Budget Set

    hoangsonww/Claude-Code-Agent-Monitor

    Define a spend budget for Claude Code and, optionally, create a cost alert rule that fires when usage crosses the limit, via POST /api/alerts/rules on the Agent Monitor dashboard.

    1.1k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Cache Efficiency

    hoangsonww/Claude-Code-Agent-Monitor

    Analyze prompt-cache effectiveness for Claude Code usage from the Agent Monitor dashboard — cache hit rate (totalcacheread / (totalcacheread + totalinput)), cachewrite vs cacheread reuse, cache-read…

    1.1k GitHub stars~966 tokensUpdated today
    Auto-check passed
  • Cost Breakdown

    hoangsonww/Claude-Code-Agent-Monitor

    Break down Claude Code costs using the Agent Monitor pricing engine.

    1.1k GitHub stars~845 tokensUpdated today
    Auto-check passed
  • Dag Map

    hoangsonww/Claude-Code-Agent-Monitor

    Render the multi-agent orchestration DAG for a session — parent→child subagent edges, tree depth, and fan-out — from the Agent Monitor workflow intelligence API.

    1.1k GitHub stars~564 tokensUpdated today
    Auto-check passed
  • Dashboard Status

    hoangsonww/Claude-Code-Agent-Monitor

    Quick dashboard health and status overview — checks the Agent Monitor API (port 4820), reports session/agent/event counts from /api/stats, confirms WebSocket connectivity, reads the redacted hook…

    1.1k GitHub stars~600 tokensUpdated today
    Auto-check passed

Questions about Regression Watch

What does Regression Watch do?

Detect quality and efficiency regressions over time using Agent Monitor data — rising error rate (APIError events), falling cache hit rate, growing compaction frequency, and climbing cost-per-session. Regression Watch is an agent skill from hoangsonww/Claude-Code-Agent-Monitor. Detect quality and efficiency regressions over time using Agent Monitor data — rising error rate (APIError events), falling cache hit rate, growing compaction frequency, and climbing cost-per-session.

When should I use Regression Watch?

Regression Watch fits situations like: checking whether things are degrading; trending in the wrong direction.

How do I install Regression Watch in Claude Code?

Run `npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill regression-watch -a claude-code`. Or copy the skill folder (plugins/ccam-insights/skills/regression-watch in hoangsonww/Claude-Code-Agent-Monitor) into .claude/skills/regression-watch in your project. Claude Code loads it when a task matches its description.

How do I install Regression Watch in Codex?

Run `npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill regression-watch -a codex`. Or copy the skill folder (plugins/ccam-insights/skills/regression-watch in hoangsonww/Claude-Code-Agent-Monitor) into .agents/skills/regression-watch in your project. Codex loads it when a task matches its description.

Can I use Regression Watch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill regression-watch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/regression-watch, .gemini/skills/regression-watch, .github/skills/regression-watch and .opencode/skills/regression-watch in your project.

What does Regression Watch need to run?

SKILL.md names no scripts, command-line tools or credentials: Regression Watch is instructions for the agent only.

Does Regression Watch access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Regression Watch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Regression Watch use?

Regression Watch is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Regression Watch use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Regression Watch?

Skills that share tags, products or a category with Regression Watch: Detecting Performance Regressions (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Detection And Monitoring (cbrock84/headcount, 2k stars), Canary Watch (affaan-m/ECC, 276k stars) and Detecting Performance Regressions (foryourhealth111-pixel/Vibe-Skills, 3.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Regression Watch?

hoangsonww (a GitHub user) maintains it in hoangsonww/Claude-Code-Agent-Monitor, which has 1,058 GitHub stars. The repository holds 78 skills in this directory. The repository was last updated on October 10, 2026.

Source: hoangsonww/Claude-Code-Agent-Monitor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.