Agent skill

Monitoring

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and…

MITAuto-check passedDevOps & Cloud

Install Monitoring

skills CLI
$ npx skills add ericrisco/rsc-harness --skill monitoring -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness monitoring --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/monitoring .claude/skills/monitoring && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
monitoring
GitHub stars
156
Token cost
~3.1k tokens
SKILL.md length
1,509 words
Files
6 (incl. scripts, references)
Skills in repo
229
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and…

  • Works in 4 steps: Trigger a real synthetic failure (stop… → Confirm the page actually lands on a… → Let the ack timeout lapse and confirm… → …
  • Setting up uptime and health monitoring
  • SKILL.md covers The one rule, The 4-layer stack, Pick a tool and Health endpoints done right, plus 6 more sections
  • Runs Shell scripts from its folder

What it does

Monitoring is an agent skill from ericrisco/rsc-harness. Use when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and readiness probes, alerting on SLO error-budget burn rather than raw counts, curing alert fatigue, on-call rotation with escalation, and a status page. NOT instrumenting logs, metrics or traces, or wiring OpenTelemetry (that is observability).

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/burn-rate-and-oncall.md`).

It sits in DevOps & Cloud, covering Site reliability engineering, Observability and Incident response. It works with OpenTelemetry. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Setting up uptime and health monitoring
  • On-call basics for a service already in production
  • So you learn it is down before customers do — health and readiness probes
  • Alerting on SLO error-budget burn rather than raw counts

Example prompts

  • “/monitoring”

Requirements

  • Python 3
  • A Bash shell
  • Docker

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Trigger a real synthetic failure (stop the service, or point a monitor at a forced-503 route).
  2. Confirm the page actually lands on a phone — not just "the rule exists in the UI."
  3. Let the ack timeout lapse and confirm escalation hops to the secondary.
  4. Restore, and confirm the resolve/all-clear notification fires too.

What it can do on your machine

Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Monitoring loads about 3.1k tokens when it runs, and up to ~5.7k if it reads all its reference files. Until then it costs about 106 tokens; SKILL.md has 1,509 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,509 words, ~3,056 tokens.

Download SKILL.mdSave it as .claude/skills/monitoring/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
monitoring
description
Use when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and readiness probes, alerting on SLO error-budget burn rather than raw counts, curing alert fatigue, on-call rotation with escalation, and a status page. NOT instrumenting logs, metrics or traces, or wiring OpenTelemetry (that is `observability`).
tags
monitoring, uptime, alerting, on-call, sre
recommends
observability, deployment, error-handling, domains-dns, scaling
origin
risco

Monitoring

You are wiring up the outside view of a service that already shipped: is it alive, is it fast, and when it breaks, does exactly one human get exactly one actionable page. This skill emits a concrete setup — a checker config, a health-endpoint contract, symptom-based alert rules, and an on-call rotation. Not telemetry instrumentation (that is ../observability/SKILL.md), not the release-gating healthcheck (that is ../deployment/SKILL.md).

The one rule

A page is justified only when there is real user impact AND a human action the system can't take itself. Everything below descends from this. Internalize the three tiers:

  • Page (wake someone): users are hurting now and a human must intervene. Checkout returns 5xx. Site unreachable. Error budget burning fast.
  • Ticket (look during business hours): degraded but not bleeding. Slow-burn budget use, cert expiring in 14 days.
  • Dashboard-only (don't notify): CPU at 80%, a single retry, a transient blip the system already healed.

If an alert doesn't map to an immediate human action, it is not a page — it's noise, and noise trains people to ignore the one page that matters.

The 4-layer stack

Don't skip layers and don't collapse them — each answers a different question.

LayerAnswersBuilt withFires when
1. External uptime probe"Is it reachable from the outside?"Uptime Kuma / UptimeRobot / Better StackURL down, TLS broken, p95 latency over budget
2. Health endpoints"Is the process alive, and are its deps reachable?"/livez + /readyz on the serviceliveness fails → restart; readiness fails → pull from rotation
3. SLO burn-rate alert"Are we spending the error budget too fast?"Prometheus/Grafana/Better Stack rulemulti-window burn rate exceeds threshold
4. On-call escalation"Who acts, and who's the backup?"PagerDuty / incident.io / Better Stacka page from layers 1–3 routes + escalates

Layer 1 catches "the whole thing is gone." Layer 2 catches "a dependency died" before users do. Layer 3 catches "we're degrading faster than we can afford." Layer 4 makes sure a human shows up.

Pick a tool

Decide by budget, team size, and self-host appetite. Pricing as of 2026-06.

ToolFree tierCheck intervalBest forPaid from
Uptime Kuma 2.1.3Fully free, self-hosteddown to ~20sfull control, you have a VPS$0 (you pay the box)
UptimeRobot50 monitors @ 5-min, 1 status page5-min free / 1-min paidno infra, want managed$7/mo
Better Stack10 monitors + incidents + logs30sincident mgmt + on-call in one$24/mo
Pingdomnonesub-minuteenterprise synthetic + RUM$15/mo

Default: Uptime Kuma if you already run a VPS (1 GB box is comfortable; needs ~400 MB RAM), UptimeRobot free tier if you don't want infra. Kuma 2.1 added Globalping worldwide probe locations (so you test from regions, not just your one box) and built-in domain-expiry monitors. Concrete docker-compose and notification wiring live in references/tool-setup.md.

Health endpoints done right

Split liveness from readiness — they trigger different machine actions.

  • /livez (liveness): is the process healthy? If it fails, the orchestrator restarts the container. Keep it dumb: just "am I running and not deadlocked." Never check the database here.
  • /readyz (readiness): are dependencies reachable so I can serve traffic? If it fails, the orchestrator pulls this instance from the load-balancer rotation but does not kill it.

The cheap-probe rule: a probe runs constantly, so it must be <100ms and must not cascade-check every downstream. Why: if /livez pings the DB and the DB is briefly slow, liveness fails, the container restarts, the restart hammers the recovering DB — a restart loop that turns a 30-second blip into an outage.

python
# Bad — one /health that cascades and returns 500 on any hiccup.
# A slow Redis takes the whole service down and triggers restart loops.
@app.get("/health")
def health():
    db.execute("SELECT 1")          # blocks
    redis.ping()                    # blocks
    requests.get(PAYMENTS_URL)      # blocks on a third party!
    return {"status": "ok"}         # 500 if ANY of these throws
python
# Good — split, cheap, correct status codes, small JSON.
@app.get("/livez")                  # liveness: process only. Restart if this fails.
def livez():
    return {"status": "alive"}      # 200, ~1ms, touches nothing downstream

@app.get("/readyz")                 # readiness: deps with short timeouts. Pull from LB if this fails.
def readyz():
    checks = {"db": ping(db, timeout=0.2), "cache": ping(redis, timeout=0.1)}
    ok = all(checks.values())
    return JSONResponse(
        {"status": "ready" if ok else "degraded", "checks": checks},
        status_code=200 if ok else 503,   # 503 so the LB pulls this instance
    )

Do not put a third-party API call in readiness — a payment provider's outage shouldn't pull all your instances from rotation and take you fully down. Degrade that path in code (../error-handling/SKILL.md), don't fail-closed on it here. Go and FastAPI handler examples are in references/tool-setup.md.

What to actually monitor

The golden checklist. Monitor the symptom users feel, not just that one URL returns 200.

  • Availability — the critical endpoint(s) reachable, from multiple regions.
  • Latency p95/p99 — averages hide the tail; alert on p95/p99 against a budget, not the mean.
  • Error rate — the 5xx-to-total ratio, because raw 5xx count says nothing without traffic volume.
  • SSL cert expiry — alert at 14 days; an expired cert is a full outage that no app metric catches.
  • Domain expiry — alert at 30 days; a lapsed domain is the most embarrassing avoidable outage.
  • The critical user journey — a synthetic check that does login → core action → result, NOT just GET /. The homepage can be 200 while checkout is broken.

The homepage being up tells you almost nothing. Probe the path that makes you money.

Alerts that don't cry wolf

Alert on symptoms, not causes. Page on "checkout error rate >2% for 5 min" (user impact), not on "CPU >80%" (a cause that may be harmless and self-resolving). High CPU with happy users is a dashboard line, not a 3am page.

Use multi-window, multi-burn-rate for SLO alerts (Google SRE workbook). Burn rate = how fast you're spending the monthly error budget. Require a long and a short window to both fire — the long window says "this is real," the short window says "this is still happening," and together they kill false positives from a single spike.

Burn rateWindowBudget spentSeverityAction
> 14.41h (+ 5m short)2% in 1hcriticalpage
> 66h (+ 30m short)5% in 6hwarningticket
> 13d (+ 6h short)10% in 3dinforeview

Routing: critical → page, warning → ticket, info → dashboard/review. Dedupe and group related alerts into one incident (10 hosts failing the same check = one page, not ten). Set maintenance windows so planned deploys don't page anyone. The full burn-rate math and a copy-paste Prometheus-style rule with a runbook_url annotation are in references/burn-rate-and-oncall.md.

Show full SKILL.md (572 more words)Show less

On-call basics

  • Rotation: weekly, with a primary and a secondary. One person can't be the single point of failure for the system that catches single points of failure.
  • Escalation policy: page primary → if no ack in 5–10 min, page secondary → then the manager. No-ack must always hop; an unacked page is a dropped page.
  • Runbook per alert: every alert links to a runbook with five fields — Symptom, Impact, First 3 checks, Mitigation, How to escalate. The person paged at 3am should not have to think from scratch. Template in references/burn-rate-and-oncall.md.
  • Status page + comms: a public status page (Kuma and Better Stack include one) so customers self-serve "is it you or me," cutting inbound during an incident.
  • Do NOT start a new on-call on Opsgenie — Atlassian is retiring it (EOL 2027-04-05; new sales ended 2025-06-04). Grafana OnCall OSS was also deprecated (folded into Grafana Cloud IRM). For a new setup use PagerDuty, incident.io, Better Stack, or Grafana Cloud IRM.

Verify it works

An untested alert is not an alert. Before you call monitoring "done," prove the wire end-to-end:

  1. Trigger a real synthetic failure (stop the service, or point a monitor at a forced-503 route).
  2. Confirm the page actually lands on a phone — not just "the rule exists in the UI."
  3. Let the ack timeout lapse and confirm escalation hops to the secondary.
  4. Restore, and confirm the resolve/all-clear notification fires too.

scripts/verify.sh enforces the structural half of this on your config: liveness split from readiness, at least one symptom/burn-rate alert with two windows, a runbook reference on every alert, and a banlist for Opsgenie-as-new-setup and homepage-only monitors. Run it in CI so config drift can't silently re-introduce a noisy or untested setup.

Anti-patterns

BadWhy it bitesDo instead
Monitor only the homepage // is 200 while checkout is broken — you learn from customerssynthetic check of the critical journey
Page on CPU/memory thresholdself-resolves, no user impact → alert fatiguepage on the symptom (error rate, latency, availability)
Alert on cause, not symptomcauses are noisy and ambiguousalert on what the user feels
Alert with no runbook linkthe paged human improvises at 3amevery alert links a 5-field runbook
Single-window threshold "by vibes"one spike pages; tuned by guessworkmulti-window multi-burn-rate against an SLO
Probe cascades every downstreamone slow dep fails the probe → restart loopcheap <100ms probe, deps in readiness with timeouts
Liveness checks the databaseslow DB → restart loop turns a blip into an outageliveness = process only; DB lives in readiness
Start new on-call on OpsgenieEOL 2027-04-05, dead endPagerDuty / incident.io / Better Stack / Grafana Cloud IRM
No secondary on-callprimary asleep/offline → page droppedprimary + secondary + manager escalation
No status pageinbound floods support during incidentspublic status page customers can self-serve
Never tested the alert"the rule exists" ≠ "the page lands"trigger a synthetic failure, confirm it pages + escalates
Page on warningstrains people to ignore pageswarning → ticket; only critical → page

References

  • references/tool-setup.md — Uptime Kuma docker-compose (rootless image, volume, port), HTTP + push-heartbeat + SSL/domain-expiry monitors, a synthetic multi-step journey, notification wiring (ntfy / Slack webhook / PagerDuty integration key), a UptimeRobot/Better Stack monitor JSON shape, and Go + FastAPI health-endpoint handlers.
  • references/burn-rate-and-oncall.md — full multi-window multi-burn-rate math, a copy-paste Prometheus-style alert rule with runbook_url, severity mapping, a sample escalation-policy YAML, and a fill-in runbook template.

Related: ../observability/SKILL.md (what the app emits), ../deployment/SKILL.md (release-gating healthcheck + rollback), ../error-handling/SKILL.md (degrade in code), ../domains-dns/SKILL.md (provision certs/DNS), ../scaling/SKILL.md (survive the load monitoring detected).

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/monitoring of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/burn-rate-and-oncall.md
  • references/tool-setup.md
  • scripts/verify.sh

Open the folder on GitHubat commit 92fde8f

Compare with similar skills

Monitoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Monitoring compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Monitoring this skillericrisco/rsc-harness156—~3.1kAutomated safety check: PassMIT
Monitoring EngineerFerroxLabs/wayland608—~3.9kAutomated safety check: PassApache-2.0
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Oma Observabilityfirst-fluke/oh-my-agent1.3k—~4.9kAutomated safety check: PassMIT
Observability Sre Triageelastic/agent-skills592—~7.4kAutomated safety check: PassApache-2.0
Observability MonitoringAnastasiyaW/codex-claude-code-config154—~4.1kAutomated safety check: PassMIT

Similar skills

  • Monitoring Engineer

    FerroxLabs/wayland

    Observability and monitoring. An agent skill from FerroxLabs/wayland.

    608 GitHub stars~3.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 5 mo ago
    DevOps & CloudAuto-check passed
  • Oma Observability

    first-fluke/oh-my-agent

    Intent-based observability + traceability router across layers, boundaries, and signals.

    1.3k GitHub stars~4.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Observability Sre Triage

    elastic/agent-skills

    Official

    Triage a degraded or suspect service end to end: read SLO status and burn rate, check active alerting rules and ML anomalies, measure throughput, latency, and error rate, assess dependency health…

    592 GitHub stars~7.4k tokensUpdated 5 days ago
    DevOps & CloudAuto-check passed
  • Observability Monitoring

    AnastasiyaW/codex-claude-code-config

    Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…

    154 GitHub stars~4.1k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • Observability Sre

    majiayu000/spellbook

    Observability and SRE expert. An agent skill from majiayu000/spellbook.

    286 GitHub stars~3.3k tokensUpdated today
    DevOps & CloudAuto-check passed

More from ericrisco/rsc-harness

All 229 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    156 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    156 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    156 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    156 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Monitoring

What does Monitoring do?

A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and…. Monitoring is an agent skill from ericrisco/rsc-harness. Use when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and readiness probes, alerting on SLO error-budget burn rather than raw counts, curing alert fatigue, on-call rotation with escalation, and a status page.

When should I use Monitoring?

Monitoring fits situations like: setting up uptime and health monitoring; on-call basics for a service already in production; so you learn it is down before customers do — health and readiness probes; alerting on SLO error-budget burn rather than raw counts.

How do I install Monitoring in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill monitoring -a claude-code`. Or copy the skill folder (skills/monitoring in ericrisco/rsc-harness) into .claude/skills/monitoring in your project. Claude Code loads it when a task matches its description.

How do I install Monitoring in Codex?

Run `npx skills add ericrisco/rsc-harness --skill monitoring -a codex`. Or copy the skill folder (skills/monitoring in ericrisco/rsc-harness) into .agents/skills/monitoring in your project. Codex loads it when a task matches its description.

Can I use Monitoring in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill monitoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/monitoring, .gemini/skills/monitoring, .github/skills/monitoring and .opencode/skills/monitoring in your project.

What does Monitoring need to run?

Going by SKILL.md and its folder, Monitoring needs a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell; Docker.

Does Monitoring access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Monitoring safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Monitoring use?

Monitoring is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Monitoring use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.

What are the alternatives to Monitoring?

Skills that share tags, products or a category with Monitoring: Monitoring Engineer (FerroxLabs/wayland, 608 stars), Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Oma Observability (first-fluke/oh-my-agent, 1.3k stars) and Observability Sre Triage (elastic/agent-skills, 592 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Monitoring?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.