Monitoring Engineer
FerroxLabs/wayland
Observability and monitoring. An agent skill from FerroxLabs/wayland.
A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and…
$ npx skills add ericrisco/rsc-harness --skill monitoring -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness monitoring --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/monitoring .claude/skills/monitoring && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "monitoring" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/monitoring into .claude/skills/monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/monitoringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill monitoring -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness monitoring --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/monitoring .agents/skills/monitoring && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "monitoring" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/monitoring into .agents/skills/monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill monitoring -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness monitoring --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/monitoring .cursor/skills/monitoring && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "monitoring" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/monitoring into .cursor/skills/monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/monitoring--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill monitoring -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness monitoring --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/monitoring .gemini/skills/monitoring && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "monitoring" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/monitoring into .gemini/skills/monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness monitoringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill monitoring -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/monitoring .github/skills/monitoring && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "monitoring" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/monitoring into .github/skills/monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill monitoring -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness monitoring --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/monitoring .opencode/skills/monitoring && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "monitoring" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/monitoring into .opencode/skills/monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
monitoringA skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and…
Monitoring is an agent skill from ericrisco/rsc-harness. Use when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and readiness probes, alerting on SLO error-budget burn rather than raw counts, curing alert fatigue, on-call rotation with escalation, and a status page. NOT instrumenting logs, metrics or traces, or wiring OpenTelemetry (that is observability).
Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/burn-rate-and-oncall.md`).
It sits in DevOps & Cloud, covering Site reliability engineering, Observability and Incident response. It works with OpenTelemetry. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Monitoring loads about 3.1k tokens when it runs, and up to ~5.7k if it reads all its reference files. Until then it costs about 106 tokens; SKILL.md has 1,509 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,509 words, ~3,056 tokens.
.claude/skills/monitoring/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.You are wiring up the outside view of a service that already shipped: is it alive, is it fast, and when it breaks, does exactly one human get exactly one actionable page. This skill emits a concrete setup — a checker config, a health-endpoint contract, symptom-based alert rules, and an on-call rotation. Not telemetry instrumentation (that is ../observability/SKILL.md), not the release-gating healthcheck (that is ../deployment/SKILL.md).
A page is justified only when there is real user impact AND a human action the system can't take itself. Everything below descends from this. Internalize the three tiers:
If an alert doesn't map to an immediate human action, it is not a page — it's noise, and noise trains people to ignore the one page that matters.
Don't skip layers and don't collapse them — each answers a different question.
| Layer | Answers | Built with | Fires when |
|---|---|---|---|
| 1. External uptime probe | "Is it reachable from the outside?" | Uptime Kuma / UptimeRobot / Better Stack | URL down, TLS broken, p95 latency over budget |
| 2. Health endpoints | "Is the process alive, and are its deps reachable?" | /livez + /readyz on the service | liveness fails → restart; readiness fails → pull from rotation |
| 3. SLO burn-rate alert | "Are we spending the error budget too fast?" | Prometheus/Grafana/Better Stack rule | multi-window burn rate exceeds threshold |
| 4. On-call escalation | "Who acts, and who's the backup?" | PagerDuty / incident.io / Better Stack | a page from layers 1–3 routes + escalates |
Layer 1 catches "the whole thing is gone." Layer 2 catches "a dependency died" before users do. Layer 3 catches "we're degrading faster than we can afford." Layer 4 makes sure a human shows up.
Decide by budget, team size, and self-host appetite. Pricing as of 2026-06.
| Tool | Free tier | Check interval | Best for | Paid from |
|---|---|---|---|---|
| Uptime Kuma 2.1.3 | Fully free, self-hosted | down to ~20s | full control, you have a VPS | $0 (you pay the box) |
| UptimeRobot | 50 monitors @ 5-min, 1 status page | 5-min free / 1-min paid | no infra, want managed | $7/mo |
| Better Stack | 10 monitors + incidents + logs | 30s | incident mgmt + on-call in one | $24/mo |
| Pingdom | none | sub-minute | enterprise synthetic + RUM | $15/mo |
Default: Uptime Kuma if you already run a VPS (1 GB box is comfortable; needs ~400 MB RAM), UptimeRobot free tier if you don't want infra. Kuma 2.1 added Globalping worldwide probe locations (so you test from regions, not just your one box) and built-in domain-expiry monitors. Concrete docker-compose and notification wiring live in references/tool-setup.md.
Split liveness from readiness — they trigger different machine actions.
/livez (liveness): is the process healthy? If it fails, the orchestrator restarts the container. Keep it dumb: just "am I running and not deadlocked." Never check the database here./readyz (readiness): are dependencies reachable so I can serve traffic? If it fails, the orchestrator pulls this instance from the load-balancer rotation but does not kill it.The cheap-probe rule: a probe runs constantly, so it must be <100ms and must not cascade-check every downstream. Why: if /livez pings the DB and the DB is briefly slow, liveness fails, the container restarts, the restart hammers the recovering DB — a restart loop that turns a 30-second blip into an outage.
# Bad — one /health that cascades and returns 500 on any hiccup.
# A slow Redis takes the whole service down and triggers restart loops.
@app.get("/health")
def health():
db.execute("SELECT 1") # blocks
redis.ping() # blocks
requests.get(PAYMENTS_URL) # blocks on a third party!
return {"status": "ok"} # 500 if ANY of these throws# Good — split, cheap, correct status codes, small JSON.
@app.get("/livez") # liveness: process only. Restart if this fails.
def livez():
return {"status": "alive"} # 200, ~1ms, touches nothing downstream
@app.get("/readyz") # readiness: deps with short timeouts. Pull from LB if this fails.
def readyz():
checks = {"db": ping(db, timeout=0.2), "cache": ping(redis, timeout=0.1)}
ok = all(checks.values())
return JSONResponse(
{"status": "ready" if ok else "degraded", "checks": checks},
status_code=200 if ok else 503, # 503 so the LB pulls this instance
)Do not put a third-party API call in readiness — a payment provider's outage shouldn't pull all your instances from rotation and take you fully down. Degrade that path in code (../error-handling/SKILL.md), don't fail-closed on it here. Go and FastAPI handler examples are in references/tool-setup.md.
The golden checklist. Monitor the symptom users feel, not just that one URL returns 200.
GET /. The homepage can be 200 while checkout is broken.The homepage being up tells you almost nothing. Probe the path that makes you money.
Alert on symptoms, not causes. Page on "checkout error rate >2% for 5 min" (user impact), not on "CPU >80%" (a cause that may be harmless and self-resolving). High CPU with happy users is a dashboard line, not a 3am page.
Use multi-window, multi-burn-rate for SLO alerts (Google SRE workbook). Burn rate = how fast you're spending the monthly error budget. Require a long and a short window to both fire — the long window says "this is real," the short window says "this is still happening," and together they kill false positives from a single spike.
| Burn rate | Window | Budget spent | Severity | Action |
|---|---|---|---|---|
| > 14.4 | 1h (+ 5m short) | 2% in 1h | critical | page |
| > 6 | 6h (+ 30m short) | 5% in 6h | warning | ticket |
| > 1 | 3d (+ 6h short) | 10% in 3d | info | review |
Routing: critical → page, warning → ticket, info → dashboard/review. Dedupe and group related alerts into one incident (10 hosts failing the same check = one page, not ten). Set maintenance windows so planned deploys don't page anyone. The full burn-rate math and a copy-paste Prometheus-style rule with a runbook_url annotation are in references/burn-rate-and-oncall.md.
references/burn-rate-and-oncall.md.An untested alert is not an alert. Before you call monitoring "done," prove the wire end-to-end:
scripts/verify.sh enforces the structural half of this on your config: liveness split from readiness, at least one symptom/burn-rate alert with two windows, a runbook reference on every alert, and a banlist for Opsgenie-as-new-setup and homepage-only monitors. Run it in CI so config drift can't silently re-introduce a noisy or untested setup.
| Bad | Why it bites | Do instead |
|---|---|---|
Monitor only the homepage / | / is 200 while checkout is broken — you learn from customers | synthetic check of the critical journey |
| Page on CPU/memory threshold | self-resolves, no user impact → alert fatigue | page on the symptom (error rate, latency, availability) |
| Alert on cause, not symptom | causes are noisy and ambiguous | alert on what the user feels |
| Alert with no runbook link | the paged human improvises at 3am | every alert links a 5-field runbook |
| Single-window threshold "by vibes" | one spike pages; tuned by guesswork | multi-window multi-burn-rate against an SLO |
| Probe cascades every downstream | one slow dep fails the probe → restart loop | cheap <100ms probe, deps in readiness with timeouts |
| Liveness checks the database | slow DB → restart loop turns a blip into an outage | liveness = process only; DB lives in readiness |
| Start new on-call on Opsgenie | EOL 2027-04-05, dead end | PagerDuty / incident.io / Better Stack / Grafana Cloud IRM |
| No secondary on-call | primary asleep/offline → page dropped | primary + secondary + manager escalation |
| No status page | inbound floods support during incidents | public status page customers can self-serve |
| Never tested the alert | "the rule exists" ≠ "the page lands" | trigger a synthetic failure, confirm it pages + escalates |
| Page on warnings | trains people to ignore pages | warning → ticket; only critical → page |
references/tool-setup.md — Uptime Kuma docker-compose (rootless image, volume, port), HTTP + push-heartbeat + SSL/domain-expiry monitors, a synthetic multi-step journey, notification wiring (ntfy / Slack webhook / PagerDuty integration key), a UptimeRobot/Better Stack monitor JSON shape, and Go + FastAPI health-endpoint handlers.references/burn-rate-and-oncall.md — full multi-window multi-burn-rate math, a copy-paste Prometheus-style alert rule with runbook_url, severity mapping, a sample escalation-policy YAML, and a fill-in runbook template.Related: ../observability/SKILL.md (what the app emits), ../deployment/SKILL.md (release-gating healthcheck + rollback), ../error-handling/SKILL.md (degrade in code), ../domains-dns/SKILL.md (provision certs/DNS), ../scaling/SKILL.md (survive the load monitoring detected).
© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in skills/monitoring of ericrisco/rsc-harness.
Open the folder on GitHubat commit 92fde8f
Monitoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Monitoring this skillericrisco/rsc-harness | 156 | — | ~3.1k | Automated safety check: Pass | MIT | |
| Monitoring EngineerFerroxLabs/wayland | 608 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | |
| Monitoring Observabilityahmedasmar/devops-claude-skills | 203 | — | ~3.9k | Automated safety check: Pass | None | |
| Oma Observabilityfirst-fluke/oh-my-agent | 1.3k | — | ~4.9k | Automated safety check: Pass | MIT | |
| Observability Sre Triageelastic/agent-skills | 592 | — | ~7.4k | Automated safety check: Pass | Apache-2.0 | |
| Observability MonitoringAnastasiyaW/codex-claude-code-config | 154 | — | ~4.1k | Automated safety check: Pass | MIT |
FerroxLabs/wayland
Observability and monitoring. An agent skill from FerroxLabs/wayland.
ahmedasmar/devops-claude-skills
Monitoring and observability strategy, implementation, and troubleshooting.
first-fluke/oh-my-agent
Intent-based observability + traceability router across layers, boundaries, and signals.
elastic/agent-skills
Triage a degraded or suspect service end to end: read SLO status and burn rate, check active alerting rules and ML anomalies, measure throughput, latency, and error rate, assess dependency health…
AnastasiyaW/codex-claude-code-config
Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…
majiayu000/spellbook
Observability and SRE expert. An agent skill from majiayu000/spellbook.
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
Works with
Categories
A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and…. Monitoring is an agent skill from ericrisco/rsc-harness. Use when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and readiness probes, alerting on SLO error-budget burn rather than raw counts, curing alert fatigue, on-call rotation with escalation, and a status page.
Monitoring fits situations like: setting up uptime and health monitoring; on-call basics for a service already in production; so you learn it is down before customers do — health and readiness probes; alerting on SLO error-budget burn rather than raw counts.
Run `npx skills add ericrisco/rsc-harness --skill monitoring -a claude-code`. Or copy the skill folder (skills/monitoring in ericrisco/rsc-harness) into .claude/skills/monitoring in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill monitoring -a codex`. Or copy the skill folder (skills/monitoring in ericrisco/rsc-harness) into .agents/skills/monitoring in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill monitoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/monitoring, .gemini/skills/monitoring, .github/skills/monitoring and .opencode/skills/monitoring in your project.
Going by SKILL.md and its folder, Monitoring needs a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell; Docker.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Monitoring is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Monitoring: Monitoring Engineer (FerroxLabs/wayland, 608 stars), Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Oma Observability (first-fluke/oh-my-agent, 1.3k stars) and Observability Sre Triage (elastic/agent-skills, 592 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.