Agent skill

Monitoring And Alerting

by rampstackco in rampstackco/claude-skills

Design and run a monitoring system for a website or web app.

MITAuto-check passedDevOps & Cloud

Install Monitoring And Alerting

skills CLI
$ npx skills add rampstackco/claude-skills --skill monitoring-and-alerting -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rampstackco/claude-skills monitoring-and-alerting --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/monitoring-and-alerting .claude/skills/monitoring-and-alerting && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
monitoring-and-alerting
GitHub stars
940
Token cost
~2.6k tokens
SKILL.md length
1,406 words
Files
3 (incl. references)
Skills in repo
103
Repo updated
First seen
Licence
MIT

At a glance

Design and run a monitoring system for a website or web app.

  • Works in 8 steps: Inventory what's already monitored → Map the system → Define the SLOs → …
  • Setting up uptime checks
  • SKILL.md covers When to use, When NOT to use, Required inputs and The framework: 4 layers, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Monitoring And Alerting is an agent skill from rampstackco/claude-skills. Design and run a monitoring system for a website or web app. Use this skill when setting up uptime checks, defining SLOs, configuring error tracking, choosing what to alert on, designing on-call rotations, or fixing alert fatigue. Triggers on monitoring, alerts, uptime, SLO, SLA, error rate, on-call, pager, alert fatigue, observability, dashboards, what should we monitor. Also triggers when an incident reveals a gap in monitoring.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `README.md` and `references/slo-design-guide.md`).

It sits in DevOps & Cloud, covering Site reliability engineering, Monitoring and alerting and Incident response. The repository describes itself as: Stack-agnostic Claude Skills covering the full website lifecycle: brand, design, content, SEO, dev, ops, growth, and research. Build, ship, audit, optimize. The licence is MIT.

When your agent uses it

  • Setting up uptime checks
  • Configuring error tracking
  • Choosing what to alert on
  • Designing on-call rotations

Example prompts

  • “/monitoring-and-alerting”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Inventory what's already monitored
  2. Map the system
  3. Define the SLOs
  4. Set up checks across the 4 layers
  5. Decide what pages and what doesn't
  6. Configure routing
  7. Build dashboards
  8. Run an alert audit

What it can do on your machine

Read from SKILL.md and the folder at commit 482c9bf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Monitoring And Alerting loads about 2.6k tokens when it runs, and up to ~4.9k if it reads all its reference files. Until then it costs about 115 tokens; SKILL.md has 1,406 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rampstackco/claude-skills at commit 482c9bf, republished under its MIT licence (© rampstackco). 1,406 words, ~2,614 tokens.

Download SKILL.mdSave it as .claude/skills/monitoring-and-alerting/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
monitoring-and-alerting
description
Design and run a monitoring system for a website or web app. Use this skill when setting up uptime checks, defining SLOs, configuring error tracking, choosing what to alert on, designing on-call rotations, or fixing alert fatigue. Triggers on monitoring, alerts, uptime, SLO, SLA, error rate, on-call, pager, alert fatigue, observability, dashboards, what should we monitor. Also triggers when an incident reveals a gap in monitoring.
category
operations
catalog_summary
SLO design, uptime checks, alert routing, on-call rotations
display_order
5

Monitoring and Alerting

Decide what to watch, what to alert on, and how to make sure the right person finds out when things break.


When to use

  • Setting up monitoring on a new site or service
  • Defining SLOs (service level objectives) and error budgets
  • Choosing which alerts page someone vs which go to a quiet channel
  • Designing or fixing on-call rotation
  • Diagnosing alert fatigue
  • Filling monitoring gaps revealed by an incident
  • Migrating monitoring vendors

When NOT to use

  • Responding to an active incident (use incident-response)
  • Writing the post-mortem (use after-action-report)
  • Designing analytics dashboards for product metrics (use analytics-strategy)
  • Performance optimization itself (use performance-optimization)

Required inputs

  • The system you're monitoring (URLs, services, dependencies)
  • Existing monitoring tools (uptime, errors, logs, APM)
  • Business hours and team timezone(s)
  • Who is on-call or available for incidents
  • Existing SLOs or success metrics, if any

The framework: 4 layers

Monitoring works in layers. Skip a layer and you'll miss a class of problems.

Layer 1: Availability

Is the site up? The simplest, most important layer.

  • HTTP checks from multiple regions (every 1-5 minutes)
  • DNS resolution checks
  • Certificate expiration checks
  • Status code checks (alert on 5xx, not just timeout)

Threshold: any sustained downtime (more than 2 consecutive failed checks) pages.

Layer 2: Correctness

The site is up, but is it serving the right thing?

  • Synthetic checks (a script that loads the homepage, clicks a button, validates expected text)
  • Critical user journeys (signup, checkout, search)
  • Content presence checks (homepage hasn't gone blank)
  • API contract checks (response shape and key fields are present)

Threshold: failures of critical-path synthetics page. Non-critical page-level synthetics alert during business hours only.

Layer 3: Performance

The site is up and correct, but is it fast enough?

  • Core Web Vitals (LCP, INP, CLS) from real users (RUM)
  • Synthetic performance (Lighthouse, WebPageTest, custom)
  • API response times (p50, p95, p99)
  • Database query times for slow queries
  • Dependency response times (third-party APIs)

Threshold: regressions from baseline (e.g., p95 doubled in 5 minutes). Don't alert on absolute thresholds without baselines.

Layer 4: Errors and anomalies

The site is up, correct, and fast for most, but errors are happening.

  • Error rate (% of requests returning 5xx)
  • Client-side error rate (uncaught JS exceptions)
  • Log error volume (unexpected spikes)
  • Anomaly detection (traffic falling off a cliff)
  • Background job failures
  • Queue depth

Threshold: rate-based, not count-based. "Error rate above 1% for 5 minutes" beats "more than 100 errors per minute."


SLOs and error budgets

A Service Level Objective is the target for reliability. Common form: "99.9% of homepage requests succeed in under 2 seconds, measured over 30 days."

The components:

  • The thing you're measuring (homepage requests)
  • The success criterion (returns 2xx in under 2 seconds)
  • The target (99.9% of them)
  • The window (over 30 days)

The error budget is the inverse: 0.1% of requests can fail. If you've used the whole budget, slow down on risky changes.

Picking SLOs

Don't aim for 100%. Don't aim for "five nines" (99.999%) unless you really need it. Each nine costs an order of magnitude more.

SLOAllowed downtime per month
99%7 hours, 18 minutes
99.9%43 minutes
99.95%21 minutes
99.99%4 minutes, 22 seconds
99.999%26 seconds

For most marketing sites, 99.9% is plenty. For SaaS, 99.95% is reasonable. Anything higher needs significant infrastructure investment.

Using error budgets

When the budget is healthy, ship aggressively. When the budget is half-spent, slow down. When the budget is exhausted, freeze risky changes until reliability recovers.

This is what makes SLOs useful: they create a feedback loop between reliability and velocity.


Workflow

Step 1: Inventory what's already monitored

What tools are in place? What checks exist? What dashboards? What alerts?

Many teams have a tangle of half-configured tools. The first job is the inventory.

Step 2: Map the system

Draw the architecture. Front-end, back-end, database, third-party APIs, queues, workers. Each box is a candidate for monitoring.

For each box, ask:

  • What does "up" mean?
  • What does "correct" mean?
  • What does "fast" mean?
  • What's the most common failure mode?
Step 3: Define the SLOs

Pick 3-5 SLOs. They should be:

  • Tied to user-visible behavior (not internal metrics)
  • Achievable with current infrastructure
  • Measured automatically
  • Reviewed at least quarterly
Step 4: Set up checks across the 4 layers

For each box, configure checks at each layer. Some boxes won't have all four; that's fine.

BoxAvailabilityCorrectnessPerformanceErrors
HomepageHTTP checkSyntheticLCP/INPJS errors
Login APIHTTP checkSynthetic flowp95 latency5xx rate
Step 5: Decide what pages and what doesn't

Three tiers:

  1. Page (wakes someone up): site down, critical flow broken, error rate spike, security incident.
  2. Notify (during business hours): non-critical synthetic failure, performance regression, slow query, dependency degradation.
  3. Log (no notification): anomalies for later review, low-priority warnings, info-level events.

Anything in tier 1 must be:

  • Actionable (the on-call can do something about it)
  • Important (it represents real impact)
  • Rare (less than 1-2 per week is the goal)

If tier 1 alerts fire frequently, alert fatigue sets in. People stop responding.

Show full SKILL.md (580 more words)Show less
Step 6: Configure routing

Where do alerts go?

  • Tier 1: paging system (e.g., PagerDuty). Do not onboard onto Opsgenie: Atlassian ended sales in June 2025 and support ends April 2027. Direct to on-call.
  • Tier 2: chat channel (Slack, Teams). Tagged with the area.
  • Tier 3: dashboard or log only.

Each tier should have a documented escalation path. If the on-call doesn't ack within 5-15 minutes, escalate.

Step 7: Build dashboards

One dashboard per audience:

  • Real-time ops dashboard: current health, recent alerts, error rates, throughput
  • SLO dashboard: SLO status and error budget consumption
  • Per-service dashboards: detail for individual services or pages
  • Executive dashboard: uptime over weeks/months, key business metrics

Dashboards are different from alerts. Alerts say "look now." Dashboards say "here's what's happening."

Step 8: Run an alert audit

Every quarter, audit:

  • Which alerts fired? Were they actionable?
  • Which alerts didn't fire when they should have?
  • Are any alerts noisy (more than once a week, low actionability)?
  • Are runbooks up to date?
  • Have SLOs been met? Any consistently breached?

Tune the system. Monitoring drifts without active maintenance.


Failure patterns

Alert on cause, not symptom. "CPU is high" is a cause. "Users are slow" is a symptom. Alert on symptoms; investigate causes.

Alert without a runbook. If the on-call doesn't know what to do, the alert is useless. Every paging alert needs a runbook (even a one-line one).

No baselines for "normal." Alerting on "more than 100 errors per minute" sounds reasonable but a busy day might exceed that without anything being wrong. Use rate-based and anomaly-based alerts.

Single-region monitoring. Your monitoring service in the same region as your site means you'll miss regional outages and you'll get woken up when monitoring itself has issues.

Monitoring the monitoring. Or rather, not. If your alerting platform is down, who tells you? Most paging services offer their own status feeds. Subscribe.

Too many tiers of severity. P0/P1/P2/P3/P4 with different SLAs becomes a sorting exercise. Three tiers (page, notify, log) is plenty.

Synthetics that don't match reality. A synthetic that hits the homepage every minute tests "is the homepage up." It doesn't test "is the actual user flow working." Build synthetics for the journeys that matter.

Static thresholds that never get tuned. Traffic grows, behavior changes, thresholds set last year are wrong. Review thresholds quarterly.

On-call rotation with no handoffs. Each new on-call has to figure out the system. Document. Run weekly handoff meetings or async updates.

Pager fatigue. If on-call is paged more than once or twice a week, something is wrong. Audit the alerts. Reduce, tune, or fix the underlying issues.


Output format

A monitoring plan includes:

  • System map: what's being monitored
  • SLOs: the 3-5 reliability targets
  • Checks per layer: availability, correctness, performance, errors
  • Alert tiering: what pages, what notifies, what logs
  • Routing: where alerts go, escalation paths
  • Dashboards: what audiences see
  • Runbooks: linked from each paging alert
  • Audit cadence: when this gets reviewed

If required data is unavailable

This skill's output depends on data, measurements, or tool results it cannot generate on its own. When a required input, tool, or data source is unavailable or unverifiable, the sanctioned output is the deliverable with the gap stated: what was needed, what was actually obtained or verified, and which parts of the output are affected. Fabricating, estimating, or interpolating a required number to complete the deliverable is never sanctioned. A stated gap is a complete answer.


Reference files

© rampstackco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/monitoring-and-alerting of rampstackco/claude-skills.

  • SKILL.md
  • README.md
  • references/slo-design-guide.md

Open the folder on GitHubat commit 482c9bf

Compare with similar skills

Monitoring And Alerting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Monitoring And Alerting compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Monitoring And Alerting this skillrampstackco/claude-skills940—~2.6kAutomated safety check: PassMIT
Alerting Irmgrafana/skills2791 repos~1.9kAutomated safety check: PassApache-2.0
SRE EngineerJeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT
Monitoringericrisco/rsc-harness167—~3.1kAutomated safety check: PassMIT
Sentry Alert TunerLeoYeAI/openclaw-master-skills2.2k—~7.3kAutomated safety check: PassMIT
Promqlgrafana/skills2791 repos~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    279 GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • SRE Engineer

    Jeffallan/claude-skills

    Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

    12k GitHub stars~1.7k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • Monitoring

    ericrisco/rsc-harness

    A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and…

    167 GitHub stars~3.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Sentry Alert Tuner

    LeoYeAI/openclaw-master-skills

    Reduce Sentry alert fatigue by surgically tuning issue grouping, fingerprint rules, severity mapping, sample rates, before-send filters, sourcemap pipelines, and release-health gates.

    2.2k GitHub stars~7.3k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Promql

    grafana/skills

    Official

    Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.

    279 GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check passed
  • Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality.

    40k GitHub starsUsed in 9 repos~607 tokens
    DevOps & CloudAuto-check passed

More from rampstackco/claude-skills

All 103 skills in this repo
  • After Action Report

    rampstackco/claude-skills

    Run a structured after-action review (postmortem, retrospective) on a launch, incident, or completed project to capture timeline, root cause analysis, contributing factors, and actionable lessons.

    940 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Analytics Strategy

    rampstackco/claude-skills

    Design measurement frameworks including event taxonomy, KPI hierarchy, dashboard architecture, attribution models, and analytics implementation strategy.

    940 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Brand Style Guide

    rampstackco/claude-skills

    Build or audit a comprehensive brand style guide that documents the full brand system including story, logo system, color, typography, imagery, voice, applications, and dos/don'ts.

    940 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Brand Voice

    rampstackco/claude-skills

    Develop or document a complete brand voice and tone system covering voice attributes, tone shifts by context, vocabulary preferences, grammar rules, and copy examples.

    940 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Content And Copy

    rampstackco/claude-skills

    Write or edit website copy, blog content, and editorial pieces with attention to voice, structure, and goal.

    940 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Content Strategy

    rampstackco/claude-skills

    Develop a content strategy covering editorial positioning, content pillars, formats, calendar, governance, and topical authority planning.

    940 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Monitoring And Alerting

What does Monitoring And Alerting do?

Design and run a monitoring system for a website or web app. Monitoring And Alerting is an agent skill from rampstackco/claude-skills. Design and run a monitoring system for a website or web app.

When should I use Monitoring And Alerting?

Monitoring And Alerting fits situations like: setting up uptime checks; configuring error tracking; choosing what to alert on; designing on-call rotations.

How do I install Monitoring And Alerting in Claude Code?

Run `npx skills add rampstackco/claude-skills --skill monitoring-and-alerting -a claude-code`. Or copy the skill folder (skills/monitoring-and-alerting in rampstackco/claude-skills) into .claude/skills/monitoring-and-alerting in your project. Claude Code loads it when a task matches its description.

How do I install Monitoring And Alerting in Codex?

Run `npx skills add rampstackco/claude-skills --skill monitoring-and-alerting -a codex`. Or copy the skill folder (skills/monitoring-and-alerting in rampstackco/claude-skills) into .agents/skills/monitoring-and-alerting in your project. Codex loads it when a task matches its description.

Can I use Monitoring And Alerting in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rampstackco/claude-skills --skill monitoring-and-alerting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/monitoring-and-alerting, .gemini/skills/monitoring-and-alerting, .github/skills/monitoring-and-alerting and .opencode/skills/monitoring-and-alerting in your project.

What does Monitoring And Alerting need to run?

SKILL.md names no scripts, command-line tools or credentials: Monitoring And Alerting is instructions for the agent only.

Does Monitoring And Alerting access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Monitoring And Alerting safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Monitoring And Alerting use?

Monitoring And Alerting is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Monitoring And Alerting use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Monitoring And Alerting?

Skills that share tags, products or a category with Monitoring And Alerting: Alerting Irm (grafana/skills, 279 stars), SRE Engineer (Jeffallan/claude-skills, 12k stars), Monitoring (ericrisco/rsc-harness, 167 stars) and Sentry Alert Tuner (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Monitoring And Alerting?

rampstackco (a GitHub organization) maintains it in rampstackco/claude-skills, which has 940 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 7, 2026.

Source: rampstackco/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.