Alerting Irm
grafana/skills
Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…
Design and run a monitoring system for a website or web app.
$ npx skills add rampstackco/claude-skills --skill monitoring-and-alerting -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install rampstackco/claude-skills monitoring-and-alerting --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/monitoring-and-alerting .claude/skills/monitoring-and-alerting && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "monitoring-and-alerting" agent skill from https://github.com/rampstackco/claude-skills/tree/main/skills/monitoring-and-alerting into .claude/skills/monitoring-and-alerting/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-and-alerting", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/rampstackco/claude-skills/tree/main/skills/monitoring-and-alertingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add rampstackco/claude-skills --skill monitoring-and-alerting -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install rampstackco/claude-skills monitoring-and-alerting --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/monitoring-and-alerting .agents/skills/monitoring-and-alerting && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "monitoring-and-alerting" agent skill from https://github.com/rampstackco/claude-skills/tree/main/skills/monitoring-and-alerting into .agents/skills/monitoring-and-alerting/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-and-alerting", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rampstackco/claude-skills --skill monitoring-and-alerting -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install rampstackco/claude-skills monitoring-and-alerting --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/monitoring-and-alerting .cursor/skills/monitoring-and-alerting && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "monitoring-and-alerting" agent skill from https://github.com/rampstackco/claude-skills/tree/main/skills/monitoring-and-alerting into .cursor/skills/monitoring-and-alerting/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-and-alerting", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/rampstackco/claude-skills.git --path skills/monitoring-and-alerting--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add rampstackco/claude-skills --skill monitoring-and-alerting -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install rampstackco/claude-skills monitoring-and-alerting --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/monitoring-and-alerting .gemini/skills/monitoring-and-alerting && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "monitoring-and-alerting" agent skill from https://github.com/rampstackco/claude-skills/tree/main/skills/monitoring-and-alerting into .gemini/skills/monitoring-and-alerting/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-and-alerting", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install rampstackco/claude-skills monitoring-and-alertingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add rampstackco/claude-skills --skill monitoring-and-alerting -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/monitoring-and-alerting .github/skills/monitoring-and-alerting && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "monitoring-and-alerting" agent skill from https://github.com/rampstackco/claude-skills/tree/main/skills/monitoring-and-alerting into .github/skills/monitoring-and-alerting/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-and-alerting", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rampstackco/claude-skills --skill monitoring-and-alerting -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install rampstackco/claude-skills monitoring-and-alerting --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/monitoring-and-alerting .opencode/skills/monitoring-and-alerting && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "monitoring-and-alerting" agent skill from https://github.com/rampstackco/claude-skills/tree/main/skills/monitoring-and-alerting into .opencode/skills/monitoring-and-alerting/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-and-alerting", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
monitoring-and-alertingDesign and run a monitoring system for a website or web app.
Monitoring And Alerting is an agent skill from rampstackco/claude-skills. Design and run a monitoring system for a website or web app. Use this skill when setting up uptime checks, defining SLOs, configuring error tracking, choosing what to alert on, designing on-call rotations, or fixing alert fatigue. Triggers on monitoring, alerts, uptime, SLO, SLA, error rate, on-call, pager, alert fatigue, observability, dashboards, what should we monitor. Also triggers when an incident reveals a gap in monitoring.
Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `README.md` and `references/slo-design-guide.md`).
It sits in DevOps & Cloud, covering Site reliability engineering, Monitoring and alerting and Incident response. The repository describes itself as: Stack-agnostic Claude Skills covering the full website lifecycle: brand, design, content, SEO, dev, ops, growth, and research. Build, ship, audit, optimize. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 482c9bf. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Monitoring And Alerting loads about 2.6k tokens when it runs, and up to ~4.9k if it reads all its reference files. Until then it costs about 115 tokens; SKILL.md has 1,406 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from rampstackco/claude-skills at commit 482c9bf, republished under its MIT licence (© rampstackco). 1,406 words, ~2,614 tokens.
.claude/skills/monitoring-and-alerting/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Decide what to watch, what to alert on, and how to make sure the right person finds out when things break.
incident-response)after-action-report)analytics-strategy)performance-optimization)Monitoring works in layers. Skip a layer and you'll miss a class of problems.
Is the site up? The simplest, most important layer.
Threshold: any sustained downtime (more than 2 consecutive failed checks) pages.
The site is up, but is it serving the right thing?
Threshold: failures of critical-path synthetics page. Non-critical page-level synthetics alert during business hours only.
The site is up and correct, but is it fast enough?
Threshold: regressions from baseline (e.g., p95 doubled in 5 minutes). Don't alert on absolute thresholds without baselines.
The site is up, correct, and fast for most, but errors are happening.
Threshold: rate-based, not count-based. "Error rate above 1% for 5 minutes" beats "more than 100 errors per minute."
A Service Level Objective is the target for reliability. Common form: "99.9% of homepage requests succeed in under 2 seconds, measured over 30 days."
The components:
The error budget is the inverse: 0.1% of requests can fail. If you've used the whole budget, slow down on risky changes.
Don't aim for 100%. Don't aim for "five nines" (99.999%) unless you really need it. Each nine costs an order of magnitude more.
| SLO | Allowed downtime per month |
|---|---|
| 99% | 7 hours, 18 minutes |
| 99.9% | 43 minutes |
| 99.95% | 21 minutes |
| 99.99% | 4 minutes, 22 seconds |
| 99.999% | 26 seconds |
For most marketing sites, 99.9% is plenty. For SaaS, 99.95% is reasonable. Anything higher needs significant infrastructure investment.
When the budget is healthy, ship aggressively. When the budget is half-spent, slow down. When the budget is exhausted, freeze risky changes until reliability recovers.
This is what makes SLOs useful: they create a feedback loop between reliability and velocity.
What tools are in place? What checks exist? What dashboards? What alerts?
Many teams have a tangle of half-configured tools. The first job is the inventory.
Draw the architecture. Front-end, back-end, database, third-party APIs, queues, workers. Each box is a candidate for monitoring.
For each box, ask:
Pick 3-5 SLOs. They should be:
For each box, configure checks at each layer. Some boxes won't have all four; that's fine.
| Box | Availability | Correctness | Performance | Errors |
|---|---|---|---|---|
| Homepage | HTTP check | Synthetic | LCP/INP | JS errors |
| Login API | HTTP check | Synthetic flow | p95 latency | 5xx rate |
Three tiers:
Anything in tier 1 must be:
If tier 1 alerts fire frequently, alert fatigue sets in. People stop responding.
Where do alerts go?
Each tier should have a documented escalation path. If the on-call doesn't ack within 5-15 minutes, escalate.
One dashboard per audience:
Dashboards are different from alerts. Alerts say "look now." Dashboards say "here's what's happening."
Every quarter, audit:
Tune the system. Monitoring drifts without active maintenance.
Alert on cause, not symptom. "CPU is high" is a cause. "Users are slow" is a symptom. Alert on symptoms; investigate causes.
Alert without a runbook. If the on-call doesn't know what to do, the alert is useless. Every paging alert needs a runbook (even a one-line one).
No baselines for "normal." Alerting on "more than 100 errors per minute" sounds reasonable but a busy day might exceed that without anything being wrong. Use rate-based and anomaly-based alerts.
Single-region monitoring. Your monitoring service in the same region as your site means you'll miss regional outages and you'll get woken up when monitoring itself has issues.
Monitoring the monitoring. Or rather, not. If your alerting platform is down, who tells you? Most paging services offer their own status feeds. Subscribe.
Too many tiers of severity. P0/P1/P2/P3/P4 with different SLAs becomes a sorting exercise. Three tiers (page, notify, log) is plenty.
Synthetics that don't match reality. A synthetic that hits the homepage every minute tests "is the homepage up." It doesn't test "is the actual user flow working." Build synthetics for the journeys that matter.
Static thresholds that never get tuned. Traffic grows, behavior changes, thresholds set last year are wrong. Review thresholds quarterly.
On-call rotation with no handoffs. Each new on-call has to figure out the system. Document. Run weekly handoff meetings or async updates.
Pager fatigue. If on-call is paged more than once or twice a week, something is wrong. Audit the alerts. Reduce, tune, or fix the underlying issues.
A monitoring plan includes:
This skill's output depends on data, measurements, or tool results it cannot generate on its own. When a required input, tool, or data source is unavailable or unverifiable, the sanctioned output is the deliverable with the gap stated: what was needed, what was actually obtained or verified, and which parts of the output are affected. Fabricating, estimating, or interpolating a required number to complete the deliverable is never sanctioned. A stated gap is a complete answer.
references/slo-design-guide.md: Detailed walkthrough of writing SLOs, error budget policies, and common SLO mistakes for web services.© rampstackco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in skills/monitoring-and-alerting of rampstackco/claude-skills.
Open the folder on GitHubat commit 482c9bf
Monitoring And Alerting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Monitoring And Alerting this skillrampstackco/claude-skills | 940 | — | ~2.6k | Automated safety check: Pass | MIT | |
| Alerting Irmgrafana/skills | 279 | 1 repos | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| SRE EngineerJeffallan/claude-skills | 12k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Monitoringericrisco/rsc-harness | 167 | — | ~3.1k | Automated safety check: Pass | MIT | |
| Sentry Alert TunerLeoYeAI/openclaw-master-skills | 2.2k | — | ~7.3k | Automated safety check: Pass | MIT | |
| Promqlgrafana/skills | 279 | 1 repos | ~1.1k | Automated safety check: Pass | Apache-2.0 |
grafana/skills
Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…
Jeffallan/claude-skills
Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.
ericrisco/rsc-harness
A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and…
LeoYeAI/openclaw-master-skills
Reduce Sentry alert fatigue by surgically tuning issue grouping, fingerprint rules, severity mapping, sample rates, before-send filters, sourcemap pipelines, and release-health gates.
grafana/skills
Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.
wshobson/agents
Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality.
rampstackco/claude-skills
Run a structured after-action review (postmortem, retrospective) on a launch, incident, or completed project to capture timeline, root cause analysis, contributing factors, and actionable lessons.
rampstackco/claude-skills
Design measurement frameworks including event taxonomy, KPI hierarchy, dashboard architecture, attribution models, and analytics implementation strategy.
rampstackco/claude-skills
Build or audit a comprehensive brand style guide that documents the full brand system including story, logo system, color, typography, imagery, voice, applications, and dos/don'ts.
rampstackco/claude-skills
Develop or document a complete brand voice and tone system covering voice attributes, tone shifts by context, vocabulary preferences, grammar rules, and copy examples.
rampstackco/claude-skills
Write or edit website copy, blog content, and editorial pieces with attention to voice, structure, and goal.
rampstackco/claude-skills
Develop a content strategy covering editorial positioning, content pillars, formats, calendar, governance, and topical authority planning.
Categories
Design and run a monitoring system for a website or web app. Monitoring And Alerting is an agent skill from rampstackco/claude-skills. Design and run a monitoring system for a website or web app.
Monitoring And Alerting fits situations like: setting up uptime checks; configuring error tracking; choosing what to alert on; designing on-call rotations.
Run `npx skills add rampstackco/claude-skills --skill monitoring-and-alerting -a claude-code`. Or copy the skill folder (skills/monitoring-and-alerting in rampstackco/claude-skills) into .claude/skills/monitoring-and-alerting in your project. Claude Code loads it when a task matches its description.
Run `npx skills add rampstackco/claude-skills --skill monitoring-and-alerting -a codex`. Or copy the skill folder (skills/monitoring-and-alerting in rampstackco/claude-skills) into .agents/skills/monitoring-and-alerting in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rampstackco/claude-skills --skill monitoring-and-alerting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/monitoring-and-alerting, .gemini/skills/monitoring-and-alerting, .github/skills/monitoring-and-alerting and .opencode/skills/monitoring-and-alerting in your project.
SKILL.md names no scripts, command-line tools or credentials: Monitoring And Alerting is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Monitoring And Alerting is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Monitoring And Alerting: Alerting Irm (grafana/skills, 279 stars), SRE Engineer (Jeffallan/claude-skills, 12k stars), Monitoring (ericrisco/rsc-harness, 167 stars) and Sentry Alert Tuner (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
rampstackco (a GitHub organization) maintains it in rampstackco/claude-skills, which has 940 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 7, 2026.
Source: rampstackco/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.