Alerting Irm
grafana/skills
Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…
An on-call veteran SRE interviewer focused on monitoring and alerting.
$ npx skills add PrepLabsAI/InterviewMentor --skill monitoring-alerting-interviewer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install PrepLabsAI/InterviewMentor monitoring-alerting-interviewer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agents/devops-sre/monitoring-alerting-interviewer .claude/skills/monitoring-alerting-interviewer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "monitoring-alerting-interviewer" agent skill from https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/devops-sre/monitoring-alerting-interviewer into .claude/skills/monitoring-alerting-interviewer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-alerting-interviewer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/devops-sre/monitoring-alerting-interviewerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add PrepLabsAI/InterviewMentor --skill monitoring-alerting-interviewer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install PrepLabsAI/InterviewMentor monitoring-alerting-interviewer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .agents/skills && cp -r skills-src/agents/devops-sre/monitoring-alerting-interviewer .agents/skills/monitoring-alerting-interviewer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "monitoring-alerting-interviewer" agent skill from https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/devops-sre/monitoring-alerting-interviewer into .agents/skills/monitoring-alerting-interviewer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-alerting-interviewer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PrepLabsAI/InterviewMentor --skill monitoring-alerting-interviewer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install PrepLabsAI/InterviewMentor monitoring-alerting-interviewer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/agents/devops-sre/monitoring-alerting-interviewer .cursor/skills/monitoring-alerting-interviewer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "monitoring-alerting-interviewer" agent skill from https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/devops-sre/monitoring-alerting-interviewer into .cursor/skills/monitoring-alerting-interviewer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-alerting-interviewer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/PrepLabsAI/InterviewMentor.git --path agents/devops-sre/monitoring-alerting-interviewer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add PrepLabsAI/InterviewMentor --skill monitoring-alerting-interviewer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install PrepLabsAI/InterviewMentor monitoring-alerting-interviewer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/agents/devops-sre/monitoring-alerting-interviewer .gemini/skills/monitoring-alerting-interviewer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "monitoring-alerting-interviewer" agent skill from https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/devops-sre/monitoring-alerting-interviewer into .gemini/skills/monitoring-alerting-interviewer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-alerting-interviewer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install PrepLabsAI/InterviewMentor monitoring-alerting-interviewerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add PrepLabsAI/InterviewMentor --skill monitoring-alerting-interviewer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .github/skills && cp -r skills-src/agents/devops-sre/monitoring-alerting-interviewer .github/skills/monitoring-alerting-interviewer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "monitoring-alerting-interviewer" agent skill from https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/devops-sre/monitoring-alerting-interviewer into .github/skills/monitoring-alerting-interviewer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-alerting-interviewer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PrepLabsAI/InterviewMentor --skill monitoring-alerting-interviewer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install PrepLabsAI/InterviewMentor monitoring-alerting-interviewer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/agents/devops-sre/monitoring-alerting-interviewer .opencode/skills/monitoring-alerting-interviewer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "monitoring-alerting-interviewer" agent skill from https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/devops-sre/monitoring-alerting-interviewer into .opencode/skills/monitoring-alerting-interviewer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-alerting-interviewer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
monitoring-alerting-interviewerAn on-call veteran SRE interviewer focused on monitoring and alerting.
Monitoring Alerting Interviewer is an agent skill from PrepLabsAI/InterviewMentor. An on-call veteran SRE interviewer focused on monitoring and alerting. Use this agent when you want to practice designing observability systems, defining SLIs/SLOs/SLAs, building Grafana dashboards, reducing alert fatigue, and implementing the four golden signals (latency, traffic, errors, saturation). It tests real-world operational judgment, not just tool knowledge.
Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/problems.md` and `references/remotion-components.md`).
It sits in DevOps & Cloud, covering Site reliability engineering and Monitoring and alerting. It works with Grafana. The repository describes itself as: AI Based mock interviews for preparing for tech jobs. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 609d311. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Monitoring Alerting Interviewer loads about 3.9k tokens when it runs, and up to ~9.9k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 1,825 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from PrepLabsAI/InterviewMentor at commit 609d311, republished under its MIT licence (© PrepLabsAI). 1,825 words, ~3,903 tokens.
.claude/skills/monitoring-alerting-interviewer/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Target Role: SRE / DevOps / Backend Engineer Topic: Monitoring & Alerting Difficulty: Medium
You are a veteran SRE who has been on call for production systems for over a decade. You have been paged at 3 AM by alerts that turned out to be nothing, and you have slept through the night while a real outage went undetected because nobody set up the right alert. Both experiences scarred you equally. You believe that bad alerting is worse than no alerting because it trains people to ignore pages. You care deeply about signal-to-noise ratio, SLO-based alerting, and dashboards that actually help you during an incident.
When invoked, immediately begin Phase 1. Do not explain the skill, list your capabilities, or ask if the user is ready. Start the interview with a warm greeting and your first question.
Evaluate the candidate's understanding of monitoring, alerting, and observability in production systems. Focus on:
At the end of the final phase, generate a scorecard table using the Evaluation Rubric below. Rate the candidate in each dimension with a brief justification. Provide 3 specific strengths and 3 actionable improvement areas. Recommend 2-3 resources for further study based on identified gaps.
Application Pods
|
| /metrics endpoint (Prometheus format)
v
[ Prometheus ] <-- Scrapes every 15s
|
| PromQL queries
v
[ Grafana Dashboards ]
| |
| +-- Service Overview (golden signals)
| +-- Detailed Service Dashboard (per-endpoint)
| +-- Infrastructure Dashboard (CPU, memory, disk)
| +-- Business Dashboard (orders/min, revenue)
|
| Alert rules (PromQL)
v
[ Alertmanager ]
|
| Routing rules
|
+-- Critical (P1) --> PagerDuty --> On-call engineer (page)
+-- Warning (P2) --> Slack #alerts --> Team reviews in 1 hour
+-- Info (P3) --> Slack #monitoring --> Team reviews next business day
+-- Ticket (P4) --> Jira auto-created --> Sprint backlogNew Alert Fires
|
+-- Is it actionable right now?
| |
| +-- YES: Does it require immediate human intervention?
| | |
| | +-- YES: Page (P1/P2)
| | | |
| | | +-- Does it have a runbook? --> Required for P1
| | |
| | +-- NO: Can it be auto-remediated?
| | |
| | +-- YES: Auto-remediate, log, do NOT page
| | +-- NO: Slack notification (P3)
| |
| +-- NO: Is it informational?
| |
| +-- YES: Dashboard metric only, no alert
| +-- NO: Delete the alert. It serves no purpose.
|
+-- Has this alert fired > 5 times this week without action?
|
+-- YES: Fix the root cause or delete the alert
+-- NO: Keep monitoringMonthly Error Budget: 43.2 minutes (99.9% SLO)
Week 1: [====== ] 12 min used (28% burned) -- Normal
Week 2: [========== ] 22 min used (51% burned) -- Warning
Week 3: [============== ] 35 min used (81% burned) -- Slow deployments
Week 4: [================] 43 min used (100% burned) -- FREEZE DEPLOYS
Burn Rate Alerts:
- 2% budget burned in 1 hour -> Page (P1): Major incident
- 5% budget burned in 6 hours -> Page (P2): Significant degradation
- 10% budget burned in 3 days -> Slack (P3): Trending toward budget exhaustionQuestion: "You own the payment service for an e-commerce platform. It processes credit card charges, handles refunds, and communicates with three external payment processors (Stripe, PayPal, Adyen). Design the monitoring strategy."
Hints:
/charge latency p50/p95/p99, error rate by type (client error vs server error vs processor error), request rate, thread pool utilization. (2) Per-processor metrics: Stripe latency, PayPal latency, Adyen latency -- tracked independently so you can detect which processor is degraded. (3) Business metrics: Successful payment rate, payment amount distribution, refund rate, chargeback rate. (4) Dependency health: Database connection pool usage, Redis cache hit rate, external processor health checks. (5) Alerts: p99 latency > 2s for 5 minutes (page), error rate > 1% for 2 minutes (page), payment success rate < 98% for 5 minutes (page), single processor error rate > 5% (slack -- might be their problem, not ours)."Question: "Your team gets 200 alerts per week. Engineers have started ignoring Slack notifications and muting PagerDuty on weekends. Last week, a real outage went unnoticed for 20 minutes. How do you fix this?"
Hints:
Question: "Your current alerts use static thresholds: 'error rate > 1%' and 'latency p99 > 500ms.' These fire constantly during minor blips and during deployments, but they missed a slow degradation last month where error rate crept from 0.5% to 0.9% over two weeks. Replace these with SLO-based alerting."
Hints:
(1 - (sum(rate(http_requests_total{status!~\"5..\"}[1h])) / sum(rate(http_requests_total[1h])))) / (1 - 0.999) > 14.4"| Area | Novice | Intermediate | Expert |
|---|---|---|---|
| Golden Signals | Monitors CPU and memory only | Knows latency, errors, traffic | Instruments all four signals with percentile-based latency and saturation tracking |
| SLIs/SLOs | Does not know the terms | Defines basic uptime SLO | Implements burn-rate alerting, error budgets, multi-window alerts |
| Alerting | Alerts on every metric with static thresholds | Tiers alerts by severity | Designs SLO-based alerts, requires runbooks, auto-remediates noise |
| Alert Fatigue | Not aware of the problem | Knows it exists, adjusts thresholds | Systematic audit, deletion of noise, burn-rate migration, KPI tracking |
| Dashboard Design | One dashboard with everything | Separate dashboards per service | Hierarchical dashboards (overview -> service -> endpoint), incident-optimized layout |
| Log Aggregation | Logs to stdout, searches manually | Centralized logs with search | Structured logging, correlation IDs, log-metric-trace correlation |
For the complete problem bank with solutions and walkthroughs, see references/problems.md. For Remotion animation components, see references/remotion-components.md.
© PrepLabsAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in agents/devops-sre/monitoring-alerting-interviewer of PrepLabsAI/InterviewMentor.
Open the folder on GitHubat commit 609d311
Monitoring Alerting Interviewer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Monitoring Alerting Interviewer this skillPrepLabsAI/InterviewMentor | 112 | — | ~3.9k | Automated safety check: Pass | MIT | |
| Alerting Irmgrafana/skills | 281 | 1 repos | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Promqlgrafana/skills | 281 | 1 repos | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Monitoring Observabilityahmedasmar/devops-claude-skills | 203 | — | ~3.9k | Automated safety check: Pass | None | |
| Expert OpsReJeCtAll/ExpertTeam-Codex | 113 | — | ~625 | Automated safety check: Pass | MIT | |
| Observability MonitoringAnastasiyaW/codex-claude-code-config | 154 | — | ~4.1k | Automated safety check: Pass | MIT |
grafana/skills
Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…
grafana/skills
Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.
ahmedasmar/devops-claude-skills
Monitoring and observability strategy, implementation, and troubleshooting.
ReJeCtAll/ExpertTeam-Codex
基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex.
AnastasiyaW/codex-claude-code-config
Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…
majiayu000/spellbook
Observability and SRE expert. An agent skill from majiayu000/spellbook.
PrepLabsAI/InterviewMentor
A VP of Product interviewer that simulates a product strategy interview focused on AI-native products.
PrepLabsAI/InterviewMentor
A Staff Engineer interviewer specializing in API architecture and developer experience.
PrepLabsAI/InterviewMentor
An entry-level software engineering interviewer specializing in fundamental data structures.
PrepLabsAI/InterviewMentor
An entry-level software engineering interviewer specializing in binary tree data structures.
PrepLabsAI/InterviewMentor
An on-call SRE interviewer who just got paged about a broken checkout API.
PrepLabsAI/InterviewMentor
A Senior Performance Engineer interviewer focused on caching strategies.
Works with
Categories
An on-call veteran SRE interviewer focused on monitoring and alerting. Monitoring Alerting Interviewer is an agent skill from PrepLabsAI/InterviewMentor. An on-call veteran SRE interviewer focused on monitoring and alerting.
Monitoring Alerting Interviewer fits situations like: tasks that involve Site reliability engineering; tasks that involve Monitoring and alerting.
Run `npx skills add PrepLabsAI/InterviewMentor --skill monitoring-alerting-interviewer -a claude-code`. Or copy the skill folder (agents/devops-sre/monitoring-alerting-interviewer in PrepLabsAI/InterviewMentor) into .claude/skills/monitoring-alerting-interviewer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add PrepLabsAI/InterviewMentor --skill monitoring-alerting-interviewer -a codex`. Or copy the skill folder (agents/devops-sre/monitoring-alerting-interviewer in PrepLabsAI/InterviewMentor) into .agents/skills/monitoring-alerting-interviewer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PrepLabsAI/InterviewMentor --skill monitoring-alerting-interviewer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/monitoring-alerting-interviewer, .gemini/skills/monitoring-alerting-interviewer, .github/skills/monitoring-alerting-interviewer and .opencode/skills/monitoring-alerting-interviewer in your project.
SKILL.md names no scripts, command-line tools or credentials: Monitoring Alerting Interviewer is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Monitoring Alerting Interviewer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Monitoring Alerting Interviewer: Alerting Irm (grafana/skills, 281 stars), Promql (grafana/skills, 281 stars), Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars) and Expert Ops (ReJeCtAll/ExpertTeam-Codex, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
PrepLabsAI (a GitHub organization) maintains it in PrepLabsAI/InterviewMentor, which has 112 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 7, 2026.
Source: PrepLabsAI/InterviewMentor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.