Agent skill

Alerting Oncall

by sickn33 in sickn33/agentic-awesome-skills

Set up alerting rules, configure on-call rotations, and manage incident response workflows.

MITAuto-check passedDevOps & Cloud

Install Alerting Oncall

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill alerting-oncall -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills alerting-oncall --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/alerting-oncall .claude/skills/alerting-oncall && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
alerting-oncall
GitHub stars
47k
Used in
1 other repo
Token cost
~2.8k tokens
SKILL.md length
91 words
Files
2 (incl. references)
Skills in repo
1,497
Repo updated
First seen
Licence
MIT

At a glance

Set up alerting rules, configure on-call rotations, and manage incident response workflows.

  • Tasks that involve Incident response
  • SKILL.md covers Prerequisites, Alerting Best Practices, Prometheus Alerting and PagerDuty Integration, plus 3 more sections
  • Reaches hooks.slack.com and events.pagerduty.com; needs SECRET_KEY
  • Tasks that involve Monitoring and alerting

What it does

Alerting Oncall is an agent skill from sickn33/agentic-awesome-skills. Set up alerting rules, configure on-call rotations, and manage incident response workflows.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/details.md`). Compatibility notes: Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and…

It sits in DevOps & Cloud, covering Incident response and Monitoring and alerting. It works with Prometheus. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve Incident response
  • Tasks that involve Monitoring and alerting

Example prompts

  • “/alerting-oncall”

Requirements

  • Docker
  • A credential in SECRET_KEY
  • Compatibility (from SKILL.md): Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled.

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml, go and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • hooks.slack.com
    • events.pagerduty.com

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • SECRET_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled.

    From compatibility in the SKILL.md frontmatter.

Context cost

Alerting Oncall loads about 2.8k tokens when it runs, and up to ~3.3k if it reads all its reference files. Until then it costs about 27 tokens; SKILL.md has 91 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its MIT licence (© sickn33). 91 words, ~2,845 tokens.

Download SKILL.mdSave it as .claude/skills/alerting-oncall/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
alerting-oncall
description
Set up alerting rules, configure on-call rotations, and manage incident response workflows.
compatibility
Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled.
category
devops
risk
critical
source
https://github.com/BagelHole/DevOps-Security-Agent-Skills
source_repo
BagelHole/DevOps-Security-Agent-Skills
source_type
community
date_added
2026-09-20
license
MIT
license_source
https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
metadata.author
devops-skills
metadata.version
1.0

Alerting & On-Call

Configure effective alerting and on-call management for production systems.

Prerequisites

  • Monitoring system (Prometheus, Datadog, etc.)
  • On-call platform (PagerDuty, Opsgenie, Grafana OnCall)
  • Communication channels (Slack, email)

Alerting Best Practices

Alert Categories
yaml
# Severity levels
critical:
  - Service completely down
  - Data loss imminent
  - Security breach
  response: Immediate page, wake people up

high:
  - Service degraded significantly
  - Error rate above SLO
  - Capacity near limit
  response: Page during business hours, notify after hours

medium:
  - Performance degradation
  - Non-critical component failure
  - Warning thresholds exceeded
  response: Notify via Slack, review next business day

low:
  - Informational alerts
  - Capacity planning triggers
  - Routine maintenance needed
  response: Email notification, weekly review
Alert Design Principles
yaml
# Good alert characteristics
alerts:
  actionable:
    - Every alert should require human action
    - Include runbook links
    - Clear remediation steps

  relevant:
    - Alert on symptoms, not causes
    - Focus on user impact
    - Avoid alerting on expected behavior

  timely:
    - Appropriate thresholds
    - Suitable evaluation windows
    - Account for normal variance

  unique:
    - No duplicate alerts
    - Proper alert grouping
    - Clear ownership

Prometheus Alerting

Alert Rules
yaml
# prometheus/rules/alerts.yml
groups:
  - name: service_alerts
    rules:
      # High-level service health
      - alert: ServiceDown
        expr: up{job="myapp"} == 0
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: "Service {{ $labels.instance }} is down"
          description: "{{ $labels.job }} on {{ $labels.instance }} has been down for more than 1 minute."
          runbook_url: "https://wiki.example.com/runbooks/service-down"

      # Error rate alert
      - alert: HighErrorRate
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
          / sum(rate(http_requests_total[5m])) by (service) > 0.05
        for: 5m
        labels:
          severity: critical
        annotations:
          summary: "High error rate for {{ $labels.service }}"
          description: "Error rate is {{ $value | humanizePercentage }} for the last 5 minutes"

      # Latency alert (SLO-based)
      - alert: HighLatency
        expr: |
          histogram_quantile(0.95, 
            sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service)
          ) > 0.5
        for: 5m
        labels:
          severity: high
        annotations:
          summary: "P95 latency above 500ms for {{ $labels.service }}"
Alertmanager Configuration
yaml
# alertmanager.yml
global:
  resolve_timeout: 5m
  slack_api_url: 'https://hooks.slack.com/services/xxx'
  pagerduty_url: 'https://events.pagerduty.com/v2/enqueue'

templates:
  - '/etc/alertmanager/templates/*.tmpl'

route:
  receiver: 'default-receiver'
  group_by: ['alertname', 'service']
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h
  
  routes:
    # Critical alerts go to PagerDuty
    - match:
        severity: critical
      receiver: 'pagerduty-critical'
      group_wait: 0s
      repeat_interval: 1h

    # High severity during business hours
    - match:
        severity: high
      receiver: 'slack-high'
      active_time_intervals:
        - business-hours

    # Route by team
    - match_re:
        team: platform.*
      receiver: 'platform-team'

receivers:
  - name: 'default-receiver'
    slack_configs:
      - channel: '#alerts'
        send_resolved: true

  - name: 'pagerduty-critical'
    pagerduty_configs:
      - service_key: 'xxx'
        severity: critical
        description: '{{ .CommonAnnotations.summary }}'
        details:
          firing: '{{ template "pagerduty.firing" . }}'

  - name: 'slack-high'
    slack_configs:
      - channel: '#alerts-high'
        title: '{{ .CommonAnnotations.summary }}'
        text: '{{ .CommonAnnotations.description }}'
        actions:
          - type: button
            text: 'Runbook'
            url: '{{ .CommonAnnotations.runbook_url }}'
          - type: button
            text: 'Dashboard'
            url: '{{ .CommonAnnotations.dashboard_url }}'

  - name: 'platform-team'
    slack_configs:
      - channel: '#platform-alerts'

time_intervals:
  - name: business-hours
    time_intervals:
      - weekdays: ['monday:friday']
        times:
          - start_time: '09:00'
            end_time: '17:00'

inhibit_rules:
  - source_match:
      severity: critical
    target_match:
      severity: high
    equal: ['service']

PagerDuty Integration

Service Configuration
yaml
# Terraform example
resource "pagerduty_service" "myapp" {
  name                    = "MyApp Production"
  description             = "Production application service"
  escalation_policy       = pagerduty_escalation_policy.default.id
  alert_creation          = "create_alerts_and_incidents"
  auto_resolve_timeout    = 14400  # 4 hours
  acknowledgement_timeout = 600    # 10 minutes

  incident_urgency_rule {
    type    = "use_support_hours"
    
    during_support_hours {
      type    = "constant"
      urgency = "high"
    }
    
    outside_support_hours {
      type    = "constant"
      urgency = "low"
    }
  }
}

resource "pagerduty_escalation_policy" "default" {
  name      = "Default Escalation"
  num_loops = 2

  rule {
    escalation_delay_in_minutes = 10
    target {
      type = "schedule_reference"
      id   = pagerduty_schedule.primary.id
    }
  }

  rule {
    escalation_delay_in_minutes = 15
    target {
      type = "user_reference"
      id   = pagerduty_user.manager.id
    }
  }
}
Schedule Configuration
yaml
resource "pagerduty_schedule" "primary" {
  name      = "Primary On-Call"
  time_zone = "America/New_York"

  layer {
    name                         = "Weekly Rotation"
    start                        = "2024-01-01T00:00:00-05:00"
    rotation_virtual_start       = "2024-01-01T00:00:00-05:00"
    rotation_turn_length_seconds = 604800  # 1 week
    users                        = [for user in pagerduty_user.oncall : user.id]
  }

  # Override layer for holidays
  layer {
    name                         = "Holiday Coverage"
    start                        = "2024-01-01T00:00:00-05:00"
    rotation_virtual_start       = "2024-01-01T00:00:00-05:00"
    rotation_turn_length_seconds = 86400
    users                        = [pagerduty_user.holiday_coverage.id]

    restriction {
      type              = "daily_restriction"
      start_time_of_day = "00:00:00"
      duration_seconds  = 86400
      start_day_of_week = 0  # Sunday
    }
  }
}

Grafana OnCall

Integration Setup
yaml
# docker-compose.yml addition
services:
  oncall:
    image: grafana/oncall
    environment:
      - SECRET_KEY=your-secret-key
      - BASE_URL=http://oncall:8080
      - GRAFANA_API_URL=http://grafana:3000
    ports:
      - "8080:8080"
Escalation Chain
yaml
# Example escalation chain structure
escalation_chains:
  - name: "Production Critical"
    steps:
      - step: 1
        type: notify
        persons:
          - "@oncall-primary"
        wait_delay: 0
        
      - step: 2
        type: notify
        persons:
          - "@oncall-secondary"
        wait_delay: 5m
        
      - step: 3
        type: notify
        persons:
          - "@engineering-manager"
        wait_delay: 10m
        
      - step: 4
        type: trigger_action
        action: "escalate_to_incident_commander"
        wait_delay: 15m

Alert Templates

Slack Alert Template
go
{{ define "slack.title" }}
[{{ .Status | toUpper }}{{ if eq .Status "firing" }}:{{ .Alerts.Firing | len }}{{ end }}] {{ .CommonLabels.alertname }}
{{ end }}

{{ define "slack.text" }}
{{ range .Alerts }}
*Alert:* {{ .Annotations.summary }}
*Severity:* {{ .Labels.severity }}
*Description:* {{ .Annotations.description }}
*Runbook:* {{ .Annotations.runbook_url }}
{{ end }}
{{ end }}
PagerDuty Details Template
go
{{ define "pagerduty.firing" }}
{{ range .Alerts.Firing }}
Alert: {{ .Labels.alertname }}
Service: {{ .Labels.service }}
Instance: {{ .Labels.instance }}
Value: {{ .Annotations.value }}
Started: {{ .StartsAt.Format "2006-01-02 15:04:05" }}
{{ end }}
{{ end }}

On-Call Best Practices

Rotation Guidelines
yaml
on_call_guidelines:
  rotation_length: 1 week
  handoff_time: "10:00 AM Monday"
  
  responsibilities:
    - Monitor alerts during shift
    - Respond within SLA (critical: 5min, high: 15min)
    - Document incidents
    - Handoff unresolved issues
    
  support:
    - Secondary on-call for backup
    - Clear escalation path
    - Manager availability for major incidents
    
  wellness:
    - Maximum 1 week on-call per month
    - Comp time after high-alert periods
    - No-interrupt recovery day after shift
Runbook Template
markdown
# Alert: High Error Rate

## Summary
Error rate has exceeded the threshold of 5% for the service.

## Impact
Users may experience errors when accessing the application.

## Investigation Steps
1. Check service logs: `kubectl logs -l app=myapp -n production`
2. Review recent deployments: `kubectl rollout history deployment/myapp`
3. Check database connectivity: `kubectl exec -it myapp -- nc -zv postgres 5432`
4. Review error traces in APM dashboard


## Contents

- [Remediation](references/details.md)
- [Escalation](references/details.md)
- [Alert Fatigue Reduction](references/details.md)
- [Common Issues](references/details.md)
- [Best Practices](references/details.md)
- [Related Skills](references/details.md)

## When to Use This Skill

Use this skill when:
- Setting up alerting rules and thresholds
- Configuring on-call rotations and schedules
- Implementing alert routing and escalation
- Reducing alert fatigue
- Managing incident response workflows

## Limitations

- Guidance executes against real environments: confirm target, blast radius, and rollback plan before applying anything.
- Never deploy to production without explicit approval. Docs-only import: upstream scripts and templates not bundled.

### Example

```bash
git status && git diff --stat
kubectl diff -f manifest.yaml

Adapted from BagelHole/DevOps-Security-Agent-Skills (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/alerting-oncall of sickn33/agentic-awesome-skills.

  • SKILL.md
  • references/details.md

Open the folder on GitHubat commit b84d35a

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Alerting Oncall next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Alerting Oncall compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Alerting Oncall this skillsickn33/agentic-awesome-skills47k1 repos~2.8kAutomated safety check: PassMIT
Alerting Irmgrafana/skills2821 repos~1.9kAutomated safety check: PassApache-2.0
SRE EngineerJeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT
Oncall Irmgrafana/skills282—~1.4kAutomated safety check: PassApache-2.0
Alerting OncallBagelHole/DevOps-Security-Agent-Skills1.2k—~3kAutomated safety check: PassMIT
Grafana Observabilityautomateyournetwork/netclaw676—~2.7kAutomated safety check: NotesApache-2.0

Similar skills

  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    282 GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • SRE Engineer

    Jeffallan/claude-skills

    Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

    12k GitHub stars~1.7k tokensUpdated 7 days ago
    DevOps & CloudAuto-check passed
  • Oncall Irm

    grafana/skills

    Official

    Route alerts, run on-call rotations, and drive incidents in Grafana IRM / OnCall — integrations (Alertmanager / Grafana Alerting / generic webhook / PagerDuty), Jinja2 routing + grouping templates…

    282 GitHub stars~1.4k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Alerting Oncall

    BagelHole/DevOps-Security-Agent-Skills

    Set up alerting rules, configure on-call rotations, and manage incident response workflows.

    1.2k GitHub stars~3k tokensUpdated 4 mo ago
    DevOps & CloudAuto-check passed
  • Grafana Observability

    automateyournetwork/netclaw

    Grafana observability platform — dashboards, Prometheus PromQL, Loki LogQL, alerting, incidents, OnCall schedules, annotations, datasource queries, panel rendering (75+ tools).

    676 GitHub stars~2.7k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • KubeEye Cluster Inspection

    kubesphere/kubesphere

    Deploys KubeEye on KubeSphere and writes InspectRule and InspectPlan resources to inspect cluster health, then retrieves the inspection reports.

    17k GitHub stars~3.6k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Works with

Categories

Questions about Alerting Oncall

What does Alerting Oncall do?

Set up alerting rules, configure on-call rotations, and manage incident response workflows. Alerting Oncall is an agent skill from sickn33/agentic-awesome-skills. Set up alerting rules, configure on-call rotations, and manage incident response workflows.

When should I use Alerting Oncall?

Alerting Oncall fits situations like: tasks that involve Incident response; tasks that involve Monitoring and alerting.

How do I install Alerting Oncall in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill alerting-oncall -a claude-code`. Or copy the skill folder (skills/alerting-oncall in sickn33/agentic-awesome-skills) into .claude/skills/alerting-oncall in your project. Claude Code loads it when a task matches its description.

How do I install Alerting Oncall in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill alerting-oncall -a codex`. Or copy the skill folder (skills/alerting-oncall in sickn33/agentic-awesome-skills) into .agents/skills/alerting-oncall in your project. Codex loads it when a task matches its description.

Can I use Alerting Oncall in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill alerting-oncall -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/alerting-oncall, .gemini/skills/alerting-oncall, .github/skills/alerting-oncall and .opencode/skills/alerting-oncall in your project.

What does Alerting Oncall need to run?

Going by SKILL.md and its folder, Alerting Oncall needs credentials named SECRET_KEY. Our summary lists: Docker; A credential in SECRET_KEY. Compatibility (from SKILL.md): Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled..

Does Alerting Oncall access the network?

SKILL.md names 3 domains. In commands or code: hooks.slack.com and events.pagerduty.com; the agent is likely to contact these when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Alerting Oncall safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Alerting Oncall use?

Alerting Oncall is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Alerting Oncall use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 499 tokens, read only when the agent opens those files.

What are the alternatives to Alerting Oncall?

Skills that share tags, products or a category with Alerting Oncall: Alerting Irm (grafana/skills, 282 stars), SRE Engineer (Jeffallan/claude-skills, 12k stars), Oncall Irm (grafana/skills, 282 stars) and Alerting Oncall (BagelHole/DevOps-Security-Agent-Skills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Alerting Oncall?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.