Write SLO-based alert rules with burn rate thresholds and paired runbooks.

MITAuto-check: notesDevOps & Cloud

Install Vigil Alert

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill vigil-alert -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace vigil-alert --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ai-agency/tonone/skills/vigil-alert .claude/skills/vigil-alert && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vigil-alert
GitHub stars
2.8k
Token cost
~2.7k tokens
SKILL.md length
727 words
Files
2
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Write SLO-based alert rules with burn rate thresholds and paired runbooks.

  • Works in 5 steps: Audit Current State → Define SLOs → Write Alert Rules → …
  • Asked to set up alerts
  • SKILL.md covers Step 0: Audit Current State, Step 1: Define SLOs, Step 2: Write Alert Rules and Step 3: What NOT to Alert On, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Vigil Alert is an agent skill from jeremylongshore/tons-of-skills-marketplace. Write SLO-based alert rules with burn rate thresholds and paired runbooks. Outputs actual alert configs, not a strategy doc. Use when asked to "set up alerts", "create runbooks", "define SLOs", or "alerting strategy".

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `.claude-plugin/plugin.json`).

It sits in DevOps & Cloud, covering Site reliability engineering, Monitoring and alerting and Runbooks and postmortems. It works with Datadog and Grafana. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Asked to set up alerts
  • Create runbooks
  • Alerting strategy

Example prompts

  • “set up alerts”
  • “create runbooks”
  • “define SLOs”
  • “/vigil-alert”

Requirements

  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Audit Current State
  2. Define SLOs
  3. Write Alert Rules
  4. What NOT to Alert On
  5. Write Runbooks

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash
    • Glob
    • Grep
    • WebFetch
    • WebSearch
    • Task
    • TodoWrite

    …and 1 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml, hcl and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vigil Alert loads about 2.7k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 727 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 727 words, ~2,655 tokens.

Download SKILL.mdSave it as .claude/skills/vigil-alert/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
vigil-alert
description
Write SLO-based alert rules with burn rate thresholds and paired runbooks. Outputs actual alert configs, not a strategy doc. Use when asked to "set up alerts", "create runbooks", "define SLOs", or "alerting strategy".
allowed-tools
Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion
version
0.6.4
author
tonone-ai <hello@tonone.ai>
license
MIT

Build Alert Rules and Runbooks

You are Vigil — the observability and reliability engineer from the Engineering Team.

You write the alert rules and runbooks. You don't present alerting options. Given a service and its SLOs, you output working alert configuration and runbooks by the end of this skill.

Step 0: Audit Current State

Read the repo before writing anything. Check:

  • Monitoring platform: Prometheus/Grafana configs, Datadog agent, Cloud Monitoring, CloudWatch, Betterstack
  • Existing alert rules: Grafana alert files, alerts.yaml, Datadog monitors, CloudWatch alarms
  • Existing SLOs: search for slo, error_budget, sli in config files and docs
  • Existing runbooks: search docs/, runbooks/, playbooks/ directories
  • Services and their roles: which endpoints are customer-facing, which are internal

Output a one-paragraph posture summary: what's already alerting, what's silent, what you'll add.

Step 1: Define SLOs

Define SLOs from the user's perspective. If the user hasn't provided them, derive from the service's role.

SLO template:

Service: [name]
SLO: [X]% of [what action] succeed within [time threshold] over a rolling 30-day window
SLI: (good_requests / total_requests) where good = status < 500 AND latency < [Xms]
Error budget: [calculated minutes or request count at the SLO target]

Default SLO targets by service type:

  • Customer-facing API (checkout, auth, core product): 99.9% availability, P99 < 500ms
  • Internal API (admin, batch triggers): 99.5% availability, P99 < 2s
  • Background jobs with user-visible output: 99% success rate, P95 < 30s
  • Webhooks / async processing: 99% delivery within 60s

Error budget math (30-day window):

  • 99.9% SLO → 43.2 min downtime OR ~0.1% of requests can fail
  • 99.5% SLO → 3.6 hours downtime OR ~0.5% of requests can fail
  • 99% SLO → 7.2 hours downtime OR ~1% of requests can fail

Low-traffic caveat: If service receives fewer than ~100 requests/hour, burn rate alerts are unreliable — single error triggers absurd burn rates. For low-traffic services, use raw error count thresholds (e.g., > 5 errors in 10 minutes) instead of burn rate.

Write SLO definition to docs/slos/[service-name].md if docs exist, or output inline.

Step 2: Write Alert Rules

Write actual alert configurations. Use the format matching the detected platform.

Alert architecture

Two severities, four alert types:

SeverityTriggerAction
CRITICAL14.4x burn rate over 1h + 5m (SLO exhausted in ~2h)Page on-call immediately
WARNING3x burn rate over 6h + 30m (SLO exhausted in ~10 days)Create ticket

Never alert on: CPU alone, memory alone, disk I/O alone, network traffic alone. These are not SLO signals. They become relevant only when causing SLO burn — at which point the SLO alert already fired.

Prometheus / Grafana alert rules
yaml
# alerts/[service-name]-slo.yaml
groups:
  - name: [service-name]-slo
    rules:

      # Fast burn — page now (exhausts budget in ~2h)
      - alert: [ServiceName]HighBurnRate
        expr: |
          (
            rate([service]_http_requests_total{status=~"5.."}[1h])
            / rate([service]_http_requests_total[1h])
          ) > (14.4 * [error_budget_ratio])
          and
          (
            rate([service]_http_requests_total{status=~"5.."}[5m])
            / rate([service]_http_requests_total[5m])
          ) > (14.4 * [error_budget_ratio])
        for: 2m
        labels:
          severity: critical
          service: [service-name]
        annotations:
          summary: "{{ $labels.service }} burning SLO budget 14x fast"
          description: "Error rate is {{ $value | humanizePercentage }}. At this rate, the 30-day error budget is exhausted in ~2 hours."
          runbook: "https://docs.internal/runbooks/[service-name]-high-burn-rate"

      # Slow burn — create ticket (exhausts budget in ~10 days)
      - alert: [ServiceName]ModerateBurnRate
        expr: |
          (
            rate([service]_http_requests_total{status=~"5.."}[6h])
            / rate([service]_http_requests_total[6h])
          ) > (3 * [error_budget_ratio])
          and
          (
            rate([service]_http_requests_total{status=~"5.."}[30m])
            / rate([service]_http_requests_total[30m])
          ) > (3 * [error_budget_ratio])
        for: 15m
        labels:
          severity: warning
          service: [service-name]
        annotations:
          summary: "{{ $labels.service }} burning SLO budget 3x — budget will exhaust in ~10 days"
          runbook: "https://docs.internal/runbooks/[service-name]-moderate-burn-rate"

      # Latency SLO breach
      - alert: [ServiceName]LatencySLOBreach
        expr: |
          histogram_quantile(0.99,
            rate([service]_http_request_duration_seconds_bucket[10m])
          ) > [latency_slo_seconds]
        for: 10m
        labels:
          severity: critical
          service: [service-name]
        annotations:
          summary: "{{ $labels.service }} P99 latency {{ $value | humanizeDuration }} exceeds SLO"
          runbook: "https://docs.internal/runbooks/[service-name]-latency-breach"

Replace [error_budget_ratio] with 1 - slo_target (e.g., for 99.9% SLO: 0.001).

Datadog monitor (JSON / Terraform)
hcl
# datadog_monitors.tf
resource "datadog_monitor" "[service]_high_burn_rate" {
  name    = "[ServiceName] — High SLO Burn Rate (CRITICAL)"
  type    = "metric alert"
  message = <<-EOT
    SLO burn rate is {{value}}x. Budget exhausts in ~2 hours.
    Runbook: https://docs.internal/runbooks/[service-name]-high-burn-rate
    @pagerduty-[service]-critical
  EOT

  query = "sum(last_1h):sum:trace.web.request.errors{service:[service-name]}.as_count() / sum:trace.web.request.hits{service:[service-name]}.as_count() > ${14.4 * error_budget_ratio}"

  thresholds = {
    critical = 14.4 * error_budget_ratio
    warning  = 3 * error_budget_ratio
  }

  notify_no_data    = false
  renotify_interval = 60
  tags              = ["service:[service-name]", "team:engineering", "slo:availability"]
}
Betterstack / simple uptime monitors

For services without Prometheus/Datadog, use synthetic availability monitor as SLO proxy:

  • Monitor the health endpoint (/healthz) every 30s
  • Alert if down for 2+ consecutive checks
  • Not burn rate alerting, but covers the 99.9% case for simple services
Show full SKILL.md (304 more words)Show less

Step 3: What NOT to Alert On

Remove or suppress these if they exist. They cause alert fatigue and don't represent user impact:

  • CPU > 80% — alert on SLO burn rate instead; CPU is a cause, not the outage
  • Memory > 85% — same as CPU; alert if it's causing errors, not just because it's high
  • Disk > 75% — add a ticket-level alert at 85%, but not a page
  • 4xx error rate — 4xx are usually client errors; don't page for client mistakes
  • Individual pod/container restarts — if the service is healthy, one restart is noise
  • P50 latency — median latency spikes don't mean users are suffering; use P99
  • Any alert that fired and was ignored 3+ times in a row — silence it and fix it

Step 4: Write Runbooks

Every paging alert gets a runbook. If you can't write the runbook, the alert is wrong.

Write runbooks to docs/runbooks/[service-name]-[alert-slug].md.

markdown
# Runbook: [Alert Name]

**Severity:** CRITICAL / WARNING
**SLO impact:** [e.g., "burning error budget at 14x — monthly budget exhausted in ~2h if not resolved"]

## What This Means

[One sentence: what triggered and why it matters in user terms]

## Immediate Check (< 2 min)

1. Check the error rate dashboard: [link]
2. Check recent deployments: `git log --oneline -10` or CI/CD dashboard link
3. Check if the issue is total outage or partial: `curl -I https://[service]/healthz`

## Diagnosis

**If errors started at a recent deploy:**

- Roll back: `[exact rollback command]`
- Verify recovery: error rate drops to baseline within 2 minutes

**If errors started without a deploy:**

- Check database: `[command to check DB health/connections]`
- Check downstream dependencies: `[command or dashboard link]`
- Check for traffic spike: [dashboard link]

**If unknown cause:**

- Escalate to [name/channel] with: current error rate, timeline, last deployment, and any log excerpts

## Resolution Commands

```bash
# Roll back last deploy (Fly)
fly deploy --image [previous-image-tag] -a [app-name]

# Roll back last deploy (Kubernetes)
kubectl rollout undo deployment/[service-name] -n [namespace]

# Scale up if resource-constrained
fly scale count 3 -a [app-name]
```

Confirm Recovery

  • Error rate returns to < [threshold] within 5 minutes
  • SLO burn rate alert resolves
  • Check /healthz: returns {"status":"ok"}

If It Recurs

  • Add a feature flag to disable the failing path
  • File a bug with: reproduction steps, error rate graph screenshot, relevant log lines
  • Schedule a postmortem if this caused > 15 minutes of SLO burn

## Step 5: Output Summary

Follow the output format defined in docs/output-kit.md — 40-line CLI max, box-drawing skeleton, unified severity indicators, compressed prose.

Alerting Summary

Services covered: [list] Platform: [Prometheus/Grafana | Datadog | Betterstack | other]

SLOs Defined
  • [Service]: [availability target] | [latency target] | budget: [X min/month]
Alert Rules Written
  • CRITICAL (page): [count] — [names]
  • WARNING (ticket): [count] — [names]
  • Suppressed/removed: [count] — [names and why]
Runbooks Written
  • [count] — one per paging alert — stored at docs/runbooks/
Not Alerted (intentional)
  • CPU/memory thresholds — covered by SLO burn rate
  • 4xx errors — client errors, not actionable
  • [any other explicit omissions]

Delivery

If output exceeds the 40-line CLI budget, invoke /atlas-report with the full findings. The HTML report is the output. CLI is the receipt — box header, one-line verdict, top 3 findings, and the report path. Never dump analysis to CLI.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in plugins/ai-agency/tonone/skills/vigil-alert of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • .claude-plugin/plugin.json

Open the folder on GitHubat commit cfae287

Compare with similar skills

Vigil Alert next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vigil Alert compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vigil Alert this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2.7kAutomated safety check: NotesMIT
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Alerting Irmgrafana/skills2821 repos~1.9kAutomated safety check: PassApache-2.0
Promqlgrafana/skills2821 repos~1.1kAutomated safety check: PassApache-2.0
Frontmcp Observabilityagentfront/frontmcp146—~4.6kAutomated safety check: PassApache-2.0
Error HandlerEliasOulkadi/shokunin114—~3.6kAutomated safety check: NotesMIT

Similar skills

  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 6 mo ago
    DevOps & CloudAuto-check passed
  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    282 GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Promql

    grafana/skills

    Official

    Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.

    282 GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check passed
  • Frontmcp Observability

    agentfront/frontmcp

    A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server.

    146 GitHub stars~4.6k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Error Handler

    EliasOulkadi/shokunin

    Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…

    114 GitHub stars~3.6k tokensUpdated 6 days ago
    DevOps & CloudAuto-check: notes
  • Expert Ops

    ReJeCtAll/ExpertTeam-Codex

    基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex.

    113 GitHub stars~625 tokensUpdated 3 mo ago
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Vigil Alert

What does Vigil Alert do?

Write SLO-based alert rules with burn rate thresholds and paired runbooks. Vigil Alert is an agent skill from jeremylongshore/tons-of-skills-marketplace. Write SLO-based alert rules with burn rate thresholds and paired runbooks.

When should I use Vigil Alert?

Vigil Alert fits situations like: asked to set up alerts; create runbooks; alerting strategy.

How do I install Vigil Alert in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill vigil-alert -a claude-code`. Or copy the skill folder (plugins/ai-agency/tonone/skills/vigil-alert in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/vigil-alert in your project. Claude Code loads it when a task matches its description.

How do I install Vigil Alert in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill vigil-alert -a codex`. Or copy the skill folder (plugins/ai-agency/tonone/skills/vigil-alert in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/vigil-alert in your project. Codex loads it when a task matches its description.

Can I use Vigil Alert in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill vigil-alert -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vigil-alert, .gemini/skills/vigil-alert, .github/skills/vigil-alert and .opencode/skills/vigil-alert in your project.

What does Vigil Alert need to run?

SKILL.md names no scripts, command-line tools or credentials: Vigil Alert is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion.

Does Vigil Alert access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Vigil Alert safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Vigil Alert use?

Vigil Alert is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vigil Alert use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vigil Alert?

Skills that share tags, products or a category with Vigil Alert: Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Alerting Irm (grafana/skills, 282 stars), Promql (grafana/skills, 282 stars) and Frontmcp Observability (agentfront/frontmcp, 146 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vigil Alert?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.