Agent skill

Prometheus Rules Review

by prometheus in prometheus/prometheus-mcp

Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML.

Apache-2.0Auto-check passedDevOps & Cloud

Install Prometheus Rules Review

skills CLI
$ npx skills add prometheus/prometheus-mcp --skill review-prometheus-rules -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install prometheus/prometheus-mcp review-prometheus-rules --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/prometheus/prometheus-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/pkg/mcp/assets/skills/review-prometheus-rules .claude/skills/review-prometheus-rules && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
review-prometheus-rules
GitHub stars
118
Token cost
~765 tokens
SKILL.md length
335 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML.

  • Reviewing alert coverage and rule health in a Prometheus setup
  • SKILL.md covers Getting oriented, Topics worth exploring and Suggesting changes
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Drafting new alerts or recording rules

What it does

The agent starts with list_rules for every rule group with expressions and evaluation health, list_alerts for what is firing, tsdb_stats for series counts and cardinality, and list_targets for scraped services that no rule covers. Runtime metrics help too: the ALERTS series reveals flapping, rate(prometheus_rule_evaluation_failures_total[5m]) catches rules that fail to evaluate, and prometheus_rule_group_last_duration_seconds close to the group's interval signals a group running out of time.

Review topics include broken or always-empty rules, recording rule candidates among repeated rate, sum and histogram_quantile expressions, alert quality such as symptom-based alerts, meaningful severity labels, sensible for durations and filled-in annotations, coverage gaps like SLOs without multi-window burn-rate alerts or critical metrics without absent() protection, and the level:metric:operations naming convention. The result is complete rule-group YAML.

When your agent uses it

  • Reviewing alert coverage and rule health in a Prometheus setup
  • Drafting new alerts or recording rules
  • Investigating flapping or noisy alerts
  • Speeding up slow dashboards with pre-computed expressions

Example prompts

  • “Audit our Prometheus rules and tell me which alerts have never fired or fail to evaluate.”
  • “The HighErrorRate alert keeps flapping, so use the ALERTS series to find out why.”
  • “Propose recording rules for the expressions our dashboards repeat, with names that follow the convention.”
  • “Check which scraped targets have no alerting rules at all and draft some.”

Requirements

  • A connected Prometheus MCP server
  • Compatibility (from SKILL.md): Requires the tools of a connected Prometheus MCP server

What it can do on your machine

Read from SKILL.md and the folder at commit 856fb45. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the tools of a connected Prometheus MCP server

    From compatibility in the SKILL.md frontmatter.

Context cost

Prometheus Rules Review loads about 765 tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 335 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~765

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from prometheus/prometheus-mcp at commit 856fb45, republished under its Apache-2.0 licence (© prometheus). 335 words, ~765 tokens.

Download SKILL.mdSave it as .claude/skills/review-prometheus-rules/SKILL.md (or your agent's skills folder).
name
review-prometheus-rules
description
Audit and improve recording and alerting rules. Use when reviewing alert coverage or rule health, drafting new alerts or recording rules, investigating flapping or noisy alerts, or speeding up slow dashboards with pre-computed expressions; produces complete rule-group YAML.
compatibility
Requires the tools of a connected Prometheus MCP server
license
Apache-2.0

Reviewing and Improving Prometheus Rules

Recording rules pre-compute expensive or repeated expressions into new series; alerting rules turn expressions into notifications. They live together in rule groups and are best designed together -- good alerts are often built on top of good recording rules. This runbook covers auditing the rules that exist, finding gaps, and drafting improvements.

Getting oriented

  • list_rules returns every rule group -- recording and alerting -- with expressions, labels, and per-rule evaluation health.
  • list_alerts shows what is firing right now; useful for separating "badly tuned" from "correctly noisy".
  • tsdb_stats reveals series counts and cardinality, which tells you whether a pre-aggregation would actually pay off.
  • list_targets can surface scraped services that no rule covers at all.
  • Prometheus's own metrics expose rule health at runtime: the ALERTS{} series records pending and firing alerts (range_query it to reconstruct an alert's history and spot flapping), rate(prometheus_rule_evaluation_failures_total[5m]) catches rules that fail to evaluate, and prometheus_rule_group_last_duration_seconds approaching the group's interval means evaluation is running out of budget.

Topics worth exploring

Treat these as starting points and follow what the data shows, not as a fixed checklist:

  • Rule health: list_rules reports evaluation health and last error per rule; broken or always-empty rules are quick wins. Test a suspect expression directly by querying.
  • Recording rule candidates: expressions repeated across alerts or dashboards, especially rate/sum/histogram_quantile combinations. Verify a candidate returns what you expect before proposing it, e.g. with range_query: sum by (job) (rate(http_requests_total{code=~"5.."}[5m])) / sum by (job) (rate(http_requests_total[5m]))
  • Alert quality: alert on symptoms rather than causes, keep severity labels meaningful, balance for durations against detection speed, and fill in annotations (summary, description, runbook link).
  • Coverage gaps: SLOs without multi-window multi-burn-rate alerts, critical metrics without absent() protection, resource exhaustion without capacity alerting.
  • Naming conventions: recording rules follow level:metric:operations (e.g. job:http_requests:rate5m); docs_search has the full convention and the rule config schema.

Suggesting changes

Propose complete rule-group YAML, and sanity-check every expression against live data with query or range_query first. Explain what each rule is for, and recommend validating with promtool and watching rule evaluation health after deployment.

© prometheus, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in pkg/mcp/assets/skills/review-prometheus-rules of prometheus/prometheus-mcp.

Open the folder on GitHubat commit 856fb45

Compare with similar skills

Prometheus Rules Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prometheus Rules Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prometheus Rules Review this skillprometheus/prometheus-mcp118—~765Automated safety check: PassApache-2.0
Archestra Dev Observabilityarchestra-ai/archestra4.4k—~1.2kAutomated safety check: PassCustom licence
Frontmcp Observabilityagentfront/frontmcp146—~4.6kAutomated safety check: PassApache-2.0
AWS Cost OperationsMicrock/ordinary-claude-skills4031 repos~2.5kAutomated safety check: PassCustom licence
Happy Infra Metrics and Grafanaslopus/happy24k—~2kAutomated safety check: NotesMIT
Syncmetapawurb/hotpath-rs1.9k—~1.2kAutomated safety check: NotesMIT

Similar skills

  • Archestra Dev Observability

    archestra-ai/archestra

    A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup.

    4.4k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Frontmcp Observability

    agentfront/frontmcp

    A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server.

    146 GitHub stars~4.6k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • AWS Cost Operations

    Microck/ordinary-claude-skills

    This skill provides AWS cost optimization, monitoring, and operational best practices with integrated MCP servers for billing analysis, cost estimation, observability, and security assessment.

    403 GitHub starsUsed in 1 repo~2.5k tokens
    DevOps & CloudAuto-check passed
  • Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.

    24k GitHub stars~2k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Syncmeta

    pawurb/hotpath-rs

    Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta).

    1.9k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • WizTelemetry Platform Service

    kubesphere/kubesphere

    Installs and configures the WizTelemetry Platform Service extension for KubeSphere, the shared API server behind its observability extensions.

    17k GitHub stars~1.8k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed

More from prometheus/prometheus-mcp

  • Prometheus System Health Check

    prometheus/prometheus-mcp

    Builds a picture of whether Prometheus itself is healthy and successfully monitoring its targets, covering readiness, firing alerts, target health and TSDB load.

    118 GitHub stars~584 tokensUpdated 3 days ago
    Auto-check passed
  • Prometheus Top Resource Consumers

    prometheus/prometheus-mcp

    Find the top CPU, memory, or disk consumers. Use for capacity reviews, noisy-neighbor hunts, and top-N questions about which jobs, pods, or instances use the…

    118 GitHub stars~569 tokensUpdated 3 days ago
    Auto-check passed
  • Prometheus Error Rate Investigator

    prometheus/prometheus-mcp

    Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in.

    118 GitHub stars~592 tokensUpdated 3 days ago
    Auto-check passed
  • Prometheus Cardinality Optimizer

    prometheus/prometheus-mcp

    Finds the metrics and labels behind a Prometheus series explosion using a connected Prometheus MCP server, then proposes relabeling, dropping or recording-rule fixes with measured impact.

    118 GitHub stars~547 tokensUpdated 3 days ago
    Auto-check passed
  • Finds where a Prometheus metric stops existing, whether at the target, the scrape, relabeling or the query, using the tools of a connected Prometheus MCP server.

    118 GitHub stars~587 tokensUpdated 3 days ago
    Auto-check passed
  • Tune Prometheus Config

    prometheus/prometheus-mcp

    Review and tune Prometheus configuration and performance. An agent skill from prometheus/prometheus-mcp.

    118 GitHub stars~724 tokensUpdated 3 days ago
    Auto-check passed

Categories

Questions about Prometheus Rules Review

What does Prometheus Rules Review do?

Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML. The agent starts with list_rules for every rule group with expressions and evaluation health, list_alerts for what is firing, tsdb_stats for series counts and cardinality, and list_targets for scraped services that no rule covers. Runtime metrics help too: the ALERTS series reveals flapping, rate(prometheus_rule_evaluation_failures_total[5m]) catches rules that fail to evaluate, and prometheus_rule_group_last_duration_seconds close to the group's interval signals a group running out of time.

When should I use Prometheus Rules Review?

Prometheus Rules Review fits situations like: reviewing alert coverage and rule health in a Prometheus setup; drafting new alerts or recording rules; investigating flapping or noisy alerts; speeding up slow dashboards with pre-computed expressions.

How do I install Prometheus Rules Review in Claude Code?

Run `npx skills add prometheus/prometheus-mcp --skill review-prometheus-rules -a claude-code`. Or copy the skill folder (pkg/mcp/assets/skills/review-prometheus-rules in prometheus/prometheus-mcp) into .claude/skills/review-prometheus-rules in your project. Claude Code loads it when a task matches its description.

How do I install Prometheus Rules Review in Codex?

Run `npx skills add prometheus/prometheus-mcp --skill review-prometheus-rules -a codex`. Or copy the skill folder (pkg/mcp/assets/skills/review-prometheus-rules in prometheus/prometheus-mcp) into .agents/skills/review-prometheus-rules in your project. Codex loads it when a task matches its description.

Can I use Prometheus Rules Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add prometheus/prometheus-mcp --skill review-prometheus-rules -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review-prometheus-rules, .gemini/skills/review-prometheus-rules, .github/skills/review-prometheus-rules and .opencode/skills/review-prometheus-rules in your project.

What does Prometheus Rules Review need to run?

SKILL.md names no scripts, command-line tools or credentials: Prometheus Rules Review is instructions for the agent only. Our summary lists: A connected Prometheus MCP server. Compatibility (from SKILL.md): Requires the tools of a connected Prometheus MCP server.

Does Prometheus Rules Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Prometheus Rules Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prometheus Rules Review use?

Prometheus Rules Review is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prometheus Rules Review use?

About 765 tokens (SKILL.md is roughly 3.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Prometheus Rules Review?

Skills that share tags, products or a category with Prometheus Rules Review: Archestra Dev Observability (archestra-ai/archestra, 4.4k stars), Frontmcp Observability (agentfront/frontmcp, 146 stars), AWS Cost Operations (Microck/ordinary-claude-skills, 403 stars) and Happy Infra Metrics and Grafana (slopus/happy, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prometheus Rules Review?

prometheus (a GitHub organization) maintains it in prometheus/prometheus-mcp, which has 118 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 4, 2026.

Source: prometheus/prometheus-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.