Agent skill

Prometheus Error Rate Investigator

by prometheus in prometheus/prometheus-mcp

Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in.

Apache-2.0Auto-check passedDevOps & Cloud

Install Prometheus Error Rate Investigator

skills CLI
$ npx skills add prometheus/prometheus-mcp --skill investigate-error-rates -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install prometheus/prometheus-mcp investigate-error-rates --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/prometheus/prometheus-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/pkg/mcp/assets/skills/investigate-error-rates .claude/skills/investigate-error-rates && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
investigate-error-rates
GitHub stars
120
Token cost
~592 tokens
SKILL.md length
257 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in.

  • Investigating a spike in 5xx or gRPC error responses
  • SKILL.md covers Getting oriented, Topics worth exploring and Reporting findings
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Confirming whether a firing error-rate alert reflects a real traffic-wide problem

What it does

The skill starts by discovering what error signals actually exist in the target system, since they vary: HTTP status counters, gRPC code counters, or dedicated errors or failures metrics, found with series or label-value queries and confirmed against metric metadata, alongside checking list_alerts and the ALERTS firing series for anything already flagged. From there it walks through computing a current error rate with rate() over raw counters, dividing by total traffic to get a ratio that separates more errors from more traffic, and running the same query over a longer range to find the inflection point against a known-good baseline window.

To localize the problem it aggregates the error query by job, instance, handler or method, using topk to keep the output readable, and cross-checks the same time window against deploys or restarts, resource saturation and latency histograms for the same services. The expected output is a clear statement of impact in terms of error ratio and affected traffic, an onset time, and the narrowest scope that explains the spike, backed by the actual queries run.

When your agent uses it

  • Investigating a spike in 5xx or gRPC error responses
  • Confirming whether a firing error-rate alert reflects a real traffic-wide problem
  • Narrowing an error spike down to a specific job, instance or route

Example prompts

  • “Our 5xx rate just spiked; find out when it started and which service it is concentrated in.”
  • “Check whether this firing alert is a traffic-wide issue or just one instance.”
  • “Compare the current gRPC error ratio against last week's baseline for the checkout service.”

Requirements

  • A connected Prometheus MCP server
  • Compatibility (from SKILL.md): Requires the tools of a connected Prometheus MCP server

What it can do on your machine

Read from SKILL.md and the folder at commit 856fb45. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the tools of a connected Prometheus MCP server

    From compatibility in the SKILL.md frontmatter.

Context cost

Prometheus Error Rate Investigator loads about 592 tokens when it runs. Until then it costs about 65 tokens; SKILL.md has 257 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~592

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from prometheus/prometheus-mcp at commit 856fb45, republished under its Apache-2.0 licence (© prometheus). 257 words, ~592 tokens.

Download SKILL.mdSave it as .claude/skills/investigate-error-rates/SKILL.md (or your agent's skills folder).
name
investigate-error-rates
description
Investigate elevated error rates and failing requests. Use when errors or 5xx/gRPC failures spike, SLOs burn, or an error alert fires; quantifies error ratios with rate(), compares to baseline, and isolates affected jobs and instances.
compatibility
Requires the tools of a connected Prometheus MCP server
license
Apache-2.0

Investigating High Error Response Rates

Figure out how bad the errors are, when they started, and where they are concentrated. Error metrics vary by system, so discover what actually exists before querying.

Getting oriented

  • Error signals are usually counters: HTTP status codes (http_requests_total{code=~"5.."}), gRPC codes (grpc_server_handled_total{grpc_code!="OK"}), or dedicated *_errors_total / *_failures_total metrics.
  • Discover what this system exposes with series or label_values on name using regex matchers, and confirm semantics with metric_metadata.
  • list_alerts shows whether an error-related alert is already firing and carries useful labels to start from.
  • The ALERTS{alertstate="firing"} series Prometheus generates is the queryable form of the same information: range_query it to see when alerts started firing and line their history up against the error timeline.

Topics worth exploring

Treat these as starting points and follow what the data shows:

  • Current error rate: counters need rate() before aggregating, e.g. with query: sum by (job) (rate(http_requests_total{code=~"5.."}[5m]))
  • Error ratio vs raw count: a ratio against total traffic distinguishes "more errors" from "more traffic": sum(rate(http_requests_total{code=~"5.."}[5m])) / sum(rate(http_requests_total[5m]))
  • When it started: range_query the same expression over hours or days and look for the inflection point; compare against a known-good baseline window.
  • Where it is concentrated: aggregate by job, instance, handler, or method; topk(10, ...) keeps output manageable.
  • What else changed at that time: deploys and restarts (process_start_time_seconds), saturation (CPU, memory, connection pools), and latency histograms for the same services.

Reporting findings

Aim for a clear statement of impact (error ratio and affected traffic), onset time, and the narrowest scope (service, instance, or route) that explains the errors, backed by the queries you ran.

© prometheus, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in pkg/mcp/assets/skills/investigate-error-rates of prometheus/prometheus-mcp.

Open the folder on GitHubat commit 856fb45

Compare with similar skills

Prometheus Error Rate Investigator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prometheus Error Rate Investigator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prometheus Error Rate Investigator this skillprometheus/prometheus-mcp120—~592Automated safety check: PassApache-2.0
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Observability MonitoringAnastasiyaW/codex-claude-code-config154—~4.1kAutomated safety check: PassMIT
Observability Sremajiayu000/spellbook287—~3.3kAutomated safety check: PassMIT
Observability Patternssoftspark/ai-toolkit179—~2.2kAutomated safety check: PassApache-2.0
Telemetrymagnus919/agent-skills116—~3.9kAutomated safety check: PassMIT

Similar skills

  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 6 mo ago
    DevOps & CloudAuto-check passed
  • Observability Monitoring

    AnastasiyaW/codex-claude-code-config

    Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…

    154 GitHub stars~4.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Observability Sre

    majiayu000/spellbook

    Observability and SRE expert. An agent skill from majiayu000/spellbook.

    287 GitHub stars~3.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Observability Patterns

    softspark/ai-toolkit

    Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI.

    179 GitHub stars~2.2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Telemetry

    magnus919/agent-skills

    Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector…

    116 GitHub stars~3.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Cloud Monitoring

    seb1n/awesome-ai-agent-skills

    Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability.

    206 GitHub stars~2.8k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed

More from prometheus/prometheus-mcp

  • Prometheus System Health Check

    prometheus/prometheus-mcp

    Builds a picture of whether Prometheus itself is healthy and successfully monitoring its targets, covering readiness, firing alerts, target health and TSDB load.

    120 GitHub stars~584 tokensUpdated 5 days ago
    Auto-check passed
  • Prometheus Top Resource Consumers

    prometheus/prometheus-mcp

    Find the top CPU, memory, or disk consumers. Use for capacity reviews, noisy-neighbor hunts, and top-N questions about which jobs, pods, or instances use the…

    120 GitHub stars~569 tokensUpdated 5 days ago
    Auto-check passed
  • Prometheus Cardinality Optimizer

    prometheus/prometheus-mcp

    Finds the metrics and labels behind a Prometheus series explosion using a connected Prometheus MCP server, then proposes relabeling, dropping or recording-rule fixes with measured impact.

    120 GitHub stars~547 tokensUpdated 5 days ago
    Auto-check passed
  • Prometheus Rules Review

    prometheus/prometheus-mcp

    Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML.

    120 GitHub stars~765 tokensUpdated 5 days ago
    Auto-check passed
  • Finds where a Prometheus metric stops existing, whether at the target, the scrape, relabeling or the query, using the tools of a connected Prometheus MCP server.

    120 GitHub stars~587 tokensUpdated 5 days ago
    Auto-check passed
  • Tune Prometheus Config

    prometheus/prometheus-mcp

    Review and tune Prometheus configuration and performance. An agent skill from prometheus/prometheus-mcp.

    120 GitHub stars~724 tokensUpdated 5 days ago
    Auto-check passed

Works with

Categories

Questions about Prometheus Error Rate Investigator

What does Prometheus Error Rate Investigator do?

Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in. The skill starts by discovering what error signals actually exist in the target system, since they vary: HTTP status counters, gRPC code counters, or dedicated errors or failures metrics, found with series or label-value queries and confirmed against metric metadata, alongside checking list_alerts and the ALERTS firing series for anything already flagged. From there it walks through computing a current error rate with rate() over raw counters, dividing by total traffic to get a ratio that separates more errors from more traffic, and running the same query over a longer range to find the inflection point against a known-good baseline window.

When should I use Prometheus Error Rate Investigator?

Prometheus Error Rate Investigator fits situations like: investigating a spike in 5xx or gRPC error responses; confirming whether a firing error-rate alert reflects a real traffic-wide problem; narrowing an error spike down to a specific job, instance or route.

How do I install Prometheus Error Rate Investigator in Claude Code?

Run `npx skills add prometheus/prometheus-mcp --skill investigate-error-rates -a claude-code`. Or copy the skill folder (pkg/mcp/assets/skills/investigate-error-rates in prometheus/prometheus-mcp) into .claude/skills/investigate-error-rates in your project. Claude Code loads it when a task matches its description.

How do I install Prometheus Error Rate Investigator in Codex?

Run `npx skills add prometheus/prometheus-mcp --skill investigate-error-rates -a codex`. Or copy the skill folder (pkg/mcp/assets/skills/investigate-error-rates in prometheus/prometheus-mcp) into .agents/skills/investigate-error-rates in your project. Codex loads it when a task matches its description.

Can I use Prometheus Error Rate Investigator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add prometheus/prometheus-mcp --skill investigate-error-rates -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/investigate-error-rates, .gemini/skills/investigate-error-rates, .github/skills/investigate-error-rates and .opencode/skills/investigate-error-rates in your project.

What does Prometheus Error Rate Investigator need to run?

SKILL.md names no scripts, command-line tools or credentials: Prometheus Error Rate Investigator is instructions for the agent only. Our summary lists: A connected Prometheus MCP server. Compatibility (from SKILL.md): Requires the tools of a connected Prometheus MCP server.

Does Prometheus Error Rate Investigator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Prometheus Error Rate Investigator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prometheus Error Rate Investigator use?

Prometheus Error Rate Investigator is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prometheus Error Rate Investigator use?

About 592 tokens (SKILL.md is roughly 2.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Prometheus Error Rate Investigator?

Skills that share tags, products or a category with Prometheus Error Rate Investigator: Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Observability Monitoring (AnastasiyaW/codex-claude-code-config, 154 stars), Observability Sre (majiayu000/spellbook, 287 stars) and Observability Patterns (softspark/ai-toolkit, 179 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prometheus Error Rate Investigator?

prometheus (a GitHub organization) maintains it in prometheus/prometheus-mcp, which has 120 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 4, 2026.

Source: prometheus/prometheus-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.