Agent skill

Prometheus System Health Check

by prometheus in prometheus/prometheus-mcp

Builds a picture of whether Prometheus itself is healthy and successfully monitoring its targets, covering readiness, firing alerts, target health and TSDB load.

Apache-2.0Auto-check passedDevOps & Cloud

Install Prometheus System Health Check

skills CLI
$ npx skills add prometheus/prometheus-mcp --skill check-system-health -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install prometheus/prometheus-mcp check-system-health --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/prometheus/prometheus-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/pkg/mcp/assets/skills/check-system-health .claude/skills/check-system-health && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
check-system-health
GitHub stars
121
Token cost
~584 tokens
SKILL.md length
264 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Builds a picture of whether Prometheus itself is healthy and successfully monitoring its targets, covering readiness, firing alerts, target health and TSDB load.

  • Answering is Prometheus healthy or is monitoring OK right now
  • SKILL.md covers Getting oriented, Topics worth exploring and Reporting findings
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Investigating why a scrape target shows as down or flapping

What it does

The skill treats healthy and ready as the first checks, confirming the server is up and able to serve queries, then layers in build_info and runtime_info for version and storage retention, list_targets for every scrape target's health, list_alerts for what is currently firing, and tsdb_stats for series counts, cardinality and ingestion load. It also suggests querying Prometheus's own self-monitoring metrics directly, such as a rate on prometheus_tsdb_head_samples_appended_total, and treating any steady rate on prometheus_rule_evaluation_failures_total or prometheus_notifications_dropped_total as a problem on its own.

From there it lists starting points to explore rather than a fixed checklist: firing alerts as the fastest pointer to a known problem, down or flapping targets with their lastError field explaining why, scrape duration close to its timeout found with a topk query, sudden growth in TSDB series counts as an early warning, Prometheus's own CPU, memory and disk headroom when node_exporter or cadvisor data is available, and data freshness checked with up or a timestamp comparison.

The result is meant to be a summary of server status, alert and target health, and TSDB load, with any issue called out alongside the specific evidence that shows it.

When your agent uses it

  • Answering is Prometheus healthy or is monitoring OK right now
  • Investigating why a scrape target shows as down or flapping
  • Checking for TSDB cardinality growth before it causes a bigger problem

Example prompts

  • “Is our monitoring healthy right now?”
  • “Why is the payments-service target showing as down?”
  • “Check for any firing alerts and summarize what needs attention.”

Requirements

  • A connected Prometheus MCP server
  • Compatibility (from SKILL.md): Requires the tools of a connected Prometheus MCP server

What it can do on your machine

Read from SKILL.md and the folder at commit 4e37ae1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the tools of a connected Prometheus MCP server

    From compatibility in the SKILL.md frontmatter.

Context cost

Prometheus System Health Check loads about 584 tokens when it runs. Until then it costs about 51 tokens; SKILL.md has 264 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~51
When it runs · the whole SKILL.md, loaded when a task matches
~584

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from prometheus/prometheus-mcp at commit 4e37ae1, republished under its Apache-2.0 licence (© prometheus). 264 words, ~584 tokens.

Download SKILL.mdSave it as .claude/skills/check-system-health/SKILL.md (or your agent's skills folder).
name
check-system-health
description
Health-check the Prometheus server and its targets. Use for status, uptime, or is-monitoring-OK questions; reviews readiness, active alerts, target health, TSDB stats, and runtime info.
compatibility
Requires the tools of a connected Prometheus MCP server
license
Apache-2.0

Understanding System Health

Build an overall picture of whether Prometheus itself is healthy and whether it is successfully monitoring what it should. A good health check covers the server, its targets, and the data being collected.

Getting oriented

  • healthy and ready report whether the server is up and able to serve queries.
  • build_info and runtime_info identify the version, storage retention, and runtime settings; flags exposes the full command line.
  • list_targets shows every scrape target and its health; list_alerts shows what is actively firing.
  • tsdb_stats summarizes series counts, cardinality, and ingestion (the load side of the picture).
  • query Prometheus's self-monitoring metrics directly: rate(prometheus_tsdb_head_samples_appended_total[5m]) shows ingestion throughput, and counters like prometheus_rule_evaluation_failures_total or prometheus_notifications_dropped_total should sit at zero -- any steady rate on them is a problem in itself.

Topics worth exploring

Treat these as starting points and follow what the data shows:

  • Firing alerts: anything already alerting is the fastest pointer to known problems; the labels identify the affected services.
  • Target health: down or flapping targets in list_targets mean missing data, and their lastError usually says why.
  • Scrape performance: how close scrapes run to their timeout, e.g. with query: topk(10, scrape_duration_seconds)
  • TSDB pressure: series counts and per-metric cardinality from tsdb_stats; sudden growth is an early warning sign.
  • Prometheus's own resources: if node_exporter or cadvisor metrics are available, check CPU, memory, and disk headroom, e.g.: process_resident_memory_bytes{job="prometheus"}
  • Data freshness: confirm critical metrics are current, e.g. query up, or time() - timestamp(<metric>) to spot staleness.

Reporting findings

Summarize server status, alert and target health, and TSDB load, and call out anything that needs attention together with the evidence that shows it.

© prometheus, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in pkg/mcp/assets/skills/check-system-health of prometheus/prometheus-mcp.

Open the folder on GitHubat commit 4e37ae1

Compare with similar skills

Prometheus System Health Check next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prometheus System Health Check compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prometheus System Health Check this skillprometheus/prometheus-mcp121—~584Automated safety check: PassApache-2.0
Happy Infra Metrics and Grafanaslopus/happy24k—~2kAutomated safety check: NotesMIT
WizTelemetry Platform Servicekubesphere/kubesphere17k—~1.8kAutomated safety check: PassCustom licence
Redis Observabilityredis/agent-skills1662 repos~911Automated safety check: PassMIT
Developing Funboost Mixinydf0509/funboost895—~2.1kAutomated safety check: PassNone
Archestra Dev Observabilityarchestra-ai/archestra4.4k—~1.2kAutomated safety check: PassCustom licence

Similar skills

  • Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.

    24k GitHub stars~2k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • WizTelemetry Platform Service

    kubesphere/kubesphere

    Installs and configures the WizTelemetry Platform Service extension for KubeSphere, the shared API server behind its observability extensions.

    17k GitHub stars~1.8k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Redis Observability

    redis/agent-skills

    Official

    Redis observability guidance — which metrics to monitor (memory, connections, hit ratio, ops/sec, rejected connections), which built-in commands to reach for during incident triage (SLOWLOG, INFO…

    166 GitHub starsUsed in 2 repos~911 tokens
    DevOps & CloudAuto-check passed
  • 当需要为 funboost 创建 Consumer 或 Publisher 的 Mixin 扩展类时使用。触发场景:添加监控、熔断、限流、链路追踪等横切关注点,编写自定义前置/后置处理钩子。关键词:mixin, consumeroverridecls, publisheroverridecls, ConsumerMixin, 自定义消费者, hook, 拦截器, 熔断器, 监控…

    895 GitHub stars~2.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Archestra Dev Observability

    archestra-ai/archestra

    A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup.

    4.4k GitHub stars~1.2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Frontmcp Observability

    agentfront/frontmcp

    A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server.

    146 GitHub stars~4.6k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from prometheus/prometheus-mcp

  • Prometheus Top Resource Consumers

    prometheus/prometheus-mcp

    Find the top CPU, memory, or disk consumers. Use for capacity reviews, noisy-neighbor hunts, and top-N questions about which jobs, pods, or instances use the…

    121 GitHub stars~569 tokensUpdated today
    Auto-check passed
  • Prometheus Error Rate Investigator

    prometheus/prometheus-mcp

    Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in.

    121 GitHub stars~592 tokensUpdated today
    Auto-check passed
  • Prometheus Cardinality Optimizer

    prometheus/prometheus-mcp

    Finds the metrics and labels behind a Prometheus series explosion using a connected Prometheus MCP server, then proposes relabeling, dropping or recording-rule fixes with measured impact.

    121 GitHub stars~547 tokensUpdated today
    Auto-check passed
  • Prometheus Rules Review

    prometheus/prometheus-mcp

    Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML.

    121 GitHub stars~765 tokensUpdated today
    Auto-check passed
  • Finds where a Prometheus metric stops existing, whether at the target, the scrape, relabeling or the query, using the tools of a connected Prometheus MCP server.

    121 GitHub stars~587 tokensUpdated today
    Auto-check passed
  • Tune Prometheus Config

    prometheus/prometheus-mcp

    Review and tune Prometheus configuration and performance. An agent skill from prometheus/prometheus-mcp.

    121 GitHub stars~724 tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Prometheus System Health Check

What does Prometheus System Health Check do?

Builds a picture of whether Prometheus itself is healthy and successfully monitoring its targets, covering readiness, firing alerts, target health and TSDB load. The skill treats healthy and ready as the first checks, confirming the server is up and able to serve queries, then layers in build_info and runtime_info for version and storage retention, list_targets for every scrape target's health, list_alerts for what is currently firing, and tsdb_stats for series counts, cardinality and ingestion load. It also suggests querying Prometheus's own self-monitoring metrics directly, such as a rate on prometheus_tsdb_head_samples_appended_total, and treating any steady rate on prometheus_rule_evaluation_failures_total or prometheus_notifications_dropped_total as a problem on its own.

When should I use Prometheus System Health Check?

Prometheus System Health Check fits situations like: answering is Prometheus healthy or is monitoring OK right now; investigating why a scrape target shows as down or flapping; checking for TSDB cardinality growth before it causes a bigger problem.

How do I install Prometheus System Health Check in Claude Code?

Run `npx skills add prometheus/prometheus-mcp --skill check-system-health -a claude-code`. Or copy the skill folder (pkg/mcp/assets/skills/check-system-health in prometheus/prometheus-mcp) into .claude/skills/check-system-health in your project. Claude Code loads it when a task matches its description.

How do I install Prometheus System Health Check in Codex?

Run `npx skills add prometheus/prometheus-mcp --skill check-system-health -a codex`. Or copy the skill folder (pkg/mcp/assets/skills/check-system-health in prometheus/prometheus-mcp) into .agents/skills/check-system-health in your project. Codex loads it when a task matches its description.

Can I use Prometheus System Health Check in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add prometheus/prometheus-mcp --skill check-system-health -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/check-system-health, .gemini/skills/check-system-health, .github/skills/check-system-health and .opencode/skills/check-system-health in your project.

What does Prometheus System Health Check need to run?

SKILL.md names no scripts, command-line tools or credentials: Prometheus System Health Check is instructions for the agent only. Our summary lists: A connected Prometheus MCP server. Compatibility (from SKILL.md): Requires the tools of a connected Prometheus MCP server.

Does Prometheus System Health Check access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Prometheus System Health Check safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prometheus System Health Check use?

Prometheus System Health Check is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prometheus System Health Check use?

About 584 tokens (SKILL.md is roughly 2.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Prometheus System Health Check?

Skills that share tags, products or a category with Prometheus System Health Check: Happy Infra Metrics and Grafana (slopus/happy, 24k stars), WizTelemetry Platform Service (kubesphere/kubesphere, 17k stars), Redis Observability (redis/agent-skills, 166 stars) and Developing Funboost Mixin (ydf0509/funboost, 895 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prometheus System Health Check?

prometheus (a GitHub organization) maintains it in prometheus/prometheus-mcp, which has 121 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 10, 2026.

Source: prometheus/prometheus-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.