Agent skill

Agent Health Monitoring

by cosmicstack-labs in cosmicstack-labs/mercury-agent-skills

Monitor AI agent health, detect anomalies, set up alerting, and maintain observability dashboards for production multi-agent systems.

MITAuto-check passedDevOps & Cloud

Install Agent Health Monitoring

skills CLI
$ npx skills add cosmicstack-labs/mercury-agent-skills --skill agent-health-monitoring -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cosmicstack-labs/mercury-agent-skills agent-health-monitoring --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cosmicstack-labs/mercury-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/categories/ai-ml/agent-health-monitoring .claude/skills/agent-health-monitoring && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-health-monitoring
GitHub stars
476
Token cost
~2.7k tokens
SKILL.md length
613 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Monitor AI agent health, detect anomalies, set up alerting, and maintain observability dashboards for production multi-agent systems.

  • Works in 6 steps: Instrument Every Agent → Implement Liveness & Readiness Probes → Set Up Anomaly Detection → …
  • Tasks that involve Monitoring and alerting
  • SKILL.md covers Overview, Core Concepts, Step-by-Step Implementation and Trigger Phrases, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agent Health Monitoring is an agent skill from cosmicstack-labs/mercury-agent-skills. Monitor AI agent health, detect anomalies, set up alerting, and maintain observability dashboards for production multi-agent systems. Covers liveness checks, performance metrics, drift detection, and incident response.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Monitoring and alerting, Anomaly detection and GitOps. The repository describes itself as: A curated registry of reusable Mercury Agent, Open Claw or Hermes Agent skills designed for real developer workflows, persistent memory, and token-efficient execution. The licence is MIT.

When your agent uses it

  • Tasks that involve Monitoring and alerting
  • Tasks that involve Anomaly detection
  • Tasks that involve GitOps

Example prompts

  • “/agent-health-monitoring”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Instrument Every Agent
  2. Implement Liveness & Readiness Probes
  3. Set Up Anomaly Detection
  4. Build the Alerting Pipeline
  5. Define Alert Rules
  6. Build the Dashboard

What it can do on your machine

Read from SKILL.md and the folder at commit 30392fb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Health Monitoring loads about 2.7k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 613 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cosmicstack-labs/mercury-agent-skills at commit 30392fb, republished under its MIT licence (© cosmicstack-labs). 613 words, ~2,683 tokens.

Download SKILL.mdSave it as .claude/skills/agent-health-monitoring/SKILL.md (or your agent's skills folder).
name
agent-health-monitoring
description
Monitor AI agent health, detect anomalies, set up alerting, and maintain observability dashboards for production multi-agent systems. Covers liveness checks, performance metrics, drift detection, and incident response.
metadata.author
cosmicstack-labs
metadata.version
1.0.0
metadata.category
ai-ml
metadata.tags
agent-monitoring, observability, alerting, health-checks, incident-response, production-agents

Agent Health Monitoring & Alerting

Overview

Production multi-agent systems fail silently. An agent that stops responding, returns empty results, or enters an infinite loop can degrade an entire workflow without triggering traditional infrastructure alerts. This skill covers how to build comprehensive health monitoring, metrics collection, and alerting for AI agent fleets.


Core Concepts

Agent Vital Signs
MetricWhat It MeasuresWhy It Matters
Response Rate% of agent invocations that return a resultDropping rate indicates crashes or context overflows
Latency (P50/P95/P99)Time from invocation to responseSpikes indicate context bloat or degraded model performance
Error Rate% of invocations with errors/tool failuresRising rate indicates systemic issues
Step CountNumber of reasoning steps per taskUnbounded growth indicates looping behavior
Tool Call Success Rate% of tool calls that succeedDrop indicates broken integrations or rate limiting
Token ConsumptionTokens used per agent runBudget anomalies indicate runaway agents
Context Utilization% of context window usedHigh utilization risks truncation and quality loss
Hallucination ScoreConfidence calibration or factuality checksDegrading accuracy undermines trust
Alert Severity Levels
LevelColorResponse TimeExamples
P0 (Critical)🔴 Red< 5 minAgent completely down, data loss, security breach
P1 (High)🟠 Orange< 15 minError rate > 20%, latency 5x baseline
P2 (Medium)🟡 Yellow< 1 hourError rate > 5%, slow degradation
P3 (Low)🔵 Blue< 24 hoursSingle agent underperforming, minor drift

Step-by-Step Implementation

Step 1: Instrument Every Agent

Wrap every agent invocation with telemetry:

python
class MonitoredAgent:
    """Agent wrapper that collects metrics on every invocation."""
    
    def __init__(self, agent, agent_name: str, metrics_client):
        self.agent = agent
        self.agent_name = agent_name
        self.metrics = metrics_client
    
    async def run(self, task: str) -> str:
        start_time = time.time()
        step_count = 0
        token_usage = 0
        
        try:
            result = await self.agent.run(task)
            
            # Collect metrics
            duration = time.time() - start_time
            self.metrics.timing(f"agent.{self.agent_name}.latency", duration)
            self.metrics.increment(f"agent.{self.agent_name}.invocations")
            self.metrics.increment(f"agent.{self.agent_name}.success")
            self.metrics.gauge(f"agent.{self.agent_name}.steps", step_count)
            
            return result
            
        except Exception as e:
            duration = time.time() - start_time
            self.metrics.increment(f"agent.{self.agent_name}.errors")
            self.metrics.timing(f"agent.{self.agent_name}.error_latency", duration)
            raise
Step 2: Implement Liveness & Readiness Probes
python
class AgentHealthProbe:
    """Kubernetes-style health probes for AI agents."""
    
    async def liveness_check(self, agent) -> bool:
        """Is the agent process alive and responding?"""
        try:
            result = await asyncio.wait_for(
                agent.run("Respond with: OK"),
                timeout=5.0
            )
            return "OK" in result
        except (asyncio.TimeoutError, Exception):
            return False
    
    async def readiness_check(self, agent) -> dict:
        """Is the agent ready to accept tasks?"""
        checks = {
            "model_available": await self._check_model(agent),
            "tools_available": await self._check_tools(agent),
            "memory_available": await self._check_memory(agent),
            "context_capacity": await self._check_context(agent),
        }
        return {
            "ready": all(checks.values()),
            "checks": checks
        }
    
    async def deep_check(self, agent) -> dict:
        """Full diagnostic: run a test task and validate output."""
        test_task = agent.config.test_prompt
        result = await agent.run(test_task)
        return {
            "passed": self._validate_output(result),
            "output_preview": result[:200],
            "latency_ms": self._last_latency
        }
Step 3: Set Up Anomaly Detection
python
class AnomalyDetector:
    """Detect unusual agent behavior using statistical methods."""

    def __init__(self, window_size: int = 100):
        self.window_size = window_size
        self.metrics_history = defaultdict(list)

    def record(self, agent_name: str, metric: str, value: float):
        self.metrics_history[f"{agent_name}:{metric}"].append(value)
        
        # Keep rolling window
        history = self.metrics_history[f"{agent_name}:{metric}"]
        if len(history) > self.window_size:
            history.pop(0)

    def is_anomalous(self, agent_name: str, metric: str, value: float, 
                     z_threshold: float = 3.0) -> tuple[bool, float]:
        """Check if a value is anomalous using z-score."""
        history = self.metrics_history.get(f"{agent_name}:{metric}", [])
        if len(history) < 10:
            return False, 0.0  # Not enough data
        
        mean = statistics.mean(history)
        stdev = statistics.stdev(history)
        if stdev == 0:
            return False, 0.0
        
        z_score = (value - mean) / stdev
        return abs(z_score) > z_threshold, z_score
Step 4: Build the Alerting Pipeline
python
class AlertManager:
    """Route alerts to the right channels based on severity."""
    
    def __init__(self):
        self.channels = {
            "p0": ["pagerduty", "slack-critical", "phone"],
            "p1": ["slack-critical", "email"],
            "p2": ["slack-warn", "email"],
            "p3": ["dashboard", "weekly-report"],
        }
    
    async def alert(self, severity: str, title: str, message: str, 
                    context: dict = None):
        """Send an alert through the appropriate channels."""
        channels = self.channels.get(severity, self.channels["p3"])
        
        for channel in channels:
            await self._send(channel, {
                "severity": severity,
                "title": title,
                "message": message,
                "context": context,
                "timestamp": datetime.now().isoformat()
            })
Step 5: Define Alert Rules
yaml
# alert-rules.yaml
rules:
  - name: agent_down
    condition: liveness_check == false
    for: 30s
    severity: P0
    message: "Agent {name} is unresponsive"

  - name: high_error_rate
    condition: error_rate > 0.20
    for: 5m
    severity: P1
    message: "Agent {name} error rate is {error_rate:.0%}"

  - name: latency_spike
    condition: p99_latency > 30s
    for: 3m
    severity: P1
    message: "Agent {name} p99 latency is {latency:.1f}s"

  - name: looping_detected
    condition: step_count > max_steps * 0.8
    for: 1m
    severity: P2
    message: "Agent {name} approaching step limit on {task_count} tasks"

  - name: budget_anomaly
    condition: token_usage > daily_budget * 0.5
    for: 1h
    severity: P2
    message: "Agent {name} used {usage} tokens in last hour (50% of daily budget)"
Step 6: Build the Dashboard

Essential dashboard panels for a multi-agent system:

PanelMetricDisplay
Agent GridLiveness per agentGreen/Red status cards
Latency HeatmapP50/P95/P99 per agentColor-coded time series
Error WaterfallError rate by agent + error typeStacked area chart
Token Burn RateTokens/min per agentLine chart with budget line
Active TasksTasks in-flight per agentGauge per agent
Top ErrorsMost frequent error messagesRanked list with count
Context Pressure% context window usedPer-agent gauge cluster
Alert TimelineAlerts over past 24hEvent timeline

Show full SKILL.md (268 more words)Show less

Trigger Phrases

PhraseAction
"Check agent health"Run liveness probes on all agents
"Show me the dashboard"Generate or link to monitoring dashboard
"Why is agent X slow?"Show latency breakdown for specific agent
"Any anomalies?"Run anomaly detection on recent metrics
"Set up alert for..."Create a new alert rule
"Agent X is down"Trigger incident response workflow
"Run a health check"Execute full liveness + readiness + deep check

Production Runbook

Incident: Agent Unresponsive
  1. Check liveness probe — is the process running?
  2. Check model endpoint — is the LLM provider healthy?
  3. Check context window — has the agent exceeded its limit?
  4. Restart agent with fresh context
  5. If recurring, set up circuit breaker
Incident: Error Rate Spike
  1. Identify error type — tool failure, model error, or parsing issue?
  2. Check recent deploys — did a prompt or tool change?
  3. Rollback if a recent change correlates
  4. Check rate limits — are external APIs throttling?
  5. Scale out if traffic increased
Incident: Token Budget Spike
  1. Identify which agent(s) are consuming
  2. Check for looping — excessive step counts
  3. Review recent tasks — unusually long inputs?
  4. Implement budget caps per task
  5. Alert the team if pattern persists

Anti-Patterns

Anti-PatternWhy It FailsFix
Monitoring only livenessAgent can be "alive" but uselessAdd readiness + deep checks
Same threshold for all agentsDifferent agents have different baselinesPer-agent dynamic thresholds
No alert deduplicationAlert fatigue leads to ignored alertsGroup by fingerprint, rate-limit
Fixing symptoms, not causesBand-aid solutions mask root issuesAlways capture root cause in alerts
No dashboardNo shared visibilityBuild and maintain a live dashboard

© cosmicstack-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in categories/ai-ml/agent-health-monitoring of cosmicstack-labs/mercury-agent-skills.

Open the folder on GitHubat commit 30392fb

Compare with similar skills

Agent Health Monitoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Health Monitoring compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Health Monitoring this skillcosmicstack-labs/mercury-agent-skills476—~2.7kAutomated safety check: PassMIT
Monitoring Observabilityyonatangross/orchestkit292—~2.2kAutomated safety check: PassMIT
Monitoringericrisco/rsc-harness180—~3.1kAutomated safety check: PassMIT
Kubernetes Network Root Cause Analysiskubeshark/kubeshark12k—~5.3kAutomated safety check: PassApache-2.0
Axiom Dashboard Builderopenclaw/clawhub9.5k—~4.9kAutomated safety check: PassMIT
Docs Corpus Auditmicrosoft/apm4k—~2.6kAutomated safety check: PassMIT

Similar skills

  • Monitoring Observability

    yonatangross/orchestkit

    Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (astype, scorecurrentspan, shouldexportspan, LangfuseMedia), and drift detection.

    292 GitHub stars~2.2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Monitoring

    ericrisco/rsc-harness

    A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and…

    180 GitHub stars~3.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Axiom Dashboard Builder

    openclaw/clawhub

    Designs and deploys Axiom dashboards through the API, choosing chart types and writing APL or metrics queries, with templates and migration notes for Splunk and Grafana.

    9.5k GitHub stars~4.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Docs Corpus Audit

    microsoft/apm

    Official

    A skill your agent uses to run a holistic regrounding pass on the entire microsoft/apm documentation corpus against current source code, page-by-page, and emit surgical fixes for stale claims.

    4k GitHub stars~2.6k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.

    24k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check: notes

More from cosmicstack-labs/mercury-agent-skills

All 12 skills in this repo
  • Before You Build

    cosmicstack-labs/mercury-agent-skills

    Use this before implementing a product, feature, SaaS, AI app, or side project to score product risk and choose the smallest validation step.

    476 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Hyperframes CLI

    cosmicstack-labs/mercury-agent-skills

    HyperFrames CLI dev loop — project scaffolding, validation (lint/inspect), browser preview with live reload, MP4/WebM rendering, and environment troubleshooting (doctor, browser, info, upgrade).

    476 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Hyperframes Media

    cosmicstack-labs/mercury-agent-skills

    Asset preprocessing for HyperFrames compositions — local text-to-speech narration (Kokoro-82M, no API key), audio/video transcription (Whisper), and background removal for transparent overlays…

    476 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Handoff Protocols

    cosmicstack-labs/mercury-agent-skills

    Design and implement agent-to-agent handoff protocols for multi-agent systems.

    476 GitHub stars~4.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Task Delegation

    cosmicstack-labs/mercury-agent-skills

    Design and operate task delegation systems for multi-agent fleets.

    476 GitHub stars~3.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Any2pdf

    cosmicstack-labs/mercury-agent-skills

    Convert Markdown to publication-quality PDF with reportlab — CJK/Latin mixed text, themes, cover pages, watermarks, callouts, formulas, and interactive theme selection

    476 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check: notes

Categories

Questions about Agent Health Monitoring

What does Agent Health Monitoring do?

Monitor AI agent health, detect anomalies, set up alerting, and maintain observability dashboards for production multi-agent systems. Agent Health Monitoring is an agent skill from cosmicstack-labs/mercury-agent-skills. Monitor AI agent health, detect anomalies, set up alerting, and maintain observability dashboards for production multi-agent systems.

When should I use Agent Health Monitoring?

Agent Health Monitoring fits situations like: tasks that involve Monitoring and alerting; tasks that involve Anomaly detection; tasks that involve GitOps.

How do I install Agent Health Monitoring in Claude Code?

Run `npx skills add cosmicstack-labs/mercury-agent-skills --skill agent-health-monitoring -a claude-code`. Or copy the skill folder (categories/ai-ml/agent-health-monitoring in cosmicstack-labs/mercury-agent-skills) into .claude/skills/agent-health-monitoring in your project. Claude Code loads it when a task matches its description.

How do I install Agent Health Monitoring in Codex?

Run `npx skills add cosmicstack-labs/mercury-agent-skills --skill agent-health-monitoring -a codex`. Or copy the skill folder (categories/ai-ml/agent-health-monitoring in cosmicstack-labs/mercury-agent-skills) into .agents/skills/agent-health-monitoring in your project. Codex loads it when a task matches its description.

Can I use Agent Health Monitoring in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cosmicstack-labs/mercury-agent-skills --skill agent-health-monitoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-health-monitoring, .gemini/skills/agent-health-monitoring, .github/skills/agent-health-monitoring and .opencode/skills/agent-health-monitoring in your project.

What does Agent Health Monitoring need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Health Monitoring is instructions for the agent only. Our summary lists: Python 3.

Does Agent Health Monitoring access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Health Monitoring safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Health Monitoring use?

Agent Health Monitoring is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Health Monitoring use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Health Monitoring?

Skills that share tags, products or a category with Agent Health Monitoring: Monitoring Observability (yonatangross/orchestkit, 292 stars), Monitoring (ericrisco/rsc-harness, 180 stars), Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars) and Axiom Dashboard Builder (openclaw/clawhub, 9.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Health Monitoring?

cosmicstack-labs (a GitHub organization) maintains it in cosmicstack-labs/mercury-agent-skills, which has 476 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on August 25, 2026.

Source: cosmicstack-labs/mercury-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.