Agent skill

Cloud Monitoring

by seb1n in seb1n/awesome-ai-agent-skills

Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability.

MITAuto-check passedDevOps & Cloud

Install Cloud Monitoring

skills CLI
$ npx skills add seb1n/awesome-ai-agent-skills --skill cloud-monitoring -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install seb1n/awesome-ai-agent-skills cloud-monitoring --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/seb1n/awesome-ai-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/devops-and-infrastructure/cloud-monitoring .claude/skills/cloud-monitoring && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cloud-monitoring
GitHub stars
206
Token cost
~2.8k tokens
SKILL.md length
879 words
Files
1
Skills in repo
101
Repo updated
First seen
Licence
MIT

At a glance

Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability.

  • Works in 6 steps: Identify Monitoring Objectives: The… → Select Monitoring Tools and… → Configure Metrics Collection and… → …
  • The user requests cloud monitoring
  • SKILL.md covers Workflow, Supported Technologies, Usage and Examples, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cloud Monitoring is an agent skill from seb1n/awesome-ai-agent-skills. Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability. Use when the user requests cloud monitoring or provides relevant inputs for this workflow.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Monitoring and alerting, Observability and Site reliability engineering. It works with Prometheus and Grafana. The repository describes itself as: 103 ready-to-use AI agent skills for Claude Code, OpenAI Codex, Gemini CLI, Cursor, GitHub Copilot, Windsurf, and other Agent Skills-compatible tools. Complete SKILL.md… The licence is MIT.

When your agent uses it

  • The user requests cloud monitoring
  • Provides relevant inputs for this workflow

Example prompts

  • “/cloud-monitoring”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Identify Monitoring Objectives: The agent works with the user to define what needs to be monitored and why. This includes identifying…
  2. Select Monitoring Tools and Instrumentation: Based on the cloud provider and application architecture, the agent recommends an appropriate…
  3. Configure Metrics Collection and Dashboards: The agent defines and deploys metric scrapers, exporters, and custom metrics. It builds…
  4. Establish Alerting Rules: The agent configures alerts that trigger on meaningful conditions — such as error budget burn rate exceeding…
  5. Set Up Log Aggregation and Trace Correlation: The agent configures centralized log collection with structured logging formats (JSON), log…
  6. Review and Iterate: The agent periodically audits alert noise levels, dashboard usage, and SLO compliance. Unused alerts are pruned…

What it can do on your machine

Read from SKILL.md and the folder at commit 75865a5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cloud Monitoring loads about 2.8k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 879 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from seb1n/awesome-ai-agent-skills at commit 75865a5, republished under its MIT licence (© seb1n). 879 words, ~2,763 tokens.

Download SKILL.mdSave it as .claude/skills/cloud-monitoring/SKILL.md (or your agent's skills folder).
name
cloud-monitoring
description
Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability. Use when the user requests cloud monitoring or provides relevant inputs for this workflow.
license
MIT
metadata.author
awesome-ai-agent-skills
metadata.version
1.0.0

Cloud Monitoring

This skill enables the agent to design and configure comprehensive monitoring and observability solutions for cloud infrastructure and applications. The agent understands the three pillars of observability — metrics, logs, and traces — and can set up dashboards, alerting rules, SLIs, SLOs, and SLAs using tools like Prometheus, Grafana, CloudWatch, Datadog, and OpenTelemetry. The agent also applies alerting best practices to minimize alert fatigue while ensuring critical issues are surfaced promptly.

Workflow

  1. Identify Monitoring Objectives: The agent works with the user to define what needs to be monitored and why. This includes identifying critical services, establishing Service Level Indicators (SLIs) such as request latency, error rate, and throughput, and setting Service Level Objectives (SLOs) that define acceptable performance thresholds. SLAs (Service Level Agreements) are documented as contractual commitments to customers.

  2. Select Monitoring Tools and Instrumentation: Based on the cloud provider and application architecture, the agent recommends an appropriate monitoring stack. This may include Prometheus for metrics collection, Grafana for visualization, Loki or CloudWatch Logs for log aggregation, and Jaeger or AWS X-Ray for distributed tracing. The agent configures OpenTelemetry SDKs in application code to emit standardized telemetry data.

  3. Configure Metrics Collection and Dashboards: The agent defines and deploys metric scrapers, exporters, and custom metrics. It builds dashboards that visualize the golden signals (latency, traffic, errors, saturation) and infrastructure metrics (CPU, memory, disk, network). Dashboards are organized by service tier so teams can quickly triage issues.

  4. Establish Alerting Rules: The agent configures alerts that trigger on meaningful conditions — such as error budget burn rate exceeding thresholds, sustained latency spikes, or pod restarts — rather than raw metric thresholds alone. Multi-window, multi-burn-rate alerting is used to balance detection speed with false-positive suppression. Alert routing is configured to send critical alerts to PagerDuty or Opsgenie and warnings to Slack.

  5. Set Up Log Aggregation and Trace Correlation: The agent configures centralized log collection with structured logging formats (JSON), log retention policies, and log-based alerts for error patterns. Distributed traces are correlated with logs and metrics using shared trace IDs so that a single alert can link directly to the relevant request trace and log entries.

  6. Review and Iterate: The agent periodically audits alert noise levels, dashboard usage, and SLO compliance. Unused alerts are pruned, thresholds are adjusted based on observed baselines, and new services are onboarded into the monitoring stack as the system evolves.

Supported Technologies

  • Metrics: Prometheus, AWS CloudWatch, Google Cloud Monitoring, Azure Monitor, Datadog, New Relic
  • Visualization: Grafana, CloudWatch Dashboards, Datadog Dashboards, Kibana
  • Logs: Loki, CloudWatch Logs, Elasticsearch/Fluentd/Kibana (EFK), Splunk
  • Traces: Jaeger, Zipkin, AWS X-Ray, Tempo, Datadog APM
  • Instrumentation: OpenTelemetry, Prometheus client libraries, StatsD
  • Alerting: Alertmanager, PagerDuty, Opsgenie, Slack Webhooks, SNS

Usage

Provide the agent with your cloud provider, the services to monitor, your preferred monitoring stack, and any existing SLOs or alerting requirements.

Example prompt:

Set up monitoring for our Kubernetes microservices on AWS.
- Use Prometheus and Grafana for metrics and dashboards
- Monitor API latency (p99 < 500ms) and error rate (< 1%)
- Send critical alerts to PagerDuty, warnings to Slack
- Aggregate logs with CloudWatch Logs

Examples

Example 1: Prometheus + Grafana with Alerting Rules

prometheus.yml — Prometheus scrape configuration:

yaml
global:
  scrape_interval: 15s
  evaluation_interval: 15s

rule_files:
  - "alert_rules.yml"

alerting:
  alertmanagers:
    - static_configs:
        - targets: ["alertmanager:9093"]

scrape_configs:
  - job_name: "node-exporter"
    static_configs:
      - targets: ["node-exporter:9100"]

  - job_name: "app"
    metrics_path: /metrics
    static_configs:
      - targets: ["app:8080"]

  - job_name: "kubernetes-pods"
    kubernetes_sd_configs:
      - role: pod
    relabel_configs:
      - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
        action: keep
        regex: true
      - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_port]
        action: replace
        target_label: __address__
        regex: (.+)
        replacement: $1

alert_rules.yml — SLO-based alerting rules:

yaml
groups:
  - name: slo-alerts
    rules:
      - alert: HighErrorRate
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[5m]))
          /
          sum(rate(http_requests_total[5m])) > 0.01
        for: 5m
        labels:
          severity: critical
        annotations:
          summary: "Error rate exceeds 1% SLO"
          description: "{{ $labels.job }} error rate is {{ $value | humanizePercentage }}"

      - alert: HighP99Latency
        expr: |
          histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le))
          > 0.5
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: "P99 latency exceeds 500ms SLO"

      - alert: PodCrashLooping
        expr: increase(kube_pod_container_status_restarts_total[1h]) > 3
        for: 5m
        labels:
          severity: critical
        annotations:
          summary: "Pod {{ $labels.pod }} is crash looping"

      - alert: HighMemoryUsage
        expr: (node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes) / node_memory_MemTotal_bytes > 0.9
        for: 15m
        labels:
          severity: warning
        annotations:
          summary: "Node memory usage above 90%"
Example 2: AWS CloudWatch Dashboard with Custom Metrics and Alarms

cloudwatch-dashboard.json — CloudFormation template for a monitoring stack:

json
{
  "AWSTemplateFormatVersion": "2010-09-09",
  "Resources": {
    "ApiDashboard": {
      "Type": "AWS::CloudWatch::Dashboard",
      "Properties": {
        "DashboardName": "api-service-dashboard",
        "DashboardBody": "{\"widgets\":[{\"type\":\"metric\",\"properties\":{\"metrics\":[[\"AWS/ApplicationELB\",\"TargetResponseTime\",\"TargetGroup\",\"my-tg\",{\"stat\":\"p99\"}],[\"AWS/ApplicationELB\",\"HTTPCode_Target_5XX_Count\",\"TargetGroup\",\"my-tg\"]],\"period\":300,\"title\":\"API Latency & Errors\"}},{\"type\":\"metric\",\"properties\":{\"metrics\":[[\"Custom/App\",\"ActiveConnections\"],[\"Custom/App\",\"QueueDepth\"]],\"period\":60,\"title\":\"Application Metrics\"}}]}"
      }
    },
    "HighLatencyAlarm": {
      "Type": "AWS::CloudWatch::Alarm",
      "Properties": {
        "AlarmName": "api-high-latency",
        "MetricName": "TargetResponseTime",
        "Namespace": "AWS/ApplicationELB",
        "Statistic": "p99",
        "Period": 300,
        "EvaluationPeriods": 3,
        "Threshold": 0.5,
        "ComparisonOperator": "GreaterThanThreshold",
        "AlarmActions": ["arn:aws:sns:us-east-1:123456789012:ops-alerts"],
        "Dimensions": [
          {"Name": "TargetGroup", "Value": "my-tg"}
        ]
      }
    },
    "HighErrorRateAlarm": {
      "Type": "AWS::CloudWatch::Alarm",
      "Properties": {
        "AlarmName": "api-high-error-rate",
        "MetricName": "HTTPCode_Target_5XX_Count",
        "Namespace": "AWS/ApplicationELB",
        "Statistic": "Sum",
        "Period": 300,
        "EvaluationPeriods": 2,
        "Threshold": 50,
        "ComparisonOperator": "GreaterThanThreshold",
        "AlarmActions": ["arn:aws:sns:us-east-1:123456789012:ops-alerts"]
      }
    }
  }
}
Show full SKILL.md (377 more words)Show less

Best Practices

  • Alert on symptoms, not causes: Alert on user-facing impact (high error rate, slow responses) rather than low-level infrastructure metrics (CPU at 80%). High CPU is only a problem if it degrades user experience.
  • Use multi-window burn rates: Instead of alerting on a single threshold, use SLO burn-rate alerts with fast (5m) and slow (1h) windows. This detects real incidents quickly while ignoring brief transient spikes.
  • Minimize alert fatigue: Every alert should be actionable. If an alert fires and the on-call engineer has no clear action to take, the alert should be removed or converted to a dashboard widget. Aim for fewer than five pages per on-call shift.
  • Implement structured logging: Use JSON-formatted logs with consistent fields (timestamp, level, service, trace_id, message) across all services. This enables efficient log querying and correlation with traces.
  • Set retention policies: Configure tiered retention — high-resolution metrics for 15 days, downsampled metrics for 1 year, logs for 30-90 days depending on compliance requirements. This controls storage costs while maintaining historical visibility.
  • Tag and label everything: Apply consistent labels (service, environment, team, version) to all metrics, logs, and traces so they can be filtered, grouped, and correlated across the observability stack.

Edge Cases

  • Metric cardinality explosion: Custom metrics with high-cardinality labels (e.g., user IDs, request URLs) can overwhelm Prometheus and cause out-of-memory crashes. Limit label values to bounded sets and use recording rules to pre-aggregate high-cardinality queries.
  • Clock skew in distributed traces: If service clocks are out of sync, trace spans may appear out of order or have negative durations. Ensure all hosts use NTP synchronization and tolerate small timing inconsistencies in trace visualization.
  • Alert storms during outages: A single infrastructure failure can trigger dozens of correlated alerts simultaneously. Configure alert grouping and inhibition rules in Alertmanager so that a parent alert (e.g., "node down") suppresses child alerts (e.g., "pod unhealthy on that node").
  • Missing metrics during deployments: Rolling deployments cause pods to restart, creating gaps in time-series data. Use rate() functions that tolerate missing scrapes and configure absent-metric alerts with appropriate for durations to avoid false positives during rollouts.
  • CloudWatch API throttling: Querying too many custom metrics or dashboards can hit CloudWatch API rate limits. Batch metric retrieval using GetMetricData instead of GetMetricStatistics and cache dashboard data on the client side.

© seb1n, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in devops-and-infrastructure/cloud-monitoring of seb1n/awesome-ai-agent-skills.

Open the folder on GitHubat commit 75865a5

Compare with similar skills

Cloud Monitoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cloud Monitoring compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cloud Monitoring this skillseb1n/awesome-ai-agent-skills206—~2.8kAutomated safety check: PassMIT
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Observability MonitoringAnastasiyaW/codex-claude-code-config154—~4.1kAutomated safety check: PassMIT
Observability Sremajiayu000/spellbook287—~3.3kAutomated safety check: PassMIT
Observability Patternssoftspark/ai-toolkit179—~2.2kAutomated safety check: PassApache-2.0
Telemetrymagnus919/agent-skills115—~3.9kAutomated safety check: PassMIT

Similar skills

  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 6 mo ago
    DevOps & CloudAuto-check passed
  • Observability Monitoring

    AnastasiyaW/codex-claude-code-config

    Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…

    154 GitHub stars~4.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Observability Sre

    majiayu000/spellbook

    Observability and SRE expert. An agent skill from majiayu000/spellbook.

    287 GitHub stars~3.3k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Observability Patterns

    softspark/ai-toolkit

    Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI.

    179 GitHub stars~2.2k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • Telemetry

    magnus919/agent-skills

    Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector…

    115 GitHub stars~3.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.

    24k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check: notes

More from seb1n/awesome-ai-agent-skills

All 101 skills in this repo
  • Agent Red Teaming

    seb1n/awesome-ai-agent-skills

    Plan, execute, document, and retest authorized security assessments of AI agents and multi-agent workflows using safe adversarial cases, synthetic identities, canaries, and evidence-based findings.

    206 GitHub stars~2.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Eu AI Act Readiness

    seb1n/awesome-ai-agent-skills

    Build a preliminary, evidence-based EU AI Act readiness assessment across AI-system inventory, territorial scope, operator roles, prohibited-practice screening, risk classification, transparency…

    206 GitHub stars~3.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Human In The Loop

    seb1n/awesome-ai-agent-skills

    Design and verify auditable human oversight, approval gates, escalation paths, and safe state transitions for AI agent workflows.

    206 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • MCP Server Building

    seb1n/awesome-ai-agent-skills

    Design, implement, harden, and verify Model Context Protocol (MCP) servers with precise tool contracts, least-privilege authorization, safe transports, structured errors, and interoperability tests.

    206 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • PDF Processing

    seb1n/awesome-ai-agent-skills

    Inspect, extract, OCR, create, merge, split, reorder, rotate, annotate, fill, redact, compress, secure, and verify PDF documents while preserving source files and visual fidelity.

    206 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Skill Supply Chain Audit

    seb1n/awesome-ai-agent-skills

    Audit agent skills, plugins, prompts, manifests, scripts, dependencies, and bundled assets for provenance, prompt-injection, permission, execution, exfiltration, persistence, and update risk.

    206 GitHub stars~2.4k tokensUpdated 2 mo ago
    Auto-check passed

Categories

Questions about Cloud Monitoring

What does Cloud Monitoring do?

Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability. Cloud Monitoring is an agent skill from seb1n/awesome-ai-agent-skills. Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability.

When should I use Cloud Monitoring?

Cloud Monitoring fits situations like: the user requests cloud monitoring; provides relevant inputs for this workflow.

How do I install Cloud Monitoring in Claude Code?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill cloud-monitoring -a claude-code`. Or copy the skill folder (devops-and-infrastructure/cloud-monitoring in seb1n/awesome-ai-agent-skills) into .claude/skills/cloud-monitoring in your project. Claude Code loads it when a task matches its description.

How do I install Cloud Monitoring in Codex?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill cloud-monitoring -a codex`. Or copy the skill folder (devops-and-infrastructure/cloud-monitoring in seb1n/awesome-ai-agent-skills) into .agents/skills/cloud-monitoring in your project. Codex loads it when a task matches its description.

Can I use Cloud Monitoring in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add seb1n/awesome-ai-agent-skills --skill cloud-monitoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cloud-monitoring, .gemini/skills/cloud-monitoring, .github/skills/cloud-monitoring and .opencode/skills/cloud-monitoring in your project.

What does Cloud Monitoring need to run?

SKILL.md names no scripts, command-line tools or credentials: Cloud Monitoring is instructions for the agent only.

Does Cloud Monitoring access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cloud Monitoring safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cloud Monitoring use?

Cloud Monitoring is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cloud Monitoring use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cloud Monitoring?

Skills that share tags, products or a category with Cloud Monitoring: Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Observability Monitoring (AnastasiyaW/codex-claude-code-config, 154 stars), Observability Sre (majiayu000/spellbook, 287 stars) and Observability Patterns (softspark/ai-toolkit, 179 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cloud Monitoring?

seb1n (a GitHub user) maintains it in seb1n/awesome-ai-agent-skills, which has 206 GitHub stars. The repository holds 101 skills in this directory. The repository was last updated on August 9, 2026.

Source: seb1n/awesome-ai-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.