Agent skill

Observability Monitoring

by AnastasiyaW in AnastasiyaW/codex-claude-code-config

Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…

MITAuto-check passedDevOps & Cloud

Install Observability Monitoring

skills CLI
$ npx skills add AnastasiyaW/codex-claude-code-config --skill observability-monitoring -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AnastasiyaW/codex-claude-code-config observability-monitoring --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AnastasiyaW/codex-claude-code-config.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/operational/observability-monitoring .claude/skills/observability-monitoring && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
observability-monitoring
GitHub stars
154
Token cost
~4.1k tokens
SKILL.md length
2,142 words
Files
3 (incl. references)
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…

  • Works in 9 steps: Freeze scope and collect live facts → Define the outcome before the metric → Map the monitored layers → …
  • Asked about monitoring
  • SKILL.md covers Operating rule, Workflow, Vendor mapping and Minimal runbook template, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Observability Monitoring is an agent skill from AnastasiyaW/codex-claude-code-config. Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting, burn-rate response, and postmortems. Use when asked about monitoring, наблюдаемость, алерты, Prometheus, Grafana, OpenTelemetry, logs, traces, profiles, service health, or incident evidence. Do not use for generic dashboard styling, frontend-only UI work, or unrelated code review.

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `agents/openai.yaml` and `references/source-notes.md`).

It sits in DevOps & Cloud, covering Site reliability engineering, Observability and Monitoring and alerting. It works with OpenTelemetry, Prometheus and Grafana. The repository describes itself as: Claude Code, Codex, and multi-agent configuration system: principles, hooks, skills, and workflow patterns for AI-assisted development. The licence is MIT.

When your agent uses it

  • Asked about monitoring
  • Incident evidence
  • Generic dashboard styling
  • Frontend-only UI work

Example prompts

  • “/observability-monitoring”

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Freeze scope and collect live facts
  2. Define the outcome before the metric
  3. Map the monitored layers
  4. Choose the right signal
  5. Design metrics without cardinality accidents
  6. Define SLI, SLO, SLA, and error budget
  7. Build actionable alerts
  8. Triage in evidence order
  9. Close the loop

What it can do on your machine

Read from SKILL.md and the folder at commit 67709af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are promql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Observability Monitoring loads about 4.1k tokens when it runs, and up to ~5.1k if it reads all its reference files. Until then it costs about 126 tokens; SKILL.md has 2,142 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~126
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AnastasiyaW/codex-claude-code-config at commit 67709af, republished under its MIT licence (© AnastasiyaW). 2,142 words, ~4,106 tokens.

Download SKILL.mdSave it as .claude/skills/observability-monitoring/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
observability-monitoring
description
Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting, burn-rate response, and postmortems. Use when asked about monitoring, наблюдаемость, алерты, Prometheus, Grafana, OpenTelemetry, logs, traces, profiles, service health, or incident evidence. Do not use for generic dashboard styling, frontend-only UI work, or unrelated code review.

Observability Monitoring

Use this skill to turn vague "is it working?" questions into evidence-backed monitoring, alerting, and incident workflows. Start from user or business impact, then move down through the system layers and choose the signal that can prove the current hypothesis.

Operating rule

Do not treat a green dashboard as proof of health. A monitoring claim is complete only when it names:

  1. the observed scope and time window;
  2. the user, business, or operator outcome being protected;
  3. the signal and exact query/probe that supports the claim;
  4. the threshold or SLO that defines bad;
  5. the next human action and its runbook/evidence link.

Keep code review, live runtime proof, UI/render proof, and release readiness as separate verdicts.

Investigation is read-only by default. A restart, alert suppression, metric-schema/label change, sampling or retention change, or vendor reconfiguration is a production mutation: involve the responsible owner or incident authority, preserve the relevant evidence first, capture the exact config/command diff, state the rollback condition, and verify the user probe plus SLI after the change.

Workflow

1. Freeze scope and collect live facts

Before changing a monitor, alert, host, or service:

  • Read the repository AGENTS.md, relevant rules, runbooks, and deployment docs.
  • Establish the actual checkout/branch, deployment, process, host, port, proxy/tunnel, and data source.
  • Record the observation time, environment, query/probe, and whether the result is current or historical.
  • Inspect the source code/config that emits or consumes the signal before changing it.
  • Never infer health from an old screenshot, a stale handoff, a single PID, or a dashboard with no successful user probe.

For infrastructure fixes, document the traffic path (DNS, proxy, tunnel, ingress, service, backend) before touching a surprising value such as 127.0.0.1, a non-default port, or a disabled check.

2. Define the outcome before the metric

Write the protected outcome in concrete terms:

  • User: can open the page, submit the request, download the result, or complete the workflow.
  • Business: orders, successful jobs, conversions, revenue, or another domain event continue to occur.
  • Operator: an on-call engineer receives a page early enough to act and can identify the next step.

Add at least one black-box or synthetic check for the user-facing path. Add real-user telemetry when the experience can vary by browser, geography, device, or network. Technical resource health is not a substitute for a user or business signal.

3. Map the monitored layers

Inspect the layers from bottom to top and state which ones are in scope:

  1. Hardware/infrastructure: CPU, memory, disk capacity and latency, network, temperature, power, GPU, host availability.
  2. Host/OS: load, processes, file descriptors, swap, service state, kernel/resource pressure.
  3. Network: reachability, packet loss, latency, port state, DNS, TLS, ingress/proxy health.
  4. Application/APM: request rate, errors, latency, saturation, dependency calls, queue depth.
  5. Databases and queues: throughput, slow queries, connection pools, replication lag, backlog, consumer health.
  6. Containers/orchestration: desired versus ready replicas, restarts, scheduling, evictions, resource limits, ephemeral identity.
  7. Business: successful workflows, jobs, orders, revenue, conversion, or domain-specific zero-activity checks.

Do not stop at the infrastructure layer when the business outcome is failing. A system with green CPU and memory can still have a broken payment, form, queue consumer, or model job.

4. Choose the right signal

Use the signal that answers the question instead of collecting everything indiscriminately:

QuestionPrimary signalPractical method
Is a resource busy, queued, or failing?MetricsUSE: utilization, saturation, errors
Is a service serving users correctly?MetricsRED: rate, errors, duration
What happened at a specific time?Structured logsSearch by timestamp, service, severity, request/trace ID
Where did a distributed request slow or fail?TracesFollow the trace across services and spans
Why is code or a process consuming resources?ProfilesInspect sampled stacks, CPU, memory, locks, or I/O
Do users actually succeed?Synthetic/RUM/business probesRun the workflow and inspect the domain result

Metrics answer aggregate how much/how often. Logs answer what happened. Traces answer where in the path. Profiles answer which code or process consumed the resource. Correlate them with stable resource identity, timestamps, and trace/request context.

5. Design metrics without cardinality accidents

Treat every unique metric-name plus label set as a time series. Before adding a label, estimate its possible values and lifetime.

  • Good metric dimensions are bounded and useful for aggregation: method, status class/code, service, region, queue, job type, or environment.
  • Keep high-cardinality or per-event values in logs/traces: user ID, request ID, email, raw URL with parameters, exception text, or arbitrary object IDs.
  • Prefer route templates over raw URLs.
  • Set explicit cardinality limits or views where the telemetry SDK supports them, and monitor overflow/dropped-attribute indicators.
  • Prefer histograms for fleet-wide aggregatable latency SLOs. Client-side summaries/quantiles generally cannot be aggregated across instances without changing their meaning; use them only when that limitation is understood. Never hide tail latency behind an average.

If a query needs a unique ID to be useful, that ID belongs in a trace or structured log field, not in a metric label.

6. Define SLI, SLO, SLA, and error budget

Use the terms precisely:

  • SLI: a measured indicator, such as successful requests divided by valid requests or p95 latency for a workflow.
  • SLO: the internal target and time window for that indicator.
  • SLA: an external agreement with consequences if the target is missed.
  • Error budget: the allowed unreliability implied by the SLO during the window.

Choose the SLI from the user-facing outcome, not from whichever metric is easiest to collect. Define the valid-request population, exclusions, time window, aggregation, and owner. Treat an SLO as a control loop: remaining budget informs release pace, risk, testing, and reliability work.

7. Build actionable alerts

Alert on symptoms that indicate user or service pain. Use dashboards and drill-downs to investigate causes.

Every paging alert must include:

  • symptom and affected scope;
  • exact query, threshold, duration, and evaluation window;
  • severity and owner;
  • a link to the dashboard/query and runbook;
  • concrete first actions, rollback criteria, and escalation path;
  • a test or smoke procedure proving the alert route works.

If no immediate action exists, make it a dashboard annotation, ticket, or recording rule instead of a page. Use a pending duration (for) to suppress short blips. If using hysteresis, define separate fire and clear thresholds; if using multi-window evaluation, define each window and its condition. Do not confuse either control with Slack/chat routing. Treat "more than two incidents per shift" as a review heuristic, not a universal law: frequent pages mean the system or alert policy needs repair.

Prefer multi-window or burn-rate alerts for SLOs when the backend supports them. A burn rate near 1 spends the budget at the planned rate; a materially higher rate requires faster response. Verify the math against the actual SLO window and alert implementation.

Portable examples

Adapt metric and label names to the live schema; these are patterns, not copy-paste production queries:

promql
# RED error ratio for one service over five minutes.
sum(rate(http_requests_total{service="checkout",code=~"5.."}[5m]))
/
sum(rate(http_requests_total{service="checkout"}[5m]))

# Aggregatable fleet-wide p95 from a histogram.
histogram_quantile(0.95,
  sum by (le, service) (
    rate(http_request_duration_seconds_bucket{service="checkout"}[5m])
  )
)

Use a synthetic/domain probe that asserts the real outcome, not only HTTP reachability. Prefer a read-only endpoint or a sandbox/test tenant. If a write path is unavoidable, require a dry-run or idempotency key, a bounded fixture, explicit cleanup, and an owner-approved change window: submit safe test request -> assert expected status and domain result/job ID -> record latency and trace ID -> clean up. Never send an arbitrary representative request into production.

A burn rate is observed error ratio / allowed error ratio, where allowed error ratio = 1 - SLO. Worked example: for an SLO of 99.9% (allowed = 0.001), a 5-minute observed error ratio of 1.0% (0.010) gives burn rate 10; if the local paging policy is the illustrative >=10 for 5m AND >=5 for 1h, page, otherwise route it according to the lower-severity policy. The query pattern is:

promql
(
  sum(rate(http_request_errors_total{service="checkout"}[5m]))
  /
  sum(rate(http_requests_total{service="checkout"}[5m]))
) / 0.001

The thresholds and windows are examples only; derive and test them from the real SLO, traffic population, exclusions, and paging budget.

Show full SKILL.md (865 more words)Show less
8. Triage in evidence order

For a live incident, follow this sequence and keep a short timeline. Apply the change-control boundary above before any mutation:

  1. Confirm the user/business symptom with a fresh probe.
  2. Bound blast radius: service, region, tenant, job type, version, and start time.
  3. Check RED for the affected service and USE for saturated dependencies/resources.
  4. Correlate logs by time, service, severity, and trace/request ID.
  5. Follow a failing or slow trace to the first bad span/dependency.
  6. Use a profile only when resource/code attribution remains unclear or the issue is intermittent.
  7. Compare deploys, config changes, feature flags, capacity, and external dependencies.
  8. Apply the smallest reversible mitigation with an explicit rollback condition.
  9. Re-run the user probe and the relevant SLI query; confirm recovery over a meaningful window.
  10. Record the evidence, unresolved uncertainty, and follow-up owner.

Do not restart or reconfigure a service merely because a check is red. First establish whether the check is stale, misrouted, a false positive, or a real symptom, and preserve the evidence needed to explain the decision.

9. Close the loop

After recovery:

  • Verify the original symptom is gone, not only that the process is alive.
  • Record detection time, acknowledgement, mitigation, recovery, and customer impact.
  • Review alert quality: actionable, correctly routed, deduplicated, and linked to evidence.
  • Write a blameless postmortem for material incidents: timeline, impact, contributing/systemic causes, what went well, what failed, and tracked prevention items.
  • Update the runbook, monitor, test, or architecture so the same failure is easier to detect and diagnose next time.

Vendor mapping

Use the actual stack discovered in the repository/runtime. Common roles from the source video are:

  • Prometheus or VictoriaMetrics: scrape and store metrics;
  • Grafana: visualize and query multiple sources;
  • Loki or Elasticsearch/OpenSearch: logs;
  • Jaeger or Tempo: traces;
  • OpenTelemetry: vendor-neutral collection and export of telemetry signals;
  • Mimir or Thanos: longer-term or larger-scale metric storage;
  • Zabbix/Nagios: host and infrastructure checks.

Historical names in the video (ping, syslog, SNMP, MRTG/RRD, Nagios, and Cacti; 00:55-03:17) explain the evolution of monitoring and are not default deployment recommendations.

These are roles, not defaults. Do not install, replace, or reconfigure a vendor component without checking the live architecture, documentation, compatibility, retention, cost, authorization, and rollback path.

Minimal runbook template

Use this shape for every new page:

text
Name:
User symptom:
SLI/query:
Trigger and duration:
Scope labels:
Severity/owner:
Dashboard and raw query:
First safe action:
Rollback condition:
Verification probe:
Escalation:
Known false positives:
Last tested:

Gotchas

  • A host can be healthy while the business workflow is broken.
  • Averages hide tail latency; inspect percentiles and failed requests.
  • A metric label creates a new series for every unique combination; raw URLs and IDs can exhaust memory.
  • A trace without stable service/resource context is hard to correlate; a log without timestamps or IDs is weak evidence.
  • A page without a concrete action trains the operator to ignore pages.
  • A synthetics-only check misses real-user variance; RUM-only telemetry can discover pain too late.
  • Profiles are an attribution signal, not a replacement for service-level metrics, logs, or traces.
  • An SLO target is not an SLA promise unless an external contract says so.
  • Auto-generated captions and copied vendor names can be wrong; verify terms before putting them into code or runbooks. See references/source-notes.md.

Troubleshooting

SymptomLikely causeFix
Dashboard is green but users failMissing black-box/business SLI, wrong route, or stale data sourceRun a fresh user probe, verify timestamps/labels, add the missing outcome signal
One deploy creates a huge series spikeHigh-cardinality label or unbounded route/IDRemove the label from metrics, normalize routes, move detail to logs/traces, cap SDK cardinality
Alert fires but nobody actsAlert is a cause guess, too sensitive, unowned, or lacks a runbookAlert on a user symptom, add owner/action/query/runbook, add for; if using hysteresis define separate fire/clear thresholds, and if using multi-window define each condition; test delivery
Error rate rises but cause is unclearLogs are unstructured or traces are not correlatedAdd structured fields and trace IDs, sample a failing trace, inspect dependency spans
Service is slow but RED looks normalLow traffic, bad aggregation, or resource contention outside the appCheck synthetic latency, USE for host/queue/storage, tail percentiles, and profiles
Dashboard is blank or staleTarget/scrape/exporter/collector/backend failure, retention gap, or clock skewCheck last-sample time, target and collector health, dropped/ingestion counters, retention, and time synchronization; do not interpret blank as zero
A query shows zero but the system may be uninstrumentedThe series is absent, filtered out, or genuinely zeroTest absent-series semantics, inspect raw labels and last sample, and add an explicit freshness/telemetry-health check
Trace is incompleteSampling, context propagation, collector drops, or backend retentionCheck trace ID propagation, sampling decision, collector drop counters, backend ingestion, and clock synchronization
Page storm during one incidentNo grouping/deduplication or too many low-value alertsGroup by incident/service, keep one page for the symptom, demote diagnostic signals
Burn-rate alert disagrees with the dashboardDifferent SLI population, window, exclusions, or recording ruleReconcile the numerator/denominator, window, and query; test against known scenarios
Restart appears to fix it but it returnsMitigation hid a systemic cause or destroyed evidenceCapture logs/metrics first, record the restart, identify the recurring trigger, add prevention

Source and current-practice notes

The reusable concepts in this skill were extracted from the supplied video and cross-checked against current primary documentation. See references/source-notes.md for the source URL, timestamp map, caption corrections, and official references.

© AnastasiyaW, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/operational/observability-monitoring of AnastasiyaW/codex-claude-code-config.

  • SKILL.md
  • agents/openai.yaml
  • references/source-notes.md

Open the folder on GitHubat commit 67709af

Compare with similar skills

Observability Monitoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Observability Monitoring compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Observability Monitoring this skillAnastasiyaW/codex-claude-code-config154—~4.1kAutomated safety check: PassMIT
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Observability Patternssoftspark/ai-toolkit179—~2.2kAutomated safety check: PassApache-2.0
Telemetrymagnus919/agent-skills115—~3.9kAutomated safety check: PassMIT
Observability Sremajiayu000/spellbook287—~3.3kAutomated safety check: PassMIT
Archestra Dev Observabilityarchestra-ai/archestra4.4k—~1.2kAutomated safety check: PassCustom licence

Similar skills

  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 6 mo ago
    DevOps & CloudAuto-check passed
  • Observability Patterns

    softspark/ai-toolkit

    Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI.

    179 GitHub stars~2.2k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • Telemetry

    magnus919/agent-skills

    Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector…

    115 GitHub stars~3.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Observability Sre

    majiayu000/spellbook

    Observability and SRE expert. An agent skill from majiayu000/spellbook.

    287 GitHub stars~3.3k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • Archestra Dev Observability

    archestra-ai/archestra

    A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup.

    4.4k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Frontmcp Observability

    agentfront/frontmcp

    A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server.

    146 GitHub stars~4.6k tokensUpdated today
    DevOps & CloudAuto-check passed

More from AnastasiyaW/codex-claude-code-config

All 50 skills in this repo
  • Bug Reproducer

    AnastasiyaW/codex-claude-code-config

    Find likely software bugs in a codebase, rank concrete bug candidates, and prove or reject them with focused regression tests before proposing a fix.

    154 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Motion Framer

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when implementing Motion or Framer Motion in React/JavaScript: interactive UI components, micro-interactions, gestures, layout or page transitions, and scroll-based animation.

    154 GitHub starsUsed in 1 repo~5.2k tokens
    Auto-check passed
  • Proof Verify

    AnastasiyaW/codex-claude-code-config

    Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work).

    154 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Workflow Orchestration

    AnastasiyaW/codex-claude-code-config

    Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).

    154 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Notebooklm Grounded Research

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when: NotebookLM, notebooklm MCP, large documentation sets, courses, books, papers, or citation-backed research are mentioned.

    154 GitHub stars~2.4k tokensUpdated today
    Auto-check: warnings
  • Deepseek Provider Contract

    AnastasiyaW/codex-claude-code-config

    Validate a proposed DeepSeek API integration before any key or project context is sent: check thinking-mode tool-call history, strict-schema assumptions, bounded output, and provider data boundaries.

    154 GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Categories

Questions about Observability Monitoring

What does Observability Monitoring do?

Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…. Observability Monitoring is an agent skill from AnastasiyaW/codex-claude-code-config. Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting, burn-rate response, and postmortems.

When should I use Observability Monitoring?

Observability Monitoring fits situations like: asked about monitoring; incident evidence; generic dashboard styling; frontend-only UI work.

How do I install Observability Monitoring in Claude Code?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill observability-monitoring -a claude-code`. Or copy the skill folder (skills/operational/observability-monitoring in AnastasiyaW/codex-claude-code-config) into .claude/skills/observability-monitoring in your project. Claude Code loads it when a task matches its description.

How do I install Observability Monitoring in Codex?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill observability-monitoring -a codex`. Or copy the skill folder (skills/operational/observability-monitoring in AnastasiyaW/codex-claude-code-config) into .agents/skills/observability-monitoring in your project. Codex loads it when a task matches its description.

Can I use Observability Monitoring in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AnastasiyaW/codex-claude-code-config --skill observability-monitoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/observability-monitoring, .gemini/skills/observability-monitoring, .github/skills/observability-monitoring and .opencode/skills/observability-monitoring in your project.

What does Observability Monitoring need to run?

SKILL.md names no scripts, command-line tools or credentials: Observability Monitoring is instructions for the agent only.

Does Observability Monitoring access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Observability Monitoring safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Observability Monitoring use?

Observability Monitoring is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Observability Monitoring use?

About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 958 tokens, read only when the agent opens those files.

What are the alternatives to Observability Monitoring?

Skills that share tags, products or a category with Observability Monitoring: Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Observability Patterns (softspark/ai-toolkit, 179 stars), Telemetry (magnus919/agent-skills, 115 stars) and Observability Sre (majiayu000/spellbook, 287 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Observability Monitoring?

AnastasiyaW (a GitHub user) maintains it in AnastasiyaW/codex-claude-code-config, which has 154 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.

Source: AnastasiyaW/codex-claude-code-config on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.