Monitoring Observability
ahmedasmar/devops-claude-skills
Monitoring and observability strategy, implementation, and troubleshooting.
Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector…
$ npx skills add magnus919/agent-skills --skill telemetry -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install magnus919/agent-skills telemetry --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/telemetry .claude/skills/telemetry && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "telemetry" agent skill from https://github.com/magnus919/agent-skills/tree/main/telemetry into .claude/skills/telemetry/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "telemetry", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/magnus919/agent-skills/tree/main/telemetryType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add magnus919/agent-skills --skill telemetry -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install magnus919/agent-skills telemetry --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/telemetry .agents/skills/telemetry && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "telemetry" agent skill from https://github.com/magnus919/agent-skills/tree/main/telemetry into .agents/skills/telemetry/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "telemetry", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill telemetry -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install magnus919/agent-skills telemetry --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/telemetry .cursor/skills/telemetry && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "telemetry" agent skill from https://github.com/magnus919/agent-skills/tree/main/telemetry into .cursor/skills/telemetry/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "telemetry", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/magnus919/agent-skills.git --path telemetry--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add magnus919/agent-skills --skill telemetry -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install magnus919/agent-skills telemetry --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/telemetry .gemini/skills/telemetry && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "telemetry" agent skill from https://github.com/magnus919/agent-skills/tree/main/telemetry into .gemini/skills/telemetry/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "telemetry", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install magnus919/agent-skills telemetryInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add magnus919/agent-skills --skill telemetry -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/telemetry .github/skills/telemetry && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "telemetry" agent skill from https://github.com/magnus919/agent-skills/tree/main/telemetry into .github/skills/telemetry/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "telemetry", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill telemetry -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install magnus919/agent-skills telemetry --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/telemetry .opencode/skills/telemetry && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "telemetry" agent skill from https://github.com/magnus919/agent-skills/tree/main/telemetry into .opencode/skills/telemetry/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "telemetry", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
telemetryOperate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector…
Telemetry is an agent skill from magnus919/agent-skills. Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector pipelines (receivers, processors, exporters, sampling, trace/span correlation); and Loki ingest, LogQL, retention, and label design — with a bundled read-only telemetry-check script for Prometheus rule sanity and scrape-target reachability. Use when running, tuning, or troubleshooting a Prometheus, OpenTelemetry Collector, or…
Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files, including scripts and reference files (for example `README.md`, `evals/evals.json` and `fixtures/prometheus-rules.yml`). Compatibility notes: The bundled telemetry-check script runs on Python 3.9+ and needs no Prometheus server for --help. Rule and scrape-config checks read local YAML/JSON files…
It sits in DevOps & Cloud, covering Monitoring and alerting, Observability and Site reliability engineering. It works with Prometheus, OpenTelemetry and Grafana. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit c545c2b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
The bundled telemetry-check script runs on Python 3.9+ and needs no Prometheus server for --help. Rule and scrape-config checks read local YAML/JSON files; scrape-target reachability probes use TCP connects only and require network access to the targets.
From compatibility in the SKILL.md frontmatter.
Telemetry loads about 3.9k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 238 tokens; SKILL.md has 1,715 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from magnus919/agent-skills at commit c545c2b, republished under its MIT licence (© magnus919). 1,715 words, ~3,894 tokens.
.claude/skills/telemetry/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.Use this skill to operate the telemetry stack — Prometheus, the OpenTelemetry Collector, and Loki — as the one deployment unit it ships as: collection and scraping, ingestion, retention, and the rules that turn raw signals into alerts. This is a tool skill for one stack of named tools. Observability strategy — SLIs, SLOs, error budgets, and what to instrument — belongs to platform-engineering; dashboards, panels, and Grafana-side alert rules, contact points, and notification policies belong to grafana. This skill owns the collection/ingest/retention layer and the Prometheus rules files that both of those skills consume.
telemetry-check script runs rule sanity and scrape-target reachability checks without changing anything.promtool rules push, a collector restart, a retention-policy change — require an explicit human directive naming the instance.prometheus.yml, collector pipelines, or credentials into chat.scripts/telemetry-check is an agent-first, read-only checker. It parses Prometheus rules files with a bundled stdlib YAML reader and runs dependency-free structural sanity checks; it extracts static targets from scrape configs and probes TCP reachability; and it emits bounded JSON. It never writes files and never sends data anywhere.
scripts/telemetry-check --help # no server needed
scripts/telemetry-check --rules rules.yml --json # rule sanity, machine-readable
scripts/telemetry-check --scrape prometheus.yml --json # probe static targets
scripts/telemetry-check --targets targets.txt --timeout 5Exit codes: 0 all checks passed, 1 issues found or a fatal error, 2 usage error. telemetry-check --rules checks structure only: exactly one of record/alert per rule, a non-empty expression with balanced delimiters, valid durations, recording-rule and label names, and string-only label values. Use promtool check rules separately for full PromQL parsing.
telemetry-check --rules and --scrape on the configs, then check the live status endpoints (/-/healthy, /api/v1/targets, collector health, Loki ready) where access exists.references/05-query-workflows.md. Define the signal, selector, UTC time window, step/limit, and expected unit; validate syntax separately from semantics; execute read-only instant then bounded range/log queries; and capture status, scope, time, limits, warnings, and cardinality evidence.scrape_configs): one job per scrape group with a deliberate scrape_interval, scrape_timeout below it, and metrics_path. Prefer static_configs for known endpoints and service discovery (*_sd_configs) for dynamic ones. Verify the running config with /api/v1/status/config and targets with /api/v1/targets?state=active.groups of record or alert rules with a PromQL expr, optional for/keep_firing_for durations, and labels/annotations. Validate every change with promtool check rules for full PromQL parsing and with the bundled telemetry-check --rules for dependency-free structural sanity before reload. Rules must be small, well-named, and reviewable — a 100-line expression is a debugging liability, not a rule.relabel_configs and metric_relabel_configs rewrite labels before ingestion. Use them to enforce label naming, drop high-cardinality or internal labels, and attach scrape metadata. Relabeling mistakes silently change series identity — verify with a targeted curl of /metrics and the target's scrapeUrl in /api/v1/targets.--storage.tsdb.retention.time and --storage.tsdb.retention.size bound local block retention; blocks are 2h by default. Retention is a capacity decision (see references/04-stack-integration-and-retention.md), not a default to leave alone. Watch prometheus_tsdb_head_series and prometheus_tsdb_compaction for cardinality and compaction pressure.--query.max-concurrency headroom and consistent external labels let you shard or deduplicate at the query layer (Thanos, Mimir, or Grafana data sources). Alerting rules must not double-fire: HA pairs need a dedup layer or consistent labeling, and rule evaluation must stay consistent across replicas. Rule evaluation state (for counters) is local to each instance.receivers → processors → exporters per signal type (metrics, logs, traces). Keep pipelines narrow and per-signal; a pipeline that mixes signals becomes un-debuggable. Each pipeline must have at least one exporter; unused receivers/exporters are dead configuration.memory_limiter processor belong before exporters; tail_sampling belongs on trace pipelines only.tail_sampling on traces decides at the batch level; probabilistic_sampler is stateless and cheaper. Sample deliberately: full traces for errors and slow paths, tail sampling for high-volume success traffic, and never sample away the error signal. Sampling must be coordinated with retention — a sampled trace is gone forever, so the decision belongs in the pipeline design, not in an emergency.trace_id and span_id in log lines and metric exemplars so LogQL and PromQL can pivot back to the trace. The collector's spanmetrics processor derives RED metrics from spans, and OTLP logs with trace context land in Loki with trace_id as a structured label for correlation. Trace context propagation is an application-level concern that backend-engineering owns; the collector side is here./loki/api/v1/push) from Promtail, the OTel Collector's loki exporter, or the Grafana Agent/Alloy. Verify ingest with loki_distributor_bytes_received_total and the ready endpoint; an ingest that silently drops (rate limits, too many outstanding requests) hides outages.{label="value"} |= "filter" | json selects streams and filters lines; label matchers are the primary cost driver. LogQL analytics (sum by (...) (rate({app="x"} |~ "error"[5m]))) work on the label index plus line filtering — design labels so the matchers you actually use are cheap.retention_period and retention_size apply per tenant; the compactor enforces them and merges index shards. Log volume is unbounded if ungoverned — set retention before rollout, track it with loki_compactor metrics, and treat log retention as a compliance decision with an owner.| json/| regexp or OTel structured metadata. Cardinality guidance: a label whose values change with every log line does not belong in the index.Retention is a stack-wide decision: Prometheus blocks (raw samples), OTel Collector buffering (in-memory queue, exporter retries), and Loki (indexed logs) each have independent retention, and the combined storage footprint is what the team pays for. Decide per component based on the question the data answers (hot metrics for alerting, samples for trends, logs for debugging and audit), set it in config, and review it on a schedule. See references/04-stack-integration-and-retention.md for the trade-off tables and the alerting rule that watches retention.
| Load when | Reference |
|---|---|
| Sources, version observations, refresh procedure | references/00-source-index.md |
| Scrape config, recording/alerting rules, relabeling, retention, HA | references/01-prometheus-operations.md |
| Collector pipelines, receivers/processors/exporters, sampling, correlation | references/02-opentelemetry-collector.md |
| Ingest, LogQL, retention, label design | references/03-loki-operations.md |
| PromQL/LogQL construction, semantic review, bounded cost, and no-data diagnosis | references/05-query-workflows.md |
| Cross-component retention decisions and stack integration | references/04-stack-integration-and-retention.md |
scripts/telemetry-check: read-only rule sanity + scrape-target reachability checker (stdlib-only, --json, --rules/--scrape/--targets, --help without a server).tests/test_telemetry_check.py: deterministic tests against fixture configs, including the read-only contract.fixtures/: prometheus-rules.yml (valid rules) and scrape-config.yml (valid scrape config) used by the tests and as starting points.references/: five dated, source-indexed references plus the source index, including the bounded PromQL/LogQL workflow.evals/evals.json: six output-quality evaluation cases for agent runs.| Claim | Minimum evidence |
|---|---|
| A rules file is structurally sound | telemetry-check --rules FILE --json exits 0 with no errors |
| A rules file is semantically valid | promtool check rules FILE exits 0 |
| A target is scrapable | telemetry-check --scrape CONFIG --json reports it reachable, and /api/v1/targets shows state="up" |
| A pipeline is live | Collector health endpoint responds and per-signal metrics (otelcol_receiver_*, otelcol_exporter_*) advance |
| Ingest is healthy | Distributor metrics advance and the ready endpoint returns 200 |
| Retention is governed | retention_period/retention_size are set explicitly, and compactor/TSDB metrics confirm the policy |
telemetry-check as anything but what it is — read-only. It has no mutation surface.© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 12 other files (scripts, references) in telemetry of magnus919/agent-skills.
Open the folder on GitHubat commit c545c2b
Telemetry next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Telemetry this skillmagnus919/agent-skills | 116 | — | ~3.9k | Automated safety check: Pass | MIT | |
| Monitoring Observabilityahmedasmar/devops-claude-skills | 203 | — | ~3.9k | Automated safety check: Pass | None | |
| Alloygrafana/skills | 281 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Observability MonitoringAnastasiyaW/codex-claude-code-config | 154 | — | ~4.1k | Automated safety check: Pass | MIT | |
| Observability Patternssoftspark/ai-toolkit | 179 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Observability Sremajiayu000/spellbook | 287 | — | ~3.3k | Automated safety check: Pass | MIT |
ahmedasmar/devops-claude-skills
Monitoring and observability strategy, implementation, and troubleshooting.
grafana/skills
Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /…
AnastasiyaW/codex-claude-code-config
Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…
softspark/ai-toolkit
Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI.
majiayu000/spellbook
Observability and SRE expert. An agent skill from majiayu000/spellbook.
archestra-ai/archestra
A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup.
magnus919/agent-skills
Organize durable agent research outputs as summaries, analysis, and evidence dossiers.
magnus919/agent-skills
Build portable, first-person colored ASCII city engines and small GIS-derived city packs.
magnus919/agent-skills
Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.
magnus919/agent-skills
A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…
magnus919/agent-skills
Use Docker Compose to define, run, debug, and harden multi-container applications.
magnus919/agent-skills
Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.
Works with
Categories
Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector…. Telemetry is an agent skill from magnus919/agent-skills. Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector pipelines (receivers, processors, exporters, sampling, trace/span correlation); and Loki ingest, LogQL, retention, and label design — with a bundled read-only telemetry-check script for Prometheus rule sanity and scrape-target reachability.
Telemetry fits situations like: troubleshooting a Prometheus; openTelemetry Collector; loki deployment; reviewing the collection/ingest/retention layer.
Run `npx skills add magnus919/agent-skills --skill telemetry -a claude-code`. Or copy the skill folder (telemetry in magnus919/agent-skills) into .claude/skills/telemetry in your project. Claude Code loads it when a task matches its description.
Run `npx skills add magnus919/agent-skills --skill telemetry -a codex`. Or copy the skill folder (telemetry in magnus919/agent-skills) into .agents/skills/telemetry in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill telemetry -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/telemetry, .gemini/skills/telemetry, .github/skills/telemetry and .opencode/skills/telemetry in your project.
Going by SKILL.md and its folder, Telemetry needs Python for the scripts in its folder. Our summary lists: Python 3. Compatibility (from SKILL.md): The bundled telemetry-check script runs on Python 3.9+ and needs no Prometheus server for --help. Rule and scrape-config checks read local YAML/JSON files; scrape-target reachability probes use TCP connects only and require network access to the targets..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Telemetry is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Telemetry: Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Alloy (grafana/skills, 281 stars), Observability Monitoring (AnastasiyaW/codex-claude-code-config, 154 stars) and Observability Patterns (softspark/ai-toolkit, 179 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 116 GitHub stars. The repository holds 130 skills in this directory. The repository was last updated on October 8, 2026.
Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.