Agent skill

Observability Hardening

by swyxio in swyxio/skills

Audit or improve observability for named production questions, incidents, or opaque operations using privacy-safe telemetry.

MITAuto-check passedDevOps & Cloud

Install Observability Hardening

skills CLI
$ npx skills add swyxio/skills --skill observability-hardening -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install swyxio/skills observability-hardening --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/observability-hardening .claude/skills/observability-hardening && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
observability-hardening
GitHub stars
176
Token cost
~931 tokens
SKILL.md length
423 words
Files
3 (incl. references)
Skills in repo
89
Repo updated
First seen
Licence
MIT

At a glance

Audit or improve observability for named production questions, incidents, or opaque operations using privacy-safe telemetry.

  • Works in 5 steps: Map what must be understood → Define telemetry contracts → Instrument high-value paths → …
  • Explicitly asks for an observability
  • SKILL.md covers Counterweight: questions…, Workflow and Quality Bar
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Observability Hardening is an agent skill from swyxio/skills. Audit or improve observability for named production questions, incidents, or opaque operations using privacy-safe telemetry. Use only when the user explicitly asks for an observability or instrumentation pass, telemetry design, dashboards or alerts, or when inability to explain production behavior is the primary problem. Do not trigger for ordinary debugging, adding one log line, generic production readiness, product analytics alone, or implementation work that merely benefits from diagnostics.

Its SKILL.md is about 930 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `agents/openai.yaml` and `references/checklist.md`).

It sits in DevOps & Cloud, covering Observability and Product analytics. The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.

When your agent uses it

  • Explicitly asks for an observability
  • Instrumentation pass
  • Telemetry design
  • Inability to explain production behavior is the primary problem

Example prompts

  • “/observability-hardening”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Map what must be understood
  2. Define telemetry contracts
  3. Instrument high-value paths
  4. Make it usable
  5. Validate

What it can do on your machine

Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Observability Hardening loads about 931 tokens when it runs, and up to ~1.4k if it reads all its reference files. Until then it costs about 131 tokens; SKILL.md has 423 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~131
When it runs · the whole SKILL.md, loaded when a task matches
~931
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 423 words, ~931 tokens.

Download SKILL.mdSave it as .claude/skills/observability-hardening/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
observability-hardening
description
Audit or improve observability for named production questions, incidents, or opaque operations using privacy-safe telemetry. Use only when the user explicitly asks for an observability or instrumentation pass, telemetry design, dashboards or alerts, or when inability to explain production behavior is the primary problem. Do not trigger for ordinary debugging, adding one log line, generic production readiness, product analytics alone, or implementation work that merely benefits from diagnostics.

Observability Hardening

Use this skill when the product works but failures are opaque. The output should make production behavior explainable without leaking private data.

Counterweight: questions before telemetry

  • Start with the operator or user question that cannot currently be answered. Add no signal without a named consumer and decision.
  • Inspect provider-native logs, metrics, traces, request IDs, and dashboards before building parallel telemetry.
  • Choose the cheapest sufficient signal. Do not add logs, metrics, traces, dashboards, alerts, analytics, and user-visible state as a bundle.
  • Instrument ownership boundaries and material state transitions, not every function or phase.
  • Prefer a saved query or clearer existing event over a new pipeline, schema, SDK, collector, or dashboard.
  • Budget cardinality, volume, retention, latency, privacy, and on-call noise. More telemetry can make diagnosis worse.
  • Add an alert only when someone owns a concrete response. Add user-visible progress only when users must wait or act.
  • Do not turn a local debugging gap into a repository-wide observability program or product analytics redesign.
  • A scoped audit may conclude existing signals are sufficient or recommend deletion of noisy telemetry.
  • Do not configure external services, create dashboards, change retention, or deploy instrumentation unless the active request explicitly authorizes it.
Show full SKILL.md (227 more words)Show less

Workflow

  1. Map what must be understood

    • Trace only the user, route, job, provider, or cost boundaries relevant to the named question.
    • Inspect current provider and application signals that may already answer it.
  2. Define telemetry contracts

    • Choose only the event, label, correlation, and redaction fields required by the question.
    • Separate product analytics from engineering diagnostics.
    • Keep user/prompt/content-heavy payloads local or redacted unless explicitly allowed.
  3. Instrument high-value paths

    • Add request or correlation IDs only where a multi-boundary question requires them.
    • Add structured transition logs or safe error classification only when needed.
    • Add only metrics required by named operational questions.
    • Add user-visible operation state only when users must wait, retry, or act.
  4. Make it usable

    • Prefer the smallest saved query, dashboard, or developer/admin view that answers a real question.
    • Define alert thresholds only for actionable failures.
    • Document the resulting query or diagnostic path when it will be reused.
  5. Validate

    • Run focused local or production-shaped checks proportional to the changed telemetry path.
    • Verify logs/events are emitted, correlated, redacted, and not duplicated.
    • Test representative failure paths.

Quality Bar

  • Reviewed critical failures can be traced across the boundaries relevant to them.
  • New signals are structured and low-cardinality where needed.
  • Sensitive content is redacted by default.
  • Alerts map to actions, not noise.
  • Reviewed long-running operations expose only the state users or operators need.

For the audit checklist, read checklist.md.

© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in observability-hardening of swyxio/skills.

  • SKILL.md
  • agents/openai.yaml
  • references/checklist.md

Open the folder on GitHubat commit 038ef34

Compare with similar skills

Observability Hardening next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Observability Hardening compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Observability Hardening this skillswyxio/skills176—~931Automated safety check: PassMIT
Add React Analyticsgotempsh/temps833—~2.7kAutomated safety check: PassApache-2.0
Vercel Optimize Auditvercel-labs/agent-skills32k8 repos~4.3kAutomated safety check: PassNone
Kubeshark Installerkubeshark/kubeshark12k—~3.6kAutomated safety check: NotesApache-2.0
Kubeshark KFL2 Filter Referencekubeshark/kubeshark12k—~3.6kAutomated safety check: PassApache-2.0
KubeSphere ServiceMesh Managerkubesphere/kubesphere17k—~2.4kAutomated safety check: PassCustom licence

Similar skills

  • Add React Analytics

    gotempsh/temps

    Add Temps analytics to React applications with comprehensive tracking capabilities including page views, custom events, scroll tracking, engagement monitoring, session recording, and Web Vitals…

    833 GitHub stars~2.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Vercel Optimize Audit

    vercel-labs/agent-skills

    Official

    Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.

    32k GitHub starsUsed in 8 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Kubeshark Installer

    kubeshark/kubeshark

    Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.

    12k GitHub stars~3.6k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Syntax reference for KFL2, the CEL-based display filter language used to search Kubernetes network traffic captured by Kubeshark, loaded before any filter is written.

    12k GitHub stars~3.6k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • KubeSphere ServiceMesh Manager

    kubesphere/kubesphere

    Installs, checks and troubleshoots the KubeSphere ServiceMesh extension (Istio, Kiali, Jaeger), including grayscale release, sidecar injection, topology and tracing issues.

    17k GitHub stars~2.4k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from swyxio/skills

All 89 skills in this repo
  • Programmatic Agents

    swyxio/skills

    Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.

    176 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Design, implement, audit, or refresh protected username and handle namespaces for public products.

    176 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • New Mac Setup

    swyxio/skills

    Fully automated new Mac setup for fullstack web developers and AI engineers.

    176 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Youtube API

    swyxio/skills

    Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…

    176 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.

    176 GitHub stars~1.5k tokensUpdated today
    Auto-check: warnings
  • Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.

    176 GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Categories

Questions about Observability Hardening

What does Observability Hardening do?

Audit or improve observability for named production questions, incidents, or opaque operations using privacy-safe telemetry. Observability Hardening is an agent skill from swyxio/skills. Audit or improve observability for named production questions, incidents, or opaque operations using privacy-safe telemetry.

When should I use Observability Hardening?

Observability Hardening fits situations like: explicitly asks for an observability; instrumentation pass; telemetry design; inability to explain production behavior is the primary problem.

How do I install Observability Hardening in Claude Code?

Run `npx skills add swyxio/skills --skill observability-hardening -a claude-code`. Or copy the skill folder (observability-hardening in swyxio/skills) into .claude/skills/observability-hardening in your project. Claude Code loads it when a task matches its description.

How do I install Observability Hardening in Codex?

Run `npx skills add swyxio/skills --skill observability-hardening -a codex`. Or copy the skill folder (observability-hardening in swyxio/skills) into .agents/skills/observability-hardening in your project. Codex loads it when a task matches its description.

Can I use Observability Hardening in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill observability-hardening -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/observability-hardening, .gemini/skills/observability-hardening, .github/skills/observability-hardening and .opencode/skills/observability-hardening in your project.

What does Observability Hardening need to run?

SKILL.md names no scripts, command-line tools or credentials: Observability Hardening is instructions for the agent only.

Does Observability Hardening access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Observability Hardening safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Observability Hardening use?

Observability Hardening is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Observability Hardening use?

About 931 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 439 tokens, read only when the agent opens those files.

What are the alternatives to Observability Hardening?

Skills that share tags, products or a category with Observability Hardening: Add React Analytics (gotempsh/temps, 833 stars), Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars), Kubeshark Installer (kubeshark/kubeshark, 12k stars) and Kubeshark KFL2 Filter Reference (kubeshark/kubeshark, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Observability Hardening?

swyxio (a GitHub user) maintains it in swyxio/skills, which has 176 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.

Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.