Agent skill

Observability And Reliability

by cbrock84 in cbrock84/headcount

Makes systems debuggable and reliably operable — instrumentation, alerting that is worth waking for, service objectives, and learning from failure.

MITAuto-check passedDevOps & Cloud

Install Observability And Reliability

skills CLI
$ npx skills add cbrock84/headcount --skill observability-and-reliability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cbrock84/headcount observability-and-reliability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cbrock84/headcount.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/technology/skills/observability-and-reliability .claude/skills/observability-and-reliability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
observability-and-reliability
GitHub stars
2k
Token cost
~931 tokens
SKILL.md length
510 words
Files
2 (incl. references)
Skills in repo
175
Repo updated
First seen
Licence
MIT

At a glance

Makes systems debuggable and reliably operable — instrumentation, alerting that is worth waking for, service objectives, and learning from failure.

  • Tasks that involve Site reliability engineering
  • SKILL.md covers Instrument for questions you…, Alert on symptoms, not causes, Objectives and error budgets and Learn from incidents, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Observability

What it does

Observability And Reliability is an agent skill from cbrock84/headcount. Makes systems debuggable and reliably operable — instrumentation, alerting that is worth waking for, service objectives, and learning from failure. Use this to instrument a service, fix alerting that is ignored, set error budgets or reliability targets, prepare for on-call, or run a blameless post-incident review.

Its SKILL.md is about 930 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/sources.md`).

It sits in DevOps & Cloud, covering Site reliability engineering, Observability and Incident response. The repository describes itself as: An agent organization structured as a company — 15+ departments, 125+ skills, each independently installable, citing the standards and regulators that settle the question. Runs… The licence is MIT.

When your agent uses it

  • Tasks that involve Site reliability engineering
  • Tasks that involve Observability
  • Tasks that involve Incident response

Example prompts

  • “Use the observability-and-reliability skill to make systems debuggable and reliably operable — instrumentation, alerting that is worth waking for…”
  • “/observability-and-reliability”

What it can do on your machine

Read from SKILL.md and the folder at commit 98d1c17. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Observability And Reliability loads about 931 tokens when it runs, and up to ~1.4k if it reads all its reference files. Until then it costs about 86 tokens; SKILL.md has 510 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~86
When it runs · the whole SKILL.md, loaded when a task matches
~931
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cbrock84/headcount at commit 98d1c17, republished under its MIT licence (© cbrock84). 510 words, ~931 tokens.

Download SKILL.mdSave it as .claude/skills/observability-and-reliability/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
observability-and-reliability
description
Makes systems debuggable and reliably operable — instrumentation, alerting that is worth waking for, service objectives, and learning from failure. Use this to instrument a service, fix alerting that is ignored, set error budgets or reliability targets, prepare for on-call, or run a blameless post-incident review.

Observability and reliability

Monitoring tells you a thing you predicted is happening. Observability lets you ask a question you did not anticipate. Production failures are mostly the unanticipated kind.

Instrument for questions you have not thought of yet

Emit structured events with enough context to slice afterwards — request identifiers, user or tenant, version, dependency, outcome, duration. Free-text logs are unsearchable at volume and become expensive noise.

Propagate a correlation identifier across every hop. Without it, a distributed system is a set of independent stories and reconstructing one request is manual archaeology.

Measure what the user experiences at the percentile they experience it. A p50 latency graph is mostly a graph of the people who were not affected.

Alert on symptoms, not causes

Alert when users are affected or imminently will be. High CPU is not an alert; requests failing or slowing is. Cause-based alerting produces pages for conditions the system handled and no page for novel failures that hurt.

Every alert must be actionable, urgent and specific. If the recipient's honest response is to look and close it, delete the alert — it is training the on-call to ignore the page, and the ignored page is eventually the real one.

Alert fatigue is the actual reliability risk in most organizations. Fewer, better alerts beat coverage.

Objectives and error budgets

Set service level objectives from what users need, then treat the remainder as a budget to spend. This converts a sterile argument between shipping and stability into arithmetic: budget remaining means ship, budget exhausted means the next work is reliability.

Keep the internal objective tighter than any external commitment made through operations:service-level-management, so you find out before the customer does.

Show full SKILL.md (232 more words)Show less

Learn from incidents

Post-incident review exists to find what made the failure possible and hard to detect, not who touched it last. Human error is a starting question, never the finding: what made the error easy, and why did nothing catch it?

Track the time to detect separately from time to resolve. Long detection is an observability defect, and it is the part that repeats.

Produce a small number of real actions with owners and dates. A review generating fifteen actions generates none.

Sources

references/sources.md in this skill lists the outside authorities that settle the questions here — what each one is authoritative for, and what you may do with it. Check them before answering on anything they cover, and cite what you used. Most are free to read and not free to reproduce; the use note on each is binding.

Tooling

Metrics and traces: Datadog, Grafana with Prometheus, New Relic, Honeycomb, and similar. Errors: Sentry, Rollbar, and similar. Logs: Elastic, OpenSearch, Loki, Splunk, and similar.

On-call and incident management: PagerDuty, Opsgenie, incident.io, FireHydrant, and similar.

Instrument with OpenTelemetry wherever you can. Vendor-specific instrumentation is the part that makes leaving expensive.

Never

  • Page a human for something they cannot act on.
  • Alert on a cause when you can alert on the symptom.
  • Report reliability as an average when users experience the tail.
  • Close an incident review with the finding that someone was careless.

© cbrock84, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/technology/skills/observability-and-reliability of cbrock84/headcount.

  • SKILL.md
  • references/sources.md

Open the folder on GitHubat commit 98d1c17

Compare with similar skills

Observability And Reliability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Observability And Reliability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Observability And Reliability this skillcbrock84/headcount2k—~931Automated safety check: PassMIT
Error HandlerEliasOulkadi/shokunin114—~3.6kAutomated safety check: NotesMIT
Observability Engineerdavila7/claude-code-templates32k8 repos~3.2kAutomated safety check: PassMIT
SRE EngineerJeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT
Incident Responderdavila7/claude-code-templates32k7 repos~2.6kAutomated safety check: PassMIT
Incident ResponderDokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI507—~706Automated safety check: PassCustom licence

Similar skills

  • Error Handler

    EliasOulkadi/shokunin

    Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…

    114 GitHub stars~3.6k tokensUpdated 3 days ago
    DevOps & CloudAuto-check: notes
  • Observability Engineer

    davila7/claude-code-templates

    Build production-ready monitoring, logging, and tracing systems.

    32k GitHub starsUsed in 8 repos~3.2k tokens
    DevOps & CloudAuto-check passed
  • SRE Engineer

    Jeffallan/claude-skills

    Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

    12k GitHub stars~1.7k tokensUpdated 5 days ago
    DevOps & CloudAuto-check passed
  • Incident Responder

    davila7/claude-code-templates

    Expert SRE incident responder specializing in rapid problem resolution, modern observability, and comprehensive incident management.

    32k GitHub starsUsed in 7 repos~2.6k tokens
    DevOps & CloudAuto-check passed
  • Incident Responder

    Dokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI

    Expert SRE incident responder specializing in rapid problem resolution.

    507 GitHub stars~706 tokensUpdated 3 mo ago
    DevOps & CloudAuto-check passed
  • It Operations

    davila7/claude-code-templates

    Manages IT infrastructure, monitoring, incident response, and service reliability.

    32k GitHub starsUsed in 1 repo~3.7k tokens
    DevOps & CloudAuto-check passed

More from cbrock84/headcount

All 175 skills in this repo
  • Agent Hierarchy

    cbrock84/headcount

    Designs orchestrator-and-subagent hierarchies for a repository — splitting agents by exclusive write surface, pairing every producer with an independent auditor, and enforcing the split with a…

    2k GitHub stars~1.2k tokensUpdated 21 days ago
    Auto-check passed
  • Access And Identity

    cbrock84/headcount

    Designs and audits who can reach what — authentication, authorization models, privileged access, service credentials, and joiner-mover-leaver process.

    2k GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check passed
  • Account Based Marketing

    cbrock84/headcount

    Concentrates marketing and sales effort on a named set of accounts rather than on volume — qualifying whether the model fits your economics at all, building the account list and the buying group…

    2k GitHub stars~1.2k tokensUpdated 21 days ago
    Auto-check passed
  • Activation

    cbrock84/headcount

    Gets new users from signup to first real value — signup flow, onboarding, time-to-value, and the early experience that determines whether someone becomes a user or a lapsed account.

    2k GitHub stars~865 tokensUpdated 21 days ago
    Auto-check passed
  • AI Research Analyst

    cbrock84/headcount

    Produces executive-level research — market sizing, competitor mapping, trend analysis, and strategic intelligence — grounded in cited sources with the confidence in each claim made explicit.

    2k GitHub stars~916 tokensUpdated 21 days ago
    Auto-check passed
  • AI Search Optimization

    cbrock84/headcount

    Optimizes for AI assistants and AI-generated answers — being retrievable, being cited, and being represented accurately when a model answers on your behalf.

    2k GitHub stars~829 tokensUpdated 21 days ago
    Auto-check passed

Categories

Questions about Observability And Reliability

What does Observability And Reliability do?

Makes systems debuggable and reliably operable — instrumentation, alerting that is worth waking for, service objectives, and learning from failure. Observability And Reliability is an agent skill from cbrock84/headcount. Makes systems debuggable and reliably operable — instrumentation, alerting that is worth waking for, service objectives, and learning from failure.

When should I use Observability And Reliability?

Observability And Reliability fits situations like: tasks that involve Site reliability engineering; tasks that involve Observability; tasks that involve Incident response.

How do I install Observability And Reliability in Claude Code?

Run `npx skills add cbrock84/headcount --skill observability-and-reliability -a claude-code`. Or copy the skill folder (plugins/technology/skills/observability-and-reliability in cbrock84/headcount) into .claude/skills/observability-and-reliability in your project. Claude Code loads it when a task matches its description.

How do I install Observability And Reliability in Codex?

Run `npx skills add cbrock84/headcount --skill observability-and-reliability -a codex`. Or copy the skill folder (plugins/technology/skills/observability-and-reliability in cbrock84/headcount) into .agents/skills/observability-and-reliability in your project. Codex loads it when a task matches its description.

Can I use Observability And Reliability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cbrock84/headcount --skill observability-and-reliability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/observability-and-reliability, .gemini/skills/observability-and-reliability, .github/skills/observability-and-reliability and .opencode/skills/observability-and-reliability in your project.

What does Observability And Reliability need to run?

SKILL.md names no scripts, command-line tools or credentials: Observability And Reliability is instructions for the agent only.

Does Observability And Reliability access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Observability And Reliability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Observability And Reliability use?

Observability And Reliability is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Observability And Reliability use?

About 931 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 443 tokens, read only when the agent opens those files.

What are the alternatives to Observability And Reliability?

Skills that share tags, products or a category with Observability And Reliability: Error Handler (EliasOulkadi/shokunin, 114 stars), Observability Engineer (davila7/claude-code-templates, 32k stars), SRE Engineer (Jeffallan/claude-skills, 12k stars) and Incident Responder (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Observability And Reliability?

cbrock84 (a GitHub user) maintains it in cbrock84/headcount, which has 2,007 GitHub stars. The repository holds 175 skills in this directory. The repository was last updated on September 17, 2026.

Source: cbrock84/headcount on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.