Agent skill

Observability Designer

by borghei in borghei/Claude-Skills

Design observability strategies: SLI/SLO frameworks, alerting, and dashboards.

MITAuto-check passedDevOps & Cloud

Install Observability Designer

skills CLI
$ npx skills add borghei/Claude-Skills --skill observability-designer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills observability-designer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/observability-designer .claude/skills/observability-designer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
observability-designer
GitHub stars
881
Token cost
~1.8k tokens
SKILL.md length
726 words
Files
16 (incl. scripts, references, assets)
Skills in repo
349
Repo updated
First seen
Licence
MIT

At a glance

Design observability strategies: SLI/SLO frameworks, alerting, and dashboards.

  • Instrumenting a production service
  • SKILL.md covers Core Capabilities, When to Use, Clarify First and Tools, plus 3 more sections
  • Runs Python scripts from its folder; calls python
  • Tuning alert rules

What it does

Observability Designer is an agent skill from borghei/Claude-Skills. Design observability strategies: SLI/SLO frameworks, alerting, and dashboards. Use when instrumenting a production service, tuning alert rules, designing Grafana dashboards, defining SLOs and error budgets, or reducing alert fatigue.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including scripts, reference files and assets (for example `README.md`, `assets/sample_alerts.json` and `assets/sample_service_api.json`).

It sits in DevOps & Cloud, covering Site reliability engineering, Monitoring and alerting and Observability. It works with Grafana. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Instrumenting a production service
  • Tuning alert rules
  • Designing Grafana dashboards
  • Defining SLOs and error budgets

Example prompts

  • “/observability-designer”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 4a698e8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Observability Designer loads about 1.8k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 64 tokens; SKILL.md has 726 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit 4a698e8, republished under its MIT licence (© borghei). 726 words, ~1,787 tokens.

Download SKILL.mdSave it as .claude/skills/observability-designer/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.
name
observability-designer
description
Design observability strategies: SLI/SLO frameworks, alerting, and dashboards. Use when instrumenting a production service, tuning alert rules, designing Grafana dashboards, defining SLOs and error budgets, or reducing alert fatigue.
license
MIT + Commons Clause
metadata.version
1.1.0
metadata.author
borghei
metadata.category
engineering
metadata.domain
observability
metadata.tier
POWERFUL
metadata.updated
2026-06-17

Observability Designer

Design production-ready observability strategies that combine the three pillars (metrics, logs, traces) with SLI/SLO frameworks, golden-signals monitoring, multi-window burn-rate alerting, and alert-noise optimization.

Core Capabilities

  • SLI/SLO frameworks — select SLIs from the golden signals, map them to Prometheus expressions, set SLO targets by criticality tier, and compute error budgets.
  • Burn-rate alerting — multi-window burn-rate rules with severity routing, hysteresis, suppression, and grouping to keep alert noise below 10%.
  • Dashboard design — Grafana specs following the Overview > Service > Component > Instance hierarchy, ≤7 panels per screen, role-based views (SRE/Dev/Exec/Ops).
  • Structured logging & tracing — JSON log format with correlation IDs, log-level discipline, and head/tail/adaptive trace sampling strategies.
  • Runbooks & validation — runbook template per critical alert; coverage validation that every T1 service has metrics, logs, traces, and a runbook.
  • Cost optimization — metric/log/trace retention tiers and cardinality management.

When to Use

  • Instrumenting a new or existing production service.
  • Defining SLOs and error budgets for a service tier.
  • Tuning alert rules or reducing alert fatigue / alert storms.
  • Designing Grafana dashboards or role-based views.
  • Choosing a trace sampling strategy or structured log schema.

Clarify First

Before designing the observability strategy, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Service type & criticality tier — api / pipeline / storage / ML and T1–T3 (sets SLO targets, error-budget math, and slo_designer flags)
  • User-facing vs internal — determines which golden signals become SLIs and how alert severity is routed
  • Primary pain: alert noise vs coverage gaps — decides whether to optimize existing alerts or design new burn-rate rules
  • Dashboard audience — SRE / Dev / Exec / Ops sets the role-based panel layout and hierarchy

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Tools

ToolPurposeCommand
slo_designer.pyGenerate SLI/SLO framework, error budgets, and burn-rate alerts from a service definitionpython scripts/slo_designer.py --service-type api --criticality high --user-facing true
alert_optimizer.pyAnalyze alert configs for noise, coverage gaps, and duplicates; emit an optimization reportpython scripts/alert_optimizer.py --input alerts.json --analyze-only
dashboard_generator.pyProduce Grafana-compatible dashboard JSON with golden signals and role-based viewspython scripts/dashboard_generator.py --service-type api --name "Payment Service"

References

Load the reference that matches the task — keep this file lean and pull detail on demand:

  • references/slo-and-alerting.md — the 8-step workflow, SLI/SLO quick reference, error-budget math, burn-rate alert windows, alert classification, alert-fatigue prevention, and golden signals. Read when designing SLOs or alerts.
  • references/dashboards-logs-traces.md — dashboard design rules, structured log format, trace sampling strategies, the runbook template, a complete worked payment-service spec, and cost optimization. Read when building dashboards, logs, traces, or runbooks.
  • references/tools-integration-and-troubleshooting.md — full per-script flag/output reference, the systems integration table (Prometheus/Grafana/Jaeger/PagerDuty), the troubleshooting table, and success-criteria targets. Read when running the scripts or diagnosing failures.
  • references/slo_cookbook.md — a practical, in-depth cookbook for defining and operating Service Level Objectives. Read when you need detailed SLO methodology beyond the quick reference.
  • references/alert_design_patterns.md — a deep guide to effective alerting patterns and anti-patterns. Read when designing a complete alerting strategy.
  • references/dashboard_best_practices.md — comprehensive dashboard design-for-insight best practices. Read when building a dashboard system from scratch.
Show full SKILL.md (229 more words)Show less

Scope & Limitations

Covers:

  • SLI/SLO framework design for request-driven, pipeline, storage, and ML services.
  • Multi-window burn-rate alert generation and alert noise optimization.
  • Grafana-compatible dashboard specification with role-based layouts (SRE, Developer, Executive, Ops).
  • Structured logging format, trace sampling strategy selection, and cost-optimization guidance.

Does NOT cover:

  • Infrastructure provisioning or Terraform/Helm configuration for Prometheus, Grafana, or Jaeger -- see ci-cd-pipeline-builder for deployment pipelines.
  • Incident response workflow orchestration or post-mortem facilitation -- see runbook-generator for runbook authoring.
  • Application Performance Management (APM) agent installation or vendor-specific SDK integration.
  • Security monitoring, SIEM rule design, or compliance audit logging -- see skill-security-auditor for security-focused analysis.

Integration Points

SkillIntegrationData Flow
runbook-generatorEvery burn-rate alert references a runbook; the runbook generator consumes alert definitions to scaffold investigation stepsAlert YAML --> runbook-generator --> Markdown runbook linked in alert annotations
ci-cd-pipeline-builderDeployment events feed into dashboard annotations and alert suppression windowsPipeline events --> Grafana annotations + Alertmanager silences
performance-profilerLatency SLI breaches trigger profiling; profiler results inform SLO target adjustmentsSLO burn-rate alert --> profiler invocation --> refined latency thresholds
database-designerDatabase SLIs (query latency, connection success rate, replication lag) align with schema-level health checksDB schema metadata --> SLI metric expressions for database-type services
tech-debt-trackerError budget depletion signals feed into tech debt prioritization as reliability investmentsError budget reports --> tech debt backlog items with SLO-linked severity
release-managerRelease readiness gates check remaining error budget before approving deploymentsError budget API --> release gate pass/fail decision

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 15 other files (scripts, references, assets) in engineering/observability-designer of borghei/Claude-Skills.

  • SKILL.md
  • README.md
  • assets/sample_alerts.json
  • assets/sample_service_api.json
  • assets/sample_service_web.json
  • expected_outputs/sample_dashboard.json
  • expected_outputs/sample_slo_framework.json
  • references/alert_design_patterns.md
  • references/dashboard_best_practices.md
  • references/dashboards-logs-traces.md
  • references/slo-and-alerting.md
  • references/slo_cookbook.md
  • references/tools-integration-and-troubleshooting.md
  • scripts/alert_optimizer.py
  • scripts/dashboard_generator.py
  • scripts/slo_designer.py

Open the folder on GitHubat commit 4a698e8

Compare with similar skills

Observability Designer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Observability Designer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Observability Designer this skillborghei/Claude-Skills881—~1.8kAutomated safety check: PassMIT
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Observability MonitoringAnastasiyaW/codex-claude-code-config154—~4.1kAutomated safety check: PassMIT
Observability Sremajiayu000/spellbook286—~3.3kAutomated safety check: PassMIT
Observability Patternssoftspark/ai-toolkit179—~2.2kAutomated safety check: PassApache-2.0
Telemetrymagnus919/agent-skills113—~3.9kAutomated safety check: PassMIT

Similar skills

  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 5 mo ago
    DevOps & CloudAuto-check passed
  • Observability Monitoring

    AnastasiyaW/codex-claude-code-config

    Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…

    154 GitHub stars~4.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Observability Sre

    majiayu000/spellbook

    Observability and SRE expert. An agent skill from majiayu000/spellbook.

    286 GitHub stars~3.3k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Observability Patterns

    softspark/ai-toolkit

    Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI.

    179 GitHub stars~2.2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Telemetry

    magnus919/agent-skills

    Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector…

    113 GitHub stars~3.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Cloud Monitoring

    seb1n/awesome-ai-agent-skills

    Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability.

    206 GitHub stars~2.8k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed

More from borghei/Claude-Skills

All 349 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    881 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    881 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    881 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    881 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    881 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    881 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Observability Designer

What does Observability Designer do?

Design observability strategies: SLI/SLO frameworks, alerting, and dashboards. Observability Designer is an agent skill from borghei/Claude-Skills. Design observability strategies: SLI/SLO frameworks, alerting, and dashboards.

When should I use Observability Designer?

Observability Designer fits situations like: instrumenting a production service; tuning alert rules; designing Grafana dashboards; defining SLOs and error budgets.

How do I install Observability Designer in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill observability-designer -a claude-code`. Or copy the skill folder (engineering/observability-designer in borghei/Claude-Skills) into .claude/skills/observability-designer in your project. Claude Code loads it when a task matches its description.

How do I install Observability Designer in Codex?

Run `npx skills add borghei/Claude-Skills --skill observability-designer -a codex`. Or copy the skill folder (engineering/observability-designer in borghei/Claude-Skills) into .agents/skills/observability-designer in your project. Codex loads it when a task matches its description.

Can I use Observability Designer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill observability-designer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/observability-designer, .gemini/skills/observability-designer, .github/skills/observability-designer and .opencode/skills/observability-designer in your project.

What does Observability Designer need to run?

Going by SKILL.md and its folder, Observability Designer needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Observability Designer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Observability Designer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Observability Designer use?

Observability Designer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Observability Designer use?

About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.

What are the alternatives to Observability Designer?

Skills that share tags, products or a category with Observability Designer: Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Observability Monitoring (AnastasiyaW/codex-claude-code-config, 154 stars), Observability Sre (majiayu000/spellbook, 286 stars) and Observability Patterns (softspark/ai-toolkit, 179 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Observability Designer?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 881 GitHub stars. The repository holds 349 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.