Agent skill

Observability Sre

by majiayu000 in majiayu000/spellbook

Observability and SRE expert. An agent skill from majiayu000/spellbook.

MITAuto-check passedDevOps & Cloud

Install Observability Sre

skills CLI
$ npx skills add majiayu000/spellbook --skill observability-sre -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/spellbook observability-sre --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/observability-sre .claude/skills/observability-sre && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
observability-sre
GitHub stars
287
Token cost
~3.3k tokens
SKILL.md length
377 words
Files
7
Skills in repo
97
Repo updated
First seen
Licence
MIT

At a glance

Observability and SRE expert. An agent skill from majiayu000/spellbook.

  • Setting up monitoring
  • SKILL.md covers Core Principles, Hard Rules (Must Follow), Quick Reference and Observability Architecture, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Managing incidents

What it does

Observability Sre is an agent skill from majiayu000/spellbook. Observability and SRE expert. Use when setting up monitoring, logging, tracing, defining SLOs, or managing incidents. Covers Prometheus, Grafana, OpenTelemetry, and incident response best practices.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files (for example `reference/extended.md`, `reference/incident-response.md` and `reference/logging.md`).

It sits in DevOps & Cloud, covering Site reliability engineering, Observability and Monitoring and alerting. It works with Prometheus, OpenTelemetry and Grafana. The repository describes itself as: Cross-runtime skills for Claude Code, Codex, and multi-agent workflows. The licence is MIT.

When your agent uses it

  • Setting up monitoring
  • Managing incidents

Example prompts

  • “/observability-sre”

Requirements

  • Node.js

What it can do on your machine

Read from SKILL.md and the folder at commit ed52af7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml, typescript and promql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Observability Sre loads about 3.3k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 377 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/spellbook at commit ed52af7, republished under its MIT licence (© majiayu000). 377 words, ~3,346 tokens.

Download SKILL.mdSave it as .claude/skills/observability-sre/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
observability-sre
description
Observability and SRE expert. Use when setting up monitoring, logging, tracing, defining SLOs, or managing incidents. Covers Prometheus, Grafana, OpenTelemetry, and incident response best practices.

Observability & Site Reliability Engineering

Core Principles

  • Three Pillars — Metrics, Logs, and Traces provide holistic visibility
  • Observability-First — Build systems that explain their own behavior
  • SLO-Driven — Define reliability targets that matter to users
  • Proactive Detection — Find issues before customers do
  • Blameless Culture — Learn from failures without blame
  • Automate Toil — Reduce repetitive operational work
  • Continuous Improvement — Each incident makes systems more resilient
  • Full-Stack Visibility — Monitor from infrastructure to business metrics

Hard Rules (Must Follow)

These rules are mandatory. Violating them means the skill is not working correctly.

Symptom-Based Alerts Only

Alert on user-facing symptoms, not internal infrastructure metrics.

yaml
# ❌ FORBIDDEN: Alerting on internal metrics
- alert: CPUHigh
  expr: cpu_usage > 70%
  # Users don't care about CPU, they care about latency

- alert: MemoryHigh
  expr: memory_usage > 80%
  # Internal metric, may not affect users

# ✅ REQUIRED: Alert on user experience
- alert: APILatencyHigh
  expr: slo:api_latency:p95 > 0.200
  annotations:
    summary: "Users experiencing slow response times"

- alert: ErrorRateHigh
  expr: slo:api_errors:rate5m > 0.001
  annotations:
    summary: "Users encountering errors"
Low Cardinality Labels

Loki/Prometheus labels must have low cardinality (<10 unique labels).

yaml
# ❌ FORBIDDEN: High cardinality labels
labels:
  user_id: "usr_123"      # Millions of values!
  order_id: "ord_456"     # Millions of values!
  request_id: "req_789"   # Every request is unique!

# ✅ REQUIRED: Low cardinality only
labels:
  namespace: "production"  # Few values
  app: "api-server"        # Few values
  level: "error"           # 5-6 values
  method: "GET"            # ~10 values

# High cardinality data goes in log body:
logger.info({
  user_id: "usr_123",      # In JSON body, not label
  order_id: "ord_456",
}, "Order processed");
SLO-Based Error Budgets

Every service must have defined SLOs with error budget tracking.

yaml
# ❌ FORBIDDEN: No SLO definition
# Just monitoring without targets

# ✅ REQUIRED: Explicit SLO with budget
# SLO: 99.9% availability
# Error Budget: 0.1% = 43.2 minutes/month downtime

groups:
  - name: slo_tracking
    rules:
      - record: slo:api_availability:ratio
        expr: sum(rate(http_requests_total{status!~"5.."}[5m])) / sum(rate(http_requests_total[5m]))

      - alert: ErrorBudgetBurnRate
        expr: slo:api_availability:ratio < 0.999
        for: 5m
        annotations:
          summary: "Burning error budget too fast"
Trace Context in Logs

All logs must include trace_id for correlation with distributed traces.

typescript
// ❌ FORBIDDEN: Logs without trace context
logger.info("Payment processed");

// ✅ REQUIRED: Include trace_id in every log
const span = trace.getActiveSpan();
logger.info({
  trace_id: span?.spanContext().traceId,
  span_id: span?.spanContext().spanId,
  order_id: "ord_123",
}, "Payment processed");

// Output includes correlation:
// {"trace_id":"abc123","span_id":"def456","order_id":"ord_123","msg":"Payment processed"}

Quick Reference

When to Use What
ScenarioTool/PatternReason
Metrics collectionPrometheus + GrafanaIndustry standard, powerful query language
Distributed tracingOpenTelemetry + Tempo/JaegerVendor-neutral, CNCF standard
Log aggregation (cost-sensitive)Grafana LokiIndexes only labels, 10x cheaper
Log aggregation (search-heavy)ELK StackFull-text search, advanced analytics
Unified observabilityElastic/Datadog/DynatraceSingle pane of glass for all telemetry
Incident managementPagerDuty/OpsgenieAlert routing, on-call scheduling
Chaos engineeringGremlin/Chaos MeshControlled failure injection
AIOps/Anomaly detectionDynatrace/DatadogAI-driven root cause analysis
Show full SKILL.md (168 more words)Show less
The Three Pillars
PillarWhatWhenTools
MetricsNumerical time-series dataReal-time monitoring, alertingPrometheus, StatsD, CloudWatch
LogsEvent records with contextDebugging, audit trailsLoki, ELK, Splunk
TracesRequest journey across servicesPerformance analysis, dependenciesOpenTelemetry, Jaeger, Zipkin

Fourth Pillar (Emerging): Continuous Profiling — Code-level performance data (CPU, memory usage at function level)


Observability Architecture

Layered Prometheus Setup
yaml
# 2025 Best Practice: Federated architecture
# Prevents metric chaos while enabling drill-down

# Layer 1: Application Prometheus
# - Detailed business logic metrics
# - High cardinality acceptable
# - Short retention (7 days)

# Layer 2: Cluster Prometheus
# - Per-environment/cluster metrics
# - Medium retention (30 days)
# - Aggregates from application level

# Layer 3: Global Prometheus
# - Cross-cluster critical metrics
# - Long retention (1 year)
# - Federation from cluster level

# Global Prometheus config
scrape_configs:
  - job_name: 'federate'
    scrape_interval: 15s
    honor_labels: true
    metrics_path: '/federate'
    params:
      'match[]':
        - '{job="kubernetes-nodes"}'
        - '{__name__=~"job:.*"}'  # Recording rules only
    static_configs:
      - targets:
        - 'cluster-prom-us-east.internal:9090'
        - 'cluster-prom-eu-west.internal:9090'
Recording Rules for Performance
yaml
# Precompute expensive queries
groups:
  - name: api_performance
    interval: 30s
    rules:
      # Request rate (requests per second)
      - record: job:api_requests:rate5m
        expr: sum(rate(http_requests_total[5m])) by (job, method, status)

      # Error rate
      - record: job:api_errors:rate5m
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[5m])) by (job)
          /
          sum(rate(http_requests_total[5m])) by (job)

      # P95 latency
      - record: job:api_latency:p95
        expr: histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (job, le))
Resource Optimization
yaml
# Increase scrape interval for high-target deployments
scrape_interval: 30s  # Default: 15s reduces load by 50%

# Use relabeling to drop unnecessary metrics
metric_relabel_configs:
  - source_labels: [__name__]
    regex: 'go_.*|process_.*'  # Drop Go runtime metrics
    action: drop

# Limit sample retention
storage:
  tsdb:
    retention.time: 15d  # Keep only 15 days locally
    retention.size: 50GB # Or max 50GB

Distributed Tracing with OpenTelemetry

Auto-Instrumentation Setup
typescript
// Node.js auto-instrumentation
import { NodeSDK } from '@opentelemetry/sdk-node';
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http';

const sdk = new NodeSDK({
  traceExporter: new OTLPTraceExporter({
    url: 'http://otel-collector:4318/v1/traces',
  }),
  instrumentations: [
    getNodeAutoInstrumentations({
      // Auto-instruments HTTP, Express, PostgreSQL, Redis, etc.
      '@opentelemetry/instrumentation-fs': { enabled: false }, // Too noisy
    }),
  ],
});

sdk.start();
Manual Instrumentation for Business Logic
typescript
import { trace, SpanStatusCode } from '@opentelemetry/api';

const tracer = trace.getTracer('payment-service', '1.0.0');

async function processPayment(orderId: string, amount: number) {
  // Create custom span for business operation
  return tracer.startActiveSpan('processPayment', async (span) => {
    try {
      // Add business context
      span.setAttributes({
        'order.id': orderId,
        'payment.amount': amount,
        'payment.currency': 'USD',
      });

      // Child span for external API call
      const paymentResult = await tracer.startActiveSpan('stripe.charge', async (childSpan) => {
        const result = await stripe.charges.create({ amount, currency: 'usd' });
        childSpan.setAttribute('stripe.charge_id', result.id);
        childSpan.setStatus({ code: SpanStatusCode.OK });
        childSpan.end();
        return result;
      });

      span.setStatus({ code: SpanStatusCode.OK });
      return paymentResult;
    } catch (error) {
      span.recordException(error);
      span.setStatus({ code: SpanStatusCode.ERROR, message: error.message });
      throw error;
    } finally {
      span.end();
    }
  });
}
Sampling Strategies
yaml
# OpenTelemetry Collector config
processors:
  # Probabilistic sampling: Keep 10% of traces
  probabilistic_sampler:
    sampling_percentage: 10

  # Tail sampling: Make decisions after seeing full trace
  tail_sampling:
    policies:
      # Always sample errors
      - name: error-traces
        type: status_code
        status_code: {status_codes: [ERROR]}

      # Always sample slow requests
      - name: slow-traces
        type: latency
        latency: {threshold_ms: 1000}

      # Sample 5% of normal traffic
      - name: normal-traces
        type: probabilistic
        probabilistic: {sampling_percentage: 5}
Context Propagation
typescript
// Ensure trace context flows across services
import { propagation, context } from '@opentelemetry/api';

// Outgoing HTTP request (automatic with auto-instrumentation)
fetch('https://api.example.com/data', {
  headers: {
    // W3C Trace Context headers injected automatically:
    // traceparent: 00-<trace-id>-<span-id>-01
    // tracestate: vendor=value
  },
});

// Manual propagation for non-HTTP (e.g., message queues)
const carrier = {};
propagation.inject(context.active(), carrier);
await publishMessage(queue, { data: payload, headers: carrier });

Structured Logging Best Practices

JSON Logging Format
typescript
// Use structured logging library
import pino from 'pino';

const logger = pino({
  level: process.env.LOG_LEVEL || 'info',
  formatters: {
    level: (label) => ({ level: label }),
  },
  timestamp: pino.stdTimeFunctions.isoTime,
  // Include trace context in logs
  mixin() {
    const span = trace.getActiveSpan();
    if (!span) return {};

    const { traceId, spanId } = span.spanContext();
    return {
      trace_id: traceId,
      span_id: spanId,
    };
  },
});

// Structured logging with context
logger.info(
  {
    user_id: '123',
    order_id: 'ord_456',
    amount: 99.99,
    payment_method: 'card',
  },
  'Payment processed successfully'
);

// Output:
// {"level":"info","time":"2025-01-15T10:30:00.000Z","trace_id":"abc123","span_id":"def456","user_id":"123","order_id":"ord_456","amount":99.99,"payment_method":"card","msg":"Payment processed successfully"}
Log Levels
typescript
// Follow standard severity levels
logger.trace({ details }, 'Low-level debugging');     // Very verbose
logger.debug({ state }, 'Debug information');          // Development
logger.info({ event }, 'Normal operation');            // Production default
logger.warn({ issue }, 'Warning condition');           // Potential issues
logger.error({ error, context }, 'Error occurred');    // Errors
logger.fatal({ critical }, 'Fatal error');             // Process crash
Grafana Loki Configuration
yaml
# Promtail config - ships logs to Loki
server:
  http_listen_port: 9080

positions:
  filename: /tmp/positions.yaml

clients:
  - url: http://loki:3100/loki/api/v1/push

scrape_configs:
  - job_name: kubernetes
    kubernetes_sd_configs:
      - role: pod
    relabel_configs:
      # Add pod labels as Loki labels (LOW cardinality only!)
      - source_labels: [__meta_kubernetes_namespace]
        target_label: namespace
      - source_labels: [__meta_kubernetes_pod_name]
        target_label: pod
      - source_labels: [__meta_kubernetes_pod_label_app]
        target_label: app
    pipeline_stages:
      # Parse JSON logs
      - json:
          expressions:
            level: level
            trace_id: trace_id
      # Extract fields as labels
      - labels:
          level:
          trace_id:
Loki Best Practices
  • Low Cardinality Labels — Use only 5-10 labels (namespace, app, level)
  • High Cardinality in Log Body — Put user_id, order_id in JSON, not labels
  • LogQL for Filtering — Use {app="api"} | json | user_id="123"
  • Retention Policy — Keep recent logs longer, compress old logs
promql
# LogQL query examples
{namespace="production", app="api"} |= "error"  # Text search

{app="api"} | json | level="error" | line_format "{{.msg}}"  # JSON parsing

rate({app="api"}[5m])  # Log rate per second

sum by (level) (count_over_time({namespace="production"}[1h]))  # Count by level

Extended Reference

Detailed material starting at ## SLO/SLI/SLA Management has been moved to reference/extended.md to keep this skill concise. Load that reference when the task requires the moved examples, command catalogs, checklists, platform details, or implementation templates.

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files in skills/observability-sre of majiayu000/spellbook.

  • SKILL.md
  • reference/extended.md
  • reference/incident-response.md
  • reference/logging.md
  • reference/monitoring.md
  • reference/tracing.md
  • templates/slo-template.md

Open the folder on GitHubat commit ed52af7

Compare with similar skills

Observability Sre next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Observability Sre compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Observability Sre this skillmajiayu000/spellbook287—~3.3kAutomated safety check: PassMIT
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Observability MonitoringAnastasiyaW/codex-claude-code-config154—~4.1kAutomated safety check: PassMIT
Observability Patternssoftspark/ai-toolkit179—~2.2kAutomated safety check: PassApache-2.0
Telemetrymagnus919/agent-skills115—~3.9kAutomated safety check: PassMIT
Archestra Dev Observabilityarchestra-ai/archestra4.4k—~1.2kAutomated safety check: PassCustom licence

Similar skills

  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 6 mo ago
    DevOps & CloudAuto-check passed
  • Observability Monitoring

    AnastasiyaW/codex-claude-code-config

    Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…

    154 GitHub stars~4.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Observability Patterns

    softspark/ai-toolkit

    Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI.

    179 GitHub stars~2.2k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Telemetry

    magnus919/agent-skills

    Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector…

    115 GitHub stars~3.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Archestra Dev Observability

    archestra-ai/archestra

    A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup.

    4.4k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Frontmcp Observability

    agentfront/frontmcp

    A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server.

    146 GitHub stars~4.6k tokensUpdated today
    DevOps & CloudAuto-check passed

More from majiayu000/spellbook

All 97 skills in this repo
  • Skill Ecosystem Doctor

    majiayu000/spellbook

    Audits and repairs how coding-agent Skills are owned, copied and exposed across runtimes, from canonical sources to quarantine and retirement.

    287 GitHub stars~3k tokensUpdated 2 days ago
    Auto-check passed
  • AGENTS.md Scaffold

    majiayu000/spellbook

    Scans a repository for real evidence and proposes, or on request writes, a small stack of root and scoped AGENTS.md files with validation commands and generated-file boundaries.

    287 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • Product Demo Builder

    majiayu000/spellbook

    Plans, produces or diagnoses evidence-backed product demo videos: script, capture plan, pacing checks and verified final media built on real product behavior.

    287 GitHub stars~3.3k tokensUpdated 2 days ago
    Auto-check passed
  • Flowguard Task Guard

    majiayu000/spellbook

    Single entry point that routes long or ambiguous agent tasks, checks live state, bounds autonomous loops and leaves a resumable handoff.

    287 GitHub stars~2.1k tokensUpdated 2 days ago
    Auto-check passed
  • npm Supply Chain Check

    majiayu000/spellbook

    Scans a repository, its lockfiles and node_modules for known malicious npm package versions and install-time indicators, using a read-only Python scanner.

    287 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • Product Manager Toolkit

    majiayu000/spellbook

    Product management helpers: a RICE scoring script, an interview transcript analyzer and PRD templates for prioritizing features, synthesizing research and writing requirements.

    287 GitHub stars~2.2k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Observability Sre

What does Observability Sre do?

Observability and SRE expert. An agent skill from majiayu000/spellbook. Observability Sre is an agent skill from majiayu000/spellbook. Observability and SRE expert.

When should I use Observability Sre?

Observability Sre fits situations like: setting up monitoring; managing incidents.

How do I install Observability Sre in Claude Code?

Run `npx skills add majiayu000/spellbook --skill observability-sre -a claude-code`. Or copy the skill folder (skills/observability-sre in majiayu000/spellbook) into .claude/skills/observability-sre in your project. Claude Code loads it when a task matches its description.

How do I install Observability Sre in Codex?

Run `npx skills add majiayu000/spellbook --skill observability-sre -a codex`. Or copy the skill folder (skills/observability-sre in majiayu000/spellbook) into .agents/skills/observability-sre in your project. Codex loads it when a task matches its description.

Can I use Observability Sre in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/spellbook --skill observability-sre -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/observability-sre, .gemini/skills/observability-sre, .github/skills/observability-sre and .opencode/skills/observability-sre in your project.

What does Observability Sre need to run?

SKILL.md names no scripts, command-line tools or credentials: Observability Sre is instructions for the agent only. Our summary lists: Node.js.

Does Observability Sre access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Observability Sre safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Observability Sre use?

Observability Sre is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Observability Sre use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Observability Sre?

Skills that share tags, products or a category with Observability Sre: Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Observability Monitoring (AnastasiyaW/codex-claude-code-config, 154 stars), Observability Patterns (softspark/ai-toolkit, 179 stars) and Telemetry (magnus919/agent-skills, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Observability Sre?

majiayu000 (a GitHub user) maintains it in majiayu000/spellbook, which has 287 GitHub stars. The repository holds 97 skills in this directory. The repository was last updated on October 8, 2026.

Source: majiayu000/spellbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.