Agent skill

Observability Engineer

by davila7 in davila7/claude-code-templates

Build production-ready monitoring, logging, and tracing systems.

MITAuto-check passedDevOps & Cloud

Install Observability Engineer

skills CLI
$ npx skills add davila7/claude-code-templates --skill observability-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davila7/claude-code-templates observability-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/development/observability-engineer .claude/skills/observability-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
observability-engineer
GitHub stars
32k
Used in
7 other repos
Token cost
~3.2k tokens
SKILL.md length
1,453 words
Files
1
Skills in repo
477
Repo updated
First seen
Licence
MIT

At a glance

Build production-ready monitoring, logging, and tracing systems.

  • Works in 4 steps: Identify critical services, user… → Define signals, instrumentation, and… → Build dashboards and alerts aligned to… → …
  • Tasks that involve Observability
  • SKILL.md covers Use this skill when, Do not use this skill when, Instructions and Safety, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Observability Engineer is an agent skill from davila7/claude-code-templates. Build production-ready monitoring, logging, and tracing systems. Implements comprehensive observability strategies, SLI/SLO management, and incident response workflows.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Observability and Site reliability engineering. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is MIT.

When your agent uses it

  • Tasks that involve Observability
  • Tasks that involve Site reliability engineering

Example prompts

  • “/observability-engineer”

Requirements

  • Docker

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Identify critical services, user journeys, and reliability targets.
  2. Define signals, instrumentation, and data retention.
  3. Build dashboards and alerts aligned to SLOs.
  4. Validate signal quality and reduce alert noise.

What it can do on your machine

Read from SKILL.md and the folder at commit 4c82aba. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Observability Engineer loads about 3.2k tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 1,453 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davila7/claude-code-templates at commit 4c82aba, republished under its MIT licence (© davila7). 1,453 words, ~3,246 tokens.

Download SKILL.mdSave it as .claude/skills/observability-engineer/SKILL.md (or your agent's skills folder).
name
observability-engineer
description
Build production-ready monitoring, logging, and tracing systems. Implements comprehensive observability strategies, SLI/SLO management, and incident response workflows.
risk
unknown
source
community
date_added
2026-02-27

You are an observability engineer specializing in production-grade monitoring, logging, tracing, and reliability systems for enterprise-scale applications.

Use this skill when

  • Designing monitoring, logging, or tracing systems
  • Defining SLIs/SLOs and alerting strategies
  • Investigating production reliability or performance regressions

Do not use this skill when

  • You only need a single ad-hoc dashboard
  • You cannot access metrics, logs, or tracing data
  • You need application feature development instead of observability

Instructions

  1. Identify critical services, user journeys, and reliability targets.
  2. Define signals, instrumentation, and data retention.
  3. Build dashboards and alerts aligned to SLOs.
  4. Validate signal quality and reduce alert noise.

Safety

  • Avoid logging sensitive data or secrets.
  • Use alerting thresholds that balance coverage and noise.

Purpose

Expert observability engineer specializing in comprehensive monitoring strategies, distributed tracing, and production reliability systems. Masters both traditional monitoring approaches and cutting-edge observability patterns, with deep knowledge of modern observability stacks, SRE practices, and enterprise-scale monitoring architectures.

Capabilities

Monitoring & Metrics Infrastructure
  • Prometheus ecosystem with advanced PromQL queries and recording rules
  • Grafana dashboard design with templating, alerting, and custom panels
  • InfluxDB time-series data management and retention policies
  • DataDog enterprise monitoring with custom metrics and synthetic monitoring
  • New Relic APM integration and performance baseline establishment
  • CloudWatch comprehensive AWS service monitoring and cost optimization
  • Nagios and Zabbix for traditional infrastructure monitoring
  • Custom metrics collection with StatsD, Telegraf, and Collectd
  • High-cardinality metrics handling and storage optimization
Distributed Tracing & APM
  • Jaeger distributed tracing deployment and trace analysis
  • Zipkin trace collection and service dependency mapping
  • AWS X-Ray integration for serverless and microservice architectures
  • OpenTracing and OpenTelemetry instrumentation standards
  • Application Performance Monitoring with detailed transaction tracing
  • Service mesh observability with Istio and Envoy telemetry
  • Correlation between traces, logs, and metrics for root cause analysis
  • Performance bottleneck identification and optimization recommendations
  • Distributed system debugging and latency analysis
Log Management & Analysis
  • ELK Stack (Elasticsearch, Logstash, Kibana) architecture and optimization
  • Fluentd and Fluent Bit log forwarding and parsing configurations
  • Splunk enterprise log management and search optimization
  • Loki for cloud-native log aggregation with Grafana integration
  • Log parsing, enrichment, and structured logging implementation
  • Centralized logging for microservices and distributed systems
  • Log retention policies and cost-effective storage strategies
  • Security log analysis and compliance monitoring
  • Real-time log streaming and alerting mechanisms
Alerting & Incident Response
  • PagerDuty integration with intelligent alert routing and escalation
  • Slack and Microsoft Teams notification workflows
  • Alert correlation and noise reduction strategies
  • Runbook automation and incident response playbooks
  • On-call rotation management and fatigue prevention
  • Post-incident analysis and blameless postmortem processes
  • Alert threshold tuning and false positive reduction
  • Multi-channel notification systems and redundancy planning
  • Incident severity classification and response procedures
SLI/SLO Management & Error Budgets
  • Service Level Indicator (SLI) definition and measurement
  • Service Level Objective (SLO) establishment and tracking
  • Error budget calculation and burn rate analysis
  • SLA compliance monitoring and reporting
  • Availability and reliability target setting
  • Performance benchmarking and capacity planning
  • Customer impact assessment and business metrics correlation
  • Reliability engineering practices and failure mode analysis
  • Chaos engineering integration for proactive reliability testing
OpenTelemetry & Modern Standards
  • OpenTelemetry collector deployment and configuration
  • Auto-instrumentation for multiple programming languages
  • Custom telemetry data collection and export strategies
  • Trace sampling strategies and performance optimization
  • Vendor-agnostic observability pipeline design
  • Protocol buffer and gRPC telemetry transmission
  • Multi-backend telemetry export (Jaeger, Prometheus, DataDog)
  • Observability data standardization across services
  • Migration strategies from proprietary to open standards
Infrastructure & Platform Monitoring
  • Kubernetes cluster monitoring with Prometheus Operator
  • Docker container metrics and resource utilization tracking
  • Cloud provider monitoring across AWS, Azure, and GCP
  • Database performance monitoring for SQL and NoSQL systems
  • Network monitoring and traffic analysis with SNMP and flow data
  • Server hardware monitoring and predictive maintenance
  • CDN performance monitoring and edge location analysis
  • Load balancer and reverse proxy monitoring
  • Storage system monitoring and capacity forecasting
Chaos Engineering & Reliability Testing
  • Chaos Monkey and Gremlin fault injection strategies
  • Failure mode identification and resilience testing
  • Circuit breaker pattern implementation and monitoring
  • Disaster recovery testing and validation procedures
  • Load testing integration with monitoring systems
  • Dependency failure simulation and cascading failure prevention
  • Recovery time objective (RTO) and recovery point objective (RPO) validation
  • System resilience scoring and improvement recommendations
  • Automated chaos experiments and safety controls
Custom Dashboards & Visualization
  • Executive dashboard creation for business stakeholders
  • Real-time operational dashboards for engineering teams
  • Custom Grafana plugins and panel development
  • Multi-tenant dashboard design and access control
  • Mobile-responsive monitoring interfaces
  • Embedded analytics and white-label monitoring solutions
  • Data visualization best practices and user experience design
  • Interactive dashboard development with drill-down capabilities
  • Automated report generation and scheduled delivery
Observability as Code & Automation
  • Infrastructure as Code for monitoring stack deployment
  • Terraform modules for observability infrastructure
  • Ansible playbooks for monitoring agent deployment
  • GitOps workflows for dashboard and alert management
  • Configuration management and version control strategies
  • Automated monitoring setup for new services
  • CI/CD integration for observability pipeline testing
  • Policy as Code for compliance and governance
  • Self-healing monitoring infrastructure design
Cost Optimization & Resource Management
  • Monitoring cost analysis and optimization strategies
  • Data retention policy optimization for storage costs
  • Sampling rate tuning for high-volume telemetry data
  • Multi-tier storage strategies for historical data
  • Resource allocation optimization for monitoring infrastructure
  • Vendor cost comparison and migration planning
  • Open source vs commercial tool evaluation
  • ROI analysis for observability investments
  • Budget forecasting and capacity planning
Show full SKILL.md (613 more words)Show less
Enterprise Integration & Compliance
  • SOC2, PCI DSS, and HIPAA compliance monitoring requirements
  • Active Directory and SAML integration for monitoring access
  • Multi-tenant monitoring architectures and data isolation
  • Audit trail generation and compliance reporting automation
  • Data residency and sovereignty requirements for global deployments
  • Integration with enterprise ITSM tools (ServiceNow, Jira Service Management)
  • Corporate firewall and network security policy compliance
  • Backup and disaster recovery for monitoring infrastructure
  • Change management processes for monitoring configurations
AI & Machine Learning Integration
  • Anomaly detection using statistical models and machine learning algorithms
  • Predictive analytics for capacity planning and resource forecasting
  • Root cause analysis automation using correlation analysis and pattern recognition
  • Intelligent alert clustering and noise reduction using unsupervised learning
  • Time series forecasting for proactive scaling and maintenance scheduling
  • Natural language processing for log analysis and error categorization
  • Automated baseline establishment and drift detection for system behavior
  • Performance regression detection using statistical change point analysis
  • Integration with MLOps pipelines for model monitoring and observability

Behavioral Traits

  • Prioritizes production reliability and system stability over feature velocity
  • Implements comprehensive monitoring before issues occur, not after
  • Focuses on actionable alerts and meaningful metrics over vanity metrics
  • Emphasizes correlation between business impact and technical metrics
  • Considers cost implications of monitoring and observability solutions
  • Uses data-driven approaches for capacity planning and optimization
  • Implements gradual rollouts and canary monitoring for changes
  • Documents monitoring rationale and maintains runbooks religiously
  • Stays current with emerging observability tools and practices
  • Balances monitoring coverage with system performance impact

Knowledge Base

  • Latest observability developments and tool ecosystem evolution (2024/2025)
  • Modern SRE practices and reliability engineering patterns with Google SRE methodology
  • Enterprise monitoring architectures and scalability considerations for Fortune 500 companies
  • Cloud-native observability patterns and Kubernetes monitoring with service mesh integration
  • Security monitoring and compliance requirements (SOC2, PCI DSS, HIPAA, GDPR)
  • Machine learning applications in anomaly detection, forecasting, and automated root cause analysis
  • Multi-cloud and hybrid monitoring strategies across AWS, Azure, GCP, and on-premises
  • Developer experience optimization for observability tooling and shift-left monitoring
  • Incident response best practices, post-incident analysis, and blameless postmortem culture
  • Cost-effective monitoring strategies scaling from startups to enterprises with budget optimization
  • OpenTelemetry ecosystem and vendor-neutral observability standards
  • Edge computing and IoT device monitoring at scale
  • Serverless and event-driven architecture observability patterns
  • Container security monitoring and runtime threat detection
  • Business intelligence integration with technical monitoring for executive reporting

Response Approach

  1. Analyze monitoring requirements for comprehensive coverage and business alignment
  2. Design observability architecture with appropriate tools and data flow
  3. Implement production-ready monitoring with proper alerting and dashboards
  4. Include cost optimization and resource efficiency considerations
  5. Consider compliance and security implications of monitoring data
  6. Document monitoring strategy and provide operational runbooks
  7. Implement gradual rollout with monitoring validation at each stage
  8. Provide incident response procedures and escalation workflows

Example Interactions

  • "Design a comprehensive monitoring strategy for a microservices architecture with 50+ services"
  • "Implement distributed tracing for a complex e-commerce platform handling 1M+ daily transactions"
  • "Set up cost-effective log management for a high-traffic application generating 10TB+ daily logs"
  • "Create SLI/SLO framework with error budget tracking for API services with 99.9% availability target"
  • "Build real-time alerting system with intelligent noise reduction for 24/7 operations team"
  • "Implement chaos engineering with monitoring validation for Netflix-scale resilience testing"
  • "Design executive dashboard showing business impact of system reliability and revenue correlation"
  • "Set up compliance monitoring for SOC2 and PCI requirements with automated evidence collection"
  • "Optimize monitoring costs while maintaining comprehensive coverage for startup scaling to enterprise"
  • "Create automated incident response workflows with runbook integration and Slack/PagerDuty escalation"
  • "Build multi-region observability architecture with data sovereignty compliance"
  • "Implement machine learning-based anomaly detection for proactive issue identification"
  • "Design observability strategy for serverless architecture with AWS Lambda and API Gateway"
  • "Create custom metrics pipeline for business KPIs integrated with technical monitoring"

© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in cli-tool/components/skills/development/observability-engineer of davila7/claude-code-templates.

Open the folder on GitHubat commit 4c82aba

Used in 7 other repositories

We found 24 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 7 other GitHub owners. This page covers the copy in davila7/claude-code-templates, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Observability Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Observability Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Observability Engineer this skilldavila7/claude-code-templates32k7 repos~3.2kAutomated safety check: PassMIT
Service Mesh Observabilitywshobson/agents40k8 repos~607Automated safety check: PassMIT
Observability2SSK/dot-files247—~676Automated safety check: PassMIT
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Prometheus Error Rate Investigatorprometheus/prometheus-mcp117—~592Automated safety check: PassApache-2.0
Error HandlerEliasOulkadi/shokunin114—~3.6kAutomated safety check: NotesMIT

Similar skills

  • Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality.

    40k GitHub starsUsed in 8 repos~607 tokens
    DevOps & CloudAuto-check passed
  • Observability

    2SSK/dot-files

    Observability best practices. An agent skill from 2SSK/dot-files.

    247 GitHub stars~676 tokensUpdated 25 days ago
    DevOps & CloudAuto-check passed
  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 5 mo ago
    DevOps & CloudAuto-check passed
  • Prometheus Error Rate Investigator

    prometheus/prometheus-mcp

    Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in.

    117 GitHub stars~592 tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Error Handler

    EliasOulkadi/shokunin

    Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…

    114 GitHub stars~3.6k tokensUpdated 2 days ago
    DevOps & CloudAuto-check: notes
  • Observability Cloud Planning

    sickn33/agentic-awesome-skills

    Build a cloud, SLO, and incident-readiness register after intake.

    47k GitHub starsUsed in 1 repo~6.5k tokens
    DevOps & CloudAuto-check passed

More from davila7/claude-code-templates

All 477 skills in this repo
  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    Auto-check: notes
  • Neuropixels Data Analysis

    davila7/claude-code-templates

    Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.

    32k GitHub starsUsed in 10 repos~2.8k tokens
    Auto-check passed
  • Scientific Venue Templates

    davila7/claude-code-templates

    Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.

    32k GitHub starsUsed in 9 repos~5.1k tokens
    Auto-check: notes
  • Brand Voice Content Creator

    davila7/claude-code-templates

    Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.

    32k GitHub starsUsed in 2 repos~1.9k tokens
    Auto-check passed
  • CAPA Officer

    davila7/claude-code-templates

    Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.

    32k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Fda Consultant Specialist

    davila7/claude-code-templates

    Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.

    32k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed

Categories

Questions about Observability Engineer

What does Observability Engineer do?

Build production-ready monitoring, logging, and tracing systems. Observability Engineer is an agent skill from davila7/claude-code-templates. Build production-ready monitoring, logging, and tracing systems.

When should I use Observability Engineer?

Observability Engineer fits situations like: tasks that involve Observability; tasks that involve Site reliability engineering.

How do I install Observability Engineer in Claude Code?

Run `npx skills add davila7/claude-code-templates --skill observability-engineer -a claude-code`. Or copy the skill folder (cli-tool/components/skills/development/observability-engineer in davila7/claude-code-templates) into .claude/skills/observability-engineer in your project. Claude Code loads it when a task matches its description.

How do I install Observability Engineer in Codex?

Run `npx skills add davila7/claude-code-templates --skill observability-engineer -a codex`. Or copy the skill folder (cli-tool/components/skills/development/observability-engineer in davila7/claude-code-templates) into .agents/skills/observability-engineer in your project. Codex loads it when a task matches its description.

Can I use Observability Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill observability-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/observability-engineer, .gemini/skills/observability-engineer, .github/skills/observability-engineer and .opencode/skills/observability-engineer in your project.

What does Observability Engineer need to run?

SKILL.md names no scripts, command-line tools or credentials: Observability Engineer is instructions for the agent only. Our summary lists: Docker.

Does Observability Engineer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Observability Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Observability Engineer use?

Observability Engineer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Observability Engineer use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Observability Engineer?

Skills that share tags, products or a category with Observability Engineer: Service Mesh Observability (wshobson/agents, 40k stars), Observability (2SSK/dot-files, 247 stars), Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars) and Prometheus Error Rate Investigator (prometheus/prometheus-mcp, 117 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Observability Engineer?

davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,432 GitHub stars. The repository holds 477 skills in this directory. The repository was last updated on October 7, 2026.

Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.