Agent skill

Tsh Implementing Observability

by TheSoftwareHouse in TheSoftwareHouse/copilot-collections

Observability patterns for logging, monitoring, alerting, and distributed tracing.

MITAuto-check passedDevOps & Cloud

Install Tsh Implementing Observability

skills CLI
$ npx skills add TheSoftwareHouse/copilot-collections --skill tsh-implementing-observability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TheSoftwareHouse/copilot-collections tsh-implementing-observability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TheSoftwareHouse/copilot-collections.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/tsh-implementing-observability .claude/skills/tsh-implementing-observability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tsh-implementing-observability
GitHub stars
284
Token cost
~2k tokens
SKILL.md length
557 words
Files
1
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Observability patterns for logging, monitoring, alerting, and distributed tracing.

  • Works in 8 steps: Discover context → Check existing… → Choose stack → Use decision matrix based… → Instrument apps → Add OpenTelemetry SDK… → …
  • Implementing metrics collection
  • SKILL.md covers When to Use, Three Pillars of Observability, Stack Detection and Solution Decision Matrix, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tsh Implementing Observability is an agent skill from TheSoftwareHouse/copilot-collections. Observability patterns for logging, monitoring, alerting, and distributed tracing. Use when implementing metrics collection, log aggregation, alerting rules, or distributed tracing across services.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Observability and Monitoring and alerting. It works with Prometheus, Kubernetes, OpenTelemetry and Datadog. The repository describes itself as: Opinionated AI-enabled workflows for product engineering. The licence is MIT.

When your agent uses it

  • Implementing metrics collection
  • Log aggregation
  • Distributed tracing across services

Example prompts

  • “/tsh-implementing-observability”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Discover context → Check existing observability setup (Prometheus, CloudWatch, etc.)
  2. Choose stack → Use decision matrix based on environment and requirements
  3. Instrument apps → Add OpenTelemetry SDK or auto-instrumentation
  4. Configure collection → Set up collectors, exporters, and storage
  5. Define SLOs → Establish SLIs, targets, and error budgets
  6. Create alerts → Implement actionable alerts with runbooks
  7. Build dashboards → Create service and infrastructure dashboards
  8. Document runbooks → Write response procedures for each alert

What it can do on your machine

Read from SKILL.md and the folder at commit 2fbe51e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tsh Implementing Observability loads about 2k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 557 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from TheSoftwareHouse/copilot-collections at commit 2fbe51e, republished under its MIT licence (© TheSoftwareHouse). 557 words, ~1,982 tokens.

Download SKILL.mdSave it as .claude/skills/tsh-implementing-observability/SKILL.md (or your agent's skills folder).
name
tsh-implementing-observability
description
Observability patterns for logging, monitoring, alerting, and distributed tracing. Use when implementing metrics collection, log aggregation, alerting rules, or distributed tracing across services.
user-invocable
false

Observability Patterns

When to Use

  • Setting up monitoring and alerting for applications
  • Implementing centralized logging
  • Adding distributed tracing to microservices
  • Designing SLOs/SLIs and error budgets
  • Creating dashboards and runbooks

Three Pillars of Observability

PillarPurposeTools
MetricsQuantitative measurements over timePrometheus, CloudWatch, Datadog, Grafana
LogsDiscrete events with contextELK, Loki, CloudWatch Logs, Splunk
TracesRequest flow across servicesJaeger, Zipkin, X-Ray, Tempo

Stack Detection

Check which observability stack the project uses:

  • prometheus.yml or ServiceMonitor → Prometheus
  • fluent-bit.conf or fluentd.conf → Fluent Bit/Fluentd
  • otel-collector-config.yaml → OpenTelemetry
  • AWS with aws_cloudwatch_* resources → CloudWatch
  • datadog-agent or DD_* env vars → Datadog

Use context7 to look up stack-specific configuration syntax.

Solution Decision Matrix

Metrics Stack
ScenarioRecommended Solution
Kubernetes-native, cost-sensitivePrometheus + Grafana
AWS-native, simple setupCloudWatch Metrics
Multi-cloud, enterpriseDatadog or New Relic
OpenTelemetry-firstPrometheus with OTLP receiver
Logging Stack
ScenarioRecommended Solution
Kubernetes, cost-sensitiveLoki + Grafana
AWS-nativeCloudWatch Logs
High volume, complex queriesElasticsearch (ELK)
Multi-cloud, managedDatadog Logs or Splunk
Tracing Stack
ScenarioRecommended Solution
Kubernetes, open-sourceJaeger or Tempo
AWS-nativeX-Ray
Multi-cloud, correlatedDatadog APM
Vendor-agnosticOpenTelemetry → any backend

Kubernetes Observability Pattern

┌─────────────────────────────────────────────────────┐
│                   Applications                      │
│  (instrumented with OpenTelemetry SDK or auto-inst) │
└──────────────────────┬──────────────────────────────┘
                       │ OTLP
                       ▼
┌─────────────────────────────────────────────────────┐
│            OpenTelemetry Collector                  │
│  (receives, processes, exports telemetry)           │
└───────┬─────────────────┬─────────────────┬─────────┘
        │                 │                 │
        ▼                 ▼                 ▼
   Prometheus          Loki             Tempo/Jaeger
   (metrics)          (logs)            (traces)
        │                 │                 │
        └────────────────┬┴─────────────────┘
                         ▼
                      Grafana
                   (visualization)

SLO/SLI Framework

Key Metrics (RED Method for Services)
MetricDescriptionExample SLI
RateRequests per secondrate(http_requests_total[5m])
ErrorsFailed requestsrate(http_requests_total{status=~"5.."}[5m])
DurationLatency distributionhistogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))
Key Metrics (USE Method for Resources)
MetricDescriptionExample
Utilization% time resource is busyCPU usage, memory usage
SaturationQueue depth, waitingPod pending, connection pool
ErrorsError countOOM kills, disk errors
SLO Definition Template
yaml
# Example: API availability SLO
slo:
  name: api-availability
  description: "API returns successful responses"
  sli:
    metric: |
      sum(rate(http_requests_total{status!~"5.."}[5m]))
      /
      sum(rate(http_requests_total[5m]))
  target: 99.9%
  window: 30d
  error_budget: 0.1%  # ~43 minutes/month downtime allowed

Alerting Strategy

Alert Severity Levels
SeverityResponseExample
CriticalPage on-call immediatelyService down, data loss risk
WarningInvestigate within hoursError rate elevated, disk 80%
InfoReview during business hoursDeployment completed, scaling event
Alert Quality Rules
  • Actionable: Every alert must have a clear response action
  • Relevant: Alert on symptoms (user impact), not causes
  • Unique: Avoid duplicate alerts for same incident
  • Timely: Alert early enough to prevent impact
Alert Template (Prometheus)
yaml
groups:
  - name: api-alerts
    rules:
      - alert: HighErrorRate
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[5m]))
          /
          sum(rate(http_requests_total[5m])) > 0.01
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "High error rate detected"
          description: "Error rate is {{ $value | humanizePercentage }} (threshold: 1%)"
          runbook_url: "https://runbooks.example.com/high-error-rate"

Structured Logging

Log Format (JSON)
json
{
  "timestamp": "2024-01-15T10:30:00Z",
  "level": "error",
  "message": "Payment processing failed",
  "service": "payment-api",
  "trace_id": "abc123",
  "span_id": "def456",
  "user_id": "user-789",
  "error": {
    "type": "PaymentGatewayError",
    "message": "Connection timeout"
  },
  "context": {
    "payment_id": "pay-123",
    "amount": 99.99
  }
}
Required Log Fields
FieldPurposeCorrelation
timestampWhen event occurredTime-based queries
levelSeverity (debug/info/warn/error)Filtering
serviceSource service nameService filtering
trace_idDistributed trace identifierCross-service correlation
messageHuman-readable descriptionSearch
Show full SKILL.md (208 more words)Show less

Process

  1. Discover context → Check existing observability setup (Prometheus, CloudWatch, etc.)
  2. Choose stack → Use decision matrix based on environment and requirements
  3. Instrument apps → Add OpenTelemetry SDK or auto-instrumentation
  4. Configure collection → Set up collectors, exporters, and storage
  5. Define SLOs → Establish SLIs, targets, and error budgets
  6. Create alerts → Implement actionable alerts with runbooks
  7. Build dashboards → Create service and infrastructure dashboards
  8. Document runbooks → Write response procedures for each alert

Checklist

  • All services emit metrics, logs, and traces
  • Trace IDs propagated across service boundaries
  • Structured logging with consistent format (JSON)
  • SLOs defined with error budgets
  • Alerts are actionable with runbook links
  • Dashboards show service health at a glance
  • Log retention policy configured
  • PII/sensitive data excluded from logs
  • On-call rotation defined for critical alerts

Anti-Patterns

Don'tDo
Alert on every metric thresholdAlert on user-impacting symptoms
Log everything at DEBUG in productionUse appropriate log levels
Unstructured log messagesStructured JSON logging
Missing trace contextPropagate trace IDs across services
Dashboards with 50+ panelsFocused dashboards per service/domain
Alerts without runbooksEvery alert links to response procedure
Store logs indefinitelyDefine retention based on compliance needs
  • tsh-implementing-kubernetes - For K8s-native observability setup
  • tsh-implementing-ci-cd - For pipeline observability integration
  • tsh-managing-secrets - For secure credential storage for observability tools

© TheSoftwareHouse, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .github/skills/tsh-implementing-observability of TheSoftwareHouse/copilot-collections.

Open the folder on GitHubat commit 2fbe51e

Compare with similar skills

Tsh Implementing Observability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tsh Implementing Observability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tsh Implementing Observability this skillTheSoftwareHouse/copilot-collections284—~2kAutomated safety check: PassMIT
Frontmcp Observabilityagentfront/frontmcp146—~4.6kAutomated safety check: PassApache-2.0
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Alloygrafana/skills279—~1.3kAutomated safety check: PassApache-2.0
ObservabilityTheBeardedBearSAS/claude-craft107—~547Automated safety check: PassMIT
Archestra Dev Observabilityarchestra-ai/archestra4.4k—~1.2kAutomated safety check: PassCustom licence

Similar skills

  • Frontmcp Observability

    agentfront/frontmcp

    A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server.

    146 GitHub stars~4.6k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 5 mo ago
    DevOps & CloudAuto-check passed
  • Alloy

    grafana/skills

    Official

    Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /…

    279 GitHub stars~1.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Observability

    TheBeardedBearSAS/claude-craft

    OpenTelemetry, distributed tracing, structured logging, metrics (Prometheus, Grafana, Datadog).

    107 GitHub stars~547 tokensUpdated 23 days ago
    DevOps & CloudAuto-check passed
  • Archestra Dev Observability

    archestra-ai/archestra

    A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup.

    4.4k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Observability Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs.

    117 GitHub stars~1.3k tokensUpdated today
    DevOps & CloudAuto-check passed

More from TheSoftwareHouse/copilot-collections

All 21 skills in this repo
  • Tsh Creating Skills

    TheSoftwareHouse/copilot-collections

    Create new skills (SKILL.md) for GitHub Copilot. An agent skill from TheSoftwareHouse/copilot-collections.

    284 GitHub stars~4.1k tokensUpdated 2 days ago
    Auto-check passed
  • Tsh Implementing Frontend

    TheSoftwareHouse/copilot-collections

    Frontend component patterns, composition, design token integration, barrel file organization, error handling, and Figma-to-code workflow.

    284 GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed
  • Tsh Implementing Terraform Modules

    TheSoftwareHouse/copilot-collections

    Build reusable Terraform modules for AWS, Azure, and GCP infrastructure following infrastructure-as-code best practices.

    284 GitHub stars~1.6k tokensUpdated 2 days ago
    Auto-check passed
  • Tsh Optimizing Frontend

    TheSoftwareHouse/copilot-collections

    Frontend rendering optimization, code splitting, memoization strategies, bundle size control, asset optimization, and memory management.

    284 GitHub stars~4.1k tokensUpdated 2 days ago
    Auto-check passed
  • Tsh Reviewing Frontend

    TheSoftwareHouse/copilot-collections

    Frontend-specific code review criteria, component anti-patterns, hooks quality, rendering correctness, accessibility and performance spot-checks, and module organization issues.

    284 GitHub stars~4.4k tokensUpdated 2 days ago
    Auto-check passed
  • Tsh Writing Hooks

    TheSoftwareHouse/copilot-collections

    Custom hook and composable patterns — naming, composition, stable return shapes, lifecycle cleanup, and testing strategies.

    284 GitHub stars~3k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Tsh Implementing Observability

What does Tsh Implementing Observability do?

Observability patterns for logging, monitoring, alerting, and distributed tracing. Tsh Implementing Observability is an agent skill from TheSoftwareHouse/copilot-collections. Observability patterns for logging, monitoring, alerting, and distributed tracing.

When should I use Tsh Implementing Observability?

Tsh Implementing Observability fits situations like: implementing metrics collection; log aggregation; distributed tracing across services.

How do I install Tsh Implementing Observability in Claude Code?

Run `npx skills add TheSoftwareHouse/copilot-collections --skill tsh-implementing-observability -a claude-code`. Or copy the skill folder (.github/skills/tsh-implementing-observability in TheSoftwareHouse/copilot-collections) into .claude/skills/tsh-implementing-observability in your project. Claude Code loads it when a task matches its description.

How do I install Tsh Implementing Observability in Codex?

Run `npx skills add TheSoftwareHouse/copilot-collections --skill tsh-implementing-observability -a codex`. Or copy the skill folder (.github/skills/tsh-implementing-observability in TheSoftwareHouse/copilot-collections) into .agents/skills/tsh-implementing-observability in your project. Codex loads it when a task matches its description.

Can I use Tsh Implementing Observability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TheSoftwareHouse/copilot-collections --skill tsh-implementing-observability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tsh-implementing-observability, .gemini/skills/tsh-implementing-observability, .github/skills/tsh-implementing-observability and .opencode/skills/tsh-implementing-observability in your project.

What does Tsh Implementing Observability need to run?

SKILL.md names no scripts, command-line tools or credentials: Tsh Implementing Observability is instructions for the agent only.

Does Tsh Implementing Observability access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tsh Implementing Observability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tsh Implementing Observability use?

Tsh Implementing Observability is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tsh Implementing Observability use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tsh Implementing Observability?

Skills that share tags, products or a category with Tsh Implementing Observability: Frontmcp Observability (agentfront/frontmcp, 146 stars), Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Alloy (grafana/skills, 279 stars) and Observability (TheBeardedBearSAS/claude-craft, 107 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tsh Implementing Observability?

TheSoftwareHouse (a GitHub organization) maintains it in TheSoftwareHouse/copilot-collections, which has 284 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 5, 2026.

Source: TheSoftwareHouse/copilot-collections on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.