Agent skill

Service Mesh Observability

by wshobson in wshobson/agents

Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality.

MITAuto-check passedDevOps & Cloud

Install Service Mesh Observability

skills CLI
$ npx skills add wshobson/agents --skill service-mesh-observability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents service-mesh-observability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/cloud-infrastructure/skills/service-mesh-observability .claude/skills/service-mesh-observability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
service-mesh-observability
GitHub stars
40k
Used in
9 other repos
Token cost
~607 tokens
SKILL.md length
166 words
Files
2 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality.

  • Works in 2 steps: Three Pillars of Observability → Golden Signals for Mesh
  • Setting up distributed tracing across mesh services
  • SKILL.md covers When to Use This Skill, Core Concepts, Templates and detailed worked… and Best Practices
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill covers observability for service meshes such as Istio and Linkerd: distributed tracing across services, mesh metrics and dashboards, dependency visualization and SLOs for service-to-service communication. The work is framed around the three pillars of observability and a golden-signals table for the mesh.

That table pairs latency (P50 and P99), traffic, errors and saturation with alert thresholds: P99 above 500ms, anomaly detection for traffic, an error rate above 1 percent and resource use above 80 percent. Other advice covers sampling every trace in development but only a small share in production, propagating trace context headers consistently, linking metrics to traces with exemplars, limiting label cardinality, using hot and cold storage tiers and watching what observability itself costs. Templates are in references/details.md.

When your agent uses it

  • Setting up distributed tracing across mesh services
  • Debugging latency or error spikes between services
  • Defining SLOs and alerts for service-to-service traffic
  • Building dashboards that show service dependencies

Example prompts

  • “Add distributed tracing to our Istio mesh and propagate trace headers through every service.”
  • “Create alerts for the four golden signals on our checkout service.”
  • “P99 latency doubled between the cart and payment services. Help me trace where the time goes.”

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Three Pillars of Observability
  2. Golden Signals for Mesh

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Service Mesh Observability loads about 607 tokens when it runs, and up to ~2.5k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 166 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~607
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 166 words, ~607 tokens.

Download SKILL.mdSave it as .claude/skills/service-mesh-observability/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
service-mesh-observability
description
Implement comprehensive observability for service meshes including distributed tracing, metrics, and visualization. Use when setting up mesh monitoring, debugging latency issues, or implementing SLOs for service communication.

Service Mesh Observability

Complete guide to observability patterns for Istio, Linkerd, and service mesh deployments.

When to Use This Skill

  • Setting up distributed tracing across services
  • Implementing service mesh metrics and dashboards
  • Debugging latency and error issues
  • Defining SLOs for service communication
  • Visualizing service dependencies
  • Troubleshooting mesh connectivity

Core Concepts

1. Three Pillars of Observability
┌─────────────────────────────────────────────────────┐
│                  Observability                       │
├─────────────────┬─────────────────┬─────────────────┤
│     Metrics     │     Traces      │      Logs       │
│                 │                 │                 │
│ • Request rate  │ • Span context  │ • Access logs   │
│ • Error rate    │ • Latency       │ • Error details │
│ • Latency P50   │ • Dependencies  │ • Debug info    │
│ • Saturation    │ • Bottlenecks   │ • Audit trail   │
└─────────────────┴─────────────────┴─────────────────┘
2. Golden Signals for Mesh
SignalDescriptionAlert Threshold
LatencyRequest duration P50, P99P99 > 500ms
TrafficRequests per secondAnomaly detection
Errors5xx error rate> 1%
SaturationResource utilization> 80%

Templates and detailed worked examples

Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.

Best Practices

Do's
  • Sample appropriately - 100% in dev, 1-10% in prod
  • Use trace context - Propagate headers consistently
  • Set up alerts - For golden signals
  • Correlate metrics/traces - Use exemplars
  • Retain strategically - Hot/cold storage tiers
Don'ts
  • Don't over-sample - Storage costs add up
  • Don't ignore cardinality - Limit label values
  • Don't skip dashboards - Visualize dependencies
  • Don't forget costs - Monitor observability costs

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/cloud-infrastructure/skills/service-mesh-observability of wshobson/agents.

  • SKILL.md
  • references/details.md

Open the folder on GitHubat commit 46891e7

Used in 9 other repositories

We found 19 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 9 other GitHub owners. This page covers the copy in wshobson/agents, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Service Mesh Observability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Service Mesh Observability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Service Mesh Observability this skillwshobson/agents40k9 repos~607Automated safety check: PassMIT
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Prometheus Error Rate Investigatorprometheus/prometheus-mcp121—~592Automated safety check: PassApache-2.0
Oma Observabilityfirst-fluke/oh-my-agent1.3k—~4.9kAutomated safety check: PassMIT
Observability Sre Triageelastic/agent-skills592—~7.4kAutomated safety check: PassApache-2.0
Observability Service Healthaspectrr/deer405—~1.2kAutomated safety check: PassMIT

Similar skills

  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 6 mo ago
    DevOps & CloudAuto-check passed
  • Prometheus Error Rate Investigator

    prometheus/prometheus-mcp

    Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in.

    121 GitHub stars~592 tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Oma Observability

    first-fluke/oh-my-agent

    Intent-based observability + traceability router across layers, boundaries, and signals.

    1.3k GitHub stars~4.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Observability Sre Triage

    elastic/agent-skills

    Official

    Triage a degraded or suspect service end to end: read SLO status and burn rate, check active alerting rules and ML anomalies, measure throughput, latency, and error rate, assess dependency health…

    592 GitHub stars~7.4k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • Assess APM service health using SLOs, alerts, ML, throughput, latency, error rate, and dependencies.

    405 GitHub stars~1.2k tokensUpdated 5 mo ago
    DevOps & CloudAuto-check passed
  • Observability Monitoring

    AnastasiyaW/codex-claude-code-config

    Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…

    154 GitHub stars~4.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 13 repos~473 tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Distributed Tracing

    wshobson/agents

    Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.

    40k GitHub starsUsed in 12 repos~527 tokens
    Auto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 6 days ago
    Auto-check passed

Categories

Questions about Service Mesh Observability

What does Service Mesh Observability do?

Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality. The skill covers observability for service meshes such as Istio and Linkerd: distributed tracing across services, mesh metrics and dashboards, dependency visualization and SLOs for service-to-service communication. The work is framed around the three pillars of observability and a golden-signals table for the mesh.

When should I use Service Mesh Observability?

Service Mesh Observability fits situations like: setting up distributed tracing across mesh services; debugging latency or error spikes between services; defining SLOs and alerts for service-to-service traffic; building dashboards that show service dependencies.

How do I install Service Mesh Observability in Claude Code?

Run `npx skills add wshobson/agents --skill service-mesh-observability -a claude-code`. Or copy the skill folder (plugins/cloud-infrastructure/skills/service-mesh-observability in wshobson/agents) into .claude/skills/service-mesh-observability in your project. Claude Code loads it when a task matches its description.

How do I install Service Mesh Observability in Codex?

Run `npx skills add wshobson/agents --skill service-mesh-observability -a codex`. Or copy the skill folder (plugins/cloud-infrastructure/skills/service-mesh-observability in wshobson/agents) into .agents/skills/service-mesh-observability in your project. Codex loads it when a task matches its description.

Can I use Service Mesh Observability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill service-mesh-observability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/service-mesh-observability, .gemini/skills/service-mesh-observability, .github/skills/service-mesh-observability and .opencode/skills/service-mesh-observability in your project.

What does Service Mesh Observability need to run?

SKILL.md names no scripts, command-line tools or credentials: Service Mesh Observability is instructions for the agent only.

Does Service Mesh Observability access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Service Mesh Observability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Service Mesh Observability use?

Service Mesh Observability is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Service Mesh Observability use?

About 607 tokens (SKILL.md is roughly 2.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.9k tokens, read only when the agent opens those files.

What are the alternatives to Service Mesh Observability?

Skills that share tags, products or a category with Service Mesh Observability: Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Prometheus Error Rate Investigator (prometheus/prometheus-mcp, 121 stars), Oma Observability (first-fluke/oh-my-agent, 1.3k stars) and Observability Sre Triage (elastic/agent-skills, 592 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Service Mesh Observability?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,314 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.