Agent skill

Agent Observability

by seb1n in seb1n/awesome-ai-agent-skills

Design privacy-aware observability for AI agents using traces, spans, structured events, metrics, cost attribution, dashboards, alerts, and investigation workflows.

MITAuto-check passedDevOps & Cloud

Install Agent Observability

skills CLI
$ npx skills add seb1n/awesome-ai-agent-skills --skill agent-observability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install seb1n/awesome-ai-agent-skills agent-observability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/seb1n/awesome-ai-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agent-engineering/agent-observability .claude/skills/agent-observability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-observability
GitHub stars
206
Token cost
~1.4k tokens
SKILL.md length
710 words
Files
4 (incl. scripts, references)
Skills in repo
91
Repo updated
First seen
Licence
MIT

At a glance

Design privacy-aware observability for AI agents using traces, spans, structured events, metrics, cost attribution, dashboards, alerts, and investigation workflows.

  • Works in 6 steps: An observability objective and system… → A trace and event schema with… → Metrics with definitions, units,… → …
  • Instrumenting an agent
  • SKILL.md covers Use when, Inputs, Output contract and Workflow, plus 4 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Agent Observability is an agent skill from seb1n/awesome-ai-agent-skills. Design privacy-aware observability for AI agents using traces, spans, structured events, metrics, cost attribution, dashboards, alerts, and investigation workflows. Use when instrumenting an agent, debugging intermittent tool or model failures, defining service-level objectives, analyzing latency or spend, auditing agent decisions, or preparing production monitoring.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/trace-schema.md` and `scripts/summarize_traces.py`).

It sits in DevOps & Cloud, covering Observability. The repository describes itself as: 103 ready-to-use AI agent skills for Claude Code, OpenAI Codex, Gemini CLI, Cursor, GitHub Copilot, Windsurf, and other Agent Skills-compatible tools. Complete SKILL.md… The licence is MIT.

When your agent uses it

  • Instrumenting an agent
  • Debugging intermittent tool
  • Defining service-level objectives
  • Analyzing latency

Example prompts

  • “/agent-observability”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. An observability objective and system boundary.
  2. A trace and event schema with identifiers, span taxonomy, attributes, and redaction rules.
  3. Metrics with definitions, units, dimensions, and ownership.
  4. Dashboard and alert specifications tied to user impact.
  5. A sampling, retention, access, and cost plan.
  6. An investigation runbook and instrumentation verification results.

What it can do on your machine

Read from SKILL.md and the folder at commit 75865a5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Observability loads about 1.4k tokens when it runs, and up to ~2.3k if it reads all its reference files. Until then it costs about 97 tokens; SKILL.md has 710 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from seb1n/awesome-ai-agent-skills at commit 75865a5, republished under its MIT licence (© seb1n). 710 words, ~1,419 tokens.

Download SKILL.mdSave it as .claude/skills/agent-observability/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
agent-observability
description
Design privacy-aware observability for AI agents using traces, spans, structured events, metrics, cost attribution, dashboards, alerts, and investigation workflows. Use when instrumenting an agent, debugging intermittent tool or model failures, defining service-level objectives, analyzing latency or spend, auditing agent decisions, or preparing production monitoring.

Agent Observability

Make agent behavior explainable from request entry through model, retrieval, tool, handoff, and response spans.

Use when

  • Add telemetry to a new or existing agent workflow.
  • Diagnose slow, costly, incorrect, looping, or failed executions.
  • Define dashboards, alerts, service-level indicators, or audit evidence.
  • Standardize traces across models, tools, and orchestration frameworks.

Inputs

Collect the workflow graph, runtime boundaries, incident questions, traffic and failure expectations, telemetry stack, data classification, retention policy, sampling limits, and owners. State what cannot be observed.

Output contract

Produce:

  1. An observability objective and system boundary.
  2. A trace and event schema with identifiers, span taxonomy, attributes, and redaction rules.
  3. Metrics with definitions, units, dimensions, and ownership.
  4. Dashboard and alert specifications tied to user impact.
  5. A sampling, retention, access, and cost plan.
  6. An investigation runbook and instrumentation verification results.

Workflow

  1. Start with operational questions such as “Which tool causes timeouts?” or “Why did cost per resolved task rise?” Do not collect fields without a decision use.
  2. Define one trace per user-visible attempt. Create spans for model calls, retrieval, tools, handoffs, approvals, retries, and final validation. Preserve parent-child relationships and propagate a correlation identifier across queues.
  3. Record stable semantic fields. Include version identifiers, status, timing, token and cost measures, retry counts, tool names, policy outcomes, and evaluation tags when available. Read trace-schema.md before defining attributes.
  4. Separate content from metadata. Default to content-free telemetry; allow prompt or response capture only through explicit authorization, redaction, access controls, and retention limits.
  5. Derive a small set of service indicators: task success, critical-policy violations, end-to-end latency, tool failure rate, escalation rate, and cost per completed task. Define denominators and treatment of cancellations and timeouts.
  6. Build dashboards from user outcome to dependency detail. Alert on actionable sustained impact, not individual noisy spans, and attach an owner and runbook.
  7. Control cardinality, sampling, and storage cost. Retain all critical failures when permitted; use head or tail sampling for normal traffic without losing rare error classes.
  8. Test trace propagation, redaction, retry linkage, clock handling, and degraded telemetry behavior before relying on the data.

Use python3 scripts/summarize_traces.py spans.jsonl for a content-free structural and health summary of normalized spans. It exits 1 for missing parents, a root count other than one, parent cycles, or disconnected components. Add --strict to also exit 1 when likely content-, personal-data-, secret-, or credential-bearing fields are found. Add --output summary.json for an atomic file write; the destination must be a new path or regular file and cannot alias the input through spelling, resolution, a symlink, or a hard link. Findings include field paths and reason classes, never suspected values.

Show full SKILL.md (272 more words)Show less

Safety and permissions

  • Never record secrets, authentication tokens, raw credentials, payment data, or unapproved personal data.
  • Do not enable production capture, change retention, export telemetry, or widen access without authorization.
  • Restrict raw traces by least privilege and log access to sensitive telemetry.
  • Do not let instrumentation failures block the user path unless a mandated audit control requires fail-closed behavior.
  • Treat traces as partial evidence: absent telemetry does not prove an action did not occur.

Verification

  • Follow a synthetic request end to end and confirm every expected span shares one trace identifier, has exactly one root, and forms one acyclic connected parent graph.
  • Trigger a tool error, timeout, retry, refusal, and approval path; verify statuses and parentage.
  • Search emitted telemetry for seeded secrets and personal-data canaries.
  • Recalculate dashboard metrics from raw spans and confirm units, denominators, and time windows.
  • Verify alerts name an owner, include diagnostic context, and avoid high-cardinality dimensions.

Failure handling

  • If trace propagation breaks, preserve local logs with correlation fields and mark cross-service conclusions as incomplete.
  • If timestamps are unreliable, prefer monotonic durations within a process and avoid false cross-host ordering.
  • If telemetry volume exceeds budget, reduce verbose attributes and normal-traffic sampling before dropping critical errors.
  • If sensitive content is found, stop capture, restrict access, follow the incident policy, and purge only with authorized retention owners.

Example

For “find why the research agent became slower after a release,” compare version-tagged end-to-end traces, slice p95 latency by retrieval, model, and browser-tool spans, inspect retries and queue delay, verify sampling did not change, correlate the regression with deployment time, and return the responsible stage, affected cohort, evidence limits, and a monitored remediation.

© seb1n, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in agent-engineering/agent-observability of seb1n/awesome-ai-agent-skills.

  • SKILL.md
  • agents/openai.yaml
  • references/trace-schema.md
  • scripts/summarize_traces.py

Open the folder on GitHubat commit 75865a5

Compare with similar skills

Agent Observability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Observability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Observability this skillseb1n/awesome-ai-agent-skills206—~1.4kAutomated safety check: PassMIT
Vercel Optimize Auditvercel-labs/agent-skills32k9 repos~4.3kAutomated safety check: PassNone
Kubeshark Installerkubeshark/kubeshark12k—~3.6kAutomated safety check: NotesApache-2.0
Kubeshark KFL2 Filter Referencekubeshark/kubeshark12k—~3.6kAutomated safety check: PassApache-2.0
KubeSphere ServiceMesh Managerkubesphere/kubesphere17k—~2.4kAutomated safety check: PassCustom licence
Kubernetes Network Root Cause Analysiskubeshark/kubeshark12k—~5.3kAutomated safety check: PassApache-2.0

Similar skills

  • Vercel Optimize Audit

    vercel-labs/agent-skills

    Official

    Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.

    32k GitHub starsUsed in 9 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Kubeshark Installer

    kubeshark/kubeshark

    Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.

    12k GitHub stars~3.6k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Syntax reference for KFL2, the CEL-based display filter language used to search Kubernetes network traffic captured by Kubeshark, loaded before any filter is written.

    12k GitHub stars~3.6k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • KubeSphere ServiceMesh Manager

    kubesphere/kubesphere

    Installs, checks and troubleshoots the KubeSphere ServiceMesh extension (Istio, Kiali, Jaeger), including grayscale release, sidecar injection, topology and tracing issues.

    17k GitHub stars~2.4k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Caveman Gateway Setup

    JuliusBrussee/caveman

    Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.

    110k GitHub starsUsed in 1 repo~2.6k tokens
    DevOps & CloudAuto-check: warnings

More from seb1n/awesome-ai-agent-skills

All 91 skills in this repo
  • Agent Red Teaming

    seb1n/awesome-ai-agent-skills

    Plan, execute, document, and retest authorized security assessments of AI agents and multi-agent workflows using safe adversarial cases, synthetic identities, canaries, and evidence-based findings.

    206 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Eu AI Act Readiness

    seb1n/awesome-ai-agent-skills

    Build a preliminary, evidence-based EU AI Act readiness assessment across AI-system inventory, territorial scope, operator roles, prohibited-practice screening, risk classification, transparency…

    206 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Human In The Loop

    seb1n/awesome-ai-agent-skills

    Design and verify auditable human oversight, approval gates, escalation paths, and safe state transitions for AI agent workflows.

    206 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • MCP Server Building

    seb1n/awesome-ai-agent-skills

    Design, implement, harden, and verify Model Context Protocol (MCP) servers with precise tool contracts, least-privilege authorization, safe transports, structured errors, and interoperability tests.

    206 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Skill Supply Chain Audit

    seb1n/awesome-ai-agent-skills

    Audit agent skills, plugins, prompts, manifests, scripts, dependencies, and bundled assets for provenance, prompt-injection, permission, execution, exfiltration, persistence, and update risk.

    206 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Spreadsheet Analysis

    seb1n/awesome-ai-agent-skills

    Inspect, profile, clean, reconcile, analyze, visualize, and verify spreadsheet data while preserving formulas, formatting, types, and source files.

    206 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Agent Observability

What does Agent Observability do?

Design privacy-aware observability for AI agents using traces, spans, structured events, metrics, cost attribution, dashboards, alerts, and investigation workflows. Agent Observability is an agent skill from seb1n/awesome-ai-agent-skills. Design privacy-aware observability for AI agents using traces, spans, structured events, metrics, cost attribution, dashboards, alerts, and investigation workflows.

When should I use Agent Observability?

Agent Observability fits situations like: instrumenting an agent; debugging intermittent tool; defining service-level objectives; analyzing latency.

How do I install Agent Observability in Claude Code?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill agent-observability -a claude-code`. Or copy the skill folder (agent-engineering/agent-observability in seb1n/awesome-ai-agent-skills) into .claude/skills/agent-observability in your project. Claude Code loads it when a task matches its description.

How do I install Agent Observability in Codex?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill agent-observability -a codex`. Or copy the skill folder (agent-engineering/agent-observability in seb1n/awesome-ai-agent-skills) into .agents/skills/agent-observability in your project. Codex loads it when a task matches its description.

Can I use Agent Observability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add seb1n/awesome-ai-agent-skills --skill agent-observability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-observability, .gemini/skills/agent-observability, .github/skills/agent-observability and .opencode/skills/agent-observability in your project.

What does Agent Observability need to run?

Going by SKILL.md and its folder, Agent Observability needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Agent Observability access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Observability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Agent Observability use?

Agent Observability is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Observability use?

About 1.4k tokens (SKILL.md is roughly 5.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 881 tokens, read only when the agent opens those files.

What are the alternatives to Agent Observability?

Skills that share tags, products or a category with Agent Observability: Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars), Kubeshark Installer (kubeshark/kubeshark, 12k stars), Kubeshark KFL2 Filter Reference (kubeshark/kubeshark, 12k stars) and KubeSphere ServiceMesh Manager (kubesphere/kubesphere, 17k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Observability?

seb1n (a GitHub user) maintains it in seb1n/awesome-ai-agent-skills, which has 206 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on August 9, 2026.

Source: seb1n/awesome-ai-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.