Agent skill

Agent Observability Spec

by mohitagw15856 in mohitagw15856/pm-claude-skills

Specify the tracing, metrics, and alerting for an AI agent or LLM feature in production.

MITAuto-check passedDevOps & Cloud

Install Agent Observability Spec

skills CLI
$ npx skills add mohitagw15856/pm-claude-skills --skill agent-observability-spec -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitagw15856/pm-claude-skills agent-observability-spec --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-observability-spec .claude/skills/agent-observability-spec && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-observability-spec
GitHub stars
1.4k
Token cost
~1.5k tokens
SKILL.md length
735 words
Files
1
Skills in repo
1,348
Repo updated
First seen
Licence
MIT

At a glance

Specify the tracing, metrics, and alerting for an AI agent or LLM feature in production.

  • Asked what to log for an LLM app
  • SKILL.md covers What This Skill Produces, Required Inputs, Trace Schema and Metrics and Alerts, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Design agent tracing

What it does

Agent Observability Spec is an agent skill from mohitagw15856/pm-claude-skills. Specify the tracing, metrics, and alerting for an AI agent or LLM feature in production. Use when asked what to log for an LLM app, design agent tracing or spans, define quality and cost monitors, or answer 'how do we know if the agent is misbehaving?'. Produces an observability spec with a trace schema, metric definitions with owners and alert thresholds, sampling and retention policy, and a privacy note for logged content.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Observability and Product metrics. The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.

When your agent uses it

  • Asked what to log for an LLM app
  • Design agent tracing
  • Define quality and cost monitors
  • Answer how do we know if the agent is misbehaving?

Example prompts

  • “how do we know if the agent is misbehaving?”
  • “/agent-observability-spec”

What it can do on your machine

Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Observability Spec loads about 1.5k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 735 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 735 words, ~1,453 tokens.

Download SKILL.mdSave it as .claude/skills/agent-observability-spec/SKILL.md (or your agent's skills folder).
name
agent-observability-spec
description
Specify the tracing, metrics, and alerting for an AI agent or LLM feature in production. Use when asked what to log for an LLM app, design agent tracing or spans, define quality and cost monitors, or answer 'how do we know if the agent is misbehaving?'. Produces an observability spec with a trace schema, metric definitions with owners and alert thresholds, sampling and retention policy, and a privacy note for logged content.

Agent Observability Spec Skill

You can't fix what you didn't record. For LLM systems the unit of observability is the trace — everything the model saw and did — because behaviour, not uptime, is what fails. This skill specifies what to capture, what to compute from it, and when to page someone.

Not quite this? Use agent-spec when you are specifying what the agent should do before it is built.

What This Skill Produces

  • A trace schema: per-request spans and the fields each must carry
  • Metric definitions across health, quality, cost, and behaviour — each with a threshold and owner
  • A sampling and retention policy that keeps cost sane and debugging possible
  • A privacy note: what logged content contains, who can see it, and how long it lives

Required Inputs

Ask for (if not already provided):

  • The system's shape — single LLM call, RAG pipeline, or multi-step tool-using agent
  • Traffic volume and cost sensitivity — full tracing at 10M req/day is a budget decision
  • What "misbehaving" means here — the two or three failure modes that matter most (wrong facts? wrong actions? cost? refusals?)
  • Existing observability stack (Datadog, Langfuse, OTel, homegrown) — spec into it, not around it

Trace Schema

Every request produces one trace; every model call, retrieval, guardrail check, and tool execution is a span. Minimum fields:

SpanMust capture
Request rootrequest id, user/session (pseudonymous), feature + prompt version, model id, total tokens, total cost, latency, terminal status
Model callfull input context (or content-addressed ref), output, finish reason, tokens in/out, cached-token share, temperature
Retrievalquery, top-k ids + scores, which chunks entered the context
Tool calltool name, arguments, result (or ref), duration, error
Guardrailcheck name, verdict, and what it did (blocked / rewrote / flagged)
User signaledits, regenerates, thumbs, abandonment — joined to the trace id

The test of the schema: an engineer can replay any incident from its trace alone (see agent-incident-postmortem).

Metrics and Alerts

Define four families; every metric gets a threshold, a window, and an owner.

  • Health — error rate, p50/p95 latency, timeout rate, provider 429/5xx rate. Page on these.
  • Cost — cost per request (p50, p99), tokens per request, cache hit rate, daily spend vs. budget (pair with llm-cost-latency-budget). Alert on p99 and daily-budget burn — cost incidents are caused by the tail, not the mean.
  • Quality proxies — format/schema violation rate, refusal rate, groundedness-check failure rate, judge score on a sampled slice, regenerate/edit rate. Alert on drift vs. a rolling baseline: absolute thresholds go stale, deltas don't.
  • Behaviour (agents) — steps per task, tool-error rate, loop detection (same tool + same args N times), unauthorised-action attempts caught by guardrails. Page on the last one.
Show full SKILL.md (307 more words)Show less

Sampling & Retention

  • Metadata for 100% of requests (ids, versions, tokens, cost, status) — this is cheap and non-negotiable.
  • Full content traces: 100% for errors, guardrail hits, and negative user signals; [1-10]% random sample for the rest, adjusted to volume.
  • Retention: full content [30-90] days, metadata [12+] months for trend baselines; incident traces pinned indefinitely.
  • Privacy: logged context contains user data — state where it lives, who has access, how deletion requests reach it, and that traces are scrubbed or access-gated before wide sharing.

Output Format

Observability Spec: [feature/agent]

System shape: [calls/pipeline/agent] · Volume: [req/day] · Stack: [tooling]

Trace schema: [the span table, tailored]

Metrics:

MetricFamilyThreshold / baselineWindowAlert → owner

Sampling & retention: [the policy]

Privacy: [content classification, access, deletion path]

Dashboards: [the 2-3 views: live health, quality drift, cost]

First incident drill: pick yesterday's worst trace and confirm it can be replayed end-to-end from the stored data.

Quality Checks

  • Any incident is replayable from its trace alone — the schema was tested against that bar
  • Every metric has a number, a window, and a named owner — no orphan dashboards
  • Quality alerts are drift-based against a rolling baseline, not absolute guesses
  • Sampling keeps 100% of error/guardrail/negative-signal traces
  • The privacy note exists and names retention and access — logged prompts are user data

Anti-Patterns

  • Do not log only inputs and outputs — without retrieval and tool spans, root cause analysis is guesswork
  • Do not alert on mean cost or mean latency — the tail is where both incidents live
  • Do not run judge-based quality scoring on 100% of traffic — sample; spend the budget on better baselines
  • Do not treat observability as launch-week scaffolding — drift metrics only work with months of baseline
  • Do not ship an agent that can take actions without logging the guardrail verdicts alongside the actions

Example Trigger Phrases

  • "What to log for an LLM app?"
  • "Design agent tracing."
  • "Define quality and cost monitors."

© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/agent-observability-spec of mohitagw15856/pm-claude-skills.

Open the folder on GitHubat commit 1cbf1f0

Compare with similar skills

Agent Observability Spec next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Observability Spec compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Observability Spec this skillmohitagw15856/pm-claude-skills1.4k—~1.5kAutomated safety check: PassMIT
Vercel Optimize Auditvercel-labs/agent-skills32k8 repos~4.3kAutomated safety check: PassNone
Kubeshark Installerkubeshark/kubeshark12k—~3.6kAutomated safety check: NotesApache-2.0
Kubeshark KFL2 Filter Referencekubeshark/kubeshark12k—~3.6kAutomated safety check: PassApache-2.0
KubeSphere ServiceMesh Managerkubesphere/kubesphere17k—~2.4kAutomated safety check: PassCustom licence
Kubernetes Network Root Cause Analysiskubeshark/kubeshark12k—~5.3kAutomated safety check: PassApache-2.0

Similar skills

  • Vercel Optimize Audit

    vercel-labs/agent-skills

    Official

    Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.

    32k GitHub starsUsed in 8 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Kubeshark Installer

    kubeshark/kubeshark

    Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.

    12k GitHub stars~3.6k tokensUpdated 2 days ago
    DevOps & CloudAuto-check: notes
  • Syntax reference for KFL2, the CEL-based display filter language used to search Kubernetes network traffic captured by Kubeshark, loaded before any filter is written.

    12k GitHub stars~3.6k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • KubeSphere ServiceMesh Manager

    kubesphere/kubesphere

    Installs, checks and troubleshoots the KubeSphere ServiceMesh extension (Istio, Kiali, Jaeger), including grayscale release, sidecar injection, topology and tracing issues.

    17k GitHub stars~2.4k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Caveman Gateway Setup

    JuliusBrussee/caveman

    Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.

    111k GitHub starsUsed in 1 repo~2.6k tokens
    DevOps & CloudAuto-check: warnings

More from mohitagw15856/pm-claude-skills

All 1,348 skills in this repo
  • Car Tco

    mohitagw15856/pm-claude-skills

    Compare the total cost of car ownership across buy-new, buy-used, lease, and keep-your-current-car — depreciation, insurance, maintenance ramp, and fuel over a real horizon, not just the monthly…

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Cs Health Scorecard

    mohitagw15856/pm-claude-skills

    Build a customer health scorecard for a specific account. An agent skill from mohitagw15856/pm-claude-skills.

    1.4k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Exit Waterfall

    mohitagw15856/pm-claude-skills

    Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Feature Prioritisation

    mohitagw15856/pm-claude-skills

    Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.

    1.4k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Fire Number

    mohitagw15856/pm-claude-skills

    Compute a financial-independence (FIRE) target and years-to-reach with every assumption labeled as an assumption — plus a sensitivity table instead of a single false-precision answer.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Freelance Rate

    mohitagw15856/pm-claude-skills

    Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.

    1.4k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Agent Observability Spec

What does Agent Observability Spec do?

Specify the tracing, metrics, and alerting for an AI agent or LLM feature in production. Agent Observability Spec is an agent skill from mohitagw15856/pm-claude-skills. Specify the tracing, metrics, and alerting for an AI agent or LLM feature in production.

When should I use Agent Observability Spec?

Agent Observability Spec fits situations like: asked what to log for an LLM app; design agent tracing; define quality and cost monitors; answer how do we know if the agent is misbehaving?.

How do I install Agent Observability Spec in Claude Code?

Run `npx skills add mohitagw15856/pm-claude-skills --skill agent-observability-spec -a claude-code`. Or copy the skill folder (skills/agent-observability-spec in mohitagw15856/pm-claude-skills) into .claude/skills/agent-observability-spec in your project. Claude Code loads it when a task matches its description.

How do I install Agent Observability Spec in Codex?

Run `npx skills add mohitagw15856/pm-claude-skills --skill agent-observability-spec -a codex`. Or copy the skill folder (skills/agent-observability-spec in mohitagw15856/pm-claude-skills) into .agents/skills/agent-observability-spec in your project. Codex loads it when a task matches its description.

Can I use Agent Observability Spec in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill agent-observability-spec -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-observability-spec, .gemini/skills/agent-observability-spec, .github/skills/agent-observability-spec and .opencode/skills/agent-observability-spec in your project.

What does Agent Observability Spec need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Observability Spec is instructions for the agent only.

Does Agent Observability Spec access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Observability Spec safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Observability Spec use?

Agent Observability Spec is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Observability Spec use?

About 1.5k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Observability Spec?

Skills that share tags, products or a category with Agent Observability Spec: Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars), Kubeshark Installer (kubeshark/kubeshark, 12k stars), Kubeshark KFL2 Filter Reference (kubeshark/kubeshark, 12k stars) and KubeSphere ServiceMesh Manager (kubesphere/kubesphere, 17k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Observability Spec?

mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,433 GitHub stars. The repository holds 1,348 skills in this directory. The repository was last updated on October 8, 2026.

Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.