Backend Dev Guidelines
langfuse/langfuse
Build or review Langfuse backend code. An agent skill from langfuse/langfuse.
A skill your agent uses when the user wants to analyze agent telemetry traces to find bugs and get fix recommendations — walks through exporting traces from a local or remote watsonx Orchestrate…
$ npx skills add IBM/ibm-watsonx-orchestrate-adk --skill telemetry-analyzer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install IBM/ibm-watsonx-orchestrate-adk telemetry-analyzer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/IBM/ibm-watsonx-orchestrate-adk.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/telemetry-analyzer .claude/skills/telemetry-analyzer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "telemetry-analyzer" agent skill from https://github.com/IBM/ibm-watsonx-orchestrate-adk/tree/main/skills/telemetry-analyzer into .claude/skills/telemetry-analyzer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "telemetry-analyzer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/IBM/ibm-watsonx-orchestrate-adk/tree/main/skills/telemetry-analyzerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add IBM/ibm-watsonx-orchestrate-adk --skill telemetry-analyzer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install IBM/ibm-watsonx-orchestrate-adk telemetry-analyzer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/IBM/ibm-watsonx-orchestrate-adk.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/telemetry-analyzer .agents/skills/telemetry-analyzer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "telemetry-analyzer" agent skill from https://github.com/IBM/ibm-watsonx-orchestrate-adk/tree/main/skills/telemetry-analyzer into .agents/skills/telemetry-analyzer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "telemetry-analyzer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add IBM/ibm-watsonx-orchestrate-adk --skill telemetry-analyzer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install IBM/ibm-watsonx-orchestrate-adk telemetry-analyzer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/IBM/ibm-watsonx-orchestrate-adk.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/telemetry-analyzer .cursor/skills/telemetry-analyzer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "telemetry-analyzer" agent skill from https://github.com/IBM/ibm-watsonx-orchestrate-adk/tree/main/skills/telemetry-analyzer into .cursor/skills/telemetry-analyzer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "telemetry-analyzer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/IBM/ibm-watsonx-orchestrate-adk.git --path skills/telemetry-analyzer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add IBM/ibm-watsonx-orchestrate-adk --skill telemetry-analyzer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install IBM/ibm-watsonx-orchestrate-adk telemetry-analyzer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/IBM/ibm-watsonx-orchestrate-adk.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/telemetry-analyzer .gemini/skills/telemetry-analyzer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "telemetry-analyzer" agent skill from https://github.com/IBM/ibm-watsonx-orchestrate-adk/tree/main/skills/telemetry-analyzer into .gemini/skills/telemetry-analyzer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "telemetry-analyzer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install IBM/ibm-watsonx-orchestrate-adk telemetry-analyzerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add IBM/ibm-watsonx-orchestrate-adk --skill telemetry-analyzer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/IBM/ibm-watsonx-orchestrate-adk.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/telemetry-analyzer .github/skills/telemetry-analyzer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "telemetry-analyzer" agent skill from https://github.com/IBM/ibm-watsonx-orchestrate-adk/tree/main/skills/telemetry-analyzer into .github/skills/telemetry-analyzer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "telemetry-analyzer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add IBM/ibm-watsonx-orchestrate-adk --skill telemetry-analyzer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install IBM/ibm-watsonx-orchestrate-adk telemetry-analyzer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/IBM/ibm-watsonx-orchestrate-adk.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/telemetry-analyzer .opencode/skills/telemetry-analyzer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "telemetry-analyzer" agent skill from https://github.com/IBM/ibm-watsonx-orchestrate-adk/tree/main/skills/telemetry-analyzer into .opencode/skills/telemetry-analyzer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "telemetry-analyzer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
telemetry-analyzerA skill your agent uses when the user wants to analyze agent telemetry traces to find bugs and get fix recommendations — walks through exporting traces from a local or remote watsonx Orchestrate…
Telemetry Analyzer is an agent skill from IBM/ibm-watsonx-orchestrate-adk, published by the product's own GitHub organization. Use when the user wants to analyze agent telemetry traces to find bugs and get fix recommendations — walks through exporting traces from a local or remote watsonx Orchestrate server, parsing raw OTel JSON or Langfuse-format trace JSON directly, and reasoning over them to identify failures and suggest fixes.
Its SKILL.md is about 10k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts (for example `README.md`, `scripts/export_traces_adk.py` and `scripts/export_traces_agentops_v3.py`).
It sits in Development, covering LLM observability and Debugging. It works with Langfuse and OpenTelemetry. The repository describes itself as: The command line client for watsonx Orchestrate's agent builder experience. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b6f9065. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 4 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Telemetry Analyzer loads about 10k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 3,534 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
d) for prerequisites (ADK installation, `.env` configuration, and environment activation) before running any scripts.Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from IBM/ibm-watsonx-orchestrate-adk at commit b6f9065, republished under its MIT licence (© IBM). 3,534 words, ~10,234 tokens.
.claude/skills/telemetry-analyzer/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.This skill exports agent telemetry traces from watsonx Orchestrate and analyzes them directly from raw OTel JSON or Langfuse-format trace JSON to produce structured bug reports with root-cause analysis and fix recommendations.
Two trace formats are supported:
- OTel JSON — produced by the classic
TracesController.export_trace_to_json()ADK path and local IBM telemetry servers. Top-level key:traceData.resourceSpans.- Langfuse JSON — produced by the
agentops-v3REST API (GET /v1/agentops-v3/traces/<id>) and theexport_traces_agentops_v3.pyscript. Top-level key:observations(array of GENERATION / CHAIN / SPAN objects).Always detect the format first before reading any span data. See the
detect_format()helper in Step 3C and the normalization table in Step 4.
Setup: See README.md for prerequisites (ADK installation,
.envconfiguration, and environment activation) before running any scripts.
TracesController) or Langfuse JSON (via export_traces_agentops_v3.py / search_traces_agentops_v3.py)Use this skill when you need to:
Ask questions like:
f492e71b957ccec2d07096ae99395d19?"_adk.py) — wxO Traces API / local IBM telemetry# Search traces and save IDs for bulk export
python scripts/search_traces_adk.py --agent-name "My Agent" --last 1h --all \
--save-ids data/trace_ids.txt
# Export all discovered traces
python scripts/export_traces_adk.py --ids-file data/trace_ids.txt
# Export a single trace by ID
python scripts/export_traces_adk.py --trace-id <32-char-hex-id>_agentops_v3.py) — agentops-v3 REST API# Search traces for an agent (last 20 minutes by default)
python scripts/search_traces_agentops_v3.py --last 20m
# Narrow or widen the time window
python scripts/search_traces_agentops_v3.py --last 2h
python scripts/search_traces_agentops_v3.py --start 2025-01-01T09:00:00Z --end 2025-01-01T21:00:00Z
# Collect every trace in the window and save IDs for bulk export
python scripts/search_traces_agentops_v3.py --last 1h --all \
--save-ids data/trace_ids.txt
# Quick count before committing to a full export
python scripts/search_traces_agentops_v3.py --last 2h --count
# Export a single trace by ID
python scripts/export_traces_agentops_v3.py --trace-id <32-char-hex-id>
# Bulk-export from a saved ID list
python scripts/export_traces_agentops_v3.py --ids-file data/trace_ids.txt--limit N, --all, and --count modes for controlling result volumedata/<trace_id>.json — OTel format via export_traces_adk.py, Langfuse format via export_traces_agentops_v3.py--with-ibm-telemetrysearch_traces_agentops_v3.py / export_traces_agentops_v3.py use the raw agentops-v3 REST API (bearer token from active env, no ADK SDK calls)Detected from normalized spans, supporting both OTel (traceData.resourceSpans[0].scopeSpans[0].spans) and Langfuse (observations[]) formats:
status.code == STATUS_CODE_ERRORfinish_reason == "length", empty answer.task output, zero output tokens (via llm.token_count.completion / gen_ai.usage.output_tokens attributes)ToolMessage content in *.tool spans, retry loops, duration_ms > 30 000 msanswer.task span, empty final answer, repeated agent.task cycles (multi-step loop), empty collaborator outputhas_error on any span, widget_result.task or collaborator.task with empty outputthread_id from the correct location:traceloop.association.properties.thread_id or thread.id span attributemetadata.attributes.thread_id or sessionId top-level fieldThe analyzer produces:
Bug Report (HTML artifact for 3+ issues, inline for fewer)
Conversational Flow Report (HTML artifact, thread-scoped)
Scripts in scripts/ (created on first use by Step 1 if not already present):
| Script | Output format | Auth mechanism | Notes |
|---|---|---|---|
search_traces_adk.py | OTel | ADK TracesController | Cursor-based pagination, agent-name filter, rate-limit retry |
export_traces_adk.py | OTel | ADK TracesController | Single or bulk export by trace ID |
search_traces_agentops_v3.py | Langfuse | Bearer token from active env | Page-based pagination, returns all traces in time window |
export_traces_agentops_v3.py | Langfuse | Bearer token from active env | GET /v1/agentops-v3/traces/<id>, no ADK SDK trace call |
Use the *_agentops_v3.py scripts when the environment uses the agentops-v3 REST API. Use the standard scripts for environments backed by the classic wxO Traces API.
File placement rules — apply throughout all steps:
- All Python scripts written to disk must be saved under
scripts/(e.g.scripts/my_helper.py).- All output files (
.jsontrace dumps,.txtID lists, etc.) must be saved underdata/(e.g.data/trace_ids.txt,data/<trace_id>.json).- Inline code snippets shown in this document are for reading and reasoning only — they are not written to disk unless explicitly instructed.
Before doing anything else, confirm the required scripts exist on disk:
scripts/search_traces_adk.py
scripts/export_traces_adk.py
scripts/search_traces_agentops_v3.py
scripts/export_traces_agentops_v3.pyIf any are missing, use write_file to create them — the canonical source for each script is in the scripts/ directory of this skill. Write each missing file to its correct path under scripts/ before continuing to Step 2.
If the user's request does not clearly state where the traces or agent reside, always use ask_followup_question to clarify before doing anything else. Do not assume a source or proceed to Step 3 without a confirmed answer to all of the following that are not already clear from context:
orchestrate env activate)?Once the source is confirmed, classify it as one of:
orchestrate env activate. Always try the _agentops_v3.py scripts first; fall back to the _adk.py scripts if the agentops-v3 search returns an error or zero traces. See Step 3A for the full fallback procedure.--with-ibm-telemetry. The Traces API is available at http://localhost:4321. Use the _adk.py scripts directly — agentops-v3 is not available on local environments.{trace_id}.json files on disk ready to analyze. Confirm they are in the data/ directory; if not, ask the user for the path before proceeding.For all remote environments, always try _agentops_v3.py first. Fall back to _adk.py only if the agentops-v3 search fails or returns zero traces.
Step A1 — Try agentops-v3:
# 1. Activate the target environment
orchestrate env activate <env-name>
# 2. Search traces via agentops-v3
python scripts/search_traces_agentops_v3.py \
--start 2025-01-01T09:00:00Z --end 2025-01-01T21:00:00Z \
--all --save-ids data/trace_ids.txtIf the command exits with a non-zero status or prints Found 0 trace(s), proceed to Step A2. Otherwise, export the discovered traces:
# 3. Export all discovered traces
python scripts/export_traces_agentops_v3.py --ids-file data/trace_ids.txtStep A2 — Fall back to ADK (only if agentops-v3 failed or returned 0 traces):
# 2. Search traces via ADK — IDs checkpointed to disk after every page
python scripts/search_traces_adk.py --agent-name "My Agent" \
--start 2025-01-01T09:00:00Z --end 2025-01-01T21:00:00Z \
--all --save-ids data/trace_ids.txt
# 3. Export all discovered traces
python scripts/export_traces_adk.py --ids-file data/trace_ids.txtsearch_traces_adk.py uses cursor-based pagination, retries automatically on 429 rate-limit responses, and handles mid-run 401 token expiry by refreshing the client. Always pass --save-ids so trace IDs are checkpointed to disk page-by-page and are not lost on timeout.
Use --limit N (default 50) for a capped sample, --all to exhaust the full window, or --count for a quick volume figure before committing to a full export.
# 1. Start the local server with IBM telemetry enabled
orchestrate server start --with-ibm-telemetry --accept-terms-and-conditions
# 2. Activate the local environment
orchestrate env activate local
# 3. Search and export — identical to Option A
python scripts/search_traces_adk.py --agent-name "My Agent" --last 30m --all \
--save-ids data/trace_ids.txt
python scripts/export_traces_adk.py --ids-file data/trace_ids.txtThe local Traces API endpoint is http://localhost:4321. The scripts detect is_local_dev() automatically and set service_names=["wxo-server"] in the filter, which is required when FORCE_SINGLE_TENANT=true.
Confirm the directory or file path with the user. Always detect the format before reading spans. Use this helper:
import json
def detect_format(raw: dict) -> str:
"""Return 'otel' or 'langfuse'."""
if isinstance(raw.get('traceData'), dict) and 'resourceSpans' in raw['traceData']:
return 'otel'
if 'resourceSpans' in raw: # traceData sometimes omitted at root
return 'otel'
if 'observations' in raw: # Langfuse / agentops-v3 REST response
return 'langfuse'
return 'langfuse' # safe default for agentops-v3 output
def get_spans(raw: dict) -> list:
"""Return a flat list of span/observation dicts regardless of format."""
fmt = detect_format(raw)
if fmt == 'otel':
td = raw.get('traceData') or raw
return [
s
for rs in td.get('resourceSpans', [])
for ss in rs.get('scopeSpans', [])
for s in ss.get('spans', [])
]
return raw.get('observations', [])Load a file and get its spans:
with open("data/<trace_id>.json") as f:
raw = json.load(f)
fmt = detect_format(raw)
spans = get_spans(raw)
print(f"Format: {fmt}, span/observation count: {len(spans)}")If the user provides a directory, glob for *.json files and skip any files not matching the 32-character hex trace ID pattern.
Single-trace shortcut: If the user has asked to analyze exactly one trace (by ID or as a single file on disk), skip Steps 4 and 5 entirely. Jump straight to Step 7 to produce the conversational flow report for the thread that trace belongs to. Only return to Step 5 (bug report) if the user explicitly asks for one after seeing the flow report.
Load each trace file using detect_format() + get_spans() from Step 3C, then normalize every span with the adapter below before running any bug-detection logic.
from datetime import datetime, timezone
def normalize_span(s: dict, fmt: str) -> dict:
"""
Return a uniform dict with these keys regardless of source format:
name, start_utc, duration_ms, status_error (bool),
input, output, attrs (dict: key -> scalar value)
"""
if fmt == 'otel':
start_ns = int(s.get('startTimeUnixNano', 0))
end_ns = int(s.get('endTimeUnixNano', 0))
start_utc = datetime.fromtimestamp(start_ns / 1e9, tz=timezone.utc)
duration_ms = (end_ns - start_ns) / 1e6
status_error = s.get('status', {}).get('code') in ('STATUS_CODE_ERROR', 2)
attrs = {}
for a in s.get('attributes', []):
v = a['value']
attrs[a['key']] = (
v.get('stringValue') or v.get('intValue') or
v.get('doubleValue') or v.get('boolValue')
)
input_val = attrs.get('traceloop.entity.input')
output_val = attrs.get('traceloop.entity.output')
else: # langfuse
start_utc = datetime.fromisoformat(
s.get('startTime', '1970-01-01T00:00:00Z').replace('Z', '+00:00'))
end_str = s.get('endTime') or s.get('startTime', '1970-01-01T00:00:00Z')
end_dt = datetime.fromisoformat(end_str.replace('Z', '+00:00'))
duration_ms = (end_dt - start_utc).total_seconds() * 1000
status_error = (s.get('level') == 'ERROR' or bool(s.get('statusMessage')))
meta = s.get('metadata') or {}
attrs = dict(meta.get('attributes', {}))
usage = s.get('usage') or {}
if usage.get('input'):
attrs['llm.usage.prompt_tokens'] = usage['input']
if usage.get('output'):
attrs['llm.usage.completion_tokens'] = usage['output']
input_val = s.get('input')
output_val = s.get('output')
return dict(
name=s.get('name', ''),
start_utc=start_utc,
duration_ms=duration_ms,
status_error=status_error,
input=input_val,
output=output_val,
attrs=attrs,
)
def get_attr(attrs: dict, key: str):
"""Read a scalar from the normalized attrs dict (works for both formats)."""
return attrs.get(key)Build the normalized span list at the start of every analysis:
fmt = detect_format(raw)
spans = get_spans(raw)
nspans = [normalize_span(s, fmt) for s in spans]
llm_spans = [n for n in nspans if n['name'] == 'WatsonxChatModel.chat']
tool_spans = [n for n in nspans if 'collaborator' in n['name'].lower()
or n['attrs'].get('openinference.span.kind') == 'TOOL']Field mapping reference — use this table when reading span data:
| Concept | OTel key (in attrs after normalization) | Langfuse source (in attrs after normalization) |
|---|---|---|
| LLM input tokens | llm.token_count.prompt or llm.usage.prompt_tokens | llm.usage.prompt_tokens (from usage.input) |
| LLM output tokens | llm.token_count.completion or llm.usage.completion_tokens | llm.usage.completion_tokens (from usage.output) |
| LLM model | llm.model_name or gen_ai.request.model | llm.model_name or model top-level field |
| Thread ID | traceloop.association.properties.thread_id or thread.id | thread_id or thread.id |
| Finish reason | llm.finish_reason or gen_ai.usage.finish_reasons | llm.finish_reason |
| Cache read tokens | gen_ai.usage.cache_read_input_tokens or llm.token_count.cache_read | gen_ai.usage.cache_read_input_tokens or usage.cacheReadInputTokens (check raw obs) |
| Cache creation tokens | gen_ai.usage.cache_creation_input_tokens or llm.token_count.cache_creation | gen_ai.usage.cache_creation_input_tokens or usage.cacheCreationInputTokens (check raw obs) |
| Span/obs input | input (normalized field) | input (normalized field) |
| Span/obs output | output (normalized field) | output (normalized field) |
| LLM reasoning | inside output JSON → additional_kwargs.reasoning | inside output dict → additional_kwargs.reasoning |
| Tool call name | span name ending in .tool | SPAN observation name (e.g. chat_with_collaborator_*) |
For each trace collect: start_utc, span names, duration_ms, status_error, collaborator tool call output, final answer output, thread_id (from attrs), and cache token counts.
Look for the following bug categories. For each issue found, record the span name, start_utc, and relevant attrs values.
status_error == True.status.code == "STATUS_CODE_ERROR" or 2.level == "ERROR" or statusMessage is non-empty.name, duration_ms, and the output field for the error message.get_attr(n['attrs'], 'llm.finish_reason') == "length" — context length exceeded (works for both formats after normalization).WatsonxChatModel.chat span with duration_ms == 0 or absent — silent LLM failure.llm.token_count.prompt / llm.token_count.completion (OTel attrs) or llm.usage.prompt_tokens / llm.usage.completion_tokens (Langfuse, auto-added by normalize_span). Zero completion tokens with non-zero prompt tokens = failed generation.get_attr(attrs, 'gen_ai.usage.cache_read_input_tokens') first; fall back to the raw Langfuse usage dict. If absent on all LLM spans, note as an observation (provider may not support caching).output field contains an empty or null content (i.e. ToolMessage.content == "").content on the immediate chat_with_collaborator_* SPAN is the normal async dispatch pattern — the real result arrives via the nested collaborator CHAIN observation. Only flag as a failure if the outer collaborator CHAIN output is also empty.duration_ms > 30 000 ms — timeout or hung call.answer (or answer.task in OTel) — agent never produced a final answer.answer span output is empty or whitespace-only.agent (or agent.task) span count > 2 — multi-step LLM loop (agent re-planning instead of delegating immediately).collaborator (or collaborator.task) span present but answer output is empty — collaborator returned nothing and supervisor silently swallowed it.answer output or output.additional_kwargs.reasoning before the user-facing reply — check both OTel and Langfuse paths.widget_result (Langfuse) or widget_result.task (OTel) with empty output — widget state machine stalled.LangGraph GENERATION (Langfuse) or LangGraph.workflow span (OTel) with duration_ms >> sum of child span durations — unexplained gap in the workflow.WatsonxChatModel.chat spans exceeds 50 000.Langfuse format caveat: Cache token attributes (
cacheReadInputTokens,cacheCreationInputTokens) are not reliably collected in Langfuse-format traces. When the source format is Langfuse (fmt == 'langfuse'), attempt both themetadata.attributesand top-levelusagefallback paths below. If both are absent after exhausting all fallbacks, setcache_unsupported = Truefor the trace and skip all cache-related analysis, stat cards, table columns, and observations for that trace. Do not flag missing cache attributes as a bug or observation for Langfuse-format traces — they are expected to be absent.
Extract per-WatsonxChatModel.chat normalized span, with a Langfuse usage dict fallback for cache fields that may not be hoisted into metadata.attributes:
def get_cache_tokens(nspan: dict, raw_obs: dict = None) -> tuple:
"""
nspan — normalized span dict (attrs already flattened from either format).
raw_obs — original Langfuse observation dict (optional), used as fallback
for cache fields that are in usage but not in metadata.attributes.
Returns (cache_read, cache_creation) as ints.
"""
attrs = nspan['attrs']
cache_read = (
get_attr(attrs, 'gen_ai.usage.cache_read_input_tokens') or
get_attr(attrs, 'llm.token_count.cache_read') or 0
)
cache_creation = (
get_attr(attrs, 'gen_ai.usage.cache_creation_input_tokens') or
get_attr(attrs, 'llm.token_count.cache_creation') or 0
)
# Langfuse-specific fallback: check top-level usage dict
if raw_obs and not cache_read and not cache_creation:
usage = raw_obs.get('usage') or {}
cache_read = usage.get('cacheReadInputTokens') or 0
cache_creation = usage.get('cacheCreationInputTokens') or 0
return int(cache_read), int(cache_creation)Aggregate across all LLM normalized spans in the trace (pass both the normalized span and the original raw observation when the source is Langfuse):
total_cache_read = sum(get_cache_tokens(n, raw_obs=spans[i] if fmt=='langfuse' else None)[0]
for i, n in enumerate(llm_spans))
total_cache_creation = sum(get_cache_tokens(n, raw_obs=spans[i] if fmt=='langfuse' else None)[1]
for i, n in enumerate(llm_spans))
# For Langfuse traces: if both totals are still 0 after all fallbacks, cache data
# is not collected by this environment — mark the trace and skip cache analysis.
cache_unsupported = (fmt == 'langfuse') and (total_cache_read == 0) and (total_cache_creation == 0)
# Effective billable input = input_tokens - total_cache_read
# Cache hit rate = total_cache_read / (input_tokens - total_cache_creation)
# only meaningful when input_tokens > total_cache_creation > 0
denom = input_tokens - total_cache_creation
cache_hit_rate = (total_cache_read / denom) if (not cache_unsupported and denom > 0) else NoneFlag the following — only when cache_unsupported is False:
cache_hit_rate < 0.20 (< 20 %) on a trace that belongs to a thread with ≥ 3 turns. The system prompt is likely not structured for prefix caching or varies per turn.cache_read_input_tokens == 0 across ≥ 3 turns in the same thread. The cache is never being warmed or the prefix changes each turn.total_cache_creation > total_cache_read * 5 across the batch. The cache is being populated but rarely re-used (cache TTL mismatch or non-repeating prefix).Skip this step when the analysis scope is a single trace. Proceed directly to Step 7 instead.
Telemetry Bug Report — <AgentName>| Field | Value |
|---|---|
| Agent ID | <uuid> |
| Agent Name | <name> |
| Environment | Instance URL or local |
| Report Date | YYYY-MM-DD (UTC) |
| Time Window (UTC) | YYYY-MM-DD HH:MM – YYYY-MM-DD HH:MM UTC |
Two rows of 4 KPI stat cards each:
.stat-bad), successful count (.stat-ok), time window duration.cache_unsupported is True for every trace, omit the cache-read and cache hit rate stat cards entirely and render only a single row of 4 cards (total traces, failed, successful, avg/p95 duration).No prose paragraphs — stat cards only.
Required CSS:
.stat-grid { display: grid; grid-template-columns: repeat(4, 1fr); gap: 10px; margin-bottom: 20px; }
.stat-card { background: #f7f8fa; border: 1px solid #e5e7eb; border-radius: 6px; padding: 12px 14px; }
.stat-val { font-size: 22px; font-weight: 700; color: #1f2328; }
.stat-lbl { font-size: 12px; color: #57606a; margin-top: 2px; }
.stat-bad .stat-val { color: #b91c1c; }
.stat-ok .stat-val { color: #15803d; }
.stat-warn .stat-val { color: #b45309; }Cache-specific stat card values to compute before rendering:
# Exclude traces where cache data is known to be uncollected (Langfuse with no cache attrs)
cache_rows = [r for r in rows if not r.get("cache_unsupported")]
if cache_rows:
total_cache_read_all = sum(r.get("cache_read_tokens", 0) for r in cache_rows)
total_cache_creation_all = sum(r.get("cache_creation_tokens", 0) for r in cache_rows)
total_input_cache = sum(r.get("input_tokens", 0) for r in cache_rows)
denom_all = total_input_cache - total_cache_creation_all
batch_cache_hit_rate = (total_cache_read_all / denom_all * 100) if denom_all > 0 else None
cache_attrs_present = total_cache_read_all > 0 or total_cache_creation_all > 0
hit_rate_display = f"{batch_cache_hit_rate:.1f}%" if cache_attrs_present else "N/A"
hit_rate_class = "stat-bad" if (cache_attrs_present and batch_cache_hit_rate is not None and batch_cache_hit_rate < 20) else "stat-ok"
show_cache_cards = True
else:
# All traces are Langfuse with no cache data — omit both cache stat cards
show_cache_cards = FalsePart 1 — Categorical summary table (group by routing pattern, error type, etc.):
| Category | Traces | LLM Calls (avg) | Tool Calls (avg) | Success Rate | Status |
|---|
Part 2 — Per-trace detail table (ordered by start time ascending):
If show_cache_cards is True (cache data available for at least some traces), use the full 7-column layout:
| Trace ID | Start Time (UTC) | Spans | LLM | In / Out / Cached Tokens | Cache Hit % | Status |
|---|
In / Out / Cached Tokens — display as three values separated by /, e.g. 12 450 / 320 / 8 100. Show — for Cached on a specific row when that trace has cache_unsupported == True.
Cache Hit % — cache_read_tokens / (input_tokens - cache_creation_tokens) * 100, formatted as 42 %. Show — for rows where cache_unsupported == True.
If show_cache_cards is False (all traces are Langfuse with no cache data), drop the "Cached Tokens" and "Cache Hit %" columns and use the 5-column layout:
| Trace ID | Start Time (UTC) | Spans | LLM | In / Out Tokens | Status |
|---|
Rules:
white-space: nowrap; font-family: monospace on the Trace ID cell.| # | Span | Type | Description | Evidence | Affected Traces |
|---|
HTML rendering rules — apply these exactly to prevent column overflow:
Table layout — always set table-layout: fixed; width: 100% and declare explicit <colgroup> widths so the browser cannot let any column expand beyond its allocation:
<table style="table-layout:fixed; width:100%">
<colgroup>
<col style="width:3%"> <!-- # -->
<col style="width:12%"> <!-- Span -->
<col style="width:7%"> <!-- Type -->
<col style="width:28%"> <!-- Description -->
<col style="width:28%"> <!-- Evidence -->
<col style="width:22%"> <!-- Affected Traces -->
</colgroup>
…
</table>Every <td> and <th> — add overflow: hidden to hard-clip any content that would otherwise bleed into the next column:
<td style="overflow:hidden; word-break:break-word">…</td>Evidence column — allow wrapping and use <br> between individual evidence items. Do not set a fixed max-width — let the percentage-based <colgroup> width scale with the container:
<th>Evidence</th>
<td style="overflow:hidden; word-break:break-word">
<code>field: value</code><br>
<code>field: value</code>
</td>Affected Traces column — use word-break: break-all (not nowrap) so long hex IDs wrap within the cell, and put each ID on its own line:
<th>Affected Traces</th>
<td style="overflow:hidden; word-break:break-all; font-family:monospace; font-size:11px">
tid1<br>tid2<br>…
</td>For Observations (prose), end each bullet with affected trace IDs in parentheses.
Cache efficiency observations — add the following bullets when applicable. Skip all cache observations entirely when show_cache_cards is False (all-Langfuse batch with no cache data):
gen_ai.usage.cache_read_input_tokens nor llm.token_count.cache_read was present on any LLM span. The provider may not support prompt caching, or the traceloop-sdk version predates cache attribute emission. (list affected trace IDs)cache_read_input_tokens == 0 across all turns in thread <thread_id> (≥ 3 turns). The cache is never warmed; the static system prompt prefix likely changes per turn or is too short to cache. (list affected trace IDs)total_cache_creation_tokens (N) is more than 5× total_cache_read_tokens (M) across the batch. The prompt-cache prefix is being written but not re-used — likely a cache TTL expiry or a prefix that varies each turn. (list affected trace IDs)For each critical issue and warning:
Issue: <description>
Root Cause: <likely cause>
Recommended Fix:
- <actionable step 1>
- <actionable step 2>Cache hit rate fix recommendation template — include when cache hit rate < 20 % or zero cache reads observed on multi-turn threads:
Issue: Low prompt-cache hit rate (<X % batch average)
Root Cause: The system prompt either varies per turn, is not positioned at the start of
the messages array, or the cache TTL expired between turns in long-gap threads.
Recommended Fix:
- Pin the static system prompt as the first message so the LLM provider can cache the
longest common prefix across turns.
- Avoid injecting dynamic content (timestamps, user IDs, session state) into the system
prompt — move that to the user turn instead.
- For Anthropic: add `cache_control: {"type": "ephemeral"}` to the last static content
block to explicitly mark the cacheable prefix boundary.
- For OpenAI: ensure repeated calls use an identical token-level prefix of ≥ 1 024 tokens;
caching is automatic but the prefix must be byte-for-byte identical.
- Verify the inter-turn gap is within the provider cache TTL: ~5 min for OpenAI and
Anthropic ephemeral caches. For long-running sessions, refresh the cache before TTL
expiry by issuing a short warm-up prompt.
- If cache attributes are absent entirely, upgrade the traceloop-sdk to a version that
emits `gen_ai.usage.cache_read_input_tokens` / `gen_ai.usage.cache_creation_input_tokens`.Render as create_html_artifact for 3+ issues; inline for fewer.
HTML report outer wrapper — use width: 80% centred with margin: 0 auto so the report occupies 80 % of the page width:
<div style="width:80%; margin:0 auto; padding:20px 0; font-family:-apple-system,'Segoe UI',system-ui,sans-serif; font-size:14px; line-height:1.6">
<!-- all report content here -->
</div>After delivering the bug report, ask:
"Would you like a conversational flow report for any specific trace? If so, provide a trace ID and I'll pull all traces under the same thread and produce a full thread-level flow report."
Use ask_followup_question with suggestions drawn from the failing trace IDs.
A conversational flow report is thread-scoped — one thread ID may span multiple traces (one per user turn).
Use detect_format() from Step 3C, then read the thread ID from the correct location:
import json
with open(f"data/{trace_id}.json") as f:
raw = json.load(f)
fmt = detect_format(raw)
thread_id = None
if fmt == 'otel':
spans = get_spans(raw)
for s in spans:
for a in s.get('attributes', []):
if a['key'] in ('traceloop.association.properties.thread_id', 'thread.id'):
thread_id = a['value']['stringValue']
break
if thread_id:
break
else: # langfuse
# 1. Check top-level metadata.attributes (most reliable)
meta_attrs = (raw.get('metadata') or {}).get('attributes', {})
thread_id = (
meta_attrs.get('thread.id') or
meta_attrs.get('thread_id') or
meta_attrs.get('langfuse.session.id') or
raw.get('sessionId') # Langfuse session == wxO thread
)
# 2. Fall back to any observation's metadata.attributes
if not thread_id:
for obs in raw.get('observations', []):
obs_attrs = (obs.get('metadata') or {}).get('attributes', {})
thread_id = obs_attrs.get('thread_id') or obs_attrs.get('thread.id')
if thread_id:
break
print(f"format: {fmt}, thread_id: {thread_id}")First scan the local cache for matching thread_id, then search the API for any traces not yet downloaded. Re-run the local scan after downloading new files. If the wxO token is expired, work with cached traces only and note how many turns were found vs. potentially missing.
Sort all matched traces by start_utc ascending. For each trace call detect_format() + normalize_span() and extract using the unified field names:
| What to extract | OTel span name | Langfuse observation name | Field on normalized span |
|---|---|---|---|
| User message | POST /orchestrate/runs or LangGraph.workflow | POST /chat/completions GENERATION | input (top-level content string) |
| LLM call count | count of WatsonxChatModel.chat spans | count of WatsonxChatModel.chat GENERATIONs | name == 'WatsonxChatModel.chat' |
| Input / output tokens | llm.token_count.prompt / .completion attrs | llm.usage.prompt_tokens / .completion_tokens attrs (auto-added) | get_attr(attrs, 'llm.usage.prompt_tokens') etc. |
| Cache tokens | gen_ai.usage.cache_read_input_tokens attr | usage.cacheReadInputTokens (raw obs fallback) | get_cache_tokens(nspan, raw_obs) |
| Collaborator routing | chat_with_collaborator_*.tool span output | SPAN obs chat_with_collaborator_* → output.content | output field |
| Collaborator steps | collaborator.task, tools.task, entrypoint.task spans | CHAIN obs collaborator, tools, LangGraph[collaborator:*] | name, duration_ms, output |
| Final answer | answer.task span output | CHAIN obs answer → output | output field |
| Error flag | status_error == True on any span | status_error == True on any obs | n['status_error'] |
| LLM reasoning | output JSON → additional_kwargs.reasoning | output dict → additional_kwargs.reasoning | parse n['output'] |
Produce a create_html_artifact with four sections:
Turn | Seq | Span name | Actor | Duration (ms) | Status.After the conversational flow report, offer:
data/{tid}.json, detect format, and show full input / output content (traceloop.entity.input / traceloop.entity.output for OTel; input / output top-level fields for Langfuse). Use normalize_span() so the same display code works for both.© IBM, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts) in skills/telemetry-analyzer of IBM/ibm-watsonx-orchestrate-adk.
Open the folder on GitHubat commit b6f9065
Telemetry Analyzer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Telemetry Analyzer this skillIBM/ibm-watsonx-orchestrate-adk | 178 | — | ~10k | Automated safety check: Notes | MIT | |
| Backend Dev Guidelineslangfuse/langfuse | 35k | — | ~1.9k | Automated safety check: Pass | Custom licence | |
| Langfuse and LLM Gateway LogsKonghaYao/peri | 223 | — | ~4.3k | Automated safety check: Notes | Apache-2.0 | |
| Agentsop Observability Setupagentsope/SkillAlchemy | 457 | — | ~4.4k | Automated safety check: Pass | MIT | |
| Olore Langfuse Latestolorehq/olore | 103 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Refactor React Effectslangfuse/langfuse | 35k | — | ~1.7k | Automated safety check: Pass | Custom licence |
langfuse/langfuse
Build or review Langfuse backend code. An agent skill from langfuse/langfuse.
KonghaYao/peri
Queries Langfuse traces, prompts, datasets and sessions, and analyzes local LLM gateway logs for requests, context growth, token use and cache hits.
agentsope/SkillAlchemy
Enhancement-overlay skill — the DECISION + WIRING layer for LM observability that the single-backend skills [[langsmith]], [[phoenix]], [[mlflow]] do NOT cover.
olorehq/olore
Local Langfuse documentation reference (latest). An agent skill from olorehq/olore.
langfuse/langfuse
Refactor avoidable React useEffect usage in Langfuse frontend code.
Arize-ai/phoenix
Create Phoenix release documentation grounded in actual code changes.
IBM/ibm-watsonx-orchestrate-adk
Analyzes IBM watsonx Orchestrate agentic workflow artefacts (JSON or Python @flow) and returns prioritised architecture recommendations grouped by impact.
IBM/ibm-watsonx-orchestrate-adk
Build MCP servers for customer care agents following Watson Orchestrate specifications.
IBM/ibm-watsonx-orchestrate-adk
A skill your agent uses when building, testing, debugging, or publishing IBM watsonx Orchestrate agents, tools, flows, connections, knowledge bases, or custom models with the orchestrate CLI or ADK…
IBM/ibm-watsonx-orchestrate-adk
Evaluate an agent instructions or agent definition for achievability and produce a structured, evidence-backed report artifact with per-dimension scores, findings, deterministic signals, and…
IBM/ibm-watsonx-orchestrate-adk
Expert guidance for creating high-level solution architecture documents from business requirements, use cases, or problem statements.
IBM/ibm-watsonx-orchestrate-adk
Expert guidance for building a Standard Operating Procedure (SOP) from a workflow diagram, Langflow JSON, n8n JSON, BPMN model or workflow description.
Works with
A skill your agent uses when the user wants to analyze agent telemetry traces to find bugs and get fix recommendations — walks through exporting traces from a local or remote watsonx Orchestrate…. Telemetry Analyzer is an agent skill from IBM/ibm-watsonx-orchestrate-adk, published by the product's own GitHub organization. Use when the user wants to analyze agent telemetry traces to find bugs and get fix recommendations — walks through exporting traces from a local or remote watsonx Orchestrate server, parsing raw OTel JSON or Langfuse-format trace JSON directly, and reasoning over them to identify failures and suggest fixes.
Telemetry Analyzer fits situations like: the user wants to analyze agent telemetry traces to find bugs and get fix recommendations — walks through exporting traces from a local; remote watsonx Orchestrate server; parsing raw OTel JSON; langfuse-format trace JSON directly.
Run `npx skills add IBM/ibm-watsonx-orchestrate-adk --skill telemetry-analyzer -a claude-code`. Or copy the skill folder (skills/telemetry-analyzer in IBM/ibm-watsonx-orchestrate-adk) into .claude/skills/telemetry-analyzer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add IBM/ibm-watsonx-orchestrate-adk --skill telemetry-analyzer -a codex`. Or copy the skill folder (skills/telemetry-analyzer in IBM/ibm-watsonx-orchestrate-adk) into .agents/skills/telemetry-analyzer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add IBM/ibm-watsonx-orchestrate-adk --skill telemetry-analyzer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/telemetry-analyzer, .gemini/skills/telemetry-analyzer, .github/skills/telemetry-analyzer and .opencode/skills/telemetry-analyzer in your project.
Going by SKILL.md and its folder, Telemetry Analyzer needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Telemetry Analyzer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 10k tokens (SKILL.md is roughly 41k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Telemetry Analyzer: Backend Dev Guidelines (langfuse/langfuse, 35k stars), Langfuse and LLM Gateway Logs (KonghaYao/peri, 223 stars), Agentsop Observability Setup (agentsope/SkillAlchemy, 457 stars) and Olore Langfuse Latest (olorehq/olore, 103 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
IBM (a GitHub organization, an official publisher) maintains it in IBM/ibm-watsonx-orchestrate-adk, which has 178 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 2, 2026.
Source: IBM/ibm-watsonx-orchestrate-adk on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.