Agent skill

Error Analysis

by yonatangross in yonatangross/orchestkit

Evals-first error analysis for LLM apps: clusters real Langfuse or JSONL traces into a human-confirmed failure taxonomy with counts, then recommends binary pass/fail evals for recurring named modes.

MITAuto-check: notesAI & LLM Engineering

Install Error Analysis

skills CLI
$ npx skills add yonatangross/orchestkit --skill error-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yonatangross/orchestkit error-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/error-analysis .claude/skills/error-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
error-analysis
GitHub stars
289
Token cost
~3.6k tokens
SKILL.md length
1,530 words
Files
5 (incl. references)
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Evals-first error analysis for LLM apps: clusters real Langfuse or JSONL traces into a human-confirmed failure taxonomy with counts, then recommends binary pass/fail evals for recurring named modes.

  • Works in 5 steps: Pull Traces → Open Coding (human writes, Claude drafts) → Axial Coding (Claude proposes, human… → …
  • Learn what to measure before writing evals
  • SKILL.md covers When to Use, Task Management (CC 2.1.16), Effort Scaling (CC 2.1.76) and Phase 1: Pull Traces, plus 8 more sections
  • Calls curl; needs LANGFUSE_PUBLIC_KEY and LANGFUSE_SECRET_KEY

What it does

Error Analysis is an agent skill from yonatangross/orchestkit. Evals-first error analysis for LLM apps: clusters real Langfuse or JSONL traces into a human-confirmed failure taxonomy with counts, then recommends binary pass/fail evals for recurring named modes. Use to learn what to measure before writing evals. Not for CI failures.

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/judge-alignment.md`, `references/langfuse-traces.md` and `references/method.md`). Compatibility notes: Claude Code 2.1.277+. Needs Langfuse credentials in env or an exported traces JSONL.

It sits in AI & LLM Engineering, covering LLM evaluation, LLM observability and Failing and flaky tests. It works with Langfuse. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.

When your agent uses it

  • Learn what to measure before writing evals
  • Tasks that involve LLM evaluation
  • Tasks that involve LLM observability

Example prompts

  • “/error-analysis”

Requirements

  • Python 3
  • A credential in LANGFUSE_PUBLIC_KEY
  • A credential in LANGFUSE_SECRET_KEY
  • Compatibility (from SKILL.md): Claude Code 2.1.277+. Needs Langfuse credentials in env or an exported traces JSONL.
  • Pre-approved tools (allowed-tools): AskUserQuestion, Bash, Read, Write, Edit, Grep, Glob, Agent, TaskCreate, TaskUpdate, TaskList, TaskGet, TaskStop, WebFetch, WebSearch

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Pull Traces
  2. Open Coding (human writes, Claude drafts)
  3. Axial Coding (Claude proposes, human confirms)
  4. Write failure-taxonomy.md
  5. Recommend Evals

What it can do on your machine

Read from SKILL.md and the folder at commit 0ef71d2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • AskUserQuestion
    • Bash
    • Read
    • Write
    • Edit
    • Grep
    • Glob
    • Agent
    • TaskCreate
    • TaskUpdate

    …and 5 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • LANGFUSE_PUBLIC_KEY
    • LANGFUSE_SECRET_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Claude Code 2.1.277+. Needs Langfuse credentials in env or an exported traces JSONL.

    From compatibility in the SKILL.md frontmatter.

Context cost

Error Analysis loads about 3.6k tokens when it runs, and up to ~7.5k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 1,530 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: AskUserQuestion, Bash, Read, Write, Edit, Grep, Glob, Agent, TaskCreate, TaskUpdate, TaskList, TaskG

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from yonatangross/orchestkit at commit 0ef71d2, republished under its MIT licence (© yonatangross). 1,530 words, ~3,637 tokens.

Download SKILL.mdSave it as .claude/skills/error-analysis/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
error-analysis
description
Evals-first error analysis for LLM apps: clusters real Langfuse or JSONL traces into a human-confirmed failure taxonomy with counts, then recommends binary pass/fail evals for recurring named modes. Use to learn what to measure before writing evals. Not for CI failures.
allowed-tools
AskUserQuestion, Bash, Read, Write, Edit, Grep, Glob, Agent, TaskCreate, TaskUpdate, TaskList, TaskGet, TaskStop, WebFetch, WebSearch
compatibility
Claude Code 2.1.277+. Needs Langfuse credentials in env or an exported traces JSONL.
license
MIT
context
fork
background
false
user-invocable
false
disable-model-invocation
false
skills
testing-llm, memory
effort
high
model
sonnet
metadata.category
workflow-automation
metadata.version
1.0.0

Error Analysis

Evals-first failure analysis for LLM applications. The method (Hamel Husain, Shreya Shankar) is qualitative research applied to traces: a human open-codes real failures, Claude axial-codes the notes into a named taxonomy, and only the recurring named modes earn an automated eval. Error analysis decides what to measure; it never writes an eval for a mode that has no name and no count.

When to Use

  • An LLM feature (chatbot, RAG, agent, extractor) misbehaves and the team is guessing which evals to write
  • You have production traces in Langfuse (or a JSONL export) and need the top failure modes with counts
  • Eval scores exist but nobody trusts them because they were never grounded in real failures
  • A judge or metric exists and its agreement with human labels is unknown

Do NOT use for Claude Code session errors (errors skill), failing CI runs (ci-debug), or writing the evaluators themselves once modes are named (testing-llm, ork:eval-runner).

Task Management (CC 2.1.16)

Multi-phase workflow: create tasks before Phase 1 and keep status current.

python
t_run  = TaskCreate(subject="Error analysis: {target}", activeForm="Running error analysis on {target}")
t_pull = TaskCreate(subject="Pull traces", activeForm="Pulling traces")
t_open = TaskCreate(subject="Open-code failure notes", activeForm="Open-coding traces")
t_axial = TaskCreate(subject="Axial-code failure taxonomy", activeForm="Axial-coding notes")
t_tax  = TaskCreate(subject="Write failure-taxonomy.md", activeForm="Writing taxonomy")
t_eval = TaskCreate(subject="Recommend evals + judge alignment", activeForm="Recommending evals")
TaskUpdate(taskId=t_open,  addBlockedBy=[t_pull])
TaskUpdate(taskId=t_axial, addBlockedBy=[t_open])
TaskUpdate(taskId=t_tax,   addBlockedBy=[t_axial])
TaskUpdate(taskId=t_eval,  addBlockedBy=[t_tax])

Effort Scaling (CC 2.1.76)

EffortTrace poolNotesOutcome
low30human notes only, no Claude draftsdraft taxonomy
medium50drafts after 10 human notestaxonomy + counts
high (default)100drafts after 30 human notes, saturation checkfull deliverable

Phase 1: Pull Traces

Goal: a working pool of ~100 diverse traces (default N=100), skewed toward failures.

Source A, Langfuse public API. Requires LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY, LANGFUSE_HOST in env. Never echo or print the secret value; reference the variable names only.

bash
mkdir -p error-analysis/traces
# Credentials ride in the Authorization header, so let curl enforce the scheme:
# --proto '=https' fails closed on every non-HTTPS URL. Do not pattern-match the
# host yourself; uppercase schemes, scheme-less names, decimal IPs, and userinfo
# tricks all bypass a regex. Relax to '=http,https' ONLY for an exact local
# endpoint (http://localhost, http://127.0.0.1, http://[::1], optional port and
# path, no userinfo). For a self-hosted internal instance, use HTTPS or an SSH
# tunnel to localhost.
proto="=https"
host_lc=$(printf '%s' "$LANGFUSE_HOST" | tr 'A-Z' 'a-z')
case "$host_lc" in
  http://*)
    authority=${host_lc#http://}
    authority=${authority%%/*}
    case "$authority" in
      *@*) printf '%s\n' 'Refusing HTTP with userinfo in the URL; use HTTPS.' >&2; exit 1 ;;
    esac
    case "$authority" in
      localhost|localhost:*|127.0.0.1|127.0.0.1:*|\[::1\]|\[::1\]:*) proto="=http,https" ;;
      *) printf '%s\n' 'Refusing plain HTTP to a non-local host; use HTTPS or an SSH tunnel to localhost.' >&2; exit 1 ;;
    esac
    ;;
esac
# Langfuse v4 removed GET /api/public/traces (404); the v2 Observations API is the
# supported read. It returns observation rows, so group by traceId client-side.
# io is required in fields or the rows come back without input/output.
curl -sS --fail-with-body --proto "${proto:-=https}" -u "$LANGFUSE_PUBLIC_KEY:$LANGFUSE_SECRET_KEY" \
  "$LANGFUSE_HOST/api/public/v2/observations?fromStartTime=$FROM&toStartTime=$TO&limit=100&fields=core,basic,trace_context,io" \
  -o error-analysis/traces/page1.json

Paginate with the cursor from each response until ~100 distinct traceId values are collected (or the cursor is exhausted), aborting on any non-2xx status so an error body is not read as the end of the cursor. Then finish the selected traces: the last trace in the pool may be missing rows, so keep paging until every selected traceId has its root observation (the isRootObservation: true row, or exactly one parentObservationId == null row as fallback; zero or several null-parent rows means root-unknown, so exclude that trace and say so in the notes), or fetch the missing rows per trace with ?traceId=<id>. Group rows by traceId and normalize each trace to one JSONL line in error-analysis/traces.jsonl: {id, timestamp, input, output, scores, observations[]}. Observation rows do not carry scores; fetch them per trace id from the Scores API v3 (GET /api/public/v3/scores?traceId=<id>) and join them in, since low scores and negative feedback are the failure signals Phase 1 prioritizes. Keep the raw tool and DB observations; Phase 5 needs them for replay. On a v3 host the legacy GET /api/public/traces list still works; the v2/v1 shapes, deprecation detail, and the JSONL fallback live in references/langfuse-traces.md.

Source B, exported JSONL. If the user hands you a file, inspect the first line with head -1 and map its fields to the same normalized shape. Do not assume field names; Langfuse, Braintrust, and home-grown loggers all differ.

Prioritize traces that carry failure signals: thumbs-down feedback, low scores, exceptions, retries, escalations. If fewer than ~30 traces show any failure signal, say so and ask whether to analyze a random sample anyway.

Phase 2: Open Coding (human writes, Claude drafts)

Open coding is a human activity. The reviewer is the benevolent dictator: one person owns the labels so the taxonomy stays coherent. Claude accelerates, it does not label unattended.

For each failing trace:

  1. Render a compact view: input, final output, and each tool or DB observation with its output.
  2. Draft ONE short free-form candidate note (a sentence, not a category). Aim at the FIRST failure in the trace; upstream failures cascade, so tagging a downstream symptom as its own mode double-counts.
  3. The human accepts, edits, or replaces the note via AskUserQuestion (options: Accept, Edit, Skip; Other captures free text).
  4. Append to error-analysis/open-coding-notes.md: trace_id | note.
  5. Label passing traces too. Judge alignment needs true negatives, so the human also labels a slice of traces with no failure signal (roughly one pass for every two failures) with the same per-mode binary label. Without labeled passes, Phase 5 can report TPR but never an honest TNR.

Cadence per the method: the human writes the first ~30 notes largely unaided (Claude drafts may be shown but the human decides), then Claude may search remaining traces for likely instances of the failure patterns seen so far and the human accepts or rejects each suggestion. Continue until theoretical saturation: new traces stop revealing new failure modes. ~100 diverse traces is the usual working pool, not a quota.

Prompt or code? When a note implicates a tool or retrieval step, replay it: re-run that tool call or DB query with the inputs recorded in the trace. Replay live only when the call is read-only; a payment, message send, or DB write replays against a sandbox or a recorded fixture, never a live service, or it is not replayed at all. Compare the replayed output with the observation recorded in the trace before assigning a verdict. If the replayed output matches the recorded one and it was wrong, the failure is code or data (the tool itself). If the replayed output matches a correct recorded output yet the final answer was wrong, the failure is prompt or context (the model had good inputs and used them badly). If the replayed output differs from the recorded one, the verdict is inconclusive until the difference is explained (stale data, drift, flaky tool), because the recorded output is what the model actually saw. Record the verdict in the note, for example prompt, code:retriever, or inconclusive:replay-diverged. This one check splits the taxonomy into things a prompt edit can fix and things it cannot.

Show full SKILL.md (593 more words)Show less

Phase 3: Axial Coding (Claude proposes, human confirms)

When open coding slows or the pool is exhausted, cluster the notes:

  1. Claude reads open-coding-notes.md and proposes 5 to 8 named failure modes. Each mode gets a name, a one-sentence definition, and the notes it absorbs.
  2. Present the proposed modes with per-mode note counts via AskUserQuestion. The human confirms, merges, splits, renames, or rejects modes. This confirmation is mandatory; an unconfirmed taxonomy is a draft.
  3. Fewer than 5 modes usually means over-merged categories that will produce vague judges. More than 8 usually means noise modes with count 1 or 2 that should fold into a sibling or a misc bucket rather than earn a name.

Phase 4: Write failure-taxonomy.md

Write error-analysis/failure-taxonomy.md with two artifacts:

Table 1, the taxonomy:

Failure modeDefinitionExample trace idsCountEval?

Table 2, counts:

ModeCount% of failuresFix surface (prompt/code/data)

Rules for the Eval? column: yes only for modes that are named, recurring (count >= ~3 or a top share of failures), and expected to persist after an obvious prompt fix. One-off bugs, upstream data issues, and modes a code assertion can check deterministically get no with a reason in Phase 5. Sort both tables by count descending.

Phase 5: Recommend Evals

For each eval: yes mode, write a recommendation in error-analysis/eval-recommendations.md:

  • Name and the failure mode it detects
  • Check: binary pass/fail only. Phrase it so a grader answers "did failure X occur, yes or no". No Likert scales; a 1-5 score hides disagreement inside the middle values and makes judge alignment unmeasurable.
  • Critique: a written paragraph covering what the eval can and cannot catch, likely false-positive sources, and the cheapest viable implementation (code assertion before LLM judge; a judge is a classifier returning pass or fail, one judge per mode, never one judge grading overall quality).
  • Judge alignment: if the check needs an LLM judge, it is untrusted until measured against the Phase 2 human labels, which must include both labeled failures and labeled passes for the mode. Split labeled examples into train/dev/test, iterate the judge prompt on dev disagreements, and report TPR (failures caught) and TNR (good outputs passed) on the untouched test set. The full protocol is references/judge-alignment.md.

Delegate execution to the eval-runner agent (ork:eval-runner) when the user wants the evals actually run against a dataset; this skill produces the taxonomy and the eval specs, the runner executes and scores them.

python
Agent(subagent_type="ork:eval-runner",
      prompt="Run the recommended evals in error-analysis/eval-recommendations.md against <dataset>. Report TPR/TNR vs the human labels in error-analysis/open-coding-notes.md.")

Output Artifacts

FileContent
error-analysis/traces.jsonlNormalized trace pool
error-analysis/open-coding-notes.mdtrace_id, note, prompt-or-code verdict
error-analysis/failure-taxonomy.mdNamed modes, definitions, examples, counts, eval yes/no
error-analysis/eval-recommendations.mdBinary eval specs, critiques, judge-alignment status

Key Decisions

DecisionRecommendation
Who labelsOne human (benevolent dictator); Claude drafts, human decides
Which failure to noteThe first one in the trace; downstream symptoms cascade from it
How many modes5 to 8, human-confirmed
Eval granularityBinary pass/fail, one judge per mode
Which modes get evalsNamed and recurring only; fix trivial prompt gaps first
Judge trustUnmeasured until TPR/TNR vs human labels on a held-out test set

References

  • references/method.md: the open/axial coding method, saturation, and the source list (hamel.dev evals FAQ, hamel.dev LLM-as-judge, Anthropic eval docs)
  • references/langfuse-traces.md: Langfuse public API trace pull, pagination, JSONL export shape
  • references/judge-alignment.md: train/dev/test splits, TPR/TNR, disagreement-driven judge iteration
  • testing-llm: evaluation frameworks, Langfuse SDK v4, metric thresholds for the evals this skill recommends
  • errors: Claude Code session errors, a different domain than LLM app traces
  • cover: generates tests; run it after the taxonomy names what to test
  • verify: grades the resulting eval suite once it exists
  • ork:eval-runner: executes the recommended evals and reports scores to Langfuse

© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in src/skills/error-analysis of yonatangross/orchestkit.

  • SKILL.md
  • references/judge-alignment.md
  • references/langfuse-traces.md
  • references/method.md
  • test-cases.json

Open the folder on GitHubat commit 0ef71d2

Compare with similar skills

Error Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Error Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Error Analysis this skillyonatangross/orchestkit289—~3.6kAutomated safety check: NotesMIT
Langfuse Codebase Navigatorlangfuse/langfuse36k—~1.4kAutomated safety check: PassCustom licence
LLM Trace Review Interfaceai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.0
Langfuse Integration Pagelangfuse/langfuse-docs246—~3.7kAutomated safety check: PassMIT
Langfuselangfuse/skills299—~2.1kAutomated safety check: NotesMIT
Add Yourself To Team Langfuselangfuse/langfuse-docs246—~548Automated safety check: PassMIT

Similar skills

  • Navigate Langfuse repositories, code areas, and agent skills.

    36k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • LLM Trace Review Interface

    ai-evals-course/evals-skills

    Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.

    1.5k GitHub stars~1.4k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Langfuse Integration Page

    langfuse/langfuse-docs

    Create a new Langfuse integration page in the langfuse-docs repo.

    246 GitHub stars~3.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Langfuse

    langfuse/skills

    Interact with Langfuse and access its documentation: tracing, monitoring, creating datasets, running experiments, and evaluating AI applications.

    299 GitHub stars~2.1k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check: notes
  • Add Yourself To Team Langfuse

    langfuse/langfuse-docs

    Add a new team member to Langfuse's canonical team data and shared team table.

    246 GitHub stars~548 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Phoenix LLM Observability

    Orchestra-Research/AI-Research-SKILLs

    Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

    13k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed

More from yonatangross/orchestkit

All 108 skills in this repo
  • API Design

    yonatangross/orchestkit

    API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.

    289 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Architecture Decision Record

    yonatangross/orchestkit

    ADR templates in the Nygard format with context, decision, consequences, and alternatives.

    289 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Audit Full

    yonatangross/orchestkit

    Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.

    289 GitHub stars~3.5k tokensUpdated today
    Auto-check: notes
  • Code Review Playbook

    yonatangross/orchestkit

    Structured review processes, conventional comments, language-specific checklists, and feedback templates.

    289 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Create PR

    yonatangross/orchestkit

    Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.

    289 GitHub stars~4.5k tokensUpdated today
    Auto-check: notes
  • Explore

    yonatangross/orchestkit

    Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.

    289 GitHub stars~3.9k tokensUpdated today
    Auto-check: notes

Works with

Questions about Error Analysis

What does Error Analysis do?

Evals-first error analysis for LLM apps: clusters real Langfuse or JSONL traces into a human-confirmed failure taxonomy with counts, then recommends binary pass/fail evals for recurring named modes. Error Analysis is an agent skill from yonatangross/orchestkit. Evals-first error analysis for LLM apps: clusters real Langfuse or JSONL traces into a human-confirmed failure taxonomy with counts, then recommends binary pass/fail evals for recurring named modes.

When should I use Error Analysis?

Error Analysis fits situations like: learn what to measure before writing evals; tasks that involve LLM evaluation; tasks that involve LLM observability.

How do I install Error Analysis in Claude Code?

Run `npx skills add yonatangross/orchestkit --skill error-analysis -a claude-code`. Or copy the skill folder (src/skills/error-analysis in yonatangross/orchestkit) into .claude/skills/error-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Error Analysis in Codex?

Run `npx skills add yonatangross/orchestkit --skill error-analysis -a codex`. Or copy the skill folder (src/skills/error-analysis in yonatangross/orchestkit) into .agents/skills/error-analysis in your project. Codex loads it when a task matches its description.

Can I use Error Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill error-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/error-analysis, .gemini/skills/error-analysis, .github/skills/error-analysis and .opencode/skills/error-analysis in your project.

What does Error Analysis need to run?

Going by SKILL.md and its folder, Error Analysis needs the command-line tools its instructions call (curl) and credentials named LANGFUSE_PUBLIC_KEY and LANGFUSE_SECRET_KEY. Our summary lists: Python 3; A credential in LANGFUSE_PUBLIC_KEY; A credential in LANGFUSE_SECRET_KEY. Its frontmatter pre-approves these tools: AskUserQuestion, Bash, Read, Write, Edit, Grep, Glob, Agent, TaskCreate, TaskUpdate, TaskList, TaskGet, TaskStop, WebFetch, WebSearch. Compatibility (from SKILL.md): Claude Code 2.1.277+. Needs Langfuse credentials in env or an exported traces JSONL..

Does Error Analysis access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Error Analysis safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Error Analysis use?

Error Analysis is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Error Analysis use?

About 3.6k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.8k tokens, read only when the agent opens those files.

What are the alternatives to Error Analysis?

Skills that share tags, products or a category with Error Analysis: Langfuse Codebase Navigator (langfuse/langfuse, 36k stars), LLM Trace Review Interface (ai-evals-course/evals-skills, 1.5k stars), Langfuse Integration Page (langfuse/langfuse-docs, 246 stars) and Langfuse (langfuse/skills, 299 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Error Analysis?

yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 289 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 7, 2026.

Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.