Official agent skill

Exploring AI Failures

by PostHog in PostHog/posthog

Find where an AI/LLM application is failing in production and surface the failure patterns, working from real traces.

OfficialCustom licenceAuto-check: warnings

Install Exploring AI Failures

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add PostHog/posthog --skill exploring-ai-failures -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PostHog/posthog exploring-ai-failures --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PostHog/posthog.git skills-src && mkdir -p .claude/skills && cp -r skills-src/products/ai_observability/skills/exploring-ai-failures .claude/skills/exploring-ai-failures && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
exploring-ai-failures
GitHub stars
40k
Token cost
~2.9k tokens
SKILL.md length
1,634 words
Files
2 (incl. references)
Skills in repo
252
Repo updated
First seen
Licence
Custom licence

At a glance

Find where an AI/LLM application is failing in production and surface the failure patterns, working from real traces.

  • Works in 4 steps: Scope to one use case → Pick which traces to read → Read a batch (this is the job) → …
  • Someone wants to understand whats going wrong with an AI feature
  • SKILL.md covers Tools, Work with the user, Step 1 — Scope to one use case and Step 2 — Pick which traces to…, plus 6 more sections
  • Reaches us.posthog.com

What it does

Exploring AI Failures is an agent skill from PostHog/posthog, published by the product's own GitHub organization. Find where an AI/LLM application is failing in production and surface the failure patterns, working from real traces. Use when someone wants to understand what's going wrong with an AI feature, find and categorize failure modes, triage errors, or investigate quality issues (wrong answers, ignored instructions, hallucinations, tool misuse) — "what's failing in my agent", "surface error patterns", "why are the responses bad", "find the common failure modes", "what should I fix next". Covers scoping to one use case…

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/finding-traces.md`).

It works with PostHog. The repository describes itself as: :hedgehog: PostHog is the leading platform for building self-driving products. Our developer tools – AI observability, analytics, session replay, flags, experiments, error…

When your agent uses it

  • Someone wants to understand whats going wrong with an AI feature
  • Find and categorize failure modes
  • Investigate quality issues (wrong answers
  • Ignored instructions

Example prompts

  • “s failing in my agent”
  • “surface error patterns”
  • “why are the responses bad”
  • “/exploring-ai-failures”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Scope to one use case
  2. Pick which traces to read
  3. Read a batch (this is the job)
  4. Rank, link, and hand back to the user

What it can do on your machine

Read from SKILL.md and the folder at commit 10f9ad7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • us.posthog.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Exploring AI Failures loads about 2.9k tokens when it runs, and up to ~3.9k if it reads all its reference files. Until then it costs about 187 tokens; SKILL.md has 1,634 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~187
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningTells the agent its actions are pre-authorized / not to stop for confirmationSKILL.md:53
    coded failure modes** — don't stop to ask permission before the reading; that reading is the core

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 1,634 words (~2,917 tokens).

“The highest-value thing you can do with production AI traffic is look at where it fails and name the patterns. The catch: most failures are silent. The model returns a clean response — HTTP 200, no exception — that is…”

— opening of SKILL.md by PostHog, Custom licence
name
exploring-ai-failures

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file (references) in products/ai_observability/skills/exploring-ai-failures of PostHog/posthog.

  • SKILL.md
  • references/finding-traces.md

Open the folder on GitHubat commit 10f9ad7

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in PostHog/posthog, which our catalogue first saw on October 8, 2026.

Compare with similar skills

Exploring AI Failures next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Exploring AI Failures compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Exploring AI Failures this skillPostHog/posthog40k—~2.9kAutomated safety check: WarnCustom licence
Opik Analytics Instrumentationcomet-ml/opik22k—~4.4kAutomated safety check: PassApache-2.0
C15tc15t/c15t1.9k1 repos~1.6kAutomated safety check: PassApache-2.0
Define Feature Flagmacro-inc/macro4.6k—~780Automated safety check: PassAGPL-3.0
Soku CLIAbout-Intelligence/soku-cli304—~2.4kAutomated safety check: PassMIT
Compare Array Bundle SizePostHog/posthog-js613—~599Automated safety check: PassCustom licence

Similar skills

  • Shows how to add product analytics events to Opik's frontend, Java backend and Python SDK, all reporting through Segment to PostHog with an opik_ name prefix.

    22k GitHub stars~4.4k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • C15t

    c15t/c15t

    Work with c15t consent management docs, APIs, and integrations for Next.js, React, and JavaScript.

    1.9k GitHub starsUsed in 1 repo~1.6k tokens
    Legal & ComplianceAuto-check passed
  • Define Feature Flag

    macro-inc/macro

    Define a frontend feature flag with defineFlag and wire its readers.

    4.6k GitHub stars~780 tokensUpdated today
    Frontend & DesignAuto-check passed
  • Soku CLI

    About-Intelligence/soku-cli

    Guides an agent through the soku command line tool for ads, GA4 and PostHog data reads, ads writes, SEO hosting, automations, files and skill management.

    304 GitHub stars~2.4k tokensUpdated today
    Marketing & SEOAuto-check passed
  • Compare Array Bundle Size

    PostHog/posthog-js

    Official

    Quickly compare the posthog-js array.js bundle size in the current working tree against a git baseline using the repository's esbuild proxy.

    613 GitHub stars~599 tokensUpdated today
    Frontend & DesignAuto-check passed
  • Telemetry Analytics

    OpenHands/OpenHands

    This skill should be used when the user asks to "add tracking", "add a PostHog event", "change telemetry consent", "instrument onboarding", "debug analytics", or changes telemetry.ts…

    90k GitHub stars~305 tokensUpdated today
    DevOps & CloudAuto-check passed

More from PostHog/posthog

All 252 skills in this repo
  • Authoring Log Alerts

    PostHog/posthog

    Official

    Author useful, low-noise log alerts on services in a PostHog project.

    40k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Official

    Operating procedure for the conflict-autoresolver agent: sweep open PostHog/posthog PRs that conflict with master, resolve the trivial conflicts (generated artifacts deterministically, source…

    40k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Help users debug PostHog Error Tracking stack-trace symbolication for any supported platform — JavaScript/TypeScript web, React Native (Hermes), Android (Proguard / R8), or iOS / macOS (dSYM).

    40k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Exploring Apm Traces

    PostHog/posthog

    Official

    Investigates distributed application performance using PostHog APM (OpenTelemetry span) data via MCP.

    40k GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Exploring LLM Traces

    PostHog/posthog

    Official

    Debug and inspect LLM/AI agent traces using PostHog's MCP tools.

    40k GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Investigate Metric

    PostHog/posthog

    Official

    Diagnose why a product metric changed (dropped, spiked, or plateaued) by orchestrating breakdowns, actors, paths, lifecycle, retention, and annotations queries.

    40k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Works with

Questions about Exploring AI Failures

What does Exploring AI Failures do?

Find where an AI/LLM application is failing in production and surface the failure patterns, working from real traces. Exploring AI Failures is an agent skill from PostHog/posthog, published by the product's own GitHub organization. Find where an AI/LLM application is failing in production and surface the failure patterns, working from real traces.

When should I use Exploring AI Failures?

Exploring AI Failures fits situations like: someone wants to understand whats going wrong with an AI feature; find and categorize failure modes; investigate quality issues (wrong answers; ignored instructions.

How do I install Exploring AI Failures in Claude Code?

Run `npx skills add PostHog/posthog --skill exploring-ai-failures -a claude-code`. Or copy the skill folder (products/ai_observability/skills/exploring-ai-failures in PostHog/posthog) into .claude/skills/exploring-ai-failures in your project. Claude Code loads it when a task matches its description.

How do I install Exploring AI Failures in Codex?

Run `npx skills add PostHog/posthog --skill exploring-ai-failures -a codex`. Or copy the skill folder (products/ai_observability/skills/exploring-ai-failures in PostHog/posthog) into .agents/skills/exploring-ai-failures in your project. Codex loads it when a task matches its description.

Can I use Exploring AI Failures in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PostHog/posthog --skill exploring-ai-failures -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/exploring-ai-failures, .gemini/skills/exploring-ai-failures, .github/skills/exploring-ai-failures and .opencode/skills/exploring-ai-failures in your project.

What does Exploring AI Failures need to run?

SKILL.md names no scripts, command-line tools or credentials: Exploring AI Failures is instructions for the agent only.

Does Exploring AI Failures access the network?

SKILL.md names 1 domain. In commands or code: us.posthog.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Exploring AI Failures safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Exploring AI Failures use?

Exploring AI Failures has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Exploring AI Failures use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 965 tokens, read only when the agent opens those files.

What are the alternatives to Exploring AI Failures?

Skills that share tags, products or a category with Exploring AI Failures: Opik Analytics Instrumentation (comet-ml/opik, 22k stars), C15t (c15t/c15t, 1.9k stars), Define Feature Flag (macro-inc/macro, 4.6k stars) and Soku CLI (About-Intelligence/soku-cli, 304 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Exploring AI Failures?

PostHog (a GitHub organization, an official publisher) maintains it in PostHog/posthog, which has 40,182 GitHub stars. The repository holds 252 skills in this directory. The repository was last updated on October 8, 2026.

Source: PostHog/posthog on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.