Official agent skill

Exploring LLM Clusters

by PostHog in PostHog/posthog-foss

Investigate AI observability clusters — understand usage patterns in AI/LLM traffic, compare cluster behavior, compute cost/latency metrics, and drill into individual traces within clusters.

OfficialMITAuto-check passedDevOps & Cloud

Install Exploring LLM Clusters

skills CLI
$ npx skills add PostHog/posthog-foss --skill exploring-llm-clusters -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PostHog/posthog-foss exploring-llm-clusters --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .claude/skills && cp -r skills-src/products/ai_observability/skills/exploring-llm-clusters .claude/skills/exploring-llm-clusters && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
exploring-llm-clusters
GitHub stars
721
Token cost
~3.1k tokens
SKILL.md length
972 words
Files
2 (incl. scripts)
Skills in repo
213
Repo updated
First seen
Licence
MIT

At a glance

Investigate AI observability clusters — understand usage patterns in AI/LLM traffic, compare cluster behavior, compute cost/latency metrics, and drill into individual traces within clusters.

  • Works in 4 steps: List recent clustering runs → Get clusters from a specific run → Compute metrics for clusters → …
  • Tasks that involve Observability
  • SKILL.md covers Tools, How clustering works, Clustering jobs and Workflow: explore clusters, plus 3 more sections
  • Runs Python scripts from its folder

What it does

Exploring LLM Clusters is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Investigate AI observability clusters — understand usage patterns in AI/LLM traffic, compare cluster behavior, compute cost/latency metrics, and drill into individual traces within clusters.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/print_clusters.py`).

It sits in DevOps & Cloud, covering Observability. It works with PostHog. The repository describes itself as: PostHog FOSS is a read-only mirror of PostHog, with all proprietary code removed. NOTE: This repo is synced automatically from the main PostHog repo. Please raise any issues and… The licence is MIT.

When your agent uses it

  • Tasks that involve Observability

Example prompts

  • “/exploring-llm-clusters”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. List recent clustering runs
  2. Get clusters from a specific run
  3. Compute metrics for clusters
  4. Drill into specific traces

What it can do on your machine

Read from SKILL.md and the folder at commit 2c48221. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Exploring LLM Clusters loads about 3.1k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 972 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from PostHog/posthog-foss at commit 2c48221, republished under its MIT licence (© PostHog). 972 words, ~3,086 tokens.

Download SKILL.mdSave it as .claude/skills/exploring-llm-clusters/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
exploring-llm-clusters
description
Investigate AI observability clusters — understand usage patterns in AI/LLM traffic, compare cluster behavior, compute cost/latency metrics, and drill into individual traces within clusters.

Exploring LLM clusters

Use this skill when investigating AI observability clusters — understanding what patterns exist in your AI/LLM traffic, comparing cluster behavior, and drilling into individual clusters.

Tools

ToolPurpose
posthog:llma-clustering-job-listList clustering job configurations for the team
posthog:llma-clustering-job-getGet a specific clustering job by ID
posthog:execute-sqlQuery cluster run events and compute metrics
posthog:query-llm-traces-listFind traces belonging to a cluster
posthog:query-llm-traceInspect a specific trace in detail

How clustering works

PostHog clusters LLM traces, individual generations, or evaluation events by embedding similarity. A Temporal workflow runs periodically or on-demand, producing cluster events stored as $ai_trace_clusters (trace-level), $ai_generation_clusters (generation-level), or $ai_evaluation_clusters (evaluation-level).

Each cluster event contains:

  • $ai_clustering_run_id — unique run identifier (format: <team_id>_<level>_<YYYYMMDD>_<HHMMSS>[_<job_id>])
  • $ai_clustering_level — "trace", "generation", or "evaluation"
  • $ai_window_start / $ai_window_end — time window of the data that was analyzed
  • $ai_total_items_analyzed — number of traces, generations, or evaluations processed
  • $ai_clusters — JSON array of cluster objects
  • $ai_clustering_params — algorithm parameters used

The analyzed window closes when a run starts, and the cluster event lands once the run finishes. So the cluster event's own timestamp is always after $ai_window_end, by anything from seconds to hours. Use the window only to bound the traces, generations, and evaluations that were analyzed. To find the cluster event itself, filter on $ai_clustering_run_id with a plain recent-time bound.

Cluster object shape (inside $ai_clusters)
json
{
  "cluster_id": 0,
  "size": 42,
  "title": "User authentication flows",
  "description": "Traces involving login, signup, and token refresh operations",
  "traces": {
    "<trace_or_generation_id>": {
      "distance_to_centroid": 0.123,
      "rank": 0,
      "x": -2.34,
      "y": 1.56,
      "timestamp": "2026-03-28T10:00:00Z",
      "trace_id": "abc-123",
      "generation_id": "gen-456"
    }
  },
  "centroid_x": -2.1,
  "centroid_y": 1.4
}
  • cluster_id: -1 is the noise/outlier cluster (items that didn't fit any cluster)
  • Items in traces are keyed by trace ID (trace-level), generation event UUID (generation-level), or evaluation event UUID (evaluation-level)
  • rank orders items by proximity to centroid (0 = closest)
  • x, y are 2D coordinates for visualization (UMAP/PCA/t-SNE reduced)

Clustering jobs

Each team can have up to 10 clustering jobs. A job defines:

  • name — human-readable label
  • analysis_level — "trace", "generation", or "evaluation"
  • event_filters — property filters scoping which items are included
  • enabled — whether the job runs on schedule

Default jobs named "Default - traces", "Default - generations", and "Default - evaluations" are auto-created and disabled when a custom job is created for the same level.

Workflow: explore clusters

Step 1 — List recent clustering runs
sql
posthog:execute-sql
SELECT
    toString(properties.$ai_clustering_run_id) AS run_id,
    toString(properties.$ai_clustering_level) AS level,
    toString(properties.$ai_clustering_job_id) AS job_id,
    toString(properties.$ai_clustering_job_name) AS job_name,
    toString(properties.$ai_window_start) AS window_start,
    toString(properties.$ai_window_end) AS window_end,
    toFloat64OrNull(toString(properties.$ai_total_items_analyzed)) AS total_items,
    timestamp
FROM events
WHERE event IN ('$ai_trace_clusters', '$ai_generation_clusters', '$ai_evaluation_clusters')
    AND timestamp >= now() - INTERVAL 14 DAY
ORDER BY timestamp DESC
LIMIT 10
Step 2 — Get clusters from a specific run
sql
posthog:execute-sql
SELECT
    toString(properties.$ai_clustering_run_id) AS run_id,
    toString(properties.$ai_clustering_level) AS level,
    toString(properties.$ai_clustering_job_id) AS job_id,
    toString(properties.$ai_clustering_job_name) AS job_name,
    toString(properties.$ai_window_start) AS window_start,
    toString(properties.$ai_window_end) AS window_end,
    toFloat64OrNull(toString(properties.$ai_total_items_analyzed)) AS total_items,
    properties.$ai_clusters AS clusters,
    properties.$ai_clustering_params AS params,
    timestamp
FROM events
WHERE event IN ('$ai_trace_clusters', '$ai_generation_clusters', '$ai_evaluation_clusters')
    AND timestamp >= now() - INTERVAL 14 DAY
    AND toString(properties.$ai_clustering_run_id) = '<run_id>'
ORDER BY timestamp DESC
LIMIT 1

Keep the lookback bound wide enough to cover the timestamp Step 1 reported for the run. Never bound this query with $ai_window_start / $ai_window_end. The cluster event is emitted after the window closes, so those bounds return zero rows.

The clusters field is a JSON array. Parse it to see cluster titles, sizes, descriptions, optional metrics, and each cluster's traces map.

Important: The clusters JSON can be very large (thousands of trace, generation, or evaluation IDs with coordinates). When the result is too large for inline display, it auto-persists to a file. Use print_clusters.py from scripts/ to get a readable summary.

Step 3 — Compute metrics for clusters

For trace-level clusters, compute cost/latency/token metrics:

sql
posthog:execute-sql
SELECT
    properties.$ai_trace_id as trace_id,
    sum(toFloat(properties.$ai_total_cost_usd)) as total_cost,
    max(toFloat(properties.$ai_latency)) as latency,
    sum(toInt(properties.$ai_input_tokens)) as input_tokens,
    sum(toInt(properties.$ai_output_tokens)) as output_tokens,
    countIf(properties.$ai_is_error = 'true') as error_count
FROM events
WHERE event IN ('$ai_generation', '$ai_embedding', '$ai_span')
    AND timestamp >= parseDateTimeBestEffort('<window_start>')
    AND timestamp <= parseDateTimeBestEffort('<window_end>')
    AND properties.$ai_trace_id IN ('<trace_id_1>', '<trace_id_2>', ...)
GROUP BY trace_id

For generation-level clusters, match by event UUID:

sql
posthog:execute-sql
SELECT
    toString(uuid) as generation_id,
    toFloat(properties.$ai_total_cost_usd) as cost,
    toFloat(properties.$ai_latency) as latency,
    toInt(properties.$ai_input_tokens) as input_tokens,
    toInt(properties.$ai_output_tokens) as output_tokens,
    if(properties.$ai_is_error = 'true', 1, 0) as is_error
FROM events
WHERE event = '$ai_generation'
    AND timestamp >= parseDateTimeBestEffort('<window_start>')
    AND timestamp <= parseDateTimeBestEffort('<window_end>')
    AND uuid IN ('<gen_uuid_1>', '<gen_uuid_2>', ...)

For evaluation-level clusters, first check each cluster's metrics field from $ai_clusters (for example pass rate, N/A rate, dominant evaluator name, and average judge cost). When you need individual evaluation rows, match by event UUID:

sql
posthog:execute-sql
SELECT
    toString(uuid) AS evaluation_id,
    toString(properties.$ai_trace_id) AS trace_id,
    toString(properties.$ai_target_event_id) AS generation_id,
    toString(properties.$ai_evaluation_name) AS evaluation_name,
    toString(properties.$ai_evaluation_result) AS evaluation_result,
    toString(properties.$ai_evaluation_reasoning) AS evaluation_reasoning,
    toFloatOrNull(toString(properties.$ai_total_cost_usd)) AS judge_cost,
    timestamp
FROM events
WHERE event = '$ai_evaluation'
    AND timestamp >= parseDateTimeBestEffort('<window_start>')
    AND timestamp <= parseDateTimeBestEffort('<window_end>')
    AND uuid IN ('<eval_uuid_1>', '<eval_uuid_2>', ...)
Step 4 — Drill into specific traces

Once you've identified interesting clusters, use the trace tools to inspect individual traces:

json
posthog:query-llm-trace
{
  "traceId": "<trace_id_from_cluster>",
  "dateRange": {"date_from": "<window_start>", "date_to": "<window_end>"}
}
When you need message content

Use events for cluster events, IDs, cost/latency/token metrics, and evaluation rows. Do not query events.properties.$ai_input, $ai_output, or $ai_output_choices when you need user messages or full model inputs/outputs — those heavy fields live on posthog.ai_events.

For a few representative examples, prefer query-llm-trace; it reads posthog.ai_events for you and returns the full event tree. For batch extraction, first get the trace IDs from the cluster, then query posthog.ai_events anchored on trace_id:

sql
posthog:execute-sql
SELECT
    trace_id,
    timestamp,
    span_id,
    event,
    model,
    input,
    output_choices
FROM posthog.ai_events
WHERE trace_id IN ('<trace_id_1>', '<trace_id_2>', ...)
ORDER BY trace_id, timestamp

posthog.ai_events has a shorter retention window than events; older clusters may still have metadata and metrics but no message content. For more detail, use the exploring LLM traces skill's event reference.

Show full SKILL.md (359 more words)Show less

Investigation patterns

"What kinds of LLM usage do we have?"
  1. List recent clustering runs (Step 1)
  2. Load the latest run's clusters (Step 2)
  3. Review cluster titles and descriptions — each represents a distinct usage pattern
  4. Compare cluster sizes to understand traffic distribution
"Which cluster is most expensive / slowest?"
  1. Load clusters from a run (Step 2)
  2. Extract trace IDs from each cluster
  3. Compute metrics per cluster (Step 3)
  4. Aggregate: avg(cost), avg(latency), sum(cost) per cluster
  5. Compare across clusters
"What's in this cluster?"
  1. Load the cluster's traces (from the traces field)
  2. Sort by rank (closest to centroid = most representative)
  3. Inspect the top 3-5 traces via query-llm-trace to understand the pattern
  4. Check the cluster title and description for the AI-generated summary
"Are there error-heavy clusters?"
  1. Compute metrics (Step 3) with error_count
  2. Calculate error rate per cluster: items_with_errors / total_items
  3. Focus on clusters with high error rates
  4. Drill into errored traces to find root causes
"How do clusters compare across runs?"
  1. List multiple runs (Step 1)
  2. Load clusters from each run
  3. Compare cluster titles — similar titles across runs indicate stable patterns
  4. Track cluster size changes to detect shifts in traffic patterns
  • Clusters overview: https://app.posthog.com/ai-observability/clusters
  • Specific run: https://app.posthog.com/ai-observability/clusters/<url_encoded_run_id>
  • Cluster detail: https://app.posthog.com/ai-observability/clusters/<url_encoded_run_id>/<cluster_id>

Always surface these links so the user can verify visually in the PostHog UI.

Tips

  • Always set a time range in SQL queries — cluster events without time bounds are slow
  • Bound a search for a cluster event by when it was emitted, not by the window it analyzed — a run's $ai_window_end is earlier than the event's own timestamp
  • Start with run listing to orient, then drill into specific clusters
  • Cluster titles and descriptions are AI-generated summaries — verify by inspecting traces
  • The noise cluster (cluster_id: -1) contains outliers that didn't fit any pattern
  • Use llma-clustering-job-list to understand what clustering configs are active
  • Trace IDs in clusters can be used directly with query-llm-trace for deep inspection
  • Message content lives on posthog.ai_events, not events.properties; use query-llm-trace unless you need custom batch SQL
  • For large clusters, inspect the top-ranked traces (closest to centroid) for representative examples

© PostHog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in products/ai_observability/skills/exploring-llm-clusters of PostHog/posthog-foss.

  • SKILL.md
  • scripts/print_clusters.py

Open the folder on GitHubat commit 2c48221

Compare with similar skills

Exploring LLM Clusters next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Exploring LLM Clusters compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Exploring LLM Clusters this skillPostHog/posthog-foss721—~3.1kAutomated safety check: PassMIT
Telemetry AnalyticsOpenHands/OpenHands90k—~305Automated safety check: PassMIT
Temps Best Practicesgotempsh/temps822—~2.9kAutomated safety check: PassApache-2.0
Posthog Operationsshepherdjerred/monorepo112—~321Automated safety check: PassGPL-3.0
AI Observability Langchain PythonJwuthri/Tracely-ai1.5k—~2.2kAutomated safety check: PassMIT
Vercel Optimize Auditvercel-labs/agent-skills32k8 repos~4.3kAutomated safety check: PassNone

Similar skills

  • Telemetry Analytics

    OpenHands/OpenHands

    This skill should be used when the user asks to "add tracking", "add a PostHog event", "change telemetry consent", "instrument onboarding", "debug analytics", or changes telemetry.ts…

    90k GitHub stars~305 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Temps Best Practices

    gotempsh/temps

    Best-practices reference for preparing and instrumenting applications on Temps.

    822 GitHub stars~2.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Posthog Operations

    shepherdjerred/monorepo

    Query or manage this repository's PostHog analytics, schema, dashboards, insights, feature flags, experiments, replay, and observability through toolkit posthog.

    112 GitHub stars~321 tokensUpdated today
    DevOps & CloudAuto-check passed
  • PostHog AI Observability integration for LangChain (Python). An agent skill from Jwuthri/Tracely-ai.

    1.5k GitHub stars~2.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Vercel Optimize Audit

    vercel-labs/agent-skills

    Official

    Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.

    32k GitHub starsUsed in 8 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Kubeshark Installer

    kubeshark/kubeshark

    Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.

    12k GitHub stars~3.6k tokensUpdated 6 days ago
    DevOps & CloudAuto-check: notes

More from PostHog/posthog-foss

All 213 skills in this repo
  • Authoring Log Alerts

    PostHog/posthog-foss

    Official

    Author useful, low-noise log alerts on services in a PostHog project.

    721 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Autoresolving PR Conflicts

    PostHog/posthog-foss

    Official

    Operating procedure for the conflict-autoresolver agent: sweep open PostHog/posthog PRs that conflict with master, resolve the trivial conflicts (generated artifacts deterministically, source…

    721 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Help users debug PostHog Error Tracking stack-trace symbolication for any supported platform — JavaScript/TypeScript web, React Native (Hermes), Android (Proguard / R8), or iOS / macOS (dSYM).

    721 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Exploring Apm Traces

    PostHog/posthog-foss

    Official

    Investigates distributed application performance using PostHog APM (OpenTelemetry span) data via MCP.

    721 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Exploring LLM Traces

    PostHog/posthog-foss

    Official

    Debug and inspect LLM/AI agent traces using PostHog's MCP tools.

    721 GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Investigate Metric

    PostHog/posthog-foss

    Official

    Diagnose why a product metric changed (dropped, spiked, or plateaued) by orchestrating breakdowns, actors, paths, lifecycle, retention, and annotations queries.

    721 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Exploring LLM Clusters

What does Exploring LLM Clusters do?

Investigate AI observability clusters — understand usage patterns in AI/LLM traffic, compare cluster behavior, compute cost/latency metrics, and drill into individual traces within clusters. Exploring LLM Clusters is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Investigate AI observability clusters — understand usage patterns in AI/LLM traffic, compare cluster behavior, compute cost/latency metrics, and drill into individual traces within clusters.

When should I use Exploring LLM Clusters?

Exploring LLM Clusters fits situations like: tasks that involve Observability.

How do I install Exploring LLM Clusters in Claude Code?

Run `npx skills add PostHog/posthog-foss --skill exploring-llm-clusters -a claude-code`. Or copy the skill folder (products/ai_observability/skills/exploring-llm-clusters in PostHog/posthog-foss) into .claude/skills/exploring-llm-clusters in your project. Claude Code loads it when a task matches its description.

How do I install Exploring LLM Clusters in Codex?

Run `npx skills add PostHog/posthog-foss --skill exploring-llm-clusters -a codex`. Or copy the skill folder (products/ai_observability/skills/exploring-llm-clusters in PostHog/posthog-foss) into .agents/skills/exploring-llm-clusters in your project. Codex loads it when a task matches its description.

Can I use Exploring LLM Clusters in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PostHog/posthog-foss --skill exploring-llm-clusters -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/exploring-llm-clusters, .gemini/skills/exploring-llm-clusters, .github/skills/exploring-llm-clusters and .opencode/skills/exploring-llm-clusters in your project.

What does Exploring LLM Clusters need to run?

Going by SKILL.md and its folder, Exploring LLM Clusters needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Exploring LLM Clusters access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Exploring LLM Clusters safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Exploring LLM Clusters use?

Exploring LLM Clusters is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Exploring LLM Clusters use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Exploring LLM Clusters?

Skills that share tags, products or a category with Exploring LLM Clusters: Telemetry Analytics (OpenHands/OpenHands, 90k stars), Temps Best Practices (gotempsh/temps, 822 stars), Posthog Operations (shepherdjerred/monorepo, 112 stars) and AI Observability Langchain Python (Jwuthri/Tracely-ai, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Exploring LLM Clusters?

PostHog (a GitHub organization, an official publisher) maintains it in PostHog/posthog-foss, which has 721 GitHub stars. The repository holds 213 skills in this directory. The repository was last updated on October 7, 2026.

Source: PostHog/posthog-foss on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.