Official agent skill

Configuring Experiment Analytics

by PostHog in PostHog/posthog-foss

Configures the analytics side of a PostHog experiment — exposure criteria (server-resolved default exposure event vs custom exposure events), primary and secondary metrics, the supported metric…

OfficialMITAuto-check passed

Install Configuring Experiment Analytics

skills CLI
$ npx skills add PostHog/posthog-foss --skill configuring-experiment-analytics -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PostHog/posthog-foss configuring-experiment-analytics --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .claude/skills && cp -r skills-src/products/experiments/skills/configuring-experiment-analytics .claude/skills/configuring-experiment-analytics && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
configuring-experiment-analytics
GitHub stars
721
Token cost
~3.7k tokens
SKILL.md length
1,749 words
Files
4 (incl. references)
Skills in repo
213
Repo updated
First seen
Licence
MIT

At a glance

Configures the analytics side of a PostHog experiment — exposure criteria (server-resolved default exposure event vs custom exposure events), primary and secondary metrics, the supported metric…

  • Works in 4 steps: Check for an existing shared metric… → Discover available events (REQUIRED… → Choose metric type → …
  • Edits a primary
  • SKILL.md covers Exposure criteria, Resolving experiments, Metrics and Interpreting results, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Configuring Experiment Analytics is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Configures the analytics side of a PostHog experiment — exposure criteria (server-resolved default exposure event vs custom exposure events), primary and secondary metrics, the supported metric types (count, sum, ratio with math and mathproperty, retention with retentionwindowstart and starthandling), multivariate user handling ("Exclude from analysis" vs "Use first seen variant"), and how to read results once the experiment is live. Use when the user adds or edits a primary or secondary metric (e.g. "add a…

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/interpreting-results.md` and `references/metric-templates.md`).

It works with PostHog. The repository describes itself as: PostHog FOSS is a read-only mirror of PostHog, with all proprietary code removed. NOTE: This repo is synced automatically from the main PostHog repo. Please raise any issues and… The licence is MIT.

When your agent uses it

  • Edits a primary
  • Secondary metric (e.g

Example prompts

  • “Exclude from analysis”
  • “Use first seen variant”
  • “add a secondary metric tracking”
  • “/configuring-experiment-analytics”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Check for an existing shared metric (REQUIRED — match by definition, not name)
  2. Discover available events (REQUIRED before building an inline metric)
  3. Choose metric type
  4. Primary vs secondary

What it can do on your machine

Read from SKILL.md and the folder at commit 2c48221. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Configuring Experiment Analytics loads about 3.7k tokens when it runs, and up to ~8.7k if it reads all its reference files. Until then it costs about 252 tokens; SKILL.md has 1,749 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~252
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PostHog/posthog-foss at commit 2c48221, republished under its MIT licence (© PostHog). 1,749 words, ~3,652 tokens.

Download SKILL.mdSave it as .claude/skills/configuring-experiment-analytics/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
configuring-experiment-analytics
description
Configures the analytics side of a PostHog experiment — exposure criteria (server-resolved default exposure event vs custom exposure events), primary and secondary metrics, the supported metric types (count, sum, ratio with `math` and `math_property`, retention with `retention_window_start` and `start_handling`), multivariate user handling ("Exclude from analysis" vs "Use first seen variant"), and how to read results once the experiment is live. Use when the user adds or edits a primary or secondary metric (e.g. "add a secondary metric tracking 'downloaded_file' per user"), sets up a ratio metric (e.g. "revenue from purchase_completed / pageviews"), sets up a retention metric (e.g. "$pageview → uploaded_file, 7-day window"), configures custom exposure (e.g. "only count users who hit /checkout"), changes multivariate handling, or asks "who is in the analysis?", "how do I measure impact?", "is this winning?", "what's the confidence level?", or "should I ship?".

Configuring experiment analytics

This skill answers: Who is included in the analysis? and How to measure impact?

Exposure criteria

Exposure criteria determine which users are counted in the experiment analysis.

Exposure event

Two options:

  1. Default exposure event — users are included when the experiment's default exposure event fires for the experiment's flag: $feature_flag_called, or $experiment_exposure for newer experiments. Which one applies is resolved server-side — read resolved_exposure_event from experiment-get rather than assuming either name (both events carry the same properties). This is the standard approach — it means a user is included only when they actually encounter the feature flag in your code.
  2. Custom exposure event — users are included when a specific custom event fires. Use this when you want tighter control over who enters the analysis (e.g., only users who actually visit the page where the experiment runs).
Multiple variant handling

When a user is exposed to multiple variants (e.g., due to flag changes or race conditions):

  • Exclude from analysis — removes these users from the analysis entirely. Cleaner data, smaller sample.
  • Use first seen variant — assigns users to the first variant they were exposed to. Keeps all users in the analysis. Note that "first seen" can introduce other biases as behavior cannot be clearly attributed to a single variant and is not recommended unless necessary.

Bias risk on uneven splits. "Exclude from analysis" combined with an uneven variant split can introduce bias — multi-variant users are dropped asymmetrically and the smaller variant loses a larger fraction of its assignments. If those users behave differently from the rest, the smaller variant's metrics will be skewed.

The right mitigation depends on experiment state:

  • Not yet launched, or only exposed to a few users so far — switch to an even variant split and use the overall rollout percentage to limit test-variant exposure. This removes the bias and preserves statistical power. See configuring-experiment-rollout.
  • Live experiment with significant exposures — changing the split mid-run reassigns users across variants, which is bad for user experience and data quality. Switch this setting to "Use first seen variant" instead — it keeps already-assigned users in their original variant (no reassignment) and removes the asymmetric exclusion.
Filter test accounts

exposure_criteria.filterTestAccounts (default: true) — excludes internal/test users from the analysis.

Resolving experiments

Metric changes require an experiment ID. If the user refers to an experiment by name or description (e.g. "add metrics to the checkout test"), load the finding-experiments skill to resolve it to a concrete ID before proceeding.

Metrics

A metric reaches an experiment one of two ways, both via experiment-update:

  • Inline metric — defined directly on the experiment. Sent in the metrics array, which replaces the entire inline list, so always get the current experiment first via experiment-get to preserve existing metrics.
  • Shared (saved) metric — a reusable metric object that can be attached to many experiments. Attached by ID via saved_metrics_ids (this list also replaces the experiment's existing saved-metric links, so resend the full set — see Step 1).

Prefer reusing a shared metric over duplicating it inline. Build a new inline metric only when no suitable shared metric already exists.

Step 1: Check for an existing shared metric (REQUIRED — match by definition, not name)

Before building any new inline metric, you MUST check whether the project already has a shared (saved) metric that measures the same thing, and reuse it. Duplicating a metric that already exists as a shared metric fragments measurement and is exactly what we want to avoid.

Reuse is decided by the metric definition — the event or action plus the metric type — not the name. Saved metrics are named by each team's own conventions, which you cannot guess, so you must compare on what each metric measures (its query), never on its title.

Workflow:

  1. Know what you're about to build first. Settle the target event(s)/action(s) and metric type (mean / funnel / ratio / retention) before searching — see Step 2 to confirm the event exists via read-data-schema. You can only recognize a duplicate once you know the concrete event/action, so this check runs after you've pinned down the event, not before.
  2. Search by the event, then compare each candidate's query. Call experiment-saved-metrics-list with ?event=<the event you're measuring> to find metrics that reference it — matched directly (an EventsNode) or via the step events of any action a metric references, so action-based metrics are found by the event their action fires on. Then for each returned row, inspect its query (not the name/description): a saved metric is a reuse match when its query measures the same event or action with the same metric_type (and compatible math) as the metric you'd otherwise build, even if its name is different.
    • Match on the event, not the action's name. An action-based metric is discoverable by the event the action fires on — pass that event, not the action's label.
    • Do not use search for this. search matches only the metric's own name / description / tags — never the underlying event or action — so it cannot find a definition match. Use search only when the user names a specific saved metric to attach (name resolution, not a definition match).
  3. If a saved metric matches the definition — confirm the match with the user by name/description, then attach it instead of building a new one:
    • Call experiment-get to read the experiment's current saved_metrics.
    • Call experiment-update with saved_metrics_ids set to the full desired set — it replaces existing links, so include the already-attached ones plus the new entry. Each entry has shape { "id": <saved-metric id>, "metadata": { "type": "primary" } } — set type to "primary" or "secondary". metadata is optional and defaults to primary.
    • Watch the id when rebuilding the set: each item in the saved_metrics you just read has a top-level id (the link id) AND a saved_metric field (the metric id). saved_metrics_ids wants the saved_metric value, not the link id — sending the link id attaches the wrong metric or fails validation.
    • You do not need to build the inline metric — the shared metric already encodes its events.
  4. If nothing in the library measures the same event/action + type — build an inline metric (Step 2+). When that inline metric is likely to be reused across experiments, offer to create it as a shared metric instead, via experiment-saved-metrics-create, then attach it as above, so the next experiment can reuse it.
Show full SKILL.md (715 more words)Show less
Step 2: Discover available events (REQUIRED before building an inline metric)

Before suggesting or building any new inline metric, you MUST call read-data-schema to discover what events actually exist in the project. Do NOT skip this step. Do NOT suggest event names based on what you think the project might track — only use events you have confirmed exist. (Attaching an existing shared metric from Step 1 does not need this — it already encodes its events.)

This applies even when:

  • The user provides event names — look them up to confirm they exist and are spelled correctly
  • The user asks "what metrics do you suggest?" — look up events first, then suggest from real data
  • The context makes certain events seem obvious — they may not exist or may be named differently

Workflow:

  1. Call read-data-schema to get the project's events
  2. Present relevant events to the user based on the experiment's hypothesis
  3. User picks which events to use for metrics
  4. Configure metrics with those confirmed event names

Legitimate exception — allow_unknown_events: true: Pass this on experiment-create / experiment-update only when the user is intentionally instrumenting an event that hasn't been ingested yet (e.g. setting up the experiment before the code change ships). Confirm this with the user — never use it as a workaround for "the event lookup didn't return what I expected".

Example:

text
User: "Let's add some metrics for the checkout experiment"

WRONG: "I'd suggest using purchase_completed as the primary metric..."
  (hallucinated event name — never seen the project's actual events)

RIGHT: *calls read-data-schema* → "Here are the events in your project
  related to checkout: `checkout_step_completed`, `payment_processed`,
  `order_confirmed`. Which of these represents a successful checkout?"
Step 3: Choose metric type

Start from a template in references/metric-templates.md: it maps common requests ("more revenue", "people come back") to a metric type and the defaults to set with it (window units, winsorization, goal).

There are four metric types. Each has kind: "ExperimentMetric":

metric_typeWhen to useRequired fields
"mean"Average of a numeric property per user (revenue, session duration, pageviews per user)source
"funnel"Conversion rate from exposure through one or more ordered actionsseries (1 or more steps)
"ratio"Rate of one event relative to anothernumerator, denominator — set math: "sum" + math_property on a side to aggregate a property; filters never aggregate
"retention"Do users come back after exposure?start_event, completion_event, retention_window_start, retention_window_end, retention_window_unit, start_handling

Funnel metrics and the implicit exposure step

Funnel metrics automatically prepend the experiment's exposure event as step_0. So a funnel with 1 step in series is a valid 2-step funnel: exposure → action. This is the correct choice for measuring "what percentage of exposed users did X?"

Examples:

  • "What % of exposed users reached /login?" → funnel with 1 step ($pageview filtered to /login)
  • "What % of exposed users completed checkout?" → funnel with 1 step (checkout_completed)
  • "What % of exposed users went cart → checkout → purchase?" → funnel with 3 steps

Mean vs funnel for the same event

  • Mean measures average count/value per user (e.g. "pageviews per user", "revenue per user").
  • Funnel measures conversion rate (e.g. "% of exposed users who purchased").

Both can reference the same event — the difference is whether you care about count/magnitude (mean) or yes/no conversion (funnel).

Retention: same vs different start/completion event

The retention window is measured from the start event, so the events you pick decide what's measured: The start occurrence never counts as its own completion (only a distinct later event does), so both shapes are valid:

  • Different start and completion events → conversion-style retention ("did they reach the target action within the window?").
  • Same event → repeat retention ("did they fire it again?"). From 0 counts a repeat from the same period onward (same-day repeats included); From ≥ 1 requires an occurrence later. Use start_handling: "first_seen". When a user says "retention of <event>" they usually mean repeat retention.

See references/metric-configuration.md for the full rendered ExperimentMetric schema (all four metric types, with required fields per type) plus WRONG/RIGHT JSON pairs for the failure modes that come up most often (ratio with is_set filter instead of math: "sum" + math_property; retention without retention_window_start / start_handling). Read it before assembling a ratio or retention payload — the required fields are authoritative.

Step 4: Primary vs secondary
  • Primary metrics — the main success criteria for the experiment. These drive the ship/end decision.
  • Secondary metrics — additional measurements for context. Useful for guardrail metrics (e.g., ensuring a conversion improvement doesn't increase error rates).

Interpreting results

See references/interpreting-results.md for guidance on reading experiment results, statistical significance, and when to ship vs end.

  • configuring-experiment-rollout — the rollout side: variant splits and traffic percentage
  • diagnosing-experiment-health — when results look biased, empty, or strange
  • analyzing-experiment-session-replays — qualitative complement — watch what each variant's users actually did

© PostHog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in products/experiments/skills/configuring-experiment-analytics of PostHog/posthog-foss.

  • SKILL.md
  • references/interpreting-results.md
  • references/metric-configuration.md.j2
  • references/metric-templates.md

Open the folder on GitHubat commit 2c48221

Compare with similar skills

Configuring Experiment Analytics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Configuring Experiment Analytics compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Configuring Experiment Analytics this skillPostHog/posthog-foss721—~3.7kAutomated safety check: PassMIT
Opik Analytics Instrumentationcomet-ml/opik22k—~4.4kAutomated safety check: PassApache-2.0
C15tc15t/c15t1.9k1 repos~1.6kAutomated safety check: PassApache-2.0
Define Feature Flagmacro-inc/macro4.6k—~780Automated safety check: PassAGPL-3.0
Soku CLIAbout-Intelligence/soku-cli304—~2.4kAutomated safety check: PassMIT
Compare Array Bundle SizePostHog/posthog-js613—~599Automated safety check: PassCustom licence

Similar skills

  • Shows how to add product analytics events to Opik's frontend, Java backend and Python SDK, all reporting through Segment to PostHog with an opik_ name prefix.

    22k GitHub stars~4.4k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • C15t

    c15t/c15t

    Work with c15t consent management docs, APIs, and integrations for Next.js, React, and JavaScript.

    1.9k GitHub starsUsed in 1 repo~1.6k tokens
    Legal & ComplianceAuto-check passed
  • Define Feature Flag

    macro-inc/macro

    Define a frontend feature flag with defineFlag and wire its readers.

    4.6k GitHub stars~780 tokensUpdated today
    Frontend & DesignAuto-check passed
  • Soku CLI

    About-Intelligence/soku-cli

    Guides an agent through the soku command line tool for ads, GA4 and PostHog data reads, ads writes, SEO hosting, automations, files and skill management.

    304 GitHub stars~2.4k tokensUpdated 2 days ago
    Marketing & SEOAuto-check passed
  • Compare Array Bundle Size

    PostHog/posthog-js

    Official

    Quickly compare the posthog-js array.js bundle size in the current working tree against a git baseline using the repository's esbuild proxy.

    613 GitHub stars~599 tokensUpdated today
    Frontend & DesignAuto-check passed
  • Telemetry Analytics

    OpenHands/OpenHands

    This skill should be used when the user asks to "add tracking", "add a PostHog event", "change telemetry consent", "instrument onboarding", "debug analytics", or changes telemetry.ts…

    90k GitHub stars~305 tokensUpdated today
    DevOps & CloudAuto-check passed

More from PostHog/posthog-foss

All 213 skills in this repo
  • Authoring Log Alerts

    PostHog/posthog-foss

    Official

    Author useful, low-noise log alerts on services in a PostHog project.

    721 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Autoresolving PR Conflicts

    PostHog/posthog-foss

    Official

    Operating procedure for the conflict-autoresolver agent: sweep open PostHog/posthog PRs that conflict with master, resolve the trivial conflicts (generated artifacts deterministically, source…

    721 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Help users debug PostHog Error Tracking stack-trace symbolication for any supported platform — JavaScript/TypeScript web, React Native (Hermes), Android (Proguard / R8), or iOS / macOS (dSYM).

    721 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Exploring Apm Traces

    PostHog/posthog-foss

    Official

    Investigates distributed application performance using PostHog APM (OpenTelemetry span) data via MCP.

    721 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Exploring LLM Traces

    PostHog/posthog-foss

    Official

    Debug and inspect LLM/AI agent traces using PostHog's MCP tools.

    721 GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Investigate Metric

    PostHog/posthog-foss

    Official

    Diagnose why a product metric changed (dropped, spiked, or plateaued) by orchestrating breakdowns, actors, paths, lifecycle, retention, and annotations queries.

    721 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Works with

Questions about Configuring Experiment Analytics

What does Configuring Experiment Analytics do?

Configures the analytics side of a PostHog experiment — exposure criteria (server-resolved default exposure event vs custom exposure events), primary and secondary metrics, the supported metric…. Configuring Experiment Analytics is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Configures the analytics side of a PostHog experiment — exposure criteria (server-resolved default exposure event vs custom exposure events), primary and secondary metrics, the supported metric types (count, sum, ratio with math and mathproperty, retention with retentionwindowstart and starthandling), multivariate user handling ("Exclude from analysis" vs "Use first seen variant"), and how to read results once the experiment is live.

When should I use Configuring Experiment Analytics?

Configuring Experiment Analytics fits situations like: edits a primary; secondary metric (e.g.

How do I install Configuring Experiment Analytics in Claude Code?

Run `npx skills add PostHog/posthog-foss --skill configuring-experiment-analytics -a claude-code`. Or copy the skill folder (products/experiments/skills/configuring-experiment-analytics in PostHog/posthog-foss) into .claude/skills/configuring-experiment-analytics in your project. Claude Code loads it when a task matches its description.

How do I install Configuring Experiment Analytics in Codex?

Run `npx skills add PostHog/posthog-foss --skill configuring-experiment-analytics -a codex`. Or copy the skill folder (products/experiments/skills/configuring-experiment-analytics in PostHog/posthog-foss) into .agents/skills/configuring-experiment-analytics in your project. Codex loads it when a task matches its description.

Can I use Configuring Experiment Analytics in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PostHog/posthog-foss --skill configuring-experiment-analytics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/configuring-experiment-analytics, .gemini/skills/configuring-experiment-analytics, .github/skills/configuring-experiment-analytics and .opencode/skills/configuring-experiment-analytics in your project.

What does Configuring Experiment Analytics need to run?

SKILL.md names no scripts, command-line tools or credentials: Configuring Experiment Analytics is instructions for the agent only.

Does Configuring Experiment Analytics access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Configuring Experiment Analytics safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Configuring Experiment Analytics use?

Configuring Experiment Analytics is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Configuring Experiment Analytics use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.1k tokens, read only when the agent opens those files.

What are the alternatives to Configuring Experiment Analytics?

Skills that share tags, products or a category with Configuring Experiment Analytics: Opik Analytics Instrumentation (comet-ml/opik, 22k stars), C15t (c15t/c15t, 1.9k stars), Define Feature Flag (macro-inc/macro, 4.6k stars) and Soku CLI (About-Intelligence/soku-cli, 304 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Configuring Experiment Analytics?

PostHog (a GitHub organization, an official publisher) maintains it in PostHog/posthog-foss, which has 721 GitHub stars. The repository holds 213 skills in this directory. The repository was last updated on October 7, 2026.

Source: PostHog/posthog-foss on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.