Official agent skill

Fixing Flaky Tests

by PostHog in PostHog/posthog-foss

Guides an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo.

OfficialMITAuto-check passedTesting & QA

Install Fixing Flaky Tests

skills CLI
$ npx skills add PostHog/posthog-foss --skill fixing-flaky-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PostHog/posthog-foss fixing-flaky-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/fixing-flaky-tests .claude/skills/fixing-flaky-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fixing-flaky-tests
GitHub stars
721
Token cost
~5.9k tokens
SKILL.md length
2,835 words
Files
1
Skills in repo
213
Repo updated
First seen
Licence
MIT

At a glance

Guides an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo.

  • Works in 8 steps: Measure the failure rate — from GitHub,… → Extract the failure from CI → Reproduce locally — before touching… → …
  • A test fails intermittently in CI but passes on rerun
  • SKILL.md covers 1. Measure the failure rate —…, 2. Extract the failure from CI, 3. Reproduce locally — before… and 4. Root-cause the flake, plus 5 more sections
  • Calls git, gh and pnpm; needs TRUNK_API_TOKEN

What it does

Fixing Flaky Tests is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Guides an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo. Use when a test fails intermittently in CI but passes on rerun or locally, when hogli ci:insights or the debugging-ci-failures skill classifies a failure as a flaky test, when given a GitHub Actions URL for a flaky job, when asked to check Trunk Flaky Tests for a test, PR, or master, or when asked to deflake, stabilize, or fix a flaky Jest, pytest, or Playwright test. Core discipline: reproduce locally…

Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Failing and flaky tests and Unit testing. It works with PostHog, Playwright, GitHub Actions and Jest. The repository describes itself as: PostHog FOSS is a read-only mirror of PostHog, with all proprietary code removed. NOTE: This repo is synced automatically from the main PostHog repo. Please raise any issues and… The licence is MIT.

When your agent uses it

  • A test fails intermittently in CI but passes on rerun
  • Hogli ci:insights
  • The debugging-ci-failures skill classifies a failure as a flaky test
  • Given a GitHub Actions URL for a flaky job

Example prompts

  • “Use the fixing-flaky-tests skill to guide an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo”
  • “/fixing-flaky-tests”

Requirements

  • A credential in TRUNK_API_TOKEN

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Measure the failure rate — from GitHub, not from a digest
  2. Extract the failure from CI
  3. Reproduce locally — before touching anything
  4. Root-cause the flake
  5. Decide the outcome — fixing is one of three
  6. Fix the root cause — never mask it
  7. Validate with an N-run loop
  8. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 2c48221. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • gh
    • pnpm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, gh and pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • TRUNK_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fixing Flaky Tests loads about 5.9k tokens when it runs. Until then it costs about 239 tokens; SKILL.md has 2,835 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~239
When it runs · the whole SKILL.md, loaded when a task matches
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PostHog/posthog-foss at commit 2c48221, republished under its MIT licence (© PostHog). 2,835 words, ~5,912 tokens.

Download SKILL.mdSave it as .claude/skills/fixing-flaky-tests/SKILL.md (or your agent's skills folder).
name
fixing-flaky-tests
description
Guides an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo. Use when a test fails intermittently in CI but passes on rerun or locally, when `hogli ci:insights` or the debugging-ci-failures skill classifies a failure as a flaky test, when given a GitHub Actions URL for a flaky job, when asked to check Trunk Flaky Tests for a test, PR, or master, or when asked to deflake, stabilize, or fix a flaky Jest, pytest, or Playwright test. Core discipline: reproduce locally before changing anything, fix the root cause (never mask it with sleeps, retries, or bigger timeouts), and prove the fix with an N-run validation loop sized to the observed failure rate. Stabilizing is not the only valid outcome — the skill also gates whether the test should exist, so deleting a test that catches nothing real, or re-leveling one that flakes because of the level it runs at, are first-class endings.

Fixing flaky tests

Before you propose a change to how the suite runs in CI, check things already tried. It records measured verdicts on test parallelism, sharding, and coverage-based selection, so a rejected approach is not rebuilt.

Three non-negotiables, in order:

  1. Reproduce before you fix. A fix for a failure you never observed is a guess. Only fall back to analytical fixes when the escalation ladder below is exhausted.
  2. Fix the root cause. Sleeps, raised timeouts, retries, and weakened assertions hide flakes; they do not fix them.
  3. Validate with an N-run loop. One green run proves nothing about an intermittent failure. Size N to the observed failure rate.

Stabilizing the test is not the only valid ending. Once you know why it flakes, step 5 asks whether it should exist. A test that catches nothing real is worth deleting, and one that flakes because of the level it runs at is worth moving down a rung.

Before any of these: measure, don't assume. Flaky-vs-deterministic, and the rate, are facts to establish from verifiable GitHub run data (step 1) — never inherited from a Slack alert, a teammate's guess, or a ci:insights label.

For triaging a red CI run (finding and classifying the failure), use the debugging-ci-failures skill first — this skill takes over once the failure is classified as a flaky test. For writing new Playwright tests that aren't flaky, use the playwright-test skill. investigating-ci-failures (green/red boundary) and diagnosing-ci-and-merge-bottlenecks (the engineering-analytics-flaky-tests tool's caveats) are product skills under products/engineering_analytics/skills/, not invocable here: read their SKILL.md at that path.

1. Measure the failure rate — from GitHub, not from a digest

The GitHub Actions API (or GitHub MCP) is the source of truth. hogli ci:insights is a digest, not an oracle — it can mislabel flaky-vs-deterministic, it lags the API until GitHub's webhook settles, and it cannot give you a rate at all: it reports absolute counts, because CI emits failures but omits ordinary passing runs, so there is no denominator. Use it to validate a hypothesis or pull historical context, never as the first move or the classification authority.

Establish the timeline yourself from raw run data:

bash
# Pass/fail history for the workflow on the branch under suspicion (master shown):
gh run list --repo PostHog/posthog --workflow=ci-backend.yml --branch master \
  --status completed --limit 60 --json conclusion,headSha,createdAt,databaseId

The run-level conclusion is not enough — a red run may be failing on a different job/test. Confirm it is the same test each time by reading the failing shard's log; and for green runs, confirm the test actually ran and passed in that shard (it may have been sharded elsewhere, not fixed):

bash
gh api repos/PostHog/posthog/actions/jobs/{job_id}/logs \
  | grep -E "<exact::test::id>|short test summary"

Read the timeline before you classify:

  • Interleaved pass/fail on adjacent commits — the same unchanged test verified passing in some runs, failing in others → genuinely flaky; continue here.
  • A long unbroken failure streak (say 30+ consecutive) is statistically incompatible with a flake — at any per-run rate below ~95%, p^30 ≈ 0. That is a deterministic regression: go to debugging-ci-failures and find the introducing commit (step 4).
  • Every rerun fails on one runner, yet master passes the same code → an environment-dependent race, not a regression: kernel clock granularity, filesystem, runner speed. Retrying on the same runner cannot help, so "failed 3/3 attempts" says nothing about whether the PR caused it. A merge-queue failure hands you the control for free: the trunk-merge/** test PR's body names the master SHA it is based on, so read master's run at that SHA (gh api repos/PostHog/posthog/commits/<sha>/check-runs) and confirm the same test ran and passed there.
  • Both at once is common: a latent flake whose rate jumped to ~100%. Find the transition (last green → first red); that boundary, not the digest's one-line verdict, is what tells you what tipped it. products/engineering_analytics/skills/investigating-ci-failures/ has the boundary query if the failure reached master.

If a failure is reported as (or you suspect it is) consistent, don't serialize — measure the rate and attempt a repro in parallel.

Record the measured rate (failures / total runs, from the run data). You need it to size the validation loop in step 7.

Then confirm it is not already handled:

bash
git log --oneline -10 -- <test_file_path>          # recently fixed already?
hogli ci:insights search "<test name or error>"    # cross-run history — corroborate against the run data, do not trust blindly

search reports two surfaces; read each for what it actually claims:

  • broken tests (recent failure fingerprints, last 2 days). A potentially_resolved state means that job's latest default-branch run is green again — weak evidence a fix landed, not proof. Confirm against the run data that it covers this failure before reporting instead of re-fixing.
  • test health (ranked by blast radius; the same spans the engineering-analytics-flaky-tests MCP tool reads). confirmed_flake is the only classification backed by proof: one commit both failed and passed the test in the same matrix job, via a re-run attempt going green or an in-job retry. A pass in a different matrix leg is not recovery. suspected_regression means no recovery was recorded — absence of proof, not proof of a regression, so treat it as real until your own run data says otherwise.
Corroborate with Trunk Flaky Tests

CI uploads test results to Trunk Flaky Tests, which tracks per-test failure history across master and PR runs. On a PR, the trunk-io bot's "Trunk Test Analytics" comment links that PR's slice (https://app.trunk.io/posthog-inc/flaky-tests/pr/<number>?repo=PostHog/posthog).

The trunk MCP server in .mcp.json queries it (tools are marked experimental by Trunk):

  • search-test (repoName: "PostHog/posthog", testNameSearch: "<test name, no filepath>") → the test case ID.
  • fix-flaky-test (repoName, testCaseId) → failure history, first-seen commit, git blame, and Trunk's root-cause investigation; createNewInvestigation: true triggers a fresh analysis (takes up to a minute).

Authenticate once via /mcp → trunk (browser OAuth); headless environments instead add an Authorization: Bearer header with a TRUNK_API_TOKEN org token to the server entry.

Two limits worth knowing before you start here. AI investigations are not enabled for this repo, so fix-flaky-test returns history or nothing, never a root cause. And lookup only goes name to ID: a bare dashboard link identifies a test you cannot name, so ask for the test name rather than guessing at the ID.

Trunk attributes each test to a team through CODEOWNERS, which cannot express the owners.yaml map. .github/scripts/trunk-codeowners.sh projects the map into a generated CODEOWNERS before each upload (hogli owners:codeowners builds the same file locally), so a test's owner in Trunk should match hogli owners:who. Where it does not, the projection dropped a spelling two teams would both claim.

Like ci:insights, this is corroboration and history, not the classification authority — flaky-vs-deterministic and the rate still come from the run data above.

2. Extract the failure from CI

bash
gh run view <run-id> --log-failed
# If the job was re-run and the latest attempt passed, the flaky failure
# lives in a previous attempt:
gh api "repos/PostHog/posthog/actions/runs/{run_id}/attempts/1/jobs"
gh api "repos/PostHog/posthog/actions/jobs/{job_id}/logs"

Capture before moving on:

  • Exact test ID (file path + test name) and the failing assertion or error.
  • Surrounding warnings — [MSW] Unhandled, async leak warnings, teardown errors — these are often the actual cause, printed before the symptom.
  • Which other tests ran in the same worker/shard before it (ordering suspects).

3. Reproduce locally — before touching anything

Escalate through these conditions until the failure appears. Stop at the first level that reproduces it; that level is your validation environment for step 7.

  1. Single run: hogli test <path>::<test> — confirms the test runs at all.

  2. Repetition loop (default N=20): catches probabilistic flakes.

  3. CI-like conditions: CI runs Jest sharded with low worker counts on contended runners — read the current flags from frontend/package.json's test script and .github/workflows/ci-frontend.yml before running. Example (with the flags as of this writing), running the test alongside its shard neighbors:

    bash
    pnpm --filter=@posthog/frontend jest <test_file> <neighbor_file> --maxWorkers=2 --forceExit

    For pytest, run the whole file or class rather than the single test, so module-level fixtures and ordering match CI.

  4. Ordering: run suspected polluter tests before the victim; reverse the order within the file. Flakes that vanish in isolation are ordering bugs.

  5. Contention: re-run the loop while something CPU-heavy runs in another shell (e.g. a parallel full-file jest run). Timeout-class flakes often only show here.

The loop harness — judge by exit code, not by grepping output:

bash
N=20; PASS=0; FAIL=0
for i in $(seq 1 $N); do
  if <test command> >/tmp/flake-run.log 2>&1; then
    PASS=$((PASS+1))
  else
    FAIL=$((FAIL+1)); cp /tmp/flake-run.log /tmp/flake-fail-$i.log; echo "run $i: FAIL"
  fi
done
echo "$PASS passed, $FAIL failed out of $N"

Two cost notes for the loop:

  • While reproducing, break after the first failure — one captured failure log is enough. Complete all N runs only when measuring the failure rate or validating in step 7.
  • The pnpm --filter=@posthog/frontend jest script runs pnpm build:products before every invocation. Inside a loop, build once, then iterate with pnpm --filter=@posthog/frontend exec jest ..., which skips the rebuild.

If nothing reproduces after the full ladder, the flake is CI-environment-specific. Before settling for that, ask what the runner has that your machine lacks: Linux mtimes from a coarse clock where macOS gives nanoseconds, a case-sensitive filesystem, a different timezone, /tmp on a different filesystem. Often you can force the CI condition instead of waiting for it — os.utime the files into the ordering the coarse clock produced, run under TZ=UTC — and then validate empirically after all. Otherwise proceed with a fix grounded in the CI evidence and root-cause analysis, and say so explicitly in the report — the validation in step 7 is then analytical, not empirical.

4. Root-cause the flake

Match the symptom to a cause class; never patch the symptom.

SymptomLikely cause class
Timeout waiting for promise/listener/elementUnawaited async work, missing mock, hidden pending request
Passes alone, fails with neighbors (or vice versa)Shared state: module cache, DB rows, global config, ordering
Fails near midnight/UTC boundaries, or on slow runnersReal clock usage — missing time_machine.travel / fake timers
Assertion on list order or generated IDsNondeterministic ordering/IDs asserted as deterministic
Query can't see just-written dataEventual consistency (ClickHouse), missing flush/commit
Only fails under --maxWorkers=2 / contentionRace condition surfaced by scheduling, too-tight timeout
Every attempt fails on one runner, passes on anotherClock-granularity race: cutoff read from an earlier write
When the cause isn't obvious, bisect

If the symptom table doesn't point at a clear cause and the test file itself is unchanged (git log -- <test_file> is stale), the trigger is elsewhere — a neighbor test, a dependency bump, or a product change. Find when it started instead of guessing:

  • Bisect the CI run history first (cheap, no local builds): from the step-1 timeline, take the last-green → first-red boundary and diff the commits in that window (git log <good>..<bad>). That short list often names the culprit outright.

  • git bisect the code when you can reproduce locally and the failure is (near-)deterministic:

    bash
    git bisect start <bad-sha> <good-sha>
    git bisect run bash -c '<repro command>'   # exit 0 = good, non-zero = bad

    Caveat: for an intermittent flake, a lucky pass at a step sends git bisect down the wrong path. Trust code-bisect only when the failure is deterministic; otherwise run the repro N times per step (fail if any iteration fails), or just use the CI run-history boundary.

PostHog-specific patterns:

Show full SKILL.md (1,151 more words)Show less
Frontend (Jest + kea)
  • Missing MSW mocks: kea logics with afterMount loaders fire API calls through the connect() chain — a logic three levels deep can trigger an unmocked fetch. Unhandled requests currently resolve with a benign empty paginated 200, so the symptom is a loader succeeding with empty or wrong data, not a network error. The [MSW] Unhandled GET ... warning in the log names the missing mock — add the useMocks entry. The unhandled-request behavior has changed before (it used to hang); if symptoms don't match, read frontend/src/mocks/jest.ts for what unmocked requests do today.
  • toFinishAllListeners() timeouts: waits for ALL kea listener promises across ALL mounted logics (3s default — LISTENER_FINISH_WAIT_TIMEOUT in kea-test-utils). Any connected logic with a pending loader blocks it. Fix the pending work; do not raise the timeout.
  • Mock URL mismatch: mocksToHandlers strips trailing slashes, but query params, @current-style segments, and :param patterns must match the real request URL. Compare against the [MSW] Unhandled line.
  • Per-test handler reset: frontend/src/mocks/jest.ts registers a global afterEach(() => mswServer.resetHandlers()). Each it.each case is a separate test, so runtime mocks must be (re-)registered in beforeEach.
  • Leaked mounts: logics mounted in beforeEach and never unmounted leak async work into later tests.
Backend (pytest)
  • DB state leakage: shared rows across tests without isolation — check fixture scope and whether the test needs @pytest.mark.django_db(transaction=True).
  • Real time: use time_machine.travel(..., tick=False); never assert on now()-derived values.
  • Timestamp-derived cutoffs: pinning or filtering "as of" the previous write's recorded time — stamp the times explicitly instead (/writing-tests, "Two writes in a row are not ordered in time").
  • ClickHouse eventual consistency: a query may not see just-inserted data — flush explicitly in the test setup rather than sleeping.

5. Decide the outcome — fixing is one of three

Once you know why it flakes, ask whether the test should exist at all, before you spend effort stabilizing it. A flaky test is the one case where cost is already proven and value is not: it has demonstrably cost reruns, wall-clock, and attention. So apply /writing-tests' gate retroactively, with more force than you would to a new test:

What realistic regression does this test catch that no existing test already catches?

Three outcomes are valid. Pick deliberately; don't default to the first.

OutcomeWhenNext
Fix itGuards a real regression at roughly the right cost.step 6
Re-level itWorth guarding, but the flake is inherent to the level it runs at.below
Delete itYou cannot name the regression it catches, or another test already catches it.below
Delete it

Recurring shapes that fail the gate:

  • Tautological smoke test. Asserts a precondition many sibling tests in the same file already need to run at all — that the chart rendered, that the list is non-empty. Every later test is a stronger version of it.
  • Duplicated coverage. A thinner second test of a path an existing test already exercises. Fold it in as a parameterized case, or drop it.
  • Third-party assertion. Flaky because it exercises a vendor's eventual consistency, scheduler, or API rather than our logic. That is the vendor's test to write, and mock-based siblings usually already cover our side.
  • Permanently gated. Skipped everywhere but CI (missing credentials, opt-in marker), so nobody develops against it and only CI pays for it.

Deletion is irreversible, so it carries a higher bar than a fix:

  • Get explicit user approval, the same gate quarantine carries. Propose the deletion and hand back; don't take it yourself.
  • Account for the coverage: name what still catches this at file:line, or say plainly what is lost and why that's acceptable. "Probably covered elsewhere" is not an answer — go read the sibling test.
  • Never delete to turn a red build green under time pressure. That is quarantine with extra steps, and it is how real regressions ship.

Then skip to step 8: there is nothing to loop.

Re-level it

Move the test down the cost ladder in /writing-tests rather than hardening it in place. A round trip through a real broker, browser, or vendor API to prove logic a direct call could prove is testing the transport, and the transport is where the nondeterminism lives. Re-leveling removes the flake by construction, so prefer it over an increasingly elaborate wait.

The replacement is a new test: run /writing-tests' gate on it, and validate it at its new level rather than against the old repro conditions, which no longer exist.

6. Fix the root cause — never mask it

Tempting masking moveDo instead
sleep(2) / setTimeout before assertingAwait the specific condition (waitFor, expectLogic, explicit flush)
Raise the test/listener timeoutFind what is hanging; the timeout is the messenger
Add retries (pytest-rerunfailures --reruns, jest.retryTimes)Reserve for genuinely nondeterministic external infra, with a comment and a linked issue — never for product code under test
Skip / quarantine the testOnly with explicit user approval, with a linked issue
Loosen the assertionMake the data deterministic (sort, freeze, seed), keep the assertion strict
Harden a test that shouldn't existGo back to step 5 — deleting or re-leveling it is the cheaper fix

Keep the fix minimal and inside the test or its fixtures when possible. If the race is in product code, the flake found a real bug — fix the product code and say so in the report.

7. Validate with an N-run loop

Run the step-3 harness on the fixed code under the same conditions that reproduced the failure (same neighbors, worker count, contention).

Size N from the observed pre-fix failure rate: if it failed about 1 in k runs, you need roughly N ≥ 3k consecutive passes for ~95% confidence the flake is gone — (1 - 1/k)^(3k) ≈ 5%. So use N = max(3k, 20). Without a usable rate estimate, run 50 and note the reduced confidence in the report. If the flake was never reproducible locally, run N = 20 as a regression check and label the validation as analytical.

Any failure in the loop → back to step 4; the root cause was wrong or incomplete. Finish with one normal run of the surrounding file/suite to confirm the fix didn't break sibling tests.

Re-leveled instead? The old repro conditions no longer apply. Loop the replacement at its own level, and confirm the flake is gone because the old test is gone, not because it got faster.

Deleted instead? Nothing to loop. Run the surrounding file/suite once to confirm nothing depended on it, and carry the coverage argument into the report.

8. Report

text
Test:            <file path>::<test name>
Observed in CI:  <measured rate from run data, e.g. 8/45 runs over 3h (gh run list); ci:insights state corroborates>
Local repro:     <command + conditions, e.g. 3/20 failures with neighbor X, maxWorkers=2 | not reproducible locally>
Root cause:      <one or two sentences>
Outcome:         fixed | re-leveled (<from> → <to>) | deleted
Change:          <what changed and why it removes the cause; for a deletion, what still covers the behavior (file:line) and what coverage is genuinely lost>
Validation:      <N>/<N> passes under repro conditions | <N>/<N> at the new level | analytical only (CI-specific) | n/a, deleted
Follow-ups:      <product bug found, related tests with the same pattern, or none>

Boundaries

  • Do not rerun CI jobs, push, or post to GitHub to "test" the fix — validate locally.
  • Do not edit .github/workflows/ as part of a flake fix.
  • Do not accept/update snapshots to make a flake pass.
  • If the same root-cause pattern clearly affects sibling tests, fix them in the same change only when mechanical; otherwise list them as follow-ups.
  • Do not delete a test without explicit user approval, plus either a named replacement for the coverage or a plain statement of what is lost. Deletion is a valid outcome, never a shortcut to green.

© PostHog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/fixing-flaky-tests of PostHog/posthog-foss.

Open the folder on GitHubat commit 2c48221

Compare with similar skills

Fixing Flaky Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fixing Flaky Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fixing Flaky Tests this skillPostHog/posthog-foss721—~5.9kAutomated safety check: PassMIT
Debug Playwright Prowquay/quay2.8k—~2.2kAutomated safety check: PassApache-2.0
Designing TestsCloudAI-X/claude-workflow-v21.4k1 repos~1.5kAutomated safety check: PassMIT
Playwright Testingchongdashu/vibejam-starter-pack149—~2.1kAutomated safety check: PassNone
Playwright Testingchongdashu/vibejam-starter-pack149—~2.2kAutomated safety check: PassNone
Testingradix-ng/primitives274—~3.3kAutomated safety check: PassMIT

Similar skills

  • Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real…

    2.8k GitHub stars~2.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Designing Tests

    CloudAI-X/claude-workflow-v2

    Designs and implements testing strategies for any codebase. An agent skill from CloudAI-X/claude-workflow-v2.

    1.4k GitHub starsUsed in 1 repo~1.5k tokens
    Testing & QAAuto-check passed
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.1k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.2k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Testing

    radix-ng/primitives

    Test Radix NG primitives across every layer and pick the RIGHT one for a change: Vitest unit (zoneless), jest-axe a11y, Playwright browser regression (apps/visual-regression), SSR…

    274 GitHub stars~3.3k tokensUpdated 8 days ago
    Testing & QAAuto-check passed
  • Handsontable Unit Testing

    handsontable/handsontable

    Conventions for Handsontable's Jest unit and TypeScript type tests: where files go, how to run them, mocking limits and when to write an E2E test instead.

    22k GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed

More from PostHog/posthog-foss

All 213 skills in this repo
  • Authoring Log Alerts

    PostHog/posthog-foss

    Official

    Author useful, low-noise log alerts on services in a PostHog project.

    721 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Autoresolving PR Conflicts

    PostHog/posthog-foss

    Official

    Operating procedure for the conflict-autoresolver agent: sweep open PostHog/posthog PRs that conflict with master, resolve the trivial conflicts (generated artifacts deterministically, source…

    721 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Help users debug PostHog Error Tracking stack-trace symbolication for any supported platform — JavaScript/TypeScript web, React Native (Hermes), Android (Proguard / R8), or iOS / macOS (dSYM).

    721 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Exploring Apm Traces

    PostHog/posthog-foss

    Official

    Investigates distributed application performance using PostHog APM (OpenTelemetry span) data via MCP.

    721 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Exploring LLM Traces

    PostHog/posthog-foss

    Official

    Debug and inspect LLM/AI agent traces using PostHog's MCP tools.

    721 GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Investigate Metric

    PostHog/posthog-foss

    Official

    Diagnose why a product metric changed (dropped, spiked, or plateaued) by orchestrating breakdowns, actors, paths, lifecycle, retention, and annotations queries.

    721 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Categories

Questions about Fixing Flaky Tests

What does Fixing Flaky Tests do?

Guides an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo. Fixing Flaky Tests is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Guides an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo.

When should I use Fixing Flaky Tests?

Fixing Flaky Tests fits situations like: A test fails intermittently in CI but passes on rerun; hogli ci:insights; the debugging-ci-failures skill classifies a failure as a flaky test; given a GitHub Actions URL for a flaky job.

How do I install Fixing Flaky Tests in Claude Code?

Run `npx skills add PostHog/posthog-foss --skill fixing-flaky-tests -a claude-code`. Or copy the skill folder (.agents/skills/fixing-flaky-tests in PostHog/posthog-foss) into .claude/skills/fixing-flaky-tests in your project. Claude Code loads it when a task matches its description.

How do I install Fixing Flaky Tests in Codex?

Run `npx skills add PostHog/posthog-foss --skill fixing-flaky-tests -a codex`. Or copy the skill folder (.agents/skills/fixing-flaky-tests in PostHog/posthog-foss) into .agents/skills/fixing-flaky-tests in your project. Codex loads it when a task matches its description.

Can I use Fixing Flaky Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PostHog/posthog-foss --skill fixing-flaky-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fixing-flaky-tests, .gemini/skills/fixing-flaky-tests, .github/skills/fixing-flaky-tests and .opencode/skills/fixing-flaky-tests in your project.

What does Fixing Flaky Tests need to run?

Going by SKILL.md and its folder, Fixing Flaky Tests needs the command-line tools its instructions call (git, gh and pnpm) and credentials named TRUNK_API_TOKEN. Our summary lists: A credential in TRUNK_API_TOKEN.

Does Fixing Flaky Tests access the network?

SKILL.md contains no URLs. Its commands use git and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Fixing Flaky Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fixing Flaky Tests use?

Fixing Flaky Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fixing Flaky Tests use?

About 5.9k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fixing Flaky Tests?

Skills that share tags, products or a category with Fixing Flaky Tests: Debug Playwright Prow (quay/quay, 2.8k stars), Designing Tests (CloudAI-X/claude-workflow-v2, 1.4k stars), Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars) and Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fixing Flaky Tests?

PostHog (a GitHub organization, an official publisher) maintains it in PostHog/posthog-foss, which has 721 GitHub stars. The repository holds 213 skills in this directory. The repository was last updated on October 7, 2026.

Source: PostHog/posthog-foss on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.