Debug Playwright Prow
quay/quay
Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real…
Guides an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo.
$ npx skills add PostHog/posthog-foss --skill fixing-flaky-tests -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install PostHog/posthog-foss fixing-flaky-tests --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/fixing-flaky-tests .claude/skills/fixing-flaky-tests && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "fixing-flaky-tests" agent skill from https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/fixing-flaky-tests into .claude/skills/fixing-flaky-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fixing-flaky-tests", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/fixing-flaky-testsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add PostHog/posthog-foss --skill fixing-flaky-tests -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install PostHog/posthog-foss fixing-flaky-tests --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/fixing-flaky-tests .agents/skills/fixing-flaky-tests && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "fixing-flaky-tests" agent skill from https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/fixing-flaky-tests into .agents/skills/fixing-flaky-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fixing-flaky-tests", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PostHog/posthog-foss --skill fixing-flaky-tests -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install PostHog/posthog-foss fixing-flaky-tests --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/fixing-flaky-tests .cursor/skills/fixing-flaky-tests && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "fixing-flaky-tests" agent skill from https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/fixing-flaky-tests into .cursor/skills/fixing-flaky-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fixing-flaky-tests", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/PostHog/posthog-foss.git --path .agents/skills/fixing-flaky-tests--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add PostHog/posthog-foss --skill fixing-flaky-tests -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install PostHog/posthog-foss fixing-flaky-tests --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/fixing-flaky-tests .gemini/skills/fixing-flaky-tests && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "fixing-flaky-tests" agent skill from https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/fixing-flaky-tests into .gemini/skills/fixing-flaky-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fixing-flaky-tests", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install PostHog/posthog-foss fixing-flaky-testsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add PostHog/posthog-foss --skill fixing-flaky-tests -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/fixing-flaky-tests .github/skills/fixing-flaky-tests && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "fixing-flaky-tests" agent skill from https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/fixing-flaky-tests into .github/skills/fixing-flaky-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fixing-flaky-tests", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PostHog/posthog-foss --skill fixing-flaky-tests -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install PostHog/posthog-foss fixing-flaky-tests --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/fixing-flaky-tests .opencode/skills/fixing-flaky-tests && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "fixing-flaky-tests" agent skill from https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/fixing-flaky-tests into .opencode/skills/fixing-flaky-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fixing-flaky-tests", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
fixing-flaky-testsGuides an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo.
Fixing Flaky Tests is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Guides an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo. Use when a test fails intermittently in CI but passes on rerun or locally, when hogli ci:insights or the debugging-ci-failures skill classifies a failure as a flaky test, when given a GitHub Actions URL for a flaky job, when asked to check Trunk Flaky Tests for a test, PR, or master, or when asked to deflake, stabilize, or fix a flaky Jest, pytest, or Playwright test. Core discipline: reproduce locally…
Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering Failing and flaky tests and Unit testing. It works with PostHog, Playwright, GitHub Actions and Jest. The repository describes itself as: PostHog FOSS is a read-only mirror of PostHog, with all proprietary code removed. NOTE: This repo is synced automatically from the main PostHog repo. Please raise any issues and… The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 2c48221. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitghpnpmFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, gh and pnpm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
TRUNK_API_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Fixing Flaky Tests loads about 5.9k tokens when it runs. Until then it costs about 239 tokens; SKILL.md has 2,835 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from PostHog/posthog-foss at commit 2c48221, republished under its MIT licence (© PostHog). 2,835 words, ~5,912 tokens.
.claude/skills/fixing-flaky-tests/SKILL.md (or your agent's skills folder).Before you propose a change to how the suite runs in CI, check things already tried. It records measured verdicts on test parallelism, sharding, and coverage-based selection, so a rejected approach is not rebuilt.
Three non-negotiables, in order:
Stabilizing the test is not the only valid ending. Once you know why it flakes, step 5 asks whether it should exist. A test that catches nothing real is worth deleting, and one that flakes because of the level it runs at is worth moving down a rung.
Before any of these: measure, don't assume. Flaky-vs-deterministic, and the rate, are facts to establish from verifiable GitHub run data (step 1) — never inherited from a Slack alert, a teammate's guess, or a ci:insights label.
For triaging a red CI run (finding and classifying the failure), use the debugging-ci-failures skill first — this skill takes over once the failure is classified as a flaky test.
For writing new Playwright tests that aren't flaky, use the playwright-test skill.
investigating-ci-failures (green/red boundary) and diagnosing-ci-and-merge-bottlenecks (the engineering-analytics-flaky-tests tool's caveats) are product skills under products/engineering_analytics/skills/, not invocable here: read their SKILL.md at that path.
The GitHub Actions API (or GitHub MCP) is the source of truth. hogli ci:insights is a digest, not an oracle — it can mislabel flaky-vs-deterministic, it lags the API until GitHub's webhook settles, and it cannot give you a rate at all: it reports absolute counts, because CI emits failures but omits ordinary passing runs, so there is no denominator. Use it to validate a hypothesis or pull historical context, never as the first move or the classification authority.
Establish the timeline yourself from raw run data:
# Pass/fail history for the workflow on the branch under suspicion (master shown):
gh run list --repo PostHog/posthog --workflow=ci-backend.yml --branch master \
--status completed --limit 60 --json conclusion,headSha,createdAt,databaseIdThe run-level conclusion is not enough — a red run may be failing on a different job/test. Confirm it is the same test each time by reading the failing shard's log; and for green runs, confirm the test actually ran and passed in that shard (it may have been sharded elsewhere, not fixed):
gh api repos/PostHog/posthog/actions/jobs/{job_id}/logs \
| grep -E "<exact::test::id>|short test summary"Read the timeline before you classify:
p^30 ≈ 0. That is a deterministic regression: go to debugging-ci-failures and find the introducing commit (step 4).master passes the same code → an environment-dependent race, not a regression: kernel clock granularity, filesystem, runner speed. Retrying on the same runner cannot help, so "failed 3/3 attempts" says nothing about whether the PR caused it. A merge-queue failure hands you the control for free: the trunk-merge/** test PR's body names the master SHA it is based on, so read master's run at that SHA (gh api repos/PostHog/posthog/commits/<sha>/check-runs) and confirm the same test ran and passed there.products/engineering_analytics/skills/investigating-ci-failures/ has the boundary query if the failure reached master.If a failure is reported as (or you suspect it is) consistent, don't serialize — measure the rate and attempt a repro in parallel.
Record the measured rate (failures / total runs, from the run data). You need it to size the validation loop in step 7.
Then confirm it is not already handled:
git log --oneline -10 -- <test_file_path> # recently fixed already?
hogli ci:insights search "<test name or error>" # cross-run history — corroborate against the run data, do not trust blindlysearch reports two surfaces; read each for what it actually claims:
potentially_resolved state means that job's latest default-branch run is green again — weak evidence a fix landed, not proof. Confirm against the run data that it covers this failure before reporting instead of re-fixing.engineering-analytics-flaky-tests MCP tool reads). confirmed_flake is the only classification backed by proof: one commit both failed and passed the test in the same matrix job, via a re-run attempt going green or an in-job retry. A pass in a different matrix leg is not recovery. suspected_regression means no recovery was recorded — absence of proof, not proof of a regression, so treat it as real until your own run data says otherwise.CI uploads test results to Trunk Flaky Tests, which tracks per-test failure history across master and PR runs.
On a PR, the trunk-io bot's "Trunk Test Analytics" comment links that PR's slice (https://app.trunk.io/posthog-inc/flaky-tests/pr/<number>?repo=PostHog/posthog).
The trunk MCP server in .mcp.json queries it (tools are marked experimental by Trunk):
search-test (repoName: "PostHog/posthog", testNameSearch: "<test name, no filepath>") → the test case ID.fix-flaky-test (repoName, testCaseId) → failure history, first-seen commit, git blame, and Trunk's root-cause investigation; createNewInvestigation: true triggers a fresh analysis (takes up to a minute).Authenticate once via /mcp → trunk (browser OAuth); headless environments instead add an Authorization: Bearer header with a TRUNK_API_TOKEN org token to the server entry.
Two limits worth knowing before you start here.
AI investigations are not enabled for this repo, so fix-flaky-test returns history or nothing, never a root cause.
And lookup only goes name to ID: a bare dashboard link identifies a test you cannot name, so ask for the test name rather than guessing at the ID.
Trunk attributes each test to a team through CODEOWNERS, which cannot express the owners.yaml map.
.github/scripts/trunk-codeowners.sh projects the map into a generated CODEOWNERS before each upload (hogli owners:codeowners builds the same file locally), so a test's owner in Trunk should match hogli owners:who.
Where it does not, the projection dropped a spelling two teams would both claim.
Like ci:insights, this is corroboration and history, not the classification authority — flaky-vs-deterministic and the rate still come from the run data above.
gh run view <run-id> --log-failed
# If the job was re-run and the latest attempt passed, the flaky failure
# lives in a previous attempt:
gh api "repos/PostHog/posthog/actions/runs/{run_id}/attempts/1/jobs"
gh api "repos/PostHog/posthog/actions/jobs/{job_id}/logs"Capture before moving on:
[MSW] Unhandled, async leak warnings, teardown errors — these are often the actual cause, printed before the symptom.Escalate through these conditions until the failure appears. Stop at the first level that reproduces it; that level is your validation environment for step 7.
Single run: hogli test <path>::<test> — confirms the test runs at all.
Repetition loop (default N=20): catches probabilistic flakes.
CI-like conditions: CI runs Jest sharded with low worker counts on contended runners — read the current flags from frontend/package.json's test script and .github/workflows/ci-frontend.yml before running.
Example (with the flags as of this writing), running the test alongside its shard neighbors:
pnpm --filter=@posthog/frontend jest <test_file> <neighbor_file> --maxWorkers=2 --forceExitFor pytest, run the whole file or class rather than the single test, so module-level fixtures and ordering match CI.
Ordering: run suspected polluter tests before the victim; reverse the order within the file. Flakes that vanish in isolation are ordering bugs.
Contention: re-run the loop while something CPU-heavy runs in another shell (e.g. a parallel full-file jest run). Timeout-class flakes often only show here.
The loop harness — judge by exit code, not by grepping output:
N=20; PASS=0; FAIL=0
for i in $(seq 1 $N); do
if <test command> >/tmp/flake-run.log 2>&1; then
PASS=$((PASS+1))
else
FAIL=$((FAIL+1)); cp /tmp/flake-run.log /tmp/flake-fail-$i.log; echo "run $i: FAIL"
fi
done
echo "$PASS passed, $FAIL failed out of $N"Two cost notes for the loop:
break after the first failure — one captured failure log is enough.
Complete all N runs only when measuring the failure rate or validating in step 7.pnpm --filter=@posthog/frontend jest script runs pnpm build:products before every invocation.
Inside a loop, build once, then iterate with pnpm --filter=@posthog/frontend exec jest ..., which skips the rebuild.If nothing reproduces after the full ladder, the flake is CI-environment-specific.
Before settling for that, ask what the runner has that your machine lacks: Linux mtimes from a coarse clock where macOS gives nanoseconds, a case-sensitive filesystem, a different timezone, /tmp on a different filesystem.
Often you can force the CI condition instead of waiting for it — os.utime the files into the ordering the coarse clock produced, run under TZ=UTC — and then validate empirically after all.
Otherwise proceed with a fix grounded in the CI evidence and root-cause analysis, and say so explicitly in the report — the validation in step 7 is then analytical, not empirical.
Match the symptom to a cause class; never patch the symptom.
| Symptom | Likely cause class |
|---|---|
| Timeout waiting for promise/listener/element | Unawaited async work, missing mock, hidden pending request |
| Passes alone, fails with neighbors (or vice versa) | Shared state: module cache, DB rows, global config, ordering |
| Fails near midnight/UTC boundaries, or on slow runners | Real clock usage — missing time_machine.travel / fake timers |
| Assertion on list order or generated IDs | Nondeterministic ordering/IDs asserted as deterministic |
| Query can't see just-written data | Eventual consistency (ClickHouse), missing flush/commit |
Only fails under --maxWorkers=2 / contention | Race condition surfaced by scheduling, too-tight timeout |
| Every attempt fails on one runner, passes on another | Clock-granularity race: cutoff read from an earlier write |
If the symptom table doesn't point at a clear cause and the test file itself is unchanged (git log -- <test_file> is stale), the trigger is elsewhere — a neighbor test, a dependency bump, or a product change. Find when it started instead of guessing:
Bisect the CI run history first (cheap, no local builds): from the step-1 timeline, take the last-green → first-red boundary and diff the commits in that window (git log <good>..<bad>). That short list often names the culprit outright.
git bisect the code when you can reproduce locally and the failure is (near-)deterministic:
git bisect start <bad-sha> <good-sha>
git bisect run bash -c '<repro command>' # exit 0 = good, non-zero = badCaveat: for an intermittent flake, a lucky pass at a step sends git bisect down the wrong path. Trust code-bisect only when the failure is deterministic; otherwise run the repro N times per step (fail if any iteration fails), or just use the CI run-history boundary.
PostHog-specific patterns:
afterMount loaders fire API calls through the connect() chain — a logic three levels deep can trigger an unmocked fetch.
Unhandled requests currently resolve with a benign empty paginated 200, so the symptom is a loader succeeding with empty or wrong data, not a network error.
The [MSW] Unhandled GET ... warning in the log names the missing mock — add the useMocks entry.
The unhandled-request behavior has changed before (it used to hang); if symptoms don't match, read frontend/src/mocks/jest.ts for what unmocked requests do today.toFinishAllListeners() timeouts: waits for ALL kea listener promises across ALL mounted logics (3s default — LISTENER_FINISH_WAIT_TIMEOUT in kea-test-utils).
Any connected logic with a pending loader blocks it.
Fix the pending work; do not raise the timeout.mocksToHandlers strips trailing slashes, but query params, @current-style segments, and :param patterns must match the real request URL.
Compare against the [MSW] Unhandled line.frontend/src/mocks/jest.ts registers a global afterEach(() => mswServer.resetHandlers()).
Each it.each case is a separate test, so runtime mocks must be (re-)registered in beforeEach.beforeEach and never unmounted leak async work into later tests.@pytest.mark.django_db(transaction=True).time_machine.travel(..., tick=False); never assert on now()-derived values./writing-tests, "Two writes in a row are not ordered in time").Once you know why it flakes, ask whether the test should exist at all, before you spend effort stabilizing it.
A flaky test is the one case where cost is already proven and value is not: it has demonstrably cost reruns, wall-clock, and attention.
So apply /writing-tests' gate retroactively, with more force than you would to a new test:
What realistic regression does this test catch that no existing test already catches?
Three outcomes are valid. Pick deliberately; don't default to the first.
| Outcome | When | Next |
|---|---|---|
| Fix it | Guards a real regression at roughly the right cost. | step 6 |
| Re-level it | Worth guarding, but the flake is inherent to the level it runs at. | below |
| Delete it | You cannot name the regression it catches, or another test already catches it. | below |
Recurring shapes that fail the gate:
Deletion is irreversible, so it carries a higher bar than a fix:
Then skip to step 8: there is nothing to loop.
Move the test down the cost ladder in /writing-tests rather than hardening it in place.
A round trip through a real broker, browser, or vendor API to prove logic a direct call could prove is testing the transport, and the transport is where the nondeterminism lives.
Re-leveling removes the flake by construction, so prefer it over an increasingly elaborate wait.
The replacement is a new test: run /writing-tests' gate on it, and validate it at its new level rather than against the old repro conditions, which no longer exist.
| Tempting masking move | Do instead |
|---|---|
sleep(2) / setTimeout before asserting | Await the specific condition (waitFor, expectLogic, explicit flush) |
| Raise the test/listener timeout | Find what is hanging; the timeout is the messenger |
Add retries (pytest-rerunfailures --reruns, jest.retryTimes) | Reserve for genuinely nondeterministic external infra, with a comment and a linked issue — never for product code under test |
| Skip / quarantine the test | Only with explicit user approval, with a linked issue |
| Loosen the assertion | Make the data deterministic (sort, freeze, seed), keep the assertion strict |
| Harden a test that shouldn't exist | Go back to step 5 — deleting or re-leveling it is the cheaper fix |
Keep the fix minimal and inside the test or its fixtures when possible. If the race is in product code, the flake found a real bug — fix the product code and say so in the report.
Run the step-3 harness on the fixed code under the same conditions that reproduced the failure (same neighbors, worker count, contention).
Size N from the observed pre-fix failure rate: if it failed about 1 in k runs, you need roughly N ≥ 3k consecutive passes for ~95% confidence the flake is gone — (1 - 1/k)^(3k) ≈ 5%.
So use N = max(3k, 20).
Without a usable rate estimate, run 50 and note the reduced confidence in the report.
If the flake was never reproducible locally, run N = 20 as a regression check and label the validation as analytical.
Any failure in the loop → back to step 4; the root cause was wrong or incomplete. Finish with one normal run of the surrounding file/suite to confirm the fix didn't break sibling tests.
Re-leveled instead? The old repro conditions no longer apply. Loop the replacement at its own level, and confirm the flake is gone because the old test is gone, not because it got faster.
Deleted instead? Nothing to loop. Run the surrounding file/suite once to confirm nothing depended on it, and carry the coverage argument into the report.
Test: <file path>::<test name>
Observed in CI: <measured rate from run data, e.g. 8/45 runs over 3h (gh run list); ci:insights state corroborates>
Local repro: <command + conditions, e.g. 3/20 failures with neighbor X, maxWorkers=2 | not reproducible locally>
Root cause: <one or two sentences>
Outcome: fixed | re-leveled (<from> → <to>) | deleted
Change: <what changed and why it removes the cause; for a deletion, what still covers the behavior (file:line) and what coverage is genuinely lost>
Validation: <N>/<N> passes under repro conditions | <N>/<N> at the new level | analytical only (CI-specific) | n/a, deleted
Follow-ups: <product bug found, related tests with the same pattern, or none>.github/workflows/ as part of a flake fix.© PostHog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/fixing-flaky-tests of PostHog/posthog-foss.
Open the folder on GitHubat commit 2c48221
Fixing Flaky Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Fixing Flaky Tests this skillPostHog/posthog-foss | 721 | — | ~5.9k | Automated safety check: Pass | MIT | |
| Debug Playwright Prowquay/quay | 2.8k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Designing TestsCloudAI-X/claude-workflow-v2 | 1.4k | 1 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Playwright Testingchongdashu/vibejam-starter-pack | 149 | — | ~2.1k | Automated safety check: Pass | None | |
| Playwright Testingchongdashu/vibejam-starter-pack | 149 | — | ~2.2k | Automated safety check: Pass | None | |
| Testingradix-ng/primitives | 274 | — | ~3.3k | Automated safety check: Pass | MIT |
quay/quay
Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real…
CloudAI-X/claude-workflow-v2
Designs and implements testing strategies for any codebase. An agent skill from CloudAI-X/claude-workflow-v2.
chongdashu/vibejam-starter-pack
Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.
chongdashu/vibejam-starter-pack
Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.
radix-ng/primitives
Test Radix NG primitives across every layer and pick the RIGHT one for a change: Vitest unit (zoneless), jest-axe a11y, Playwright browser regression (apps/visual-regression), SSR…
handsontable/handsontable
Conventions for Handsontable's Jest unit and TypeScript type tests: where files go, how to run them, mocking limits and when to write an E2E test instead.
PostHog/posthog-foss
Author useful, low-noise log alerts on services in a PostHog project.
PostHog/posthog-foss
Operating procedure for the conflict-autoresolver agent: sweep open PostHog/posthog PRs that conflict with master, resolve the trivial conflicts (generated artifacts deterministically, source…
PostHog/posthog-foss
Help users debug PostHog Error Tracking stack-trace symbolication for any supported platform — JavaScript/TypeScript web, React Native (Hermes), Android (Proguard / R8), or iOS / macOS (dSYM).
PostHog/posthog-foss
Investigates distributed application performance using PostHog APM (OpenTelemetry span) data via MCP.
PostHog/posthog-foss
Debug and inspect LLM/AI agent traces using PostHog's MCP tools.
PostHog/posthog-foss
Diagnose why a product metric changed (dropped, spiked, or plateaued) by orchestrating breakdowns, actors, paths, lifecycle, retention, and annotations queries.
Categories
Guides an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo. Fixing Flaky Tests is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Guides an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo.
Fixing Flaky Tests fits situations like: A test fails intermittently in CI but passes on rerun; hogli ci:insights; the debugging-ci-failures skill classifies a failure as a flaky test; given a GitHub Actions URL for a flaky job.
Run `npx skills add PostHog/posthog-foss --skill fixing-flaky-tests -a claude-code`. Or copy the skill folder (.agents/skills/fixing-flaky-tests in PostHog/posthog-foss) into .claude/skills/fixing-flaky-tests in your project. Claude Code loads it when a task matches its description.
Run `npx skills add PostHog/posthog-foss --skill fixing-flaky-tests -a codex`. Or copy the skill folder (.agents/skills/fixing-flaky-tests in PostHog/posthog-foss) into .agents/skills/fixing-flaky-tests in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PostHog/posthog-foss --skill fixing-flaky-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fixing-flaky-tests, .gemini/skills/fixing-flaky-tests, .github/skills/fixing-flaky-tests and .opencode/skills/fixing-flaky-tests in your project.
Going by SKILL.md and its folder, Fixing Flaky Tests needs the command-line tools its instructions call (git, gh and pnpm) and credentials named TRUNK_API_TOKEN. Our summary lists: A credential in TRUNK_API_TOKEN.
SKILL.md contains no URLs. Its commands use git and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Fixing Flaky Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.9k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Fixing Flaky Tests: Debug Playwright Prow (quay/quay, 2.8k stars), Designing Tests (CloudAI-X/claude-workflow-v2, 1.4k stars), Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars) and Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
PostHog (a GitHub organization, an official publisher) maintains it in PostHog/posthog-foss, which has 721 GitHub stars. The repository holds 213 skills in this directory. The repository was last updated on October 7, 2026.
Source: PostHog/posthog-foss on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.