Agent skill

Hardening Flaky E2E Tests

by Comfy-Org in Comfy-Org/ComfyUI_frontend

Diagnoses and fixes flaky Playwright e2e tests by replacing race-prone patterns with retry-safe alternatives.

GPL-3.0Auto-check passedTesting & QA

Install Hardening Flaky E2E Tests

skills CLI
$ npx skills add Comfy-Org/ComfyUI_frontend --skill hardening-flaky-e2e-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Comfy-Org/ComfyUI_frontend hardening-flaky-e2e-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Comfy-Org/ComfyUI_frontend.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/hardening-flaky-e2e-tests .claude/skills/hardening-flaky-e2e-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hardening-flaky-e2e-tests
GitHub stars
2.1k
Token cost
~2.8k tokens
SKILL.md length
806 words
Files
1
Skills in repo
22
Repo updated
First seen
Licence
GPL-3.0

At a glance

Diagnoses and fixes flaky Playwright e2e tests by replacing race-prone patterns with retry-safe alternatives.

  • Works in 7 steps: Gather CI Evidence → Classify the Flake → Apply the Transform → …
  • Triaging CI flakes
  • SKILL.md covers Workflow and Local Noise — Do Not Fix
  • Calls gh, git and pnpm

What it does

Hardening Flaky E2E Tests is an agent skill from Comfy-Org/ComfyUI_frontend. Diagnoses and fixes flaky Playwright e2e tests by replacing race-prone patterns with retry-safe alternatives. Use when triaging CI flakes, hardening spec files, fixing timing races, or asked to stabilize browser tests. Triggers on: flaky, flake, harden, stabilize, race condition in e2e, intermittent failure.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering End-to-end testing, Failing and flaky tests and Async programming. It works with Playwright. The repository describes itself as: Official front-end implementation of ComfyUI. The licence is GPL-3.0.

When your agent uses it

  • Triaging CI flakes
  • Hardening spec files
  • Fixing timing races
  • Asked to stabilize browser tests

Example prompts

  • “Use the hardening-flaky-e2e-tests skill to diagnose and fixes flaky Playwright e2e tests by replacing race-prone patterns with retry-safe alternatives”
  • “/hardening-flaky-e2e-tests”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Gather CI Evidence
  2. Classify the Flake
  3. Apply the Transform
  4. Keep Changes Narrow
  5. Verify Narrowly
  6. Watch CI E2E Runs
  7. Pre-merge Checklist

What it can do on your machine

Read from SKILL.md and the folder at commit f9be289. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • git
    • pnpm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh, git and pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hardening Flaky E2E Tests loads about 2.8k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 806 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Comfy-Org/ComfyUI_frontend at commit f9be289, republished under its GPL-3.0 licence (© Comfy-Org). 806 words, ~2,782 tokens.

Download SKILL.mdSave it as .claude/skills/hardening-flaky-e2e-tests/SKILL.md (or your agent's skills folder).
name
hardening-flaky-e2e-tests
description
Diagnoses and fixes flaky Playwright e2e tests by replacing race-prone patterns with retry-safe alternatives. Use when triaging CI flakes, hardening spec files, fixing timing races, or asked to stabilize browser tests. Triggers on: flaky, flake, harden, stabilize, race condition in e2e, intermittent failure.

Hardening Flaky E2E Tests

Fix flaky Playwright specs by identifying race-prone patterns and replacing them with retry-safe alternatives. This skill covers diagnosis, pattern matching, and mechanical transforms — not writing new tests (see writing-playwright-tests for that). The general rules behind every transform here (wait on the real readiness boundary, never sleep, never weaken an assertion, check history before removing a guard) are in docs/guidance/testing-principles.md.

Workflow

1. Gather CI Evidence
bash
gh run list --workflow=ci-test.yaml --limit=5
gh run download <run-id> -n playwright-report
  • Open report.json and search for "status": "flaky" entries.
  • Collect file paths, test titles, and error messages.
  • Do NOT trust green checks alone — flaky tests that passed on retry still need fixing.
  • Use error-context.md, traces, and page snapshots before editing code.
  • Pull the newest run after each push instead of assuming the flaky set is unchanged.
2. Classify the Flake

Read the failing assertion and match it against the pattern table. Most flakes fall into one of these categories:

#PatternSignature in CodeFix
1Snapshot-then-assertexpect(await evaluate()).toBe(x)await expect.poll(() => evaluate()).toBe(x)
2Immediate countconst n = await loc.count(); expect(n).toBe(3)await expect(loc).toHaveCount(3)
3nextFrame after menu clickclickMenuItem(x); nextFrame()clickMenuItem(x); contextMenu.waitForHidden()
4Tight poll timeoutexpect.poll(..., { timeout: 250 })≥2000 ms; prefer default 5000 ms
5Immediate evaluate after mutationsetSetting(k, v); expect(await evaluate()).toBe(x)await expect.poll(() => evaluate()).toBe(x)
6Screenshot without readinessloadWorkflow(); nextFrame(); toHaveScreenshot()Assert the expected node state with a retrying assertion first
7Non-deterministic node ordergetNodeRefsByType('X')[0] with >1 matchgetNodeRefById(id) or guard toHaveLength(1)
8Fake readiness helperHelper clicks but doesn't assert stateRemove; poll the actual value
9Immediate graph state after dropexpect(await getLinkCount()).toBe(1)await expect.poll(() => getLinkCount()).toBe(1)
10Immediate boundingBox/layout readconst box = await loc.boundingBox(); expect(box!.width)await expect.poll(() => loc.boundingBox().then(b => b?.width))
3. Apply the Transform
Rule: Choose the Smallest Correct Assertion
  • Locator state → use built-in retrying assertions: toBeVisible(), toHaveText(), toHaveCount(), toHaveClass()
  • Single async value → expect.poll(() => asyncFn()).toBe(expected)
  • Multiple assertions that must settle together → expect(async () => { ... }).toPass()
  • Never use waitForTimeout() to hide a race.
typescript
// ✅ Single value — use expect.poll
await expect
  .poll(() => comfyPage.page.evaluate(() => window.app!.graph.links.length))
  .toBe(3)

// ✅ Locator count — use toHaveCount
await expect(comfyPage.page.locator('.dom-widget')).toHaveCount(2)

// ✅ Multiple conditions — use toPass
await expect(async () => {
  expect(await node1.getValue()).toBe('foo')
  expect(await node2.getValue()).toBe('bar')
}).toPass({ timeout: 5000 })
Rule: Wait for the Real Readiness Boundary

Visible is not always ready. Prefer user-facing assertions when possible; poll internal state only when there is no UI surface to assert on.

Vue-node tests must use @vue-nodes; the fixture enables the renderer and waits for initial readiness. Never manually set Comfy.VueNodes.Enabled or call comfyPage.vueNodes.waitForNodes() in tests.

Common readiness boundaries:

After this action...Wait for...
Canvas interaction (drag, click node)await comfyPage.nextFrame()
Menu item clickawait contextMenu.waitForHidden()
Workflow loadawait comfyPage.workflow.loadWorkflow(...) (built-in wait)
Settings writePoll the setting value with expect.poll()
Node pin/bypass/collapse toggleawait expect.poll(() => nodeRef.isPinned()).toBe(true)
Graph mutation (add/remove node, link)Poll link/node count
Clipboard writePoll pasted value
ScreenshotAssert expected node state with a retrying assertion
Rule: Expose Locators for Retrying Assertions

When a helper returns a count via await loc.count(), callers can't use toHaveCount(). Expose the underlying Locator as a getter so callers choose between:

typescript
// Helper exposes locator
get domWidgets(): Locator {
  return this.page.locator('.dom-widget')
}

// Caller uses retrying assertion
await expect(comfyPage.domWidgets).toHaveCount(2)

Replace count methods with locator getters so callers can use retrying assertions directly.

Rule: Fix Check-then-Act Races in Helpers
typescript
// ❌ Race: count can change between check and waitFor
const count = await locator.count()
if (count > 0) {
  await locator.waitFor({ state: 'hidden' })
}

// ✅ Direct: waitFor handles both cases
await locator.waitFor({ state: 'hidden' })
Show full SKILL.md (331 more words)Show less
Rule: Remove force:true from Clicks

force: true bypasses actionability checks, hiding real animation/visibility races. Remove it and fix the underlying timing issue.

typescript
// ❌ Hides the race
await closeButton.click({ force: true })

// ✅ Surfaces the real issue — fix with proper wait
await closeButton.click()
await dialog.waitForHidden()
Rule: Handle Non-deterministic Element Order

When getNodeRefsByType returns multiple nodes, the order is not guaranteed. Don't use index [0] blindly.

typescript
// ❌ Assumes order
const node = (await comfyPage.nodeOps.getNodeRefsByType('CLIPTextEncode'))[0]

// ✅ Find by ID or proximity
const nodes = await comfyPage.nodeOps.getNodeRefsByType('CLIPTextEncode')
let target = nodes[0]
for (const n of nodes) {
  const pos = await n.getPosition()
  if (Math.abs(pos.y - expectedY) < minDist) target = n
}

Or guard the assumption:

typescript
const nodes = await comfyPage.nodeOps.getNodeRefsByType('CLIPTextEncode')
expect(nodes).toHaveLength(1)
const node = nodes[0]
Rule: Use toPass for Timing-sensitive Dismiss Guards

Some UI elements (e.g. LiteGraph's graphdialog) have built-in dismiss delays. Retry the entire dismiss action:

typescript
// ✅ Retry click+assert together
await expect(async () => {
  await comfyPage.canvas.click({ position: { x: 10, y: 10 } })
  await expect(dialog).toBeHidden({ timeout: 500 })
}).toPass({ timeout: 5000 })
4. Keep Changes Narrow
  • Shared helpers should drive setup to a stable boundary.
  • Do not encode one-spec timing assumptions into generic helpers.
  • If a race only matters to one spec, prefer a local wait in that spec.
  • If a helper fails before the real test begins, remove or relax the brittle precondition and let downstream UI interaction prove readiness.
5. Verify Narrowly
bash
# Targeted rerun with repetition
pnpm test:browser:local -- browser_tests/tests/myFile.spec.ts --repeat-each 10

# Single test by line number (avoids grep quoting issues on Windows)
pnpm test:browser:local -- browser_tests/tests/myFile.spec.ts:42
  • Use --repeat-each 10 for targeted flake verification (use 20 for single test cases).
  • Verify with the smallest command that exercises the flaky path.
6. Watch CI E2E Runs

After pushing, use gh to monitor the E2E workflow:

bash
# Find the run for the current branch
gh run list --workflow="CI: Tests E2E" --branch=$(git branch --show-current) --limit=1

# Watch it live (blocks until complete, streams logs)
gh run watch <run-id>

# One-liner: find and watch the latest E2E run for the current branch
gh run list --workflow="CI: Tests E2E" --branch=$(git branch --show-current) --limit=1 --json databaseId --jq ".[0].databaseId" | xargs gh run watch

On Windows (PowerShell):

powershell
# One-liner equivalent
gh run watch (gh run list --workflow="CI: Tests E2E" --branch=$(git branch --show-current) --limit=1 --json databaseId --jq ".[0].databaseId")

After the run completes:

bash
# Download the Playwright report artifact
gh run download <run-id> -n playwright-report

# View the run summary in browser
gh run view <run-id> --web

Also watch the unit test workflow in parallel if you changed helpers:

bash
gh run list --workflow="CI: Tests Unit" --branch=$(git branch --show-current) --limit=1
7. Pre-merge Checklist

Before merging a flaky-test fix, confirm:

  • The latest CI artifact was inspected directly
  • The root cause is stated as a race or readiness mismatch
  • The fix waits on the real readiness boundary
  • The assertion primitive matches the job (poll vs toHaveCount vs toPass)
  • The fix stays local unless a shared helper truly owns the race
  • Local verification uses a targeted rerun
  • No behavioral changes to the test — only timing/retry strategy updated

Local Noise — Do Not Fix

These are local distractions, not CI root causes:

  • Missing local input fixture files required by the test path
  • Missing local models directory
  • Teardown EPERM while restoring the local browser-test user data directory
  • Local screenshot baseline differences on Windows

Rules:

  • First confirm whether it blocks the exact flaky path under investigation.
  • Do not commit temporary local assets used only for verification.
  • Do not commit local screenshot baselines.

© Comfy-Org, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/hardening-flaky-e2e-tests of Comfy-Org/ComfyUI_frontend.

Open the folder on GitHubat commit f9be289

Compare with similar skills

Hardening Flaky E2E Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hardening Flaky E2E Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hardening Flaky E2E Tests this skillComfy-Org/ComfyUI_frontend2.1k—~2.8kAutomated safety check: PassGPL-3.0
Cucumber and Playwright E2E Testslanggenius/dify158k—~682Automated safety check: PassCustom licence
Studio E2E Testssupabase/supabase111k—~2.8kAutomated safety check: PassApache-2.0
Debug Playwrightquay/quay2.8k—~1.2kAutomated safety check: PassApache-2.0
Playwright Testingchongdashu/vibejam-starter-pack149—~2.1kAutomated safety check: PassNone
Playwright Testingchongdashu/vibejam-starter-pack149—~2.2kAutomated safety check: PassNone

Similar skills

  • Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions.

    158k GitHub stars~682 tokensUpdated today
    Testing & QAAuto-check passed
  • Studio E2E Tests

    supabase/supabase

    Official

    Write and run Playwright E2E tests for Supabase Studio (e2e/studio).

    111k GitHub stars~2.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Debug Playwright E2E test failures from GitHub Actions CI runs.

    2.8k GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.1k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.2k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • E2E

    sendou-ink/sendou.ink

    Run, debug, and manage Playwright e2e tests. An agent skill from sendou-ink/sendou.ink.

    297 GitHub stars~2.1k tokensUpdated yesterday
    Testing & QAAuto-check: notes

More from Comfy-Org/ComfyUI_frontend

All 22 skills in this repo
  • Adding Deprecation Warnings

    Comfy-Org/ComfyUI_frontend

    Adds deprecation warnings for renamed or removed properties/APIs.

    2.1k GitHub stars~775 tokensUpdated today
    Auto-check passed
  • Agent Integration Replay

    Comfy-Org/ComfyUI_frontend

    Replay recorded agent conversations as Playwright tests against the real chat panel and canvas.

    2.1k GitHub stars~805 tokensUpdated today
    Auto-check: notes
  • Codegen Transform

    Comfy-Org/ComfyUI_frontend

    Transforms raw Playwright codegen output into ComfyUI convention-compliant tests.

    2.1k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Comment Sicko

    Comfy-Org/ComfyUI_frontend

    Dispatches the comment-sicko subagent to hunt gratuitous comments in a PR/diff, triages its raw findings, and posts a polite, professional writeup.

    2.1k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Perf Fix With Proof

    Comfy-Org/ComfyUI_frontend

    Ships performance fixes with CI-proven improvement using stacked PRs.

    2.1k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Publishing A New Package

    Comfy-Org/ComfyUI_frontend

    Publishes a new package from this monorepo to npm under @comfyorg and proves it is consumable from another repo.

    2.1k GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Hardening Flaky E2E Tests

What does Hardening Flaky E2E Tests do?

Diagnoses and fixes flaky Playwright e2e tests by replacing race-prone patterns with retry-safe alternatives. Hardening Flaky E2E Tests is an agent skill from Comfy-Org/ComfyUI_frontend. Diagnoses and fixes flaky Playwright e2e tests by replacing race-prone patterns with retry-safe alternatives.

When should I use Hardening Flaky E2E Tests?

Hardening Flaky E2E Tests fits situations like: triaging CI flakes; hardening spec files; fixing timing races; asked to stabilize browser tests.

How do I install Hardening Flaky E2E Tests in Claude Code?

Run `npx skills add Comfy-Org/ComfyUI_frontend --skill hardening-flaky-e2e-tests -a claude-code`. Or copy the skill folder (.claude/skills/hardening-flaky-e2e-tests in Comfy-Org/ComfyUI_frontend) into .claude/skills/hardening-flaky-e2e-tests in your project. Claude Code loads it when a task matches its description.

How do I install Hardening Flaky E2E Tests in Codex?

Run `npx skills add Comfy-Org/ComfyUI_frontend --skill hardening-flaky-e2e-tests -a codex`. Or copy the skill folder (.claude/skills/hardening-flaky-e2e-tests in Comfy-Org/ComfyUI_frontend) into .agents/skills/hardening-flaky-e2e-tests in your project. Codex loads it when a task matches its description.

Can I use Hardening Flaky E2E Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Comfy-Org/ComfyUI_frontend --skill hardening-flaky-e2e-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hardening-flaky-e2e-tests, .gemini/skills/hardening-flaky-e2e-tests, .github/skills/hardening-flaky-e2e-tests and .opencode/skills/hardening-flaky-e2e-tests in your project.

What does Hardening Flaky E2E Tests need to run?

Going by SKILL.md and its folder, Hardening Flaky E2E Tests needs the command-line tools its instructions call (gh, git and pnpm).

Does Hardening Flaky E2E Tests access the network?

SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Hardening Flaky E2E Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Hardening Flaky E2E Tests use?

Hardening Flaky E2E Tests is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hardening Flaky E2E Tests use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hardening Flaky E2E Tests?

Skills that share tags, products or a category with Hardening Flaky E2E Tests: Cucumber and Playwright E2E Tests (langgenius/dify, 158k stars), Studio E2E Tests (supabase/supabase, 111k stars), Debug Playwright (quay/quay, 2.8k stars) and Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hardening Flaky E2E Tests?

Comfy-Org (a GitHub organization) maintains it in Comfy-Org/ComfyUI_frontend, which has 2,056 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 9, 2026.

Source: Comfy-Org/ComfyUI_frontend on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.