Agent skill

Writing E2E Tests

by shift-editor in shift-editor/shift

Canonical rules for writing, rewriting, or reviewing Shift desktop Playwright E2E tests and visual goldens under apps/desktop/e2e/.

Apache-2.0Auto-check passedTesting & QA

Install Writing E2E Tests

skills CLI
$ npx skills add shift-editor/shift --skill writing-e2e-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install shift-editor/shift writing-e2e-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/shift-editor/shift.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/writing-e2e-tests .claude/skills/writing-e2e-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
writing-e2e-tests
GitHub stars
343
Token cost
~3.4k tokens
SKILL.md length
1,546 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Canonical rules for writing, rewriting, or reviewing Shift desktop Playwright E2E tests and visual goldens under apps/desktop/e2e/.

  • Works in 3 steps: Assert the semantic precondition that… → Park the pointer off the captured… → Capture the smallest element that shows…
  • Change a .spec.ts
  • SKILL.md covers Choose the layer before…, Waiting: never sleep, Oracles: assert what the… and Goldens: one contract, one…, plus 6 more sections
  • Calls pnpm and node

What it does

Writing E2E Tests is an agent skill from shift-editor/shift. Canonical rules for writing, rewriting, or reviewing Shift desktop Playwright E2E tests and visual goldens under apps/desktop/e2e/. Use whenever you add or change a .spec.ts, a fixture, EditorDriver, a screenshot baseline, or Playwright project membership, and whenever you investigate a flaky or failing E2E test. Covers waits, oracles, goldens, fixtures, projects, and flake verification.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering End-to-end testing. It works with Playwright. The repository describes itself as: A cross-platform font editor built in Rust and TypeScript. The licence is Apache-2.0.

When your agent uses it

  • Change a .spec.ts
  • A screenshot baseline
  • Playwright project membership
  • Whenever you investigate a flaky

Example prompts

  • “/writing-e2e-tests”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Assert the semantic precondition that makes the image meaningful: the point count, selection, tool state, outline targets, handle states…
  2. Park the pointer off the captured element unless hover is the contract.
  3. Capture the smallest element that shows the contract. Do not use a full-page golden to protect a toolbar button.

What it can do on your machine

Read from SKILL.md and the folder at commit e7dacfa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Writing E2E Tests loads about 3.4k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 1,546 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from shift-editor/shift at commit e7dacfa, republished under its Apache-2.0 licence (© shift-editor). 1,546 words, ~3,400 tokens.

Download SKILL.mdSave it as .claude/skills/writing-e2e-tests/SKILL.md (or your agent's skills folder).
name
writing-e2e-tests
description
Canonical rules for writing, rewriting, or reviewing Shift desktop Playwright E2E tests and visual goldens under `apps/desktop/e2e/`. Use whenever you add or change a `.spec.ts`, a fixture, `EditorDriver`, a screenshot baseline, or Playwright project membership, and whenever you investigate a flaky or failing E2E test. Covers waits, oracles, goldens, fixtures, projects, and flake verification.

/writing-e2e-tests — How desktop E2E tests are written

An E2E test is worth its cost only when it fails for exactly one reason: the user-visible behavior it names is broken. It must not fail because the machine was slow, the theme changed, a sidebar moved, or a retry happened to pass.

Read /writing-tests first for the general rules (observable state, no mocks, migration ledgers). Read apps/desktop/e2e/README.md for the full fixture and project reference. This skill is the checklist that turns those into a test that stays green for the right reasons.

Choose the layer before writing anything

The behavior is…Owning test
Pure computation (layout math, latch rules, range selection)Pure unit test — extract the function first
Observable through editor state after a click, drag, or keyTestEditor tool or command test
A real browser or Electron contract (DOM events, focus, native menus, dialogs, windows, processes, files)Semantic Playwright test
Appearance: colours, strokes, handles, layoutOne focused golden, after a semantic precondition
Hardware rendering, GPU residency, WebGPU presentationgpu project test
Save, quit, crash, recovery, file activation, cross-OS pathsplatform project test

If most assertions in a draft E2E test read editor state after page.evaluate, the behavior probably belongs in TestEditor. Keep the E2E test as a thin proof that the real UI reaches it.

Waiting: never sleep

waitForTimeout is banned. Wait for the condition that makes the next step valid:

Before you…Wait for
Read geometry, project coordinates, or send pointer input after setup changed geometry or cameraeditor.waitForCanvasRender()
Assert after raw page.mouse moveseditor.flushPointerMoves(), then waitForCanvasRender() or waitForIdle()
Act on an editor routewaitForEditorReady() / openCatalogGlyph() (works for authored and preview sessions)
Click a catalog cell by coordinatesclickFirstCatalogGlyph() — it waits for a laid-out, settled Grid
Act on a workspacewaitForWorkspaceReady()
Assert after quit or closedirtyDocumentDecisions() / dirtyDocumentRequests() — main-process records, not elapsed time
Compare Grid frames after an editPublished Grid state (data-grid-readiness, data-atlas-build-count, active location), armed after setup settles

When elapsed time is itself the contract (debounce, momentum, idle timeouts), put the rule in a pure function driven by timestamps and unit-test it. If the E2E test must remain, dispatch events inside the page and wait on the page clock (performance.now()) that stamps them. glyph-view.spec.ts zoom momentum is the template.

Negative assertions ("nothing else happened") need a positive anchor: wait for the event that would have triggered the unwanted outcome, then assert the outcome is absent.

Oracles: assert what the editor decided, not what pixels look like

Never count or sample hard-coded colours from a canvas. Colour oracles break under themes, alpha, anti-aliasing, device scale, and renderer changes, and many pass when the feature is broken (an empty canvas has zero "wrong" pixels).

Use published state instead:

  • geometry: editor.outline(), pointPosition(), selectionIds(), selectionBounds();
  • interaction: toolState(), hoverId(), activeSnapGuides();
  • rendering decisions: visibleOutlines(nodeId) and handleStates(node) on the glyph node definition;
  • UI state: roles, accessible names, aria-pressed, aria-checked, aria-disabled, data-* state published for tests;
  • persistence: savedGlyphNames() for .shift files and exportedGlyphNames() for exported fonts — never existsSync, byte inequality, or size > 0.

If the state you need is not published, publish it from the production code that renders it (same computation, not a parallel reimplementation), add a unit test for it, then assert it. That is how visibleOutlines and handleStates were added.

Assertions must be exact enough to fail on the wrong answer:

  • assert the exact set or order, not count() > N or not.toEqual(before);
  • derive expectations from the fixture or the UI (for example, read the ordered category buttons and compute the expected range) instead of hard-coding one neighbour;
  • run the mental-deletion check: if the feature were a no-op, would this assertion fail?

Goldens: one contract, one focused capture

All goldens go through fixtures/snapshots.ts; scripts/check-e2e-projects.mjs rejects toHaveScreenshot or toMatchSnapshot anywhere else.

  • expectCanvasSnapshot(editor, name) — the editor canvas stack, exact comparison, after the canvas has rendered.
  • expectPanelSnapshot(locator, name) — one panel, menu, toolbar, or dialog.
  • expectPageSnapshot(page, name) — full window; use sparingly, only for overall chrome.

Before every golden:

  1. Assert the semantic precondition that makes the image meaningful: the point count, selection, tool state, outline targets, handle states, zoom, active theme (data-color-theme), or devicePixelRatio.
  2. Park the pointer off the captured element unless hover is the contract.
  3. Capture the smallest element that shows the contract. Do not use a full-page golden to protect a toolbar button.

Rules:

  • One golden per distinct visual contract. Do not add near-duplicate images; if two captures are byte-identical, one of them is redundant.
  • Canvas goldens use maxDiffPixels: 0 with threshold: 0.02 on every host: enough for antialiasing differences between hosts (≤ 0.009), far below a real colour-token change (~0.07). Never raise either to make a mismatch pass; the default threshold of 0.2 silently accepts token changes.
  • Interface goldens (expectPanelSnapshot, expectPageSnapshot) are exact but compared on CI only, with baselines generated on the runner via the ci: update visual snapshots label. Never commit a locally generated interface baseline. Prefer semantic assertions and keep interface goldens few.
  • Goldens never pass on retry. The helpers throw on a retry attempt, so explain the first failure.
  • Use canvas-local positions (dragCanvas(), canvasPagePoint()) so layout changes cannot redraw geometry.
  • Theme goldens select the theme through the product and assert data-color-theme first. HiDPI goldens use test.use({ deviceScaleFactor: 2 }) and assert devicePixelRatio first. Add one golden per palette branch or rendering path, not one per theme.

Updating baselines: only for an intentional appearance change. Run the focused spec with --update-snapshots -g "<title>", open every changed PNG and describe what changed, then rerun without update mode and with --retries=0 --repeat-each=3. Canvas baselines may be generated locally. Interface baselines are generated only by the ci: update visual snapshots label; inspect the runner-generated images before merging.

Show full SKILL.md (614 more words)Show less

Fixtures and isolation

  • Use the shared editor fixture; construct EditorDriver directly only for additional windows.
  • Relaunch through the relaunch fixture, never raw electron.launch. The process registry owns every launched tree, attaches diagnostics, and kills it even when the test times out.
  • Every launch environment comes from shiftTestEnvironment(). Never spread process.env into a launch.
  • Script native dialogs with the fixture options (dirtyDocumentChoice(s), saveShiftPaths, openFontPath). An unexpected renderer dialog fails the test; opt in with allowRendererDialogs only when the dialog is the behavior.
  • Put shared setup in fixtures or EditorDriver (openScratchGlyph(), commitInputValue(), toolButton()), not copied helpers in specs. Do not grow EditorDriver with one-off helpers.
  • Prefer getByRole() / getByLabel() and locators from fixtures/appLocators.ts. No parent traversal (locator("..")), no styling classes, no toHaveCSS() unless the style value is the product contract.

Structure

  • One behavior per test. Split long flows when an early failure would hide unrelated coverage.
  • Do not accumulate state across unrelated steps; start each test from a fixture.
  • Setup may use page.evaluate or insertContent; the behavior under test must go through the user surface.
  • Name tests after the behavior: "shift-selects glyph categories as a visible range", not "category test".

Projects

Membership is explicit in apps/desktop/playwright.config.ts. Add a new spec to the right list — visual, platform, gpu, or perf — and run node scripts/check-e2e-projects.mjs. A spec that needs native lifecycle behavior on Windows and Linux belongs in platform. A spec whose only GPU dependency is incidental belongs in visual.

Large-data performance regressions

  • The perf project runs nightly or manually, not in the normal correctness suite. Protect a large-list UI change with a small visual interaction test that reaches an offscreen row, preserves selection and keyboard focus, and bounds the number of mounted rows.
  • In a real-app performance test, time the whole user-visible operation through confirmed edit settlement and canvas readiness with the relevant panels mounted. Timing only a local mutation, Rust apply, or IPC handler misses renderer main-thread stalls. Keep hardware-sensitive time ceilings in perf, not visual; verify the test fails for the original regression rather than only when it times out.
  • Use EditorDriver for navigation and edit settling. Keep the performance fixture responsible for the hardware launch and generated data, not editor behavior.

Proving a test is not flaky

Before committing a new or changed E2E test:

sh
pnpm test:e2e:visual e2e/<spec>.spec.ts -g "<title>" --repeat-each=10 --retries=0

For timing-sensitive flows, repeat under CPU load (for example, several yes > /dev/null & processes, killed afterwards), because CI runners are slower than development machines. For platform behavior, rely on the Windows and Linux CI jobs and say so in the pull request.

When a test fails or flakes:

  1. Read the trace and electron-diagnostics before changing anything.
  2. Classify it: product regression, test oracle, or infrastructure. Do not call a single failure flaky without a retry pass or a rerun on the same commit.
  3. Fix the oracle or the wait. Never add a sleep, a retry, or tolerance, and never regenerate a baseline without understanding the diff.

The CI "E2E Report" job lists tests that passed only on retry. Treat every entry as a bug to fix, not noise.

Review checklist

  • The layer is right; logic that TestEditor can observe is tested there.
  • No waitForTimeout, no fixed delays, and no polling without a semantic condition.
  • No colour counting, pixel sampling, byte inequality, or "something changed" oracle.
  • Assertions are exact and fail when the feature is a no-op.
  • Every golden goes through fixtures/snapshots.ts, captures one contract at the smallest element, and follows a semantic precondition.
  • Changed baselines were inspected and verified with --retries=0.
  • Launches, relaunches, and dialogs go through fixtures.
  • The spec belongs to the intended project, and the project check passes.
  • The test passed --repeat-each=10 --retries=0, and the pull request lists the exact E2E commands and any coverage not run.

© shift-editor, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .codex/skills/writing-e2e-tests of shift-editor/shift.

Open the folder on GitHubat commit e7dacfa

Compare with similar skills

Writing E2E Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Writing E2E Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Writing E2E Tests this skillshift-editor/shift343—~3.4kAutomated safety check: PassApache-2.0
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
playwright-cli Browser Automationgithub/gh-aw5.3k23 repos~2.8kAutomated safety check: PassMIT
Write and Verify Playwright Testsappsmithorg/appsmith41k—~2.9kAutomated safety check: NotesApache-2.0
Cucumber and Playwright E2E Testslanggenius/dify158k—~682Automated safety check: PassCustom licence
E2E Testinglangflow-ai/langflow156k—~3.3kAutomated safety check: PassMIT

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • Official

    Drives a real browser from the command line with playwright-cli to open pages, interact, mock requests, save state and work with Playwright tests.

    5.3k GitHub starsUsed in 23 repos~2.8k tokens
    Testing & QAAuto-check passed
  • Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.

    41k GitHub stars~2.9k tokensUpdated today
    Testing & QAAuto-check: notes
  • Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions.

    158k GitHub stars~682 tokensUpdated today
    Testing & QAAuto-check passed
  • E2E Testing

    langflow-ai/langflow

    Write and review Playwright E2E tests for Langflow. An agent skill from langflow-ai/langflow.

    156k GitHub stars~3.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Handsontable Playwright E2E Tests

    handsontable/handsontable

    Guides writing and changing Playwright end-to-end tests for Handsontable using page objects, data-testid hooks and deterministic waits.

    22k GitHub stars~1.7k tokensUpdated today
    Testing & QAAuto-check passed

More from shift-editor/shift

All 14 skills in this repo
  • Shift Commit Rules

    shift-editor/shift

    Rules for writing git commits in the Shift font editor repo: Conventional Commits subjects, user-facing changelog wording, concise subjects and logical commit boundaries.

    343 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check: notes
  • Dead Code Removal with Knip

    shift-editor/shift

    Finds unused files, exports and class members with Knip, then verifies each candidate through reference tracing before removing anything, never using knip --fix.

    343 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Shift Subsystem Docs

    shift-editor/shift

    Updates or creates DOCS.md files for Shift subsystems, recording the architecture invariants and constraints that cannot be learned from reading the source.

    343 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Adversarial Docs Audit

    shift-editor/shift

    Fact-checks DOCS.md files against the source code, testing each concrete claim and sorting it as true, false, stale or unverifiable.

    343 GitHub stars~818 tokensUpdated yesterday
    Auto-check passed
  • Shift Issue Writer

    shift-editor/shift

    Sets the rules for finding, writing and updating Shift GitHub issues: search for duplicates first, use outcome-focused titles and testable acceptance criteria.

    343 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Shift JSDoc Contracts

    shift-editor/shift

    Guides writing JSDoc for Shift exported APIs as a stable caller contract, covering ownership, lifetime, side effects and nullability that TypeScript types cannot express.

    343 GitHub stars~3.7k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Writing E2E Tests

What does Writing E2E Tests do?

Canonical rules for writing, rewriting, or reviewing Shift desktop Playwright E2E tests and visual goldens under apps/desktop/e2e/. Writing E2E Tests is an agent skill from shift-editor/shift. Canonical rules for writing, rewriting, or reviewing Shift desktop Playwright E2E tests and visual goldens under apps/desktop/e2e/.

When should I use Writing E2E Tests?

Writing E2E Tests fits situations like: change a .spec.ts; A screenshot baseline; playwright project membership; whenever you investigate a flaky.

How do I install Writing E2E Tests in Claude Code?

Run `npx skills add shift-editor/shift --skill writing-e2e-tests -a claude-code`. Or copy the skill folder (.codex/skills/writing-e2e-tests in shift-editor/shift) into .claude/skills/writing-e2e-tests in your project. Claude Code loads it when a task matches its description.

How do I install Writing E2E Tests in Codex?

Run `npx skills add shift-editor/shift --skill writing-e2e-tests -a codex`. Or copy the skill folder (.codex/skills/writing-e2e-tests in shift-editor/shift) into .agents/skills/writing-e2e-tests in your project. Codex loads it when a task matches its description.

Can I use Writing E2E Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add shift-editor/shift --skill writing-e2e-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/writing-e2e-tests, .gemini/skills/writing-e2e-tests, .github/skills/writing-e2e-tests and .opencode/skills/writing-e2e-tests in your project.

What does Writing E2E Tests need to run?

Going by SKILL.md and its folder, Writing E2E Tests needs the command-line tools its instructions call (pnpm and node).

Does Writing E2E Tests access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Writing E2E Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Writing E2E Tests use?

Writing E2E Tests is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Writing E2E Tests use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Writing E2E Tests?

Skills that share tags, products or a category with Writing E2E Tests: Web Application Testing (anthropics/skills, 180k stars), playwright-cli Browser Automation (github/gh-aw, 5.3k stars), Write and Verify Playwright Tests (appsmithorg/appsmith, 41k stars) and Cucumber and Playwright E2E Tests (langgenius/dify, 158k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Writing E2E Tests?

shift-editor (a GitHub organization) maintains it in shift-editor/shift, which has 343 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 6, 2026.

Source: shift-editor/shift on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.