Agent skill

Opik E2E Test Writer

by comet-ml in comet-ml/opik

Adds end-to-end Playwright tests to the Opik suite: scopes the feature, explores the live UI, writes a page object and spec, and runs it locally until green.

Apache-2.0Auto-check passedTesting & QA

Install Opik E2E Test Writer

skills CLI
$ npx skills add comet-ml/opik --skill writing-e2e-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install comet-ml/opik writing-e2e-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/comet-ml/opik.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/writing-e2e-tests .claude/skills/writing-e2e-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
writing-e2e-tests
GitHub stars
22k
Token cost
~2.8k tokens
SKILL.md length
1,243 words
Files
2
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Adds end-to-end Playwright tests to the Opik suite: scopes the feature, explores the live UI, writes a page object and spec, and runs it locally until green.

  • Works in 5 steps: Scope (gate, lightweight) → Analyze the feature and frontend code → Discover the live UI (gate, lightweight) → …
  • Adding an e2e test for a new Opik page or feature
  • SKILL.md covers Where tests live, Tooling — already set up, Conventions and The loop, plus 2 more sections
  • Calls docker, npx and python3

What it does

Built for the Opik repository. You name a feature, page or branch, and the skill runs a fixed loop in tests_end_to_end/e2e/: analyze the feature and frontend code, explore the running UI with the Playwright MCP, write a Page Object Model and a spec, and run it locally until it passes. For a branch or PR it reads the diff to find what changed.

The suite layout is documented: specs grouped by feature folder, page objects, fixtures that seed and tear down entities such as project, dataset, trace, experiment and test suite, SDK clients for seeding, and a FastAPI bridge that Playwright starts automatically. A conventions.md file sets rules like wrapping steps in test.step, asserting through the UI first and seeding only through the public SDK.

When your agent uses it

  • Adding an e2e test for a new Opik page or feature
  • Covering a branch's changes with a Playwright spec
  • Writing a Page Object Model for an existing page
  • Turning a manual UI flow into a regression test

Example prompts

  • “Add an e2e test for the experiments comparison page.”
  • “Write a test for the feature I just built in this branch.”
  • “Cover the dataset items flow with an e2e test.”

Requirements

  • A checkout of the Opik repository
  • Playwright, plus uv to run the SDK bridge

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Scope (gate, lightweight)
  2. Analyze the feature and frontend code
  3. Discover the live UI (gate, lightweight)
  4. Write the POM + spec
  5. Run until green

What it can do on your machine

Read from SKILL.md and the folder at commit f217a86. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • npx
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker and npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Opik E2E Test Writer loads about 2.8k tokens when it runs. Until then it costs about 126 tokens; SKILL.md has 1,243 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~126
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from comet-ml/opik at commit f217a86, republished under its Apache-2.0 licence (© comet-ml). 1,243 words, ~2,798 tokens.

Download SKILL.mdSave it as .claude/skills/writing-e2e-tests/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
writing-e2e-tests
description
Use when a developer wants to add, write, or create an end-to-end test for an Opik feature, page, or branch — e.g. "add an e2e test for the experiments comparison page", "write a test for the feature I just built", "e2e test for this branch", "cover the dataset items flow with a test". Runs the full loop in tests_end_to_end/e2e/ — analyze the feature and frontend code, explore the live UI with the Playwright MCP, write the Page Object Model + spec, and run it locally until green.

Writing E2E Tests

This skill is how we add an end-to-end test to the Opik E2E suite. You give it a feature, page, or branch; it runs a proven loop end-to-end and leaves you with a working, locally-verified Playwright test.

Announce at start: "I'm using the writing-e2e-tests skill to add an E2E test for X."

Where tests live

The suite is at tests_end_to_end/e2e/. Inside it:

  • Specs: tests/<feature>/<name>.spec.ts — one feature directory per page family (datasets, trace-explore, experiments, test-suites, online-evaluation, …).
  • Page Object Models: pom/<name>.page.ts — one class per page, methods for the interactions a test needs.
  • Fixtures: fixtures/<name>.fixture.ts — seed entities (project, dataset, trace, experiment, testSuite) and tear them down. Composed in a chain; re-exported from fixtures/index.ts.
  • SDK clients: core/sdk/ — sdkClient.python (HTTP wrapper over the bridge) and sdkClient.typescript (direct new Opik({...})) for seeding. core/backend/ holds the typed REST client for inspection + teardown.
  • Bridge: services/opik-sdk-driver/ — a FastAPI app (run with uv) wrapping the Python SDK, exposing routes the TS clients call. Playwright's webServer directive auto-spawns it during a test run; you don't start it by hand.

Specs and POMs import through path aliases: import { test, expect } from '@e2e/fixtures' and import { LogsPage } from '@e2e/pom/logs.page'.

Tooling — already set up

  • The Playwright MCP (live-UI exploration) and the playwright-test MCP (browser_generate_locator) are already configured in the repo's .mcp.json. No setup step.
  • Tests run via the plain Playwright CLI from tests_end_to_end/e2e/. The webServer directive spawns the bridge automatically.

Conventions

Read conventions.md before writing any POM or spec. It carries the rules that keep tests legible and stable: mandatory test.step() wrapping, UI-first assertions, selector preference, public-SDK-only seeding, fixture seed shapes, and the tag taxonomy. They aren't optional polish — each prevents a class of failure.

The loop

dot
digraph writing_e2e {
    rankdir=TB;
    "1. Scope (GATE)" [shape=box];
    "2. Analyze feature + FE code" [shape=box];
    "3. Discover live UI (GATE)" [shape=box];
    "4. Write POM + spec" [shape=box];
    "5. Run until green" [shape=box];
    "Green?" [shape=diamond];

    "1. Scope (GATE)" -> "2. Analyze feature + FE code";
    "2. Analyze feature + FE code" -> "3. Discover live UI (GATE)";
    "3. Discover live UI (GATE)" -> "4. Write POM + spec";
    "4. Write POM + spec" -> "5. Run until green";
    "5. Run until green" -> "Green?";
    "Green?" -> "4. Write POM + spec" [label="no — fix"];
    "Green?" -> "done" [label="yes"];
}
Step 1 — Scope (gate, lightweight)

Work out, from the request:

  • What flow / feature the test covers, and which page it lives on. If the dev pointed at a branch or PR, read the diff to find what changed.
  • The target. Default is local OSS at http://localhost:5173 (OPIK_DEPLOYMENT=oss, workspace default) — the natural target for "test the feature I just built." Only use another target if the dev asks.
  • Tags — pick a tier (@t1-smoke / @t2-cuj / @t3-nightly) and a feature tag, per conventions.md.

Run the safety check (below) before any seeding. Then confirm the scope in one short message — feature, page, target, tags — and proceed. Don't write a formal spec document.

Step 2 — Analyze the feature and frontend code

Before touching the browser:

  • Read the page's FE source under apps/opik-frontend/src/v2/pages/<Page>/ — the route it renders at, the components it composes, and any data-testid attributes already present. The route shape is what your POM's goto() will use.
  • Identify the entity preconditions: what must exist for the page to render real data (an empty project shows only the empty state). Decide how to seed it — which fixture fits, or which bridge/SDK call. Seed via the SDK/bridge, never by click-creating through the UI.
  • Check fixtures/ for an existing fixture that already seeds the shape you need; reuse it before writing a new one.
Step 3 — Discover the live UI (gate, lightweight)

Invoke the playwright-pom-discovery skill (via the Skill tool). It walks the live page with the Playwright MCP: seed state, navigate authed, snapshot the accessibility tree, enumerate data-testids, pick the most stable selector for each element you'll target, and flag any element that has no stable selector (needs a FE data-testid added in this change).

When discovery is done, report a short summary — the selectors you'll use per element, and any missing testids you'll add — and confirm before writing code. Don't write anything under pom/ before this step.

Step 4 — Write the POM + spec
  • Write or extend the POM in pom/<name>.page.ts using the selectors from discovery. Each method wraps its body in test.step() and returns through the callback (see conventions.md).
  • Write the spec in tests/<feature>/<name>.spec.ts: tier + @area: on the describe block, a @cap: per test, coarse test.step() phases, UI-first assertions.
  • Take the @area:/@cap: values from tests_end_to_end/coverage/taxonomy.yaml — never invent them. Grep it for the feature first; the capability you're covering is usually already reserved as covered: false. Update the taxonomy in the same change: add the spec to the area's specs: list and flip each covered capability to covered: true with its tier. See the Tags section of conventions.md.
  • If discovery flagged a missing/brittle selector, add a descriptive data-testid to the FE component in the same change.
Show full SKILL.md (529 more words)Show less
Rebuilding the FE after adding a data-testid

The local OSS deployment serves the frontend from a Docker image — file changes to apps/opik-frontend/ are not picked up automatically. After adding a data-testid, you must rebuild and restart the container before the test can find it.

From deployment/docker-compose/:

bash
# 1. Build a new image from the updated source
docker compose --profile opik build frontend

# 2. Recreate the container using the locally built image
#    (pull_policy defaults to "always" — override it so Docker uses the local build)
docker stop opik-frontend-1 && docker rm opik-frontend-1
OPIK_FRONTEND_PULL_POLICY=never docker compose --profile opik up -d --no-deps frontend

Verify the new data-testid is live before running the test:

bash
docker exec opik-frontend-1 sh -c 'grep -r "your-testid" /usr/share/nginx/html/ | wc -l'
# should print a non-zero number

Network note: if the rebuilt container loses connectivity to the backend (502 errors), the container may have ended up on the wrong Docker network. Fix it:

bash
docker network disconnect opik-opik_default opik-frontend-1
docker network connect opik-opik_default opik-frontend-1
Step 5 — Run until green

From tests_end_to_end/e2e/:

bash
npx playwright test tests/<feature>/<name>.spec.ts --reporter=list

The bridge auto-spawns (you'll see its startup line in the output). If a test fails, read the failure trace (npx playwright show-trace) rather than adjusting selectors blindly — see "verify the test render before blaming the backend" in conventions.md. Fix and re-run until green. Report the actual run output.

Three checks before you call it done — each catches something the single-spec run can't:

bash
# 1. Tags resolve in the taxonomy (this is the CI `tag-lint` job — run it locally, it's instant)
python3 tests_end_to_end/coverage/tag_lint.py --taxonomy tests_end_to_end/coverage/taxonomy.yaml --estate tests_end_to_end

# 2. If you touched a shared POM, the whole feature directory still passes
npx playwright test tests/<feature>/ --reporter=list

# 3. Types
npx tsc --noEmit

"My spec passes" is not "I didn't break anything": a shared POM is used by sibling specs, and tag-lint failures never surface in a Playwright run at all.

Safety: verify local config before seeding

The Python SDK behind the bridge reads ~/.opik.config. If it points at a cloud environment, seeding would create real data there. Before any seed against a local target:

bash
cat ~/.opik.config

If url_override is anything other than http://localhost:5173/api, back it up and point it local:

bash
cp ~/.opik.config ~/.opik.config.bak 2>/dev/null || true
cat > ~/.opik.config << 'EOF'
[opik]
url_override = http://localhost:5173/api
workspace = default
EOF

When the work is done, remind the dev to restore: cp ~/.opik.config.bak ~/.opik.config. If it already points local, skip this.

Anti-patterns

SymptomWhat you skipped
"Let me read the FE source to find the selector"Discovery — snapshot the rendered DOM. What renders is the only source of truth for selectors.
"I'll explore the empty page and figure out the rows later"Seeding — an empty-state-only POM never exercises the row template or open-detail actions.
"I'll write the POM and find out if it works when the whole suite runs"Run-until-green in isolation — iterate on the one spec, don't debug it inside a full suite run.
"page.locator('tbody tr:nth-child(3)') is fine"Flagging the missing testid — brittle structural selectors are the top source of flake; add a data-testid.
"I'll create the dataset through the UI so the page has data"SDK/bridge seeding — UI-create is what the test exercises, not how you set up.
"@cap:traces.bulk-delete-traces describes what my test does"Checking the taxonomy — @area:/@cap: values are a fixed vocabulary in taxonomy.yaml, not free text. CI's tag-lint rejects invented names, and the capability you want is usually already reserved there.
"I'll clean up in a try/finally at the end of the test"Fixture teardown — it already runs on pass, fail and timeout, and keeps cleanup out of the assertions (where it also escapes test.step() and never reaches the trace). Add a fixture, or a register-callback one when the id only appears mid-test.
"Adding a fixture would touch files outside my ticket"Nothing — fixtures/ and core/backend/client.ts are where seeding and teardown belong. Reuse first, but adding one is normal, not scope creep.
"The spec passes, so I'm done"tag_lint.py, the feature-directory run, and tsc — a green single spec hides tag errors and sibling breakage from a shared POM.

© comet-ml, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/writing-e2e-tests of comet-ml/opik.

  • SKILL.md
  • conventions.md

Open the folder on GitHubat commit f217a86

Compare with similar skills

Opik E2E Test Writer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Opik E2E Test Writer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Opik E2E Test Writer this skillcomet-ml/opik22k—~2.8kAutomated safety check: PassApache-2.0
Playwright E2E Testsonyx-dot-app/onyx32k1 repos~2.8kAutomated safety check: NotesCustom licence
E2E VerificationChorus-AIDLC/Chorus1.2k—~1.5kAutomated safety check: NotesAGPL-3.0
Playwright Testingchongdashu/vibejam-starter-pack149—~2.2kAutomated safety check: PassNone
Frontend Playwright E2Eansible/ansible-ui113—~2.5kAutomated safety check: NotesApache-2.0
Dev Webtestclassmethod/tsumiki974—~4.2kAutomated safety check: PassMIT

Similar skills

  • Playwright E2E Tests

    onyx-dot-app/onyx

    Write and maintain Playwright end-to-end tests for the Onyx application.

    32k GitHub starsUsed in 1 repo~2.8k tokens
    Testing & QAAuto-check: notes
  • E2E Verification

    Chorus-AIDLC/Chorus

    A skill your agent uses when manually verifying a Chorus frontend change in a real browser — finding local login credentials, driving the running dev server with the Playwright MCP, logging in…

    1.2k GitHub stars~1.5k tokensUpdated today
    Testing & QAAuto-check: notes
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.2k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Frontend Playwright E2E

    ansible/ansible-ui

    Write, run, and debug Playwright E2E / integration / live tests.

    113 GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check: notes
  • Dev Webtest

    classmethod/tsumiki

    This skill should be used when the user asks to "dev-webtest", "Webテスト", "画面の動作確認", "E2Eテスト", "web test", "visual check", "モンキーテスト", "アクセシビリティチェック", "レスポンシブテスト", "フォームテスト".

    974 GitHub stars~4.2k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Agentic Browser Testing

    petrkindlmann/qa-skills

    Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.

    163 GitHub stars~4.5k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed

More from comet-ml/opik

All 19 skills in this repo
  • Checklist for wiring a new linter into Opik's Code Quality pipeline: the four files to edit, the silent-failure gotchas and the pass/fail verification loop.

    22k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Shows how to add product analytics events to Opik's frontend, Java backend and Python SDK, all reporting through Segment to PostHog with an opik_ name prefix.

    22k GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests.

    22k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Rules for writing PR descriptions, changelog entries and feature documentation in the Opik repository, including the exact headings that CI requires.

    22k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Turns a code change into one committed, passing Playwright end-to-end spec by resolving the change scope and handing authoring to a companion skill.

    22k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Starts, rebuilds, and troubleshoots the Opik local dev stack, including an optional Comet Platform integration mode for the Opik team.

    22k GitHub stars~734 tokensUpdated today
    Auto-check passed

Categories

Questions about Opik E2E Test Writer

What does Opik E2E Test Writer do?

Adds end-to-end Playwright tests to the Opik suite: scopes the feature, explores the live UI, writes a page object and spec, and runs it locally until green. Built for the Opik repository. You name a feature, page or branch, and the skill runs a fixed loop in tests_end_to_end/e2e/: analyze the feature and frontend code, explore the running UI with the Playwright MCP, write a Page Object Model and a spec, and run it locally until it passes.

When should I use Opik E2E Test Writer?

Opik E2E Test Writer fits situations like: adding an e2e test for a new Opik page or feature; covering a branch's changes with a Playwright spec; writing a Page Object Model for an existing page; turning a manual UI flow into a regression test.

How do I install Opik E2E Test Writer in Claude Code?

Run `npx skills add comet-ml/opik --skill writing-e2e-tests -a claude-code`. Or copy the skill folder (.agents/skills/writing-e2e-tests in comet-ml/opik) into .claude/skills/writing-e2e-tests in your project. Claude Code loads it when a task matches its description.

How do I install Opik E2E Test Writer in Codex?

Run `npx skills add comet-ml/opik --skill writing-e2e-tests -a codex`. Or copy the skill folder (.agents/skills/writing-e2e-tests in comet-ml/opik) into .agents/skills/writing-e2e-tests in your project. Codex loads it when a task matches its description.

Can I use Opik E2E Test Writer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add comet-ml/opik --skill writing-e2e-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/writing-e2e-tests, .gemini/skills/writing-e2e-tests, .github/skills/writing-e2e-tests and .opencode/skills/writing-e2e-tests in your project.

What does Opik E2E Test Writer need to run?

Going by SKILL.md and its folder, Opik E2E Test Writer needs the command-line tools its instructions call (docker, npx and python3). Our summary lists: A checkout of the Opik repository; Playwright, plus uv to run the SDK bridge.

Does Opik E2E Test Writer access the network?

SKILL.md contains no URLs. Its commands use docker and npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Opik E2E Test Writer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Opik E2E Test Writer use?

Opik E2E Test Writer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Opik E2E Test Writer use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Opik E2E Test Writer?

Skills that share tags, products or a category with Opik E2E Test Writer: Playwright E2E Tests (onyx-dot-app/onyx, 32k stars), E2E Verification (Chorus-AIDLC/Chorus, 1.2k stars), Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars) and Frontend Playwright E2E (ansible/ansible-ui, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Opik E2E Test Writer?

comet-ml (a GitHub organization) maintains it in comet-ml/opik, which has 22,412 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 7, 2026.

Source: comet-ml/opik on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.