Agent skill

Explore Feature E2E Test

by comet-ml in comet-ml/opik

Turns a code change into one committed, passing Playwright end-to-end spec by resolving the change scope and handing authoring to a companion skill.

Apache-2.0Auto-check passedTesting & QA

Install Explore Feature E2E Test

skills CLI
$ npx skills add comet-ml/opik --skill explore-feature -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install comet-ml/opik explore-feature --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/comet-ml/opik.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/explore-feature .claude/skills/explore-feature && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
explore-feature
GitHub stars
22k
Token cost
~3.4k tokens
SKILL.md length
1,762 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Turns a code change into one committed, passing Playwright end-to-end spec by resolving the change scope and handing authoring to a companion skill.

  • Works in 4 steps: Resolve scope (gate) → Local-run gate → Delegate authoring to writing-e2e-tests → …
  • Adding an end-to-end test for a change you just made
  • SKILL.md covers What this does — and doesn't, The loop, Phase 1 — Resolve scope (gate) and Phase 2 — Local-run gate, plus 3 more sections
  • Calls git, npx and npm

What it does

This is a thin orchestrator with two phases of its own: resolve what to cover, and confirm the result. Authoring is delegated to the `writing-e2e-tests` skill, which analyzes the front end, discovers the live UI, writes the page objects and spec, and runs it green. The goal is a cheap, per-PR happy-path test of a single flow that passes locally within minutes, not deep bug-hunting or CI wiring.

In the first phase the scope is normalized into a ScopeSpec from a local diff, a PR number or URL read through `gh pr view`, or several PRs combined. Local-diff mode captures committed, unstaged and other working-tree changes against the merge base, and stops if nothing changed. The output is an ordinary spec under `tests/` with a tier tag, an area tag and a capability tag per test, picked up by the same suites as every other spec.

When your agent uses it

  • Adding an end-to-end test for a change you just made
  • Covering the main flow of a pull request with one Playwright spec
  • Writing a test for uncommitted work before opening a PR

Example prompts

  • “Explore this feature and add an e2e test for my branch.”
  • “Add a test for my PR and make sure it passes locally.”
  • “Cover the feature in the PR I linked with a Playwright spec.”

Requirements

  • A Playwright test setup with the `writing-e2e-tests` skill available
  • The GitHub CLI or GitHub MCP when scoping from a PR

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Resolve scope (gate)
  2. Local-run gate
  3. Delegate authoring to writing-e2e-tests
  4. Confirm

What it can do on your machine

Read from SKILL.md and the folder at commit f217a86. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • npx
    • npm
    • python3
    • gh
    • docker
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, npx, npm, gh and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Explore Feature E2E Test loads about 3.4k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 1,762 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from comet-ml/opik at commit f217a86, republished under its Apache-2.0 licence (© comet-ml). 1,762 words, ~3,350 tokens.

Download SKILL.mdSave it as .claude/skills/explore-feature/SKILL.md (or your agent's skills folder).
name
explore-feature
description
Use when a developer wants an e2e test covering a change they just made — e.g. "explore this feature", "add a test for my PR", "cover the feature in PR

Explore Feature

Turns "here's my change, cover it" into one committed, locally-green Playwright spec.

It is a thin orchestrator: it owns two phases — resolve what to cover, and confirm the result — and delegates the actual authoring (analyze FE, discover live UI, write POM/spec, run green) to the writing-e2e-tests skill.

Announce at start: "I'm using the explore-feature skill to add an e2e test for X."

What this does — and doesn't

  • Does: resolve the change surface → pick the one flow worth covering → hand it to writing-e2e-tests → confirm the spec is tagged, in the taxonomy, and green.
  • Doesn't: deep bug-hunting or edge-case coverage (that's a separate QA activity); CI wiring.
  • Cheap per-PR happy-path only. One flow, green locally in minutes.

The output is an ordinary spec: tests/<area>/<name>.spec.ts, a tier tag, an @area:, a @cap: per test — exactly what every other spec in the estate looks like, and picked up by the same suites. There is no separate lane and no version stamp.

The loop

dot
digraph explore_feature {
    rankdir=TB;
    "1. Resolve scope (GATE)" [shape=box];
    "2. Local-run gate (GATE)" [shape=box];
    "3. Delegate authoring to writing-e2e-tests" [shape=box];
    "4. Confirm tagged + green" [shape=box];

    "1. Resolve scope (GATE)" -> "2. Local-run gate (GATE)";
    "2. Local-run gate (GATE)" -> "3. Delegate authoring to writing-e2e-tests";
    "3. Delegate authoring to writing-e2e-tests" -> "4. Confirm tagged + green";
}

Phase 1 — Resolve scope (gate)

Normalize whatever the dev pointed at into one ScopeSpec before authoring anything.

Input modes (auto-detect from the argument; ask if ambiguous):

ModeTriggerResolve
Local diffno arg / dirty tree / "my branch" / "my changes"see Local-diff mode below — this is the default when no PR is named, incl. "I haven't opened a PR yet"
A PR#<n> or a PR URLPR diff + linked ticket via gh pr view <n> --json … (or the GitHub MCP)
Multi-PR / multi-ticketa listunion of the diffs

Local-diff mode — a dev running this on their branch before (or without) a PR. Capture the full change surface, not just committed work — pre-PR work is often uncommitted:

bash
base=$(git merge-base origin/main HEAD)
git diff --name-only "$base"...HEAD          # committed on the branch
git diff --name-only HEAD                     # unstaged working-tree changes
git diff --name-only --cached                 # staged but uncommitted
git ls-files --others --exclude-standard      # new untracked files

Union those for changedFiles. If all four are empty, there's nothing to cover — say so and stop.

Produce the ScopeSpec:

  • tickets[] — for naming and the PR description, if there are any. A ticket key in the branch name (andreic/OPIK-1234-… → OPIK-1234) is the usual source. Never invent one; the spec name should describe the behaviour anyway, not the ticket.
  • changedFiles[] — the FE/BE change surface.
  • area — the taxonomy area the change belongs to, from tests_end_to_end/coverage/taxonomy.yaml. This decides the directory: tests/<area>/.
  • targetPath = tests_end_to_end/e2e/tests/<area>/<name>.spec.ts — new, or an existing spec in that directory to extend. Prefer extending: a new test() in the area's existing spec beats a new file when the setup is the same.
  • capabilities[] — the @cap: keys this will cover, grepped from the taxonomy. They usually already exist as covered: false. If nothing fits, add the entry — a kebab-case name for the user-facing capability.
  • tier — @t1-smoke for fast deterministic core checks, @t2-cuj for multi-step journeys and anything destructive, @t3-nightly for slower/broader. Tier is chosen by how often it should run, not by importance; anything spending real LLM budget is not t1.
  • happyPath — the one end-to-end flow to cover. Multi-PR → the combined assembled-feature flow, as one test.

Three things to resolve while shaping the happy path — each caught a false or unbuildable test in piloting:

  • Fix PRs — cover the repro, not the easy path. For a fix:, the happy path must exercise the exact condition the bug needed. If the state can be reached two ways and only one triggered the bug (e.g. a trace shows the bug only when source=sdk via manual reference-linking, not via evaluate()), seeding the easy way makes the test pass against the pre-fix code too — a vacuous test. Identify the repro condition from the PR's root-cause description and seed that shape.
  • N equivalent surfaces — cover the most representative one. If the change fixes the same behavior in several places (e.g. experiment "Go to logs", a shared sidebar, and a Playground cell link), don't try to cover them all — pick the single most representative entry point and tell the dev which ones you left. Keeps it cheap.
  • Seeding is part of the deliverable, not a precondition. Work out early how the happy path's state gets created — an existing fixture, an SDK client, or the bridge (services/opik-sdk-driver). If the shape the repro needs isn't reachable through the current surface (e.g. the bridge only exposes evaluate() but the bug needs a manual client.trace(source=...) + ExperimentItemReferences shape), that seeding support is yours to add as part of authoring — extend the bridge route / add a fixture / use the SDK client directly — then write the test on top of it. Adding a fixture is normal, not scope creep. The only real stop is if the state cannot be produced through any public SDK / bridgeable path at all (rare) — then flag it, because it likely means the feature isn't end-to-end testable yet.

Two gates here, before expensive authoring:

  1. Scope gate — state the resolved happy path + repro seed shape + target path + tier + the @cap: keys back to the dev and get a yes. If seeding the repro needs new bridge/fixture support, say so here so the dev knows this work also touches services/opik-sdk-driver or the fixtures. Multi-PR especially: "One test in <area>/<name>.spec.ts covering X→Y→Z, @t2-cuj. OK?"
  2. Skip check — if the change is pure refactor / infra / docs with no user-facing behavior, say so and stop. Note two cases that are user-facing even though they look like config: a capability-map / constants change that adds a user-visible option (e.g. a new model in a dropdown → happy path: "open the page, the option is selectable"), and a backend-dominant change whose only visible effect is subtle (e.g. a trace that should not appear in a default list) — find the user-observable effect and cover that, don't skip. The opposite case also happens: a perf / internals change with no behavior delta by design (e.g. swapping a slow probe for a fast one, same rendered result). There's no PR-specific happy path — so either cover a generic regression on the affected page and say so, or skip-with-a-note if the suite already covers that page. Don't dress a generic regression up as PR-specific coverage. Two things to get right here: (a) skip vs cover turns on coverage of the specific state/decision the change governs, not the page as a whole — grep the existing suite (tests_end_to_end/e2e/tests) for that exact state. A page whose populated path is covered but whose empty/onboarding path (the branch a probe like this actually drives) is not is not "already covered" — cover the uncovered half. Skip-with-a-note only when the specific state is genuinely already asserted somewhere. (b) This is a fix: PR, so the "repro condition" mandate seems to apply — but a no-behavior-delta fix has no repro that renders differently pre/post. The perf-fix escape hatch overrides the repro mandate: say "N/A — no behavior delta; generic regression" and label it generic. A generic test that passes on both the pre- and post-fix build is correct, not a bug — say so when you report back.
Show full SKILL.md (673 more words)Show less

Phase 2 — Local-run gate

Before authoring can be verified, confirm the dev has a local stack with their changes:

  1. Probe for a running stack — the frontend (http://localhost:5173, or :5174 for FE-from-source) and the backend, the way the suite reaches it: GET <baseUrl>/api/is-alive/ver (e.g. http://localhost:5173/api/is-alive/ver). A standard opik.sh compose stack does not expose the backend on a bare :8080 — the FE proxies /api to it, and that proxied path returning a {"version": …} JSON is the real "backend is up" signal. A FE that answers on / but 000s on /api/is-alive/ver is a half-up stack: every seeded test fails on the first API call for env reasons, not the feature. Require the /api health check to pass, not just "/ answers on 5173." Note the returned version — it tells you which build is running (see step 2).

  2. Gate the dev: confirm the running stack actually contains their changes. A stale prebuilt opik.sh stack won't show new data-testids — if the feature adds testids, the dev must be on FE-from-source :5174 (dev-runner --restart). Verify the change is actually in the served build, not just that a FE answers: for a FE-only PR merged to main, curl http://localhost:5174/src/<changed-file> and grep for a symbol the PR added (Vite serves source), and/or check git merge-base --is-ancestor <merge-sha> HEAD. "The dev server is up" is not "the fix is present."

  3. If nothing is running / it's the wrong stack: offer to spin it (local-dev / dev-runner) or ask the dev to bring it up with their changes, then proceed. Never silently run against a stack lacking the feature — that produces false-green or false-missing-testid results.

    Worktree gotcha (FE-from-source against a prebuilt backend). dev-runner.sh is worktree-aware: in a worktree it offsets every port from a per-worktree hash and starts its own JAR-mode backend against a fresh, empty DB — so --restart there does not reuse the healthy opik.sh docker DB, and FE-from-source may not land on :5174. When you need FE-with-the-fix on top of an existing seeded docker backend, the reliable path is to run the Vite dev server directly with pinned ports and point its /api proxy at the running backend:

    • The opik.sh backend container usually publishes only its internal port to a random host port (docker port opik-<proj>-backend-1), not :8080. Vite's /api proxy strips /api and needs a bare backend, so bridge the container's app port to host :8080 on the compose network, e.g. docker run -d --name opik-be-8080 --network <compose_net> -p 8080:8080 alpine/socat tcp-listen:8080,fork,reuseaddr tcp-connect:opik-<proj>-backend-1:8080.
    • Then cd apps/opik-frontend && npm ci && VITE_DEV_PORT=5174 VITE_BACKEND_PORT=8080 npm run start.
    • Point the suite at it: OPIK_BASE_URL=http://localhost:5174 OPIK_DEPLOYMENT=oss. OSS needs no auth. Tear down the socat container when done.

Phase 3 — Delegate authoring to writing-e2e-tests

Invoke the writing-e2e-tests skill to do the analyze → discover-live-UI → write → run-green loop. Hand it the ScopeSpec: the target path, the tier + @area: + @cap: tags, the happy path and its seed shape, and "verify green against the dev's local stack."

That skill owns the conventions — test.step() wrapping, UI-first assertions, selector preference, SDK-only seeding, fixture-owned teardown, and the taxonomy update. Don't restate them here; read .agents/skills/writing-e2e-tests/conventions.md if you need them.

Phase 4 — Confirm

  • The spec exists at targetPath, tagged with one tier + @area:<area>, with a @cap: per test.
  • Those @area:/@cap: values resolve in tests_end_to_end/coverage/taxonomy.yaml, and the taxonomy was updated in the same change (spec added to specs:, covered capabilities flipped to covered: true with the tier).
  • tag_lint.py reports 0 problem(s) — this is the CI tag-lint job, so a miss here is a red build:
    bash
    python3 tests_end_to_end/coverage/tag_lint.py --taxonomy tests_end_to_end/coverage/taxonomy.yaml --estate tests_end_to_end
  • It runs green locally, and so does the rest of its feature directory — a shared POM or fixture is used by sibling specs:
    bash
    cd tests_end_to_end/e2e && npx playwright test tests/<area>/ --reporter=list
    npx tsc --noEmit
  • Report the committed spec path back to the dev, plus anything you deliberately left uncovered (the other N surfaces, edge cases) so they know the boundary.

Ownership

QA owns this skill. When a generated test misses something, the fix lands in this skill's files — this is the feedback loop. Edit in .agents/skills/explore-feature/, then make claude to mirror for local testing.

© comet-ml, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/explore-feature of comet-ml/opik.

Open the folder on GitHubat commit f217a86

Compare with similar skills

Explore Feature E2E Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Explore Feature E2E Test compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Explore Feature E2E Test this skillcomet-ml/opik22k—~3.4kAutomated safety check: PassApache-2.0
Record E2E Giflablup/backend.ai-webui133—~907Automated safety check: NotesLGPL-3.0
Write and Verify Playwright Testsappsmithorg/appsmith41k—~2.9kAutomated safety check: NotesApache-2.0
Cherry Studio Regression TestsCherryHQ/cherry-studio52k—~1.2kAutomated safety check: PassAGPL-3.0
Kane CLI Browser TestingLambdaTest/kane-cli247—~8.4kAutomated safety check: PassApache-2.0
Dev Webtest Planclassmethod/tsumiki974—~4.2kAutomated safety check: PassMIT

Similar skills

  • Record E2E Gif

    lablup/backend.ai-webui

    Record Playwright e2e tests as one GIF per test case (video → ffmpeg palette GIF) and return a markdown table for a PR description.

    133 GitHub stars~907 tokensUpdated today
    Testing & QAAuto-check: notes
  • Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.

    41k GitHub stars~2.9k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • Cherry Studio Regression Tests

    CherryHQ/cherry-studio

    Runs Cherry Studio's critical-path regression suite as deterministic Playwright E2E tests through a GitHub workflow on macOS and Windows runners.

    52k GitHub stars~1.2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Kane CLI Browser Testing

    LambdaTest/kane-cli

    Drives a real browser through the kane-cli tool and designs requirement-linked test suites from a PRD or a plain description, with mobile and cloud-grid runs.

    247 GitHub stars~8.4k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Dev Webtest Plan

    classmethod/tsumiki

    This skill should be used when the user asks to "dev-webtest-plan", "Webテスト計画を生成", "テスト計画を作成", "webtest plan", "E2Eテスト計画", "画面テスト計画", "generate webtest plan", "create test plan from requirements"…

    974 GitHub stars~4.2k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Official

    Writes UI tests that reproduce a GitHub issue in .NET MAUI and keeps iterating until the tests actually fail, proving they catch the bug.

    23k GitHub stars~3k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from comet-ml/opik

All 19 skills in this repo
  • Checklist for wiring a new linter into Opik's Code Quality pipeline: the four files to edit, the silent-failure gotchas and the pass/fail verification loop.

    22k GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Shows how to add product analytics events to Opik's frontend, Java backend and Python SDK, all reporting through Segment to PostHog with an opik_ name prefix.

    22k GitHub stars~4.4k tokensUpdated yesterday
    Auto-check passed
  • Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests.

    22k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Rules for writing PR descriptions, changelog entries and feature documentation in the Opik repository, including the exact headings that CI requires.

    22k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Starts, rebuilds, and troubleshoots the Opik local dev stack, including an optional Comet Platform integration mode for the Opik team.

    22k GitHub stars~734 tokensUpdated yesterday
    Auto-check passed
  • Specifies how to instrument an opik-backend pipeline with per-stage OpenTelemetry metrics for throughput, latency, errors and queue delay by workspace.

    22k GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Explore Feature E2E Test

What does Explore Feature E2E Test do?

Turns a code change into one committed, passing Playwright end-to-end spec by resolving the change scope and handing authoring to a companion skill. This is a thin orchestrator with two phases of its own: resolve what to cover, and confirm the result. Authoring is delegated to the `writing-e2e-tests` skill, which analyzes the front end, discovers the live UI, writes the page objects and spec, and runs it green.

When should I use Explore Feature E2E Test?

Explore Feature E2E Test fits situations like: adding an end-to-end test for a change you just made; covering the main flow of a pull request with one Playwright spec; writing a test for uncommitted work before opening a PR.

How do I install Explore Feature E2E Test in Claude Code?

Run `npx skills add comet-ml/opik --skill explore-feature -a claude-code`. Or copy the skill folder (.agents/skills/explore-feature in comet-ml/opik) into .claude/skills/explore-feature in your project. Claude Code loads it when a task matches its description.

How do I install Explore Feature E2E Test in Codex?

Run `npx skills add comet-ml/opik --skill explore-feature -a codex`. Or copy the skill folder (.agents/skills/explore-feature in comet-ml/opik) into .agents/skills/explore-feature in your project. Codex loads it when a task matches its description.

Can I use Explore Feature E2E Test in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add comet-ml/opik --skill explore-feature -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/explore-feature, .gemini/skills/explore-feature, .github/skills/explore-feature and .opencode/skills/explore-feature in your project.

What does Explore Feature E2E Test need to run?

Going by SKILL.md and its folder, Explore Feature E2E Test needs the command-line tools its instructions call (git, npx, npm, python3, gh and docker). Our summary lists: A Playwright test setup with the `writing-e2e-tests` skill available; The GitHub CLI or GitHub MCP when scoping from a PR.

Does Explore Feature E2E Test access the network?

SKILL.md contains no URLs. Its commands use git, npx, npm, gh and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Explore Feature E2E Test safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Explore Feature E2E Test use?

Explore Feature E2E Test is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Explore Feature E2E Test use?

About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Explore Feature E2E Test?

Skills that share tags, products or a category with Explore Feature E2E Test: Record E2E Gif (lablup/backend.ai-webui, 133 stars), Write and Verify Playwright Tests (appsmithorg/appsmith, 41k stars), Cherry Studio Regression Tests (CherryHQ/cherry-studio, 52k stars) and Kane CLI Browser Testing (LambdaTest/kane-cli, 247 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Explore Feature E2E Test?

comet-ml (a GitHub organization) maintains it in comet-ml/opik, which has 22,412 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 7, 2026.

Source: comet-ml/opik on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.