Agent skill

Debugging Opik E2E Tests

by comet-ml in comet-ml/opik

Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests.

Apache-2.0Auto-check passedTesting & QA

Install Debugging Opik E2E Tests

skills CLI
$ npx skills add comet-ml/opik --skill debugging-e2e-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install comet-ml/opik debugging-e2e-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/comet-ml/opik.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/debugging-e2e-tests .claude/skills/debugging-e2e-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debugging-e2e-tests
GitHub stars
22k
Token cost
~1.8k tokens
SKILL.md length
855 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests.

  • Works in 5 steps: Resolve the entry point → Gather evidence → Classify → …
  • A pull request check shows an Opik E2E test failing and you need to know why
  • SKILL.md covers What this does — and doesn't, Where the evidence lives, Tooling and The loop, plus 1 more section
  • Calls gh and npx

What it does

Give the agent a failure from a CI check, an Allure TestOps launch, a test name or a local run and it collects the evidence first: Playwright traces and videos, the HTML report and Allure results, taken from CI artifacts or the local test-results folder. It uses the allure-testops MCP to find the launch, list test results with their flaky flags and pull each test's pass and fail history, and uses gh to find the failed job and download artifacts.

The trace is opened with Playwright's show-trace to see the failing step, DOM snapshot, console and network at that moment. From the history and the trace the agent classifies the failure as a real regression or a flake and proposes a fix with the evidence cited. It is read-only: it does not edit tests or rerun the suite, and applying the fix is handed to the writing-e2e-tests skill or a separate request.

When your agent uses it

  • A pull request check shows an Opik E2E test failing and you need to know why
  • Deciding whether a test such as dataset-crud-smoke is flaky or genuinely broken
  • The nightly E2E suite turned red and you want the first failing test explained
  • Reading Playwright traces and Allure history for a failed run

Example prompts

  • “Why did dataset-crud-smoke fail on my pull request's CI run?”
  • “Is dataset-crud-smoke flaky? Check its history in TestOps.”
  • “The nightly e2e suite went red; find the first real regression and propose a fix.”

Requirements

  • The allure-testops MCP server
  • GitHub CLI (gh)
  • Playwright, to open traces

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Resolve the entry point
  2. Gather evidence
  3. Classify
  4. Diagnose
  5. Report + propose (no edits)

What it can do on your machine

Read from SKILL.md and the folder at commit f217a86. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh and npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debugging Opik E2E Tests loads about 1.8k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 855 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from comet-ml/opik at commit f217a86, republished under its Apache-2.0 licence (© comet-ml). 855 words, ~1,795 tokens.

Download SKILL.mdSave it as .claude/skills/debugging-e2e-tests/SKILL.md (or your agent's skills folder).
name
debugging-e2e-tests
description
Use when an Opik E2E test has failed and a developer wants it investigated — e.g. "why did this e2e test fail?", "investigate the failing run on my PR", "is dataset-crud-smoke flaky?", "the nightly e2e suite went red". Takes a failure from a CI check, a TestOps launch, a test name, or a local run; gathers the trace and history, classifies regression vs. flake, and proposes a fix. Read-only — it diagnoses and proposes, it does not edit tests.

Debugging E2E Tests

This skill investigates a failed test in the Opik E2E suite (tests_end_to_end/e2e/). You give it a failure from wherever you noticed it; it gathers the evidence, decides whether it's a real regression or a flake, and proposes a fix.

Announce at start: "I'm using the debugging-e2e-tests skill to investigate X."

What this does — and doesn't

  • It diagnoses and proposes, grounded in cited evidence (the trace, the error, the history). It is read-only: it does not edit tests, and it does not re-run the suite as part of investigating.
  • To apply a proposed fix, hand off to the writing-e2e-tests skill (or just say "apply it") — that's a separate, deliberate act with its own run-until-green loop.

Where the evidence lives

  • Local run — traces under tests_end_to_end/e2e/test-results/ (retained on failure), Allure results under allure-results/.
  • CI run — the suite uploads three artifacts per run (7-day retention): test-results-v2 (Playwright traces + videos), playwright-report-v2 (the HTML report), allure-results-v2. Download with gh run download <run-id> -n test-results-v2 -D <dir>.
  • Allure TestOps (comet.testops.cloud, project id 1) — results stream live during CI. Launches are named Opik v2 … <tier> - <run_id> (the trailing number is the GitHub Actions run id; the env segment varies — E2E, Post-Merge, Local, staging, production).

Tooling

  • allure-testops MCP (already connected) — the richest source. Validated calls:
    • list_launches(projectId: 1, search: "<run_id or name fragment>", sort: ["createdDate,DESC"]) or search_launches(rql: …) — find the launch.
    • list_test_results(launchId) — per-test name, fullName (spec path + line, e.g. datasets/dataset-crud-smoke.spec.ts:8:7), status, a TestOps-computed flaky flag, muted/known, tags, jobRun.url (the GitHub Actions run), and the result id. Use search to filter to the failing test.
    • get_test_result_history(id) — the pass/fail timeline for that test across recent launches. This is the flake signal.
  • gh — gh run view <run-id> to find the failed job; gh run download <run-id> -n test-results-v2 -D <dir> for the trace artifact. A launch's jobRun.url gives you the run id.
  • npx playwright show-trace <trace.zip> (from tests_end_to_end/e2e/) — open the trace to see the exact step that failed, the DOM snapshot, and console/network at that moment.
  • git — diff the suspected change against the failing test's code path.

The loop

dot
digraph debugging_e2e {
    rankdir=TB;
    "1. Resolve entry point" [shape=box];
    "2. Gather evidence" [shape=box];
    "3. Classify" [shape=box];
    "4. Diagnose" [shape=box];
    "5. Report + propose (no edits)" [shape=box];

    "1. Resolve entry point" -> "2. Gather evidence";
    "2. Gather evidence" -> "3. Classify";
    "3. Classify" -> "4. Diagnose";
    "4. Diagnose" -> "5. Report + propose (no edits)";
}
Step 1 — Resolve the entry point

Normalize whatever you were given into "a failed test + where its evidence lives":

  • A red CI check / Actions run — take the run id. gh run view <run-id> for the failed job; find the matching launch via list_launches(projectId: 1, search: "<run-id>"); the trace is in the test-results-v2 artifact (gh run download).
  • A TestOps launch — query it directly: list_test_results(launchId), filter to the failed results.
  • A test name — list_test_results with search across a recent launch, or search launches, to find the result id; then pull its history.
  • A local failure — use the local test-results/ trace and allure-results/ directly; TestOps may have nothing for an uncommitted local run, which is fine.
Step 2 — Gather evidence
  • The failed assertion and error message (from the trace, the report, or the TestOps result).
  • The trace: npx playwright show-trace on the retained/downloaded .zip. Read the failing step, the DOM snapshot at that point, and console/network around it.
  • Screenshot / video if present (only-on-failure / retain-on-failure).
  • The test's history via get_test_result_history(id), plus the TestOps flaky flag on the result. Skip history gracefully when TestOps isn't reachable (e.g. a purely local run) and fall back to trace + diff reasoning.
Show full SKILL.md (324 more words)Show less
Step 3 — Classify

Decide: real regression, flake, or environment / selector drift.

  • History when available: a clean pass streak that broke right after a related change → lean regression. Intermittent pass/fail with no related change, or a TestOps flaky: true → lean flake.
  • Diff correlation: does a recent change touch the code path the failed assertion exercises (the page/component, the POM method, the fixture)? If yes → regression is likely. If the failed area is untouched → flake or environment is likely.
  • Default to "flake / uncertain" when history is intermittent and no related diff exists — don't over-call a regression without evidence.
Step 4 — Diagnose

Root cause, grounded in cited evidence (the specific trace step, the error, the history pattern) — not speculation. Apply the suite's lenses:

  • Verify the test render before blaming the backend. A "X didn't appear" failure is often a DOM race (a loading spinner still up, an eventually-consistent write not yet landed), not a backend regression. Check the trace's DOM snapshot at the failing step.
  • Selector drift — the FE changed an accessible name / removed a data-testid, so a locator no longer resolves.
  • Eventually-consistent state — async scoring/ingestion that needed a poll, not a fixed wait.
  • Fixture seed-shape mismatch — the page rendered an empty/partial state because the seed didn't match what the assertion expects.
Step 5 — Report + propose (no edits)

Produce:

  • Verdict — classification (regression / flake / environment-or-selector) + a confidence level.
  • Evidence — the trace step, the error, the history pattern, and the correlated change (if any), each cited.
  • Proposed fix — specific. For a regression: the code/selector/poll change to make. For a flake: a poll instead of a fixed wait, a quarantine, or "no code fix — known flaky, retry."

Do not edit anything. If the developer wants the fix applied, hand off to writing-e2e-tests.

Boundaries

  • Read-only: no test edits, no investigation-driven re-runs.
  • Works from all four entry points; degrades gracefully without TestOps (local failures use the trace + diff alone).
  • Distinct from authoring: writing-e2e-tests makes a new test; this explains a red one.

© comet-ml, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/debugging-e2e-tests of comet-ml/opik.

Open the folder on GitHubat commit f217a86

Compare with similar skills

Debugging Opik E2E Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debugging Opik E2E Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debugging Opik E2E Tests this skillcomet-ml/opik22k—~1.8kAutomated safety check: PassApache-2.0
Debug E2E Testbitovi/ai-enablement-prompts121—~677Automated safety check: PassMIT
Debug Playwrightquay/quay2.8k—~1.2kAutomated safety check: PassApache-2.0
Playwright Testingchongdashu/vibejam-starter-pack149—~2.2kAutomated safety check: PassNone
Playwright Healing AgentTahanima/playwright-java-test-automation-architecture113—~210Automated safety check: PassMIT
Web Testing with Playwright and Vitestwithkynam/vibecode-pro-max-kit1.1k—~892Automated safety check: PassApache-2.0

Similar skills

  • Debug E2E Test

    bitovi/ai-enablement-prompts

    Debug and fix failing Playwright E2E tests. An agent skill from bitovi/ai-enablement-prompts.

    121 GitHub stars~677 tokensUpdated 26 days ago
    Testing & QAAuto-check passed
  • Debug Playwright E2E test failures from GitHub Actions CI runs.

    2.8k GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.2k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Playwright Healing Agent

    Tahanima/playwright-java-test-automation-architecture

    Expert in diagnosing Playwright failures. Uses the @playwright MCP to verify live DOM state.

    113 GitHub stars~210 tokensUpdated 7 mo ago
    Testing & QAAuto-check passed
  • Web Testing with Playwright and Vitest

    withkynam/vibecode-pro-max-kit

    Covers web testing from unit to E2E, load, visual, accessibility and security checks, with Playwright, Vitest and k6 guides plus a Playwright setup script.

    1.1k GitHub stars~892 tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Pester Failure Analysis

    PowerShell/PowerShell

    Investigates failing Pester tests in PowerShell CI jobs by following a six-step workflow from pull request status to documented fix recommendations.

    56k GitHub stars~5.1k tokensUpdated today
    Testing & QAAuto-check passed

More from comet-ml/opik

All 19 skills in this repo
  • Checklist for wiring a new linter into Opik's Code Quality pipeline: the four files to edit, the silent-failure gotchas and the pass/fail verification loop.

    22k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Shows how to add product analytics events to Opik's frontend, Java backend and Python SDK, all reporting through Segment to PostHog with an opik_ name prefix.

    22k GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Rules for writing PR descriptions, changelog entries and feature documentation in the Opik repository, including the exact headings that CI requires.

    22k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Turns a code change into one committed, passing Playwright end-to-end spec by resolving the change scope and handing authoring to a companion skill.

    22k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Starts, rebuilds, and troubleshoots the Opik local dev stack, including an optional Comet Platform integration mode for the Opik team.

    22k GitHub stars~734 tokensUpdated today
    Auto-check passed
  • Specifies how to instrument an opik-backend pipeline with per-stage OpenTelemetry metrics for throughput, latency, errors and queue delay by workspace.

    22k GitHub stars~3.2k tokensUpdated today
    Auto-check passed

Categories

Questions about Debugging Opik E2E Tests

What does Debugging Opik E2E Tests do?

Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests. Give the agent a failure from a CI check, an Allure TestOps launch, a test name or a local run and it collects the evidence first: Playwright traces and videos, the HTML report and Allure results, taken from CI artifacts or the local test-results folder. It uses the allure-testops MCP to find the launch, list test results with their flaky flags and pull each test's pass and fail history, and uses gh to find the failed job and download artifacts.

When should I use Debugging Opik E2E Tests?

Debugging Opik E2E Tests fits situations like: A pull request check shows an Opik E2E test failing and you need to know why; deciding whether a test such as dataset-crud-smoke is flaky or genuinely broken; the nightly E2E suite turned red and you want the first failing test explained; reading Playwright traces and Allure history for a failed run.

How do I install Debugging Opik E2E Tests in Claude Code?

Run `npx skills add comet-ml/opik --skill debugging-e2e-tests -a claude-code`. Or copy the skill folder (.agents/skills/debugging-e2e-tests in comet-ml/opik) into .claude/skills/debugging-e2e-tests in your project. Claude Code loads it when a task matches its description.

How do I install Debugging Opik E2E Tests in Codex?

Run `npx skills add comet-ml/opik --skill debugging-e2e-tests -a codex`. Or copy the skill folder (.agents/skills/debugging-e2e-tests in comet-ml/opik) into .agents/skills/debugging-e2e-tests in your project. Codex loads it when a task matches its description.

Can I use Debugging Opik E2E Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add comet-ml/opik --skill debugging-e2e-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debugging-e2e-tests, .gemini/skills/debugging-e2e-tests, .github/skills/debugging-e2e-tests and .opencode/skills/debugging-e2e-tests in your project.

What does Debugging Opik E2E Tests need to run?

Going by SKILL.md and its folder, Debugging Opik E2E Tests needs the command-line tools its instructions call (gh and npx). Our summary lists: The allure-testops MCP server; GitHub CLI (gh); Playwright, to open traces.

Does Debugging Opik E2E Tests access the network?

SKILL.md contains no URLs. Its commands use gh and npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Debugging Opik E2E Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Debugging Opik E2E Tests use?

Debugging Opik E2E Tests is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debugging Opik E2E Tests use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Debugging Opik E2E Tests?

Skills that share tags, products or a category with Debugging Opik E2E Tests: Debug E2E Test (bitovi/ai-enablement-prompts, 121 stars), Debug Playwright (quay/quay, 2.8k stars), Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars) and Playwright Healing Agent (Tahanima/playwright-java-test-automation-architecture, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debugging Opik E2E Tests?

comet-ml (a GitHub organization) maintains it in comet-ml/opik, which has 22,412 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 7, 2026.

Source: comet-ml/opik on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.