Agent skill

Opik Visual Regression Tests

by comet-ml in comet-ml/opik

Adds a screenshot-comparison test for an Opik UI page, following the project's page-object pattern and deterministic seeding rules.

Apache-2.0Auto-check passedTesting & QA

Install Opik Visual Regression Tests

skills CLI
$ npx skills add comet-ml/opik --skill writing-visual-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install comet-ml/opik writing-visual-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/comet-ml/opik.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/writing-visual-tests .claude/skills/writing-visual-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
writing-visual-tests
GitHub stars
22k
Token cost
~8.5k tokens
SKILL.md length
4,072 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Adds a screenshot-comparison test for an Opik UI page, following the project's page-object pattern and deterministic seeding rules.

  • Works in 7 steps: Scope. Which page/panel, which states… → Seed data. Add a test-helper-service… → Page object(s). Follow the pattern above… → …
  • Adding a screenshot test for a new Opik page or panel state
  • SKILL.md covers Where tests live, Page object pattern, Safety: verify local config… and Local environment, plus 11 more sections
  • Calls npx, docker and python

What it does

Unlike Opik's functional end-to-end suite, a visual test's assertion is the screenshot itself, so seeding has to produce pixel-deterministic state rather than merely correct data; a typed client over a Flask test-helper service wraps the real Python SDK for that seeding, and the project explicitly bans seeding by clicking through the UI. Fixed-name projects such as visual-project and visual-sidebar-project, created and torn down in global setup, keep screenshots identical across runs, and each new group of screenshots gets its own project rather than sharing one.

Before any file is written, the plan names the page, states, new seed route, spec file, and resulting screenshot names, and waits for explicit approval unless the request already settles every one of those choices.

When your agent uses it

  • Adding a screenshot test for a new Opik page or panel state
  • Screenshotting every tab or state of an existing panel
  • Regenerating or reviewing visual-test baselines after a UI change

Example prompts

  • “Add a visual test for the trace sidebar's empty state.”
  • “Screenshot each tab of the dataset panel for regression coverage.”
  • “Plan a visual test for the new project settings page before writing it.”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Scope. Which page/panel, which states (tabs, empty vs. populated, with/without optional data). Check whether an existing spec's beforeAll…
  2. Seed data. Add a test-helper-service route if no existing one produces the shape you need (see "Adding a seed endpoint" below), plus a…
  3. Page object(s). Follow the pattern above exactly; reuse an existing page object if the page already has one.
  4. Spec. One test() per screenshot, unique names (see "Unique screenshot names").
  5. Generate/refresh baselines for the whole suite — every time, whether or not baselines already exist (see "Baselines are local and…
  6. Stress-test for stability — do not skip this. Run the full-suite compare pass (no --update-snapshots) 3 times in a row
  7. Final full run, then work through the "Cleanup checklist" above — one run without SKIP_TEARDOWN first (so global-teardown.ts actually…

What it can do on your machine

Read from SKILL.md and the folder at commit f217a86. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • docker
    • python
    • curl
    • npm
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, docker, curl, npm and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Opik Visual Regression Tests loads about 8.5k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 4,072 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~8.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from comet-ml/opik at commit f217a86, republished under its Apache-2.0 licence (© comet-ml). 4,072 words, ~8,476 tokens.

Download SKILL.mdSave it as .claude/skills/writing-visual-tests/SKILL.md (or your agent's skills folder).
name
writing-visual-tests
description
Use when a developer wants to add a visual regression (screenshot) test for an Opik UI page or panel — e.g. "add a visual test for the trace sidebar", "screenshot each tab of the dataset panel", "visual regression test for the new empty state". Covers the page-object pattern, seeding via the test-helper-service, unique screenshot naming, per-test masking, baseline (re)generation, and the local-run-until-stable loop in tests_end_to_end/visual-tests/.

Writing Visual Tests

This skill adds a screenshot-comparison test to Opik's visual regression suite. Unlike the functional E2E suite (tests_end_to_end/e2e/, see writing-e2e-tests), this suite's assertion is the screenshot — so seeding must produce deterministic pixels, not just correct data.

Announce at start: "I'm using the writing-visual-tests skill to add a visual test for X."

Plan before implementing. Before writing or editing any file, work through step 1 ("Scope") of "The loop" below and present the user a short plan: which page/panel and states you'll cover, what new seed route/client method (if any) you'll add, which spec file and project the new test(s) will live in (and why, per "One project per screenshot group"), and the resulting screenshot name(s). Wait for the user's explicit approval before starting step 2 (seeding) or any other implementation work. Only skip this pause if the user's request already specifies all of these choices unambiguously.

Where tests live

The suite is at tests_end_to_end/visual-tests/:

  • Specs: tests/<area>.spec.ts — e.g. visual-comparison.spec.ts (happy-path pages), empty-states.spec.ts, trace-sidebar.spec.ts. One test.describe per spec, one test.beforeAll that seeds data and resolves the project ID.
  • Page objects: page-objects/<page>.page.ts — see "Page object pattern" below; always follow it, don't freehand a different shape.
  • Seeding: helpers/test-helper-client.ts — a typed HTTP client over the Flask test-helper-service (../test-helper-service/routes/*.py), which wraps the real Python SDK. Add a new route + a new client method together; never seed by clicking through the UI.
  • Screenshot util: tests/utils/screenshot.ts — screenshot(page, name, extraMasks?). Read this whole file before writing a new spec.
  • Global setup/teardown: global-setup.ts / global-teardown.ts — create/delete the fixed-name projects (visual-project, visual-empty-project, visual-sidebar-project, …) used across specs. Fixed names, no timestamp suffix, so screenshots are identical across runs. Give each new portion of screenshots its own project — see "One project per screenshot group" below.
  • Baselines: screenshots/baseline/*.png — see "Baselines are local and gitignored" below before you assume one exists.

Page object pattern

Every existing page object (projects.page.ts, experiments.page.ts, datasets.page.ts, test-suites.page.ts, logs.page.ts, …) follows the same shape. New ones must match it exactly — don't introduce test.step() wrapping (that's the e2e/ suite's convention, not this one), don't return locators from goto(), don't add a constructor signature that differs from the others.

ts
import { Page } from '@playwright/test';
import { BasePage } from './base.page';

export class WidgetsPage extends BasePage {
  constructor(page: Page, baseUrl: string, workspace: string) {
    super(page, baseUrl, workspace);
  }

  async goto(projectId: string): Promise<void> {
    await this.page.goto(this.url(`projects/${projectId}/widgets`));
    await this.page.waitForLoadState('load');
    await this.dismissWelcomeDialogIfPresent();
  }

  // Races the populated state against the empty state so the same method
  // works whether the seeded project has data or not — never a bare
  // unconditional wait, which hangs forever if the match text is wrong.
  async waitForReady(expectedCellText: string): Promise<void> {
    await this.page.getByRole('heading', { name: 'Widgets', exact: true }).waitFor({ state: 'visible', timeout: 10000 });
    await Promise.race([
      this.page.locator('tbody tr td').filter({ hasText: expectedCellText }).first().waitFor({ state: 'visible', timeout: 10000 }),
      this.page.getByRole('heading', { name: /no widgets yet/i }).waitFor({ state: 'visible', timeout: 8000 }),
    ]);
  }

  async waitForEmpty(): Promise<void> {
    await this.page.getByRole('heading', { name: 'Widgets', exact: true }).waitFor({ state: 'visible', timeout: 10000 });
    await this.page.getByRole('heading', { name: /no widgets yet/i }).waitFor({ state: 'visible', timeout: 20000 });
  }
}

Rules that fall out of this:

  • Extend BasePage and take the exact (page, baseUrl, workspace) constructor, even if a given page doesn't need workspace yet — consistency lets every spec construct page objects identically.
  • goto() always: navigate via this.url(...) (never a raw string), waitForLoadState('load'), then dismissWelcomeDialogIfPresent().
  • waitForReady(expectedCellText) takes the text to match, doesn't hardcode it — callers pass whatever they actually seeded. See "Waiting: for content, not just structure" below for what's safe to pass.
  • waitForEmpty() is a separate method, reused by empty-states.spec.ts — don't fold empty-state handling into waitForReady as a flag.
  • Add new interaction methods (click, open, switch-tab) alongside goto/waitForReady, not in the spec file.
Exception: panels and non-route components

A side panel (e.g. the trace/span detail panel) isn't its own route — it opens over an existing page. Don't force it into BasePage. Give it just a Page and a root locator scoped to the panel's own testid, then scope every other locator/action under root:

ts
import { Page, Locator } from '@playwright/test';

export class TraceDetailsPanelPage {
  constructor(private page: Page) {}

  get root(): Locator {
    return this.page.getByTestId('traces');
  }

  async waitForLoaded(): Promise<void> {
    await this.root.getByRole('button', { name: 'Close' }).waitFor({ state: 'visible', timeout: 20000 });
    // ...
  }
}

Construct it from the page the panel opens over, in the spec: new TraceDetailsPanelPage(page), right after the page object that opened it.

Safety: verify local config before seeding

The Python SDK behind test-helper-service reads ~/.opik.config. If it points at a cloud environment, seeding creates real data there.

bash
grep url_override ~/.opik.config

If url_override is anything other than http://localhost:5173/api/, back it up and point it local:

bash
cp ~/.opik.config ~/.opik.config.bak-visual-test
cat > ~/.opik.config << 'EOF'
[opik]
url_override = http://localhost:5173/api/
workspace = default
EOF

When the work is done, restore it and stop test-helper-service — see the "Cleanup checklist" below before doing either: both are machine-wide, and blindly restoring/killing can break a concurrent run.

Local environment

Use ./opik.sh (repo root) for the Docker stack — check what's already running before touching it:

bash
docker ps --format '{{.Names}}\t{{.Status}}'   # look for opik-opik-frontend-1, opik-opik-backend-1, etc.
./opik.sh --verify                             # only if you're unsure containers are healthy

If containers are already up (healthy, running), leave them alone — don't --stop/--clean/restart something a teammate or another task left running. Only run ./opik.sh (no flags) to start what's missing.

A rebuilt stack is not a fresh one. ./opik.sh --build rebuilds images but reuses existing Docker volumes (MySQL, ClickHouse, …) unless you explicitly wipe them — so a container that just restarted can still be serving months of accumulated local data: projects, feedback definitions, environments, AI provider keys, whatever previous sessions (yours or a teammate's) left behind. If your test's scope includes a default or empty state — anything you'd normally get "for free" on a brand-new workspace — don't trust what a long-lived instance currently shows. See "Know your defaults before you seed or delete" below before writing any seed/cleanup code against it. If a genuinely fresh stack (fresh volumes) is feasible and not disruptive to concurrent work, prefer it for this kind of test; if not, fall back to reasoning from the Liquibase migrations directly rather than from the current DB contents.

Start the test helper service (check port 5555 is free first):

bash
lsof -ti:5555                                  # empty output = free, safe to start
cd tests_end_to_end/test-helper-service
OPIK_BASE_URL=http://localhost:5173 OPIK_WORKSPACE=default TEST_HELPER_PORT=5555 \
  nohup .venv/bin/python app.py > /tmp/test-helper-service.log 2>&1 &
disown
curl -s http://localhost:5555/health

Restart it (kill $(lsof -ti:5555), relaunch) whenever you change a route file — Flask's dev server doesn't hot-reload here. Same caution as everywhere else this port comes up: confirm the PID is actually yours before killing it (see "Cleanup checklist" below).

Cleanup checklist

Everything this suite generates is gitignored (root .gitignore entries: tests_end_to_end/visual-tests/test-results/, visual-report/, .auth/, .test-state.json, screenshots/; repo-wide: allure-results/) — nothing here risks a bad commit, but it does accumulate on disk across sessions. Work through this list before ending a visual-testing session, not just the config restore.

~/.opik.config and the test-helper-service process are machine-wide, not scoped to your run. If another agent, terminal, or teammate could be running visual/E2E tests concurrently on the same machine, check before you touch either — restoring the config or killing the process out from under a run in progress breaks it:

bash
lsof -ti:5555                      # a PID here isn't necessarily "your" leftover — it may be someone else's active run
ps -p <pid> -o lstart,command      # when did it start, relative to when you started your own work?

If in doubt, ask rather than kill. This isn't hypothetical — it happened during this skill's own authoring session: a test-helper-service instance was found still listening after "cleanup," and it turned out to belong to a different run in progress, not a leftover.

Also note: playwright.config.ts's webServer directive auto-spawns test-helper-service on port 5555 if nothing answers /health when a test starts (reuseExistingServer: !process.env.CI) — so an instance can appear that you never manually started, including right after your own "final run." Don't assume every process on that port traces back to a command you ran.

ArtifactWhat it isAction
~/.opik.configPoints the SDK at an environmentConfirm no concurrent run needs it local (see above), then restore from your backup (cp ~/.opik.config.bak-visual-test ~/.opik.config) and delete the backup file.
test-helper-service process (port 5555)Flask bridge — may be one you started, or one webServer auto-spawned, or someone else'sConfirm it's actually idle/yours (see above) before kill $(lsof -ti:5555).
Docker containers (opik-opik-*)The stack under testLeave alone if they were already running when you started (the common case) — don't --stop/--clean a stack you didn't start. Only stop containers you personally started for this session.
test-results/Playwright's per-test output (failure screenshots, videos, traces)Regenerated every run — safe to delete: rm -rf test-results.
screenshots/comparison/Side-by-side copy written on every comparison run (i.e. whenever SKIP_TEARDOWN is not 1)Regenerated every run — safe to delete: rm -rf screenshots/comparison.
visual-report/The HTML report (npm run test:report)Safe to delete: rm -rf visual-report.
allure-results/Allure reporter outputSafe to delete: rm -rf allure-results.
.test-state.jsonExperiment-ID bridge for non-SKIP_TEARDOWN runsRemoved automatically by global-teardown.ts on a normal run; delete by hand only if you aborted mid-run without ever finishing a non-SKIP_TEARDOWN pass.
.auth/Playwright storage state written by global-setup.tsNot cleaned up by global-teardown.ts — it only removes .test-state.json and server-side data. Delete by hand: rm -rf .auth.
screenshots/baseline/*.pngYour actual baselinesNot cruft — don't reflexively delete. Gitignored and per-machine by design (see "Baselines are local and gitignored"); leave them for the next person/session to reuse, unless you specifically want to force a clean re-baseline.
Seeded project (visual-project / visual-empty-project / visual-sidebar-project) in the running Opik instanceServer-side data, not a local fileDeleted by global-teardown.ts on a plain npx playwright test run (not SKIP_TEARDOWN=1). Do a final run without SKIP_TEARDOWN before you finish (see step 7) so this actually happens — don't leave it to the next person. If teardown never ran (e.g. an interrupted SKIP_TEARDOWN=1 session), delete any of the three by hand via the running instance's UI/API.

One-shot cleanup for the local-only report/log artifacts (keeps baselines):

bash
cd tests_end_to_end/visual-tests
rm -rf test-results screenshots/comparison visual-report allure-results

Baselines are local and gitignored

screenshots/ is gitignored (see .gitignore) — baseline PNGs are never committed. There is no CI job for this suite; it's a local, on-demand check. Two consequences:

  • On a fresh clone, or after git clean, no baselines exist at all — every screenshot name in every spec needs generating before you can run a compare pass.
  • Adding a new screenshot name always needs a fresh baseline, regardless of whether other baselines in the same directory already exist from a previous session.

Always (re)generate baselines for the whole suite, not just the spec you touched, before judging a compare run — via --update-snapshots (see the loop below). Don't assume "the baseline is already there"; check, or just regenerate, it's cheap:

bash
ls tests_end_to_end/visual-tests/screenshots/baseline/ | grep <your-prefix>

Regenerate whenever any of these change, not just on first add: the screenshot name, the masks passed to it, a wait added/removed before it, or the seeded data shape.

The loop

  1. Scope. Which page/panel, which states (tabs, empty vs. populated, with/without optional data). Check whether an existing spec's beforeAll already seeds a project/entity you can extend, or whether this needs its own.
  2. Seed data. Add a test-helper-service route if no existing one produces the shape you need (see "Adding a seed endpoint" below), plus a TestHelperClient method.
  3. Page object(s). Follow the pattern above exactly; reuse an existing page object if the page already has one.
  4. Spec. One test() per screenshot, unique names (see "Unique screenshot names").
  5. Generate/refresh baselines for the whole suite — every time, whether or not baselines already exist (see "Baselines are local and gitignored" above). Run all spec files, not just the one you touched: adding or changing a test can shift shared page chrome, and stale baselines in untouched files are otherwise only caught by accident.
    bash
    cd tests_end_to_end/visual-tests
    SKIP_TEARDOWN=1 OPIK_BASE_URL=http://localhost:5173 npx playwright test --update-snapshots --reporter=list
    SKIP_TEARDOWN=1 keeps the seeded projects alive for a fast next run; the plain npx playwright test invocation still always cleans and recreates them at the start (global-setup), so each run starts from a known state. If the full-suite run surfaces failures in files you didn't touch, don't reflexively blame your change — check whether those baselines simply predate a recent, unrelated app change (compare ls -la screenshots/baseline/ timestamps against recent frontend commits) before deciding whether to regenerate them or investigate further; confirm with the user before regenerating baselines outside the scope of your own change.
  6. Stress-test for stability — do not skip this. Run the full-suite compare pass (no --update-snapshots) 3 times in a row:
    bash
    for i in 1 2 3; do
      echo "=== Run $i ==="
      SKIP_TEARDOWN=1 OPIK_BASE_URL=http://localhost:5173 npx playwright test --reporter=list
    done
    A single green run proves nothing — real timestamps, real durations, and font/layout jitter only show up over several runs. If anything fails, open the *-diff.png attachment under test-results/ before touching anything — it tells you exactly which pixels moved. Don't guess. If you change the spec, masks, or seed data in response, go back to step 5 and regenerate the baseline before stress-testing again — a stale baseline will just fail for a different, misleading reason.

    macOS gotcha: timeout is not a built-in command (no coreutils by default) — a loop like timeout 60 npx playwright test ... silently no-ops. Don't wrap runs in timeout. Shell-tool gotcha: 3 sequential runs of the full suite easily exceed a shell tool's default ~2-minute timeout, cutting the loop off mid-run. That's a harness timeout, not a test failure — don't read it as one. Pass an explicit longer timeout for the whole loop (e.g. 5+ minutes), or split into separate invocations and resume numbering (for i in 2 3; do ...) if one gets cut off.

  7. Final full run, then work through the "Cleanup checklist" above — one run without SKIP_TEARDOWN first (so global-teardown.ts actually deletes the seeded projects server-side), then every item in the checklist. Config restore and stopping test-helper-service are only two of several things left behind.

Unique screenshot names

All specs share one flat directory, screenshots/baseline/ (see snapshotPathTemplate in playwright.config.ts) — there is no per-spec-file subfolder. A name collision with another spec silently compares your new test against someone else's baseline (or overwrites theirs).

  • Prefix every screenshot name with a short, spec-unique code, e.g. visual-comparison.spec.ts uses 01-, 02-…; empty-states.spec.ts uses E01-, E02-…; trace-sidebar.spec.ts uses S01-, S02-…. Pick a prefix letter/word not already in use.
  • Before naming, check for collisions:
    bash
    grep -rn "screenshot(page, '" tests_end_to_end/visual-tests/tests/*.spec.ts
  • Keep the rest of the name descriptive (S02-trace-sidebar-details, not S02) — the file is the only durable record of what a screenshot covers once screenshots/ is gitignored.

Masks: scope them to your test, don't touch the shared list

tests/utils/screenshot.ts has two tiers, not one shared list:

  • baseMasks(page) (internal, always applied by screenshot()) — only page chrome that renders on virtually every screenshot and can legitimately vary between the two environments being compared: the breadcrumb (workspace/project name) and generic timestamp elements.
  • tableMasks(page) (exported) — masks for populated data tables: relative/absolute date cells, UUID columns, duration cells, pagination "Showing X-Y of Z". Only screenshots of an actual table with rows need this (see visual-comparison.spec.ts's tests 01-06, which pass tableMasks(page) explicitly). Empty-state screenshots have no rows and the trace sidebar has no table at all, so neither passes it.

Do not add a mask to baseMasks for something specific to your new page; it silently changes masking for every other visual test in the suite. If your new screenshot is a populated table, pass the existing tableMasks(page) — don't reinvent the same regexes. If it's something narrower (a single stat row, one specific element), pass a test-specific mask through screenshot()'s third argument instead:

ts
// tests/utils/screenshot.ts
export async function screenshot(page: Page, name: string, extraMasks: Locator[] = []) { ... }
ts
// your spec
await screenshot(page, 'S02-trace-sidebar-details', [panel.statsRowMask]);
// or, for a table page:
await screenshot(page, '07-widgets-page', tableMasks(page));

Expose the locator as a getter on the relevant page object (see TraceDetailsPanelPage.statsRowMask) so the masking rationale lives next to the DOM it targets, not buried in the spec.

Only promote a mask into baseMasks or tableMasks if the pattern is genuinely generic and reusable across many pages (e.g. "any td matching a relative-date regex") — not a one-off element on the one page you're testing.

Mask the stable-sized container, not the small dynamic element

If the dynamic content sits inline in a flex/wrap row next to other elements (a stats bar, a badge row), masking just the small element isn't enough — its rendered width still varies by a pixel run to run (real timestamp/duration text, font sub-pixel rendering), which shifts every sibling after it even though they're masked-adjacent, not masked themselves. Symptom: a screenshot diff of only a few hundred/thousand pixels, always right next to a masked box, that comes and goes across runs.

Fix: mask the smallest ancestor with a stable box (e.g. one that's w-full or otherwise sized independent of its children) so internal jitter never leaks past the mask edge:

ts
// Masks the whole stats row (created-at, duration, score counts) as one block,
// because that row is `w-full` — its own box doesn't move even when the text inside it does.
get statsRowMask(): Locator {
  return this.page.getByTestId('data-viewer-created-at').locator('xpath=..');
}
Show full SKILL.md (1,674 more words)Show less

Adding a seed endpoint

Follow the existing pattern in test-helper-service/routes/*.py: a Flask blueprint route using get_opik_client(), validate_required_fields(), success_response(). Then add a matching method to TestHelperClient in helpers/test-helper-client.ts.

Make the seeded data deterministic — this is the #1 source of visual-test flakiness:

  • Don't rely on @track/opik_context decorator timing if the trace/span's duration will be visible on screen. Real wall-clock execution time varies run to run, even if it always rounds to "0s" — the rendered text can still differ by a sub-pixel width and shift siblings. Instead call client.trace(...) / trace.span(...) directly with an identical start_time and end_time, so duration is exactly zero every time:
    python
    now = datetime.datetime.now(datetime.timezone.utc)
    trace = client.trace(..., start_time=now, end_time=now)
    span = trace.span(..., start_time=now, end_time=now)
  • Log at most one feedback score per trace/span in a screenshot-visible table unless you've confirmed the UI sorts them deterministically. Two scores logged in the same call have no guaranteed render order (no explicit sort key), so a table showing both can silently swap row order between runs.
  • To attach a prompt to a trace/span for the Prompts tab, don't rely on the @track context helpers — build the metadata directly so it works with plain client.trace():
    python
    prompt = client.create_prompt(name=..., prompt=...)
    client.trace(..., metadata={"opik_prompts": [prompt.__internal_api__to_info_dict__()]})
  • Attachments: resolve paths via the existing resolve_attachment_path() helper (relative to tests_end_to_end/), and check in a small fixture file under visual-tests/fixtures/ if one doesn't already exist for your case.

Know your defaults before you seed or delete

Before you write a test that asserts an "empty" or "default" state, find out what a genuinely fresh workspace actually contains for that entity — don't infer it from what a long-lived local instance currently shows, and don't assume "I deleted it and now it's empty" proves the state is reachable elsewhere.

Backend migrations can seed real per-workspace defaults that are visually indistinguishable from leftover test/dev artifacts:

bash
grep -rl "INSERT" apps/opik-backend/src/main/resources/liquibase/**/*.sql | xargs grep -l "<table_name>"

What you find splits into two cases, and they need different handling:

  • Hard-protected default (backend blocks deletion). Some seeded rows are permanently undeletable by design — e.g. the "User feedback" feedback definition, guarded in FeedbackDefinitionService.java (containsNameByIds(... USER_FEEDBACK) → 409 on delete). If you hit a 409 trying to clear one, that "empty" screenshot is not reachable on any real Opik instance. Don't chase it by deleting whatever it's referencing (e.g. cascading into trace feedback scores to free it up) — drop that screenshot, tell the user why, and note the gap in the PR description. Confirm this with the user before dropping it; don't decide unilaterally.
  • Deletable-but-still-a-default. Other rows are seeded the same way (a migration backfills them for every workspace) but have no delete guard — e.g. environments gets development/staging/production from a real migration (...seed_default_environments.sql), not from anyone clicking around. Deleting these locally only fixes your machine. If your test's cleanup routine (global-setup.ts/global-teardown.ts) doesn't also delete these known default names on every run, the test only passes on the instance you personally vandalized — it fails on a teammate's clone, a fresh volume, or CI, either by rendering extra rows in a populated screenshot or by never reaching the empty state at all.

Before deleting anything on a shared/long-lived local instance to "get to empty," ask whether it's a real default (migration-seeded, potentially relied on elsewhere) or actual scratch data (e.g. an old e2e run's leftover config) — check for the entity's name in tests_end_to_end/e2e/** seeding helpers too, since some things that look like defaults (e.g. a lone "openai" AI provider key on an otherwise-default instance) are really just residue from a different test suite (ensureProviderConfigured) rather than anything the backend seeds. If deleting something that turns out to matter beyond your own session (shared instance, other suites depending on it), stop and confirm with the user first rather than deleting and hoping it doesn't matter.

Waiting: for content, not just structure

waitForReady() methods should wait for the specific value you seeded to be visible (e.g. a known input string), not just a generic loading state. Many Opik tables don't render a "Name" column by default, so waitFor({ hasText: entityName }) can time out even though the row is there — match on whatever text is actually visible in the default column set (e.g. the input/output preview), not the entity's name.

For any panel with lazy-loaded content (a trace/span side panel's Input/Output, media, etc.), wait for its loading placeholder to clear before switching tabs, and again after switching tabs, before screenshotting:

ts
await this.root.getByText('Loading', { exact: true }).waitFor({ state: 'hidden', timeout: 15000 });

One project per screenshot group

Don't seed a new spec's data into an existing spec's shared project (e.g. visual-project) just because it's already there and already has a beforeAll you could piggyback on. Every spec file that seeds its own traces/threads/datasets should create and use its own dedicated project in global-setup.ts/global-teardown.ts (see visual-sidebar-project for trace-sidebar.spec.ts).

Reusing a shared project causes two distinct problems, both discovered the hard way while adding trace-sidebar.spec.ts:

  • Cross-spec pollution. Data seeded by spec B's beforeAll becomes visible in spec A's table screenshot (an extra row, a shifted count/chart) even though spec A never changed — because both specs' entities live in the same project. This produces a confusing, flaky-looking diff in a test you didn't touch.
  • Non-idempotent-seed collisions compound. Combined with the hook-restart gotcha below, two specs sharing one project multiplies the chance that a retried beforeAll collides with another spec's still-present data (dataset/name conflicts, unexpected table rows) — independent of any bug in your own spec.

Cost is low — global-setup.ts already loops over project creation — so default to a new project per spec unless it's an empty-state test explicitly reusing visual-empty-project on purpose.

This rule is about other specs' projects, not your own. Adding a new test() to a spec file that already owns a project (e.g. adding test 07 to visual-comparison.spec.ts, which already owns visual-project) correctly reuses that same project — that's the same screenshot group, not a new one. Don't manufacture a new project just because you're adding a test.

But within that reuse, isolate new seeded entities from ones an existing baseline already depends on. If your new test needs data the existing beforeAll doesn't already produce, add a new, separate trace/thread/dataset — don't extend an entity that an earlier, already-baselined test in the same file screenshots. Extending it (e.g. adding a sibling span to a trace whose tree is shown in an earlier screenshot) changes what that earlier screenshot renders (an extra tree row, a shifted count) and silently breaks its baseline. This is the same "cross-pollution" failure mode as above, just within one spec file instead of across two — and it's easy to miss because it's your own change causing it, not another spec's. Give the new state its own uniquely-named trace/entity in the same project instead.

If unsure whether "same project, new test" or "same project, new isolated entity" is right for what you're adding, it usually comes down to: does an already-passing test in this file screenshot the shared context (a tree, a table, a count) that your new seed data would also appear in? If yes, isolate with a new entity; if no (e.g. a standalone page-level screenshot with its own query), reusing the existing seed is fine.

Known Playwright gotcha: hook re-run on failure

If a test in a describe block fails, Playwright restarts the worker process before the next test in that file — which re-runs test.beforeAll. If your beforeAll seeds data (as ours does), one real failure produces a second copy of the seeded entity before the next test runs, which can cascade into unrelated-looking failures in every subsequent test in the file (growing diffs, extra table rows). When diagnosing a multi-test failure, always look at the first failure in isolation — it's usually the only real bug; the rest may just be contamination from the worker restart.

Anti-patterns

SymptomWhat you skipped
"The baseline folder already has PNGs, I don't need --update-snapshots"Baselines are gitignored and per-machine — a name you just added has no baseline yet, regardless of what else is in the folder.
"It passed once, ship it"The stress-test loop — real timestamps and font jitter are intermittent by nature; one green run doesn't prove stability.
"I'll match the row by trace/dataset name"Checking what's actually rendered — many tables don't show a Name column; match visible cell text instead.
"I'll add this mask to baseMasks(), it's easier"Scoping — that changes every other visual test's masking. Use tableMasks(page) if it's a table page, or screenshot()'s extraMasks param otherwise.
"I'll mask just the timestamp div"Checking whether it sits in a flex row with siblings — mask the stable-sized parent instead, or the siblings will still jitter.
"The decorator's duration always shows 0s, it's fine"Determinism — "usually 0s" still varies by a sub-pixel width; force it to exactly 0 with matching start_time/end_time.
"All 4 tests failed, must be a bigger bug"The hook-restart gotcha — check the first failure alone before assuming later ones are independent.
"I'll just seed my new spec into visual-project, it's already there"One project per screenshot group — reusing another spec's project pollutes its table screenshots with your seeded data.
"I'll add my new span/row to the trace an earlier test in this file already screenshots"Isolating new entities within a reused project — extending a shared entity shifts what an already-baselined screenshot in the same file renders (extra tree row, changed count), breaking it. Seed a new, separate entity instead.
"I'll wrap this panel in BasePage"The panel exception — a non-route panel takes just Page + a root testid locator, not the full BasePage constructor.
"I stopped the test-helper-service, I'm done"The rest of the cleanup checklist — config restore, a non-SKIP_TEARDOWN final run (server-side project deletion), and the local report/log directories (test-results/, visual-report/, allure-results/, screenshots/comparison/).
"I deleted the leftover rows, now the tab is empty, ship it"Checking whether what you deleted was a real per-workspace default (migration-seeded, e.g. environments' development/staging/production) rather than test residue — if it's a real default, your cleanup routine must delete it by name on every run, or the empty state only exists on the one instance you manually cleared.
"The rebuild finished, the stack is fresh"--build rebuilds images, not volumes — a "fresh" container can still be serving months of accumulated local data. Don't infer default/empty state from a long-lived instance's current contents; check the Liquibase migrations.

© comet-ml, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/writing-visual-tests of comet-ml/opik.

Open the folder on GitHubat commit f217a86

Compare with similar skills

Opik Visual Regression Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Opik Visual Regression Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Opik Visual Regression Tests this skillcomet-ml/opik22k—~8.5kAutomated safety check: PassApache-2.0
Drive MiMo CodeXiaomiMiMo/MiMo-Code14k—~3.9kAutomated safety check: PassMIT
Glance TestDebugBase/glance156—~827Automated safety check: PassMIT
Playwright Testingchongdashu/vibejam-starter-pack149—~2.2kAutomated safety check: PassNone
Handsontable Visual Test Demoshandsontable/handsontable22k—~1.3kAutomated safety check: PassCustom licence
Playwright Expertcin12211/orca-q223—~1.3kAutomated safety check: PassMIT

Similar skills

  • Drive MiMo Code

    XiaomiMiMo/MiMo-Code

    Lets one MiMoCode process drive another, headless with JSON events or interactively through tmux, to test behavior and visual regressions with parseable evidence.

    14k GitHub stars~3.9k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Glance Test

    DebugBase/glance

    Run E2E browser tests on any web application using Glance MCP.

    156 GitHub stars~827 tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.2k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Handsontable Visual Test Demos

    handsontable/handsontable

    Explains how to add or change the demo pages that Handsontable's visual regression suite photographs, including per-feature routes in the js demo and the shared grid.

    22k GitHub stars~1.3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Playwright Expert

    cin12211/orca-q

    Playwright E2E testing expert for browser automation, cross-browser testing, visual regression, network interception, and CI integration.

    223 GitHub stars~1.3k tokensUpdated 16 days ago
    Testing & QAAuto-check passed
  • Test Specialist

    travisjneuman/.claude

    Test-writing patterns for JS/TS, Python, Go, and Rust (unit, integration, E2E, visual regression).

    101 GitHub stars~3.7k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from comet-ml/opik

All 19 skills in this repo
  • Checklist for wiring a new linter into Opik's Code Quality pipeline: the four files to edit, the silent-failure gotchas and the pass/fail verification loop.

    22k GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Shows how to add product analytics events to Opik's frontend, Java backend and Python SDK, all reporting through Segment to PostHog with an opik_ name prefix.

    22k GitHub stars~4.4k tokensUpdated yesterday
    Auto-check passed
  • Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests.

    22k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Rules for writing PR descriptions, changelog entries and feature documentation in the Opik repository, including the exact headings that CI requires.

    22k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Turns a code change into one committed, passing Playwright end-to-end spec by resolving the change scope and handing authoring to a companion skill.

    22k GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Starts, rebuilds, and troubleshoots the Opik local dev stack, including an optional Comet Platform integration mode for the Opik team.

    22k GitHub stars~734 tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Opik Visual Regression Tests

What does Opik Visual Regression Tests do?

Adds a screenshot-comparison test for an Opik UI page, following the project's page-object pattern and deterministic seeding rules. Unlike Opik's functional end-to-end suite, a visual test's assertion is the screenshot itself, so seeding has to produce pixel-deterministic state rather than merely correct data; a typed client over a Flask test-helper service wraps the real Python SDK for that seeding, and the project explicitly bans seeding by clicking through the UI. Fixed-name projects such as visual-project and visual-sidebar-project, created and torn down in global setup, keep screenshots identical across runs, and each new group of screenshots gets its own project rather than sharing one.

When should I use Opik Visual Regression Tests?

Opik Visual Regression Tests fits situations like: adding a screenshot test for a new Opik page or panel state; screenshotting every tab or state of an existing panel; regenerating or reviewing visual-test baselines after a UI change.

How do I install Opik Visual Regression Tests in Claude Code?

Run `npx skills add comet-ml/opik --skill writing-visual-tests -a claude-code`. Or copy the skill folder (.agents/skills/writing-visual-tests in comet-ml/opik) into .claude/skills/writing-visual-tests in your project. Claude Code loads it when a task matches its description.

How do I install Opik Visual Regression Tests in Codex?

Run `npx skills add comet-ml/opik --skill writing-visual-tests -a codex`. Or copy the skill folder (.agents/skills/writing-visual-tests in comet-ml/opik) into .agents/skills/writing-visual-tests in your project. Codex loads it when a task matches its description.

Can I use Opik Visual Regression Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add comet-ml/opik --skill writing-visual-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/writing-visual-tests, .gemini/skills/writing-visual-tests, .github/skills/writing-visual-tests and .opencode/skills/writing-visual-tests in your project.

What does Opik Visual Regression Tests need to run?

Going by SKILL.md and its folder, Opik Visual Regression Tests needs the command-line tools its instructions call (npx, docker, python, curl, npm and git).

Does Opik Visual Regression Tests access the network?

SKILL.md contains no URLs. Its commands use npx, docker, curl, npm and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Opik Visual Regression Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Opik Visual Regression Tests use?

Opik Visual Regression Tests is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Opik Visual Regression Tests use?

About 8.5k tokens (SKILL.md is roughly 34k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Opik Visual Regression Tests?

Skills that share tags, products or a category with Opik Visual Regression Tests: Drive MiMo Code (XiaomiMiMo/MiMo-Code, 14k stars), Glance Test (DebugBase/glance, 156 stars), Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars) and Handsontable Visual Test Demos (handsontable/handsontable, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Opik Visual Regression Tests?

comet-ml (a GitHub organization) maintains it in comet-ml/opik, which has 22,412 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 7, 2026.

Source: comet-ml/opik on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.