Agent skill

Opik Verify

by comet-ml in comet-ml/opik-mcp

Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost…

Apache-2.0Auto-check: notesTesting & QA

Install Opik Verify

skills CLI
$ npx skills add comet-ml/opik-mcp --skill opik-verify -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install comet-ml/opik-mcp opik-verify --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/opik_mcp/skills/opik-verify .claude/skills/opik-verify && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
opik-verify
GitHub stars
219
Token cost
~2.8k tokens
SKILL.md length
1,384 words
Files
14 (incl. references)
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost…

  • Works in 6 steps: Load the policy → Resolve the two runs → Read both runs, item by item → …
  • Is this safe to ship
  • SKILL.md covers Inputs, The policy, Activation — the only in-scope… and Blockers, plus 4 more sections
  • Runs Python scripts from its folder

What it does

Opik Verify is an agent skill from comet-ml/opik-mcp. Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost budgets, flakiness, evidence size, and whether the judge is validated. Reads two experiments on an Opik test suite via the SDK (or the MCP when connected) and returns a verdict with every criterion shown pass/fail. Use for "is this safe to ship", "can I merge this", "go/no-go on this change", "gate this release", "should I…

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files, including reference files (for example `evals/HARNESS.md`, `evals/cases.yaml` and `evals/fixtures/gate/opik-release-policy.yaml`). Compatibility notes: Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a test suite with a…

It sits in Testing & QA, covering MCP servers, Test generation and Failing and flaky tests. It works with Model Context Protocol. The repository describes itself as: Model Context Protocol (MCP) server for Opik, the open-source LLM observability and evaluation platform, built by Comet. Read traces, log scores, and manage prompts from Claude… The licence is Apache-2.0.

When your agent uses it

  • Is this safe to ship
  • Can I merge this
  • Go/no-go on this change
  • Gate this release

Example prompts

  • “is this safe to ship”
  • “can I merge this”
  • “go/no-go on this change”
  • “/opik-verify”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a test suite with a baseline and a candidate experiment (from the compare skill). Install the `opik` skill alongside this one — it holds the shared test-suite and experiment references; without it, this skill falls back to the public docs.
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, Write

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Load the policy
  2. Resolve the two runs
  3. Read both runs, item by item
  4. Evaluate every criterion, in this order
  5. Decide
  6. Report, and record only on request

What it can do on your machine

Read from SKILL.md and the folder at commit e0c2057. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • comet.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a test suite with a baseline and a candidate experiment (from the compare skill). Install the `opik` skill alongside this one — it holds the shared test-suite and experiment references; without it, this skill falls back to the public docs.

    From compatibility in the SKILL.md frontmatter.

Context cost

Opik Verify loads about 2.8k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 163 tokens; SKILL.md has 1,384 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~163
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from comet-ml/opik-mcp at commit e0c2057, republished under its Apache-2.0 licence (© comet-ml). 1,384 words, ~2,759 tokens.

Download SKILL.mdSave it as .claude/skills/opik-verify/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
opik-verify
description
Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost budgets, flakiness, evidence size, and whether the judge is validated. Reads two experiments on an Opik test suite via the SDK (or the MCP when connected) and returns a verdict with every criterion shown pass/fail. Use for "is this safe to ship", "can I merge this", "go/no-go on this change", "gate this release", "should I roll this out". Not for producing the numbers (use compare), building an evaluation (use evaluate), or deploying anything.
allowed-tools
Read, Grep, Glob, Bash, Write
compatibility
Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a test suite with a baseline and a candidate experiment (from the compare skill). Install the `opik` skill alongside this one — it holds the shared test-suite and experiment references; without it, this skill falls back to the public docs.
metadata.last_updated
2026-09-17
metadata.source_commit
2.0.0
metadata.argument-hint
[suite, or baseline and candidate experiment ids; optional --policy path]

Verify — Ship or Hold, Against a Policy You Can Read

Definition of done: one verdict — ship, hold, needs_review, or insufficient_evidence — computed from a declared policy over the baseline-vs-candidate numbers, with every criterion listed with its threshold, the observed value, and pass/fail, the cases behind any failure named, and the compare-view link. The policy is either the repo's opik-release-policy.yaml or the documented defaults, and the report says which. If the two runs can't be read or aren't comparable, stop at the first genuine blocker and return exactly one next step. "Looks good to me" is not a verdict; a verdict without its criteria is not one either.

Operate: apply the policy mechanically, show your arithmetic, refuse to ship on a judge nobody validated, and change no application code. The only file this skill may write is the policy file, and only when the user says so. It never deploys.

Inputs

The entry point is /opik-verify right after /opik-compare (its baseline and candidate), /opik-verify <suite> (the two most recent runs on the suite), or /opik-verify <baseline-id> <candidate-id>. Infer the rest; treat these as optional overrides:

  • policy (default: opik-release-policy.yaml at the repo root or under .opik/, else the defaults below) · which experiments (default: as above) · --record (default: off — write the verdict into the candidate experiment's config).

Ask only at a genuine, non-inferable blocker (see Blockers).

The policy

Every key is optional; missing keys take the defaults listed in references/policy.md. Say in the report which source applied.

judge_validated: false is the human-review gate: until someone has confirmed the judge agrees with people (/opik-evaluate's validate-evaluator reference), a passing run yields needs_review, not ship. Flip it to true in the file once that is done — deliberately a human edit, never something this skill sets on its own.

Activation — the only in-scope work

1. Load the policy

Look for opik-release-policy.yaml at the repo root, then .opik/. Parse it; unknown keys → Blocker (name the key). No file → defaults, and say so. Never invent thresholds not in the file or the defaults.

2. Resolve the two runs

Take them from /opik-compare's output when it just ran. Otherwise, the SDK read in references/sdk-reads.md (Resolve the two runs). Skip a failed-judge run (a run whose judge had no credential is not a candidate — /opik-compare explains how it happens). scoring_failed does not survive the read path; the read-back signal is: every item failed and every assertion reason mentions a missing credential or an LLM infrastructure error. Say which run you skipped and why. When the hosted MCP is connected, list('experiment', name=…) shows each run's averages and pass rate to pick from; the item-level read below stays on the SDK.

3. Read both runs, item by item

An experiment holds one item per run: with runs_per_item: 3 a dataset item appears three times, same dataset_item_id, different trace_id. Group — a dict keyed on dataset_item_id silently keeps one run and loses the counts. The SDK read, grouped per item, with the pass thresholds taken from the suite: references/sdk-reads.md (Read both runs). Comparability first: same dataset_version_id, same item set, same judge model (experiment config). Different → Blocker ("rerun the candidate on suite version X with judge Y, then /opik-verify") — a verdict on non-comparable runs is not a verdict.

4. Evaluate every criterion, in this order

Compute all of them even after the first failure — the report shows the whole table.

  1. Evidence size — scored items (items with assertions) ≥ min_items. Below → the verdict is insufficient_evidence regardless of the rest, unless a gate criterion (2–7) also failed — then it is hold: a known safety regression outranks thin evidence. Still report every criterion.
  2. Regressions — items passed in baseline and not in candidate. An item is flaky when runs_passed is strictly between 0 and runs_total in either run (the counts from step 3), or it flips between two runs of the same code if you have them. Under flaky_policy: exclude a flaky item is dropped from the regression count and listed separately with its counts; under count it stays in. When every item has runs_total == 1, flakiness is not observable — report the flaky check as not_evaluated, state that exclude excluded nothing, and suggest runs_per_item: 3 on the suite if the user wants the protection. Count ≤ max_regressions.
  3. Safety — any regression whose data.tags intersects safety_tags → fail, no exceptions, no exclusions.
  4. Pass rate — candidate pass_rate vs baseline, or vs the number given.
  5. Subgroups — when subgroup_key is set, pass rate per value of that key must not fall.
  6. Latency — candidate p90 duration ≤ baseline p90 × (1 + latency_p90_max_increase), from the experiments' duration percentiles (p50/p90/p99 on the experiment record) (or per-item duration from the REST experiment items).
  7. Cost — candidate mean total_estimated_cost per item ≤ baseline × (1 + cost_per_item_max_increase). Skip and say "no cost data" when neither run carries costs. Aggregates lag. Right after a run finishes, the experiment record's duration can read 0.0 and total_estimated_cost_avg None for a few seconds while the backend aggregates (observed). A zero or missing aggregate on one side is not data — re-read after a short wait, or compute p90 and mean cost from the per-item duration / total_estimated_cost fields on the REST experiment items; never let a 0.0 pass or fail the gate.
  8. Evidence strength — a paired sign test on the flips: with f fixes and r regressions, the two-sided binomial p-value under 50/50. Report it; it is not a gate. With f + r < 6 say "too few flips to call it more than noise".
  9. Judge — judge_validated from the policy. False → cap the verdict at needs_review.
  10. Attribution — flips whose reason reads as judge hesitation on an unchanged output (see /opik-compare step 5.6) are listed for the human under needs_review, never silently counted either way.
Show full SKILL.md (447 more words)Show less
5. Decide

Precedence, top to bottom — the first line that applies wins:

  • Any of criteria 2–7 failed → hold (even when criterion 1 also failed).
  • Criterion 1 failed → insufficient_evidence.
  • All gates pass but judge_validated: false, or attribution flagged items → needs_review, naming exactly what a person should look at.
  • Otherwise → ship.

Never round a hold up because the deltas are "mostly positive"; never round a ship down because of a hunch. The policy is the judgment; changing it is the user's move.

6. Report, and record only on request

The table (criterion · threshold · observed · pass/fail), the regressions named with their assertion and trace link, the compare URL with both ids, the policy source, and one next step. With --record, write the verdict into the candidate experiment's config — read the existing config first and merge, update_experiment replaces it: references/sdk-reads.md (Record the verdict). Offer — do not do — writing opik-release-policy.yaml with the defaults when no file existed, so the next verdict is reproducible.

Blockers

Stop at the earliest blocker and return exactly one next step:

  • "Run opik configure, then rerun /opik-verify."
  • "Suite <name> has fewer than two comparable runs — run /opik-compare <suite> first."
  • "Baseline and candidate are on different suite versions (v3 vs v4) — rerun the candidate on v3, or re-baseline on v4, then /opik-verify."
  • "opik-release-policy.yaml has an unknown key <key> — fix or remove it."
  • "The candidate run's judge failed (every item scoring_failed) — set the judge's provider key and rerun /opik-compare."

Output

User-facing: the verdict in one line, the criteria table, the regressions (case, assertion, why, link), the policy source, the compare link, and the single next step. Not a narrative, not JSON.

Underneath (for composition / evals), one shape, with its invariants: references/output-shape.md.

Examples

Worked runs (ship, hold, needs review, insufficient evidence): references/examples.md.

Anti-patterns

A verdict without the criteria table; thresholds pulled from thin air rather than the file or the defaults; shipping on an unvalidated judge; treating a flaky item as a regression (or a regression as flaky) without the run data to say so; comparing runs on different suite versions; averaging away a safety regression; a p-value presented as a gate on six flips; rounding hold to ship because the aggregate went up; writing the policy file or recording the verdict without being asked; editing application code; deploying.

References

Test-suite and experiment detail live in the opik skill, installed beside this one — paths relative to this file: ../opik/references/evaluation-test-suites.md (execution policies, runs_passed/runs_total, versions, get_test_suite_experiments), ../opik/references/evaluation-datasets.md (experiments, OQL). The numbers this skill judges come from ../opik-compare/SKILL.md; judge validation is ../opik-evaluate/references/validate-evaluator.md. If your host lays skills out differently, locate the opik skill's references/ directory.

If the opik skill isn't installed, say so in the report and use https://www.comet.com/docs/opik/ rather than working from memory.

© comet-ml, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (references) in src/opik_mcp/skills/opik-verify of comet-ml/opik-mcp.

  • SKILL.md
  • evals/.gitignore
  • evals/HARNESS.md
  • evals/cases.yaml
  • evals/fixtures/gate/opik-release-policy.yaml
  • evals/fixtures/gate/pyproject.toml
  • evals/fixtures/gate/seed.py
  • evals/grader.py
  • evals/metrics.py
  • evals/run_evals.py
  • references/examples.md
  • references/output-shape.md
  • references/policy.md
  • references/sdk-reads.md

Open the folder on GitHubat commit e0c2057

Compare with similar skills

Opik Verify next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Opik Verify compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Opik Verify this skillcomet-ml/opik-mcp219—~2.8kAutomated safety check: NotesApache-2.0
Ue Test AuthoringJasonMa0012/MooaToon749—~2.1kAutomated safety check: NotesCustom licence
Azsdk Common Pipeline AnalysisAzure/azure-sdk-tools134—~1.2kAutomated safety check: PassMIT
Chatgpt App Submissionnteract/semiotic2.7k—~2.8kAutomated safety check: PassApache-2.0
Create Test Runjeremylongshore/tons-of-skills-marketplace2.8k—~3kAutomated safety check: PassMIT
Zizkadb TestZIZKA-AI-SL/ZizkaDB123—~358Automated safety check: PassCustom licence

Similar skills

  • Ue Test Authoring

    JasonMa0012/MooaToon

    A skill your agent uses when writing or modifying UE automated tests (Automation, CQTest, Functional, Gauntlet, LowLevel) with Rider MCP available.

    749 GitHub stars~2.1k tokensUpdated 20 days ago
    Testing & QAAuto-check: notes
  • Azsdk Common Pipeline Analysis

    Azure/azure-sdk-tools

    Official

    Analyze Azure SDK CI/CD pipeline failures into a structured diagnosis, and define the required output format.

    134 GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Chatgpt App Submission

    nteract/semiotic

    Inspect a ChatGPT Apps MCP server codebase and generate chatgpt-app-submission.json with app info suggestions, tool hint justifications, test cases, and negative test cases, then report review-check…

    2.7k GitHub stars~2.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Create Test Run

    jeremylongshore/tons-of-skills-marketplace

    Create a Kobiton test run from a test case or suite, then offer to monitor it.

    2.8k GitHub stars~3k tokensUpdated today
    Testing & QAAuto-check passed
  • Zizkadb Test

    ZIZKA-AI-SL/ZizkaDB

    Run the full ZizkaDB test suite across all layers — lint, Python unit tests, SDK tests, MCP tests, TypeScript tests, and dashboard build verification.

    123 GitHub stars~358 tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Testrail

    alirezarezvani/claude-skills

    Sync tests with TestRail. An agent skill from alirezarezvani/claude-skills.

    28k GitHub starsUsed in 1 repo~966 tokens
    Testing & QAAuto-check passed

More from comet-ml/opik-mcp

All 10 skills in this repo
  • Opik

    comet-ml/opik-mcp

    Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST).

    219 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Opik Compare

    comet-ml/opik-mcp

    Run a candidate against the baseline over an Opik test suite and read the numbers back — which cases broke, which got fixed, the per-metric deltas, worst rows, and whether the two runs are…

    219 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Opik Diagnose

    comet-ml/opik-mcp

    Surface the Opik traces worth a developer's attention, ranked by signal — Diagnostics issues first, then errors, failed tool calls, latency, regressions, and low online-eval scores.

    219 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes
  • Opik Evaluate

    comet-ml/opik-mcp

    Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link.

    219 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check: notes
  • Opik Instrument

    comet-ml/opik-mcp

    Add Opik tracing to an existing app and verify a real trace lands.

    219 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes
  • Opik Optimize

    comet-ml/opik-mcp

    Improve a prompt with the Opik Agent Optimizer — resolve the prompt, a dataset, and a metric, pick the algorithm, run a bounded optimization, check the gain on held-out data, and save the winner as…

    219 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes

Questions about Opik Verify

What does Opik Verify do?

Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost…. Opik Verify is an agent skill from comet-ml/opik-mcp. Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost budgets, flakiness, evidence size, and whether the judge is validated.

When should I use Opik Verify?

Opik Verify fits situations like: is this safe to ship; can I merge this; go/no-go on this change; gate this release.

How do I install Opik Verify in Claude Code?

Run `npx skills add comet-ml/opik-mcp --skill opik-verify -a claude-code`. Or copy the skill folder (src/opik_mcp/skills/opik-verify in comet-ml/opik-mcp) into .claude/skills/opik-verify in your project. Claude Code loads it when a task matches its description.

How do I install Opik Verify in Codex?

Run `npx skills add comet-ml/opik-mcp --skill opik-verify -a codex`. Or copy the skill folder (src/opik_mcp/skills/opik-verify in comet-ml/opik-mcp) into .agents/skills/opik-verify in your project. Codex loads it when a task matches its description.

Can I use Opik Verify in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add comet-ml/opik-mcp --skill opik-verify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/opik-verify, .gemini/skills/opik-verify, .github/skills/opik-verify and .opencode/skills/opik-verify in your project.

What does Opik Verify need to run?

Going by SKILL.md and its folder, Opik Verify needs Python for the scripts in its folder. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, Write. Compatibility (from SKILL.md): Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a test suite with a baseline and a candidate experiment (from the compare skill). Install the `opik` skill alongside this one — it holds the shared test-suite and experiment references; without it, this skill falls back to the public docs..

Does Opik Verify access the network?

SKILL.md names 1 domain. As links in the text: comet.com. This is read from the text; nothing was executed.

Is Opik Verify safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Opik Verify use?

Opik Verify is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Opik Verify use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.

What are the alternatives to Opik Verify?

Skills that share tags, products or a category with Opik Verify: Ue Test Authoring (JasonMa0012/MooaToon, 749 stars), Azsdk Common Pipeline Analysis (Azure/azure-sdk-tools, 134 stars), Chatgpt App Submission (nteract/semiotic, 2.7k stars) and Create Test Run (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Opik Verify?

comet-ml (a GitHub organization) maintains it in comet-ml/opik-mcp, which has 219 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 7, 2026.

Source: comet-ml/opik-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.