Ue Test Authoring
JasonMa0012/MooaToon
A skill your agent uses when writing or modifying UE automated tests (Automation, CQTest, Functional, Gauntlet, LowLevel) with Rider MCP available.
Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost…
$ npx skills add comet-ml/opik-mcp --skill opik-verify -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install comet-ml/opik-mcp opik-verify --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/opik_mcp/skills/opik-verify .claude/skills/opik-verify && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "opik-verify" agent skill from https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opik-verify into .claude/skills/opik-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opik-verify", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opik-verifyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add comet-ml/opik-mcp --skill opik-verify -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install comet-ml/opik-mcp opik-verify --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .agents/skills && cp -r skills-src/src/opik_mcp/skills/opik-verify .agents/skills/opik-verify && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "opik-verify" agent skill from https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opik-verify into .agents/skills/opik-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opik-verify", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add comet-ml/opik-mcp --skill opik-verify -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install comet-ml/opik-mcp opik-verify --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/src/opik_mcp/skills/opik-verify .cursor/skills/opik-verify && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "opik-verify" agent skill from https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opik-verify into .cursor/skills/opik-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opik-verify", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/comet-ml/opik-mcp.git --path src/opik_mcp/skills/opik-verify--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add comet-ml/opik-mcp --skill opik-verify -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install comet-ml/opik-mcp opik-verify --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/src/opik_mcp/skills/opik-verify .gemini/skills/opik-verify && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "opik-verify" agent skill from https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opik-verify into .gemini/skills/opik-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opik-verify", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install comet-ml/opik-mcp opik-verifyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add comet-ml/opik-mcp --skill opik-verify -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .github/skills && cp -r skills-src/src/opik_mcp/skills/opik-verify .github/skills/opik-verify && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "opik-verify" agent skill from https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opik-verify into .github/skills/opik-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opik-verify", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add comet-ml/opik-mcp --skill opik-verify -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install comet-ml/opik-mcp opik-verify --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/src/opik_mcp/skills/opik-verify .opencode/skills/opik-verify && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "opik-verify" agent skill from https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opik-verify into .opencode/skills/opik-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opik-verify", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
opik-verifyDecide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost…
Opik Verify is an agent skill from comet-ml/opik-mcp. Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost budgets, flakiness, evidence size, and whether the judge is validated. Reads two experiments on an Opik test suite via the SDK (or the MCP when connected) and returns a verdict with every criterion shown pass/fail. Use for "is this safe to ship", "can I merge this", "go/no-go on this change", "gate this release", "should I…
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files, including reference files (for example `evals/HARNESS.md`, `evals/cases.yaml` and `evals/fixtures/gate/opik-release-policy.yaml`). Compatibility notes: Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a test suite with a…
It sits in Testing & QA, covering MCP servers, Test generation and Failing and flaky tests. It works with Model Context Protocol. The repository describes itself as: Model Context Protocol (MCP) server for Opik, the open-source LLM observability and evaluation platform, built by Comet. Read traces, log scores, and manage prompts from Claude… The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e0c2057. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGrepGlobBashWriteFrom allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
comet.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a test suite with a baseline and a candidate experiment (from the compare skill). Install the `opik` skill alongside this one — it holds the shared test-suite and experiment references; without it, this skill falls back to the public docs.
From compatibility in the SKILL.md frontmatter.
Opik Verify loads about 2.8k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 163 tokens; SKILL.md has 1,384 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, Grep, Glob, Bash, WriteAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from comet-ml/opik-mcp at commit e0c2057, republished under its Apache-2.0 licence (© comet-ml). 1,384 words, ~2,759 tokens.
.claude/skills/opik-verify/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.Definition of done: one verdict — ship, hold, needs_review, or insufficient_evidence — computed from a declared policy over the baseline-vs-candidate numbers, with every criterion listed with its threshold, the observed value, and pass/fail, the cases behind any failure named, and the compare-view link. The policy is either the repo's opik-release-policy.yaml or the documented defaults, and the report says which. If the two runs can't be read or aren't comparable, stop at the first genuine blocker and return exactly one next step. "Looks good to me" is not a verdict; a verdict without its criteria is not one either.
Operate: apply the policy mechanically, show your arithmetic, refuse to ship on a judge nobody validated, and change no application code. The only file this skill may write is the policy file, and only when the user says so. It never deploys.
The entry point is /opik-verify right after /opik-compare (its baseline and candidate), /opik-verify <suite> (the two most recent runs on the suite), or /opik-verify <baseline-id> <candidate-id>. Infer the rest; treat these as optional overrides:
opik-release-policy.yaml at the repo root or under .opik/, else the defaults below) · which experiments (default: as above) · --record (default: off — write the verdict into the candidate experiment's config).Ask only at a genuine, non-inferable blocker (see Blockers).
Every key is optional; missing keys take the defaults listed in references/policy.md. Say in the report which source applied.
judge_validated: false is the human-review gate: until someone has confirmed the judge agrees with people (/opik-evaluate's validate-evaluator reference), a passing run yields needs_review, not ship. Flip it to true in the file once that is done — deliberately a human edit, never something this skill sets on its own.
Look for opik-release-policy.yaml at the repo root, then .opik/. Parse it; unknown keys → Blocker (name the key). No file → defaults, and say so. Never invent thresholds not in the file or the defaults.
Take them from /opik-compare's output when it just ran. Otherwise, the SDK read in references/sdk-reads.md (Resolve the two runs).
Skip a failed-judge run (a run whose judge had no credential is not a candidate — /opik-compare explains how it happens). scoring_failed does not survive the read path; the read-back signal is: every item failed and every assertion reason mentions a missing credential or an LLM infrastructure error. Say which run you skipped and why. When the hosted MCP is connected, list('experiment', name=…) shows each run's averages and pass rate to pick from; the item-level read below stays on the SDK.
An experiment holds one item per run: with runs_per_item: 3 a dataset item appears three times, same dataset_item_id, different trace_id. Group — a dict keyed on dataset_item_id silently keeps one run and loses the counts.
The SDK read, grouped per item, with the pass thresholds taken from the suite: references/sdk-reads.md (Read both runs).
Comparability first: same dataset_version_id, same item set, same judge model (experiment config). Different → Blocker ("rerun the candidate on suite version X with judge Y, then /opik-verify") — a verdict on non-comparable runs is not a verdict.
Compute all of them even after the first failure — the report shows the whole table.
min_items. Below → the verdict is insufficient_evidence regardless of the rest, unless a gate criterion (2–7) also failed — then it is hold: a known safety regression outranks thin evidence. Still report every criterion.passed in baseline and not in candidate. An item is flaky when runs_passed is strictly between 0 and runs_total in either run (the counts from step 3), or it flips between two runs of the same code if you have them. Under flaky_policy: exclude a flaky item is dropped from the regression count and listed separately with its counts; under count it stays in. When every item has runs_total == 1, flakiness is not observable — report the flaky check as not_evaluated, state that exclude excluded nothing, and suggest runs_per_item: 3 on the suite if the user wants the protection. Count ≤ max_regressions.data.tags intersects safety_tags → fail, no exceptions, no exclusions.pass_rate vs baseline, or vs the number given.subgroup_key is set, pass rate per value of that key must not fall.latency_p90_max_increase), from the experiments' duration percentiles (p50/p90/p99 on the experiment record) (or per-item duration from the REST experiment items).total_estimated_cost per item ≤ baseline × (1 + cost_per_item_max_increase). Skip and say "no cost data" when neither run carries costs.
Aggregates lag. Right after a run finishes, the experiment record's duration can read 0.0 and total_estimated_cost_avg None for a few seconds while the backend aggregates (observed). A zero or missing aggregate on one side is not data — re-read after a short wait, or compute p90 and mean cost from the per-item duration / total_estimated_cost fields on the REST experiment items; never let a 0.0 pass or fail the gate.f fixes and r regressions, the two-sided binomial p-value under 50/50. Report it; it is not a gate. With f + r < 6 say "too few flips to call it more than noise".judge_validated from the policy. False → cap the verdict at needs_review.reason reads as judge hesitation on an unchanged output (see /opik-compare step 5.6) are listed for the human under needs_review, never silently counted either way.Precedence, top to bottom — the first line that applies wins:
hold (even when criterion 1 also failed).insufficient_evidence.judge_validated: false, or attribution flagged items → needs_review, naming exactly what a person should look at.ship.Never round a hold up because the deltas are "mostly positive"; never round a ship down because of a hunch. The policy is the judgment; changing it is the user's move.
The table (criterion · threshold · observed · pass/fail), the regressions named with their assertion and trace link, the compare URL with both ids, the policy source, and one next step. With --record, write the verdict into the candidate experiment's config — read the existing config first and merge, update_experiment replaces it: references/sdk-reads.md (Record the verdict).
Offer — do not do — writing opik-release-policy.yaml with the defaults when no file existed, so the next verdict is reproducible.
Stop at the earliest blocker and return exactly one next step:
opik configure, then rerun /opik-verify."<name> has fewer than two comparable runs — run /opik-compare <suite> first."/opik-verify."opik-release-policy.yaml has an unknown key <key> — fix or remove it."scoring_failed) — set the judge's provider key and rerun /opik-compare."User-facing: the verdict in one line, the criteria table, the regressions (case, assertion, why, link), the policy source, the compare link, and the single next step. Not a narrative, not JSON.
Underneath (for composition / evals), one shape, with its invariants: references/output-shape.md.
Worked runs (ship, hold, needs review, insufficient evidence): references/examples.md.
A verdict without the criteria table; thresholds pulled from thin air rather than the file or the defaults; shipping on an unvalidated judge; treating a flaky item as a regression (or a regression as flaky) without the run data to say so; comparing runs on different suite versions; averaging away a safety regression; a p-value presented as a gate on six flips; rounding hold to ship because the aggregate went up; writing the policy file or recording the verdict without being asked; editing application code; deploying.
Test-suite and experiment detail live in the opik skill, installed beside this one — paths relative to this file: ../opik/references/evaluation-test-suites.md (execution policies, runs_passed/runs_total, versions, get_test_suite_experiments), ../opik/references/evaluation-datasets.md (experiments, OQL). The numbers this skill judges come from ../opik-compare/SKILL.md; judge validation is ../opik-evaluate/references/validate-evaluator.md. If your host lays skills out differently, locate the opik skill's references/ directory.
If the opik skill isn't installed, say so in the report and use https://www.comet.com/docs/opik/ rather than working from memory.
© comet-ml, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 13 other files (references) in src/opik_mcp/skills/opik-verify of comet-ml/opik-mcp.
Open the folder on GitHubat commit e0c2057
Opik Verify next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Opik Verify this skillcomet-ml/opik-mcp | 219 | — | ~2.8k | Automated safety check: Notes | Apache-2.0 | |
| Ue Test AuthoringJasonMa0012/MooaToon | 749 | — | ~2.1k | Automated safety check: Notes | Custom licence | |
| Azsdk Common Pipeline AnalysisAzure/azure-sdk-tools | 134 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Chatgpt App Submissionnteract/semiotic | 2.7k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Create Test Runjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~3k | Automated safety check: Pass | MIT | |
| Zizkadb TestZIZKA-AI-SL/ZizkaDB | 123 | — | ~358 | Automated safety check: Pass | Custom licence |
JasonMa0012/MooaToon
A skill your agent uses when writing or modifying UE automated tests (Automation, CQTest, Functional, Gauntlet, LowLevel) with Rider MCP available.
Azure/azure-sdk-tools
Analyze Azure SDK CI/CD pipeline failures into a structured diagnosis, and define the required output format.
nteract/semiotic
Inspect a ChatGPT Apps MCP server codebase and generate chatgpt-app-submission.json with app info suggestions, tool hint justifications, test cases, and negative test cases, then report review-check…
jeremylongshore/tons-of-skills-marketplace
Create a Kobiton test run from a test case or suite, then offer to monitor it.
ZIZKA-AI-SL/ZizkaDB
Run the full ZizkaDB test suite across all layers — lint, Python unit tests, SDK tests, MCP tests, TypeScript tests, and dashboard build verification.
alirezarezvani/claude-skills
Sync tests with TestRail. An agent skill from alirezarezvani/claude-skills.
comet-ml/opik-mcp
Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST).
comet-ml/opik-mcp
Run a candidate against the baseline over an Opik test suite and read the numbers back — which cases broke, which got fixed, the per-metric deltas, worst rows, and whether the two runs are…
comet-ml/opik-mcp
Surface the Opik traces worth a developer's attention, ranked by signal — Diagnostics issues first, then errors, failed tool calls, latency, regressions, and low online-eval scores.
comet-ml/opik-mcp
Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link.
comet-ml/opik-mcp
Add Opik tracing to an existing app and verify a real trace lands.
comet-ml/opik-mcp
Improve a prompt with the Opik Agent Optimizer — resolve the prompt, a dataset, and a metric, pick the algorithm, run a bounded optimization, check the gain on held-out data, and save the winner as…
Works with
Categories
Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost…. Opik Verify is an agent skill from comet-ml/opik-mcp. Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost budgets, flakiness, evidence size, and whether the judge is validated.
Opik Verify fits situations like: is this safe to ship; can I merge this; go/no-go on this change; gate this release.
Run `npx skills add comet-ml/opik-mcp --skill opik-verify -a claude-code`. Or copy the skill folder (src/opik_mcp/skills/opik-verify in comet-ml/opik-mcp) into .claude/skills/opik-verify in your project. Claude Code loads it when a task matches its description.
Run `npx skills add comet-ml/opik-mcp --skill opik-verify -a codex`. Or copy the skill folder (src/opik_mcp/skills/opik-verify in comet-ml/opik-mcp) into .agents/skills/opik-verify in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add comet-ml/opik-mcp --skill opik-verify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/opik-verify, .gemini/skills/opik-verify, .github/skills/opik-verify and .opencode/skills/opik-verify in your project.
Going by SKILL.md and its folder, Opik Verify needs Python for the scripts in its folder. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, Write. Compatibility (from SKILL.md): Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a test suite with a baseline and a candidate experiment (from the compare skill). Install the `opik` skill alongside this one — it holds the shared test-suite and experiment references; without it, this skill falls back to the public docs..
SKILL.md names 1 domain. As links in the text: comet.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Opik Verify is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Opik Verify: Ue Test Authoring (JasonMa0012/MooaToon, 749 stars), Azsdk Common Pipeline Analysis (Azure/azure-sdk-tools, 134 stars), Chatgpt App Submission (nteract/semiotic, 2.7k stars) and Create Test Run (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
comet-ml (a GitHub organization) maintains it in comet-ml/opik-mcp, which has 219 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 7, 2026.
Source: comet-ml/opik-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.