Agent skill

Opik Online Eval

by comet-ml in comet-ml/opik-mcp

Take a judge live on production traffic — create an Opik online evaluation rule (LLM-as-judge or Python metric) on a project with sampling, filters, variable mapping, and a cost cap, then confirm…

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Opik Online Eval

skills CLI
$ npx skills add comet-ml/opik-mcp --skill opik-online-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install comet-ml/opik-mcp opik-online-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/opik_mcp/skills/opik-online-eval .claude/skills/opik-online-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
opik-online-eval
GitHub stars
220
Token cost
~3k tokens
SKILL.md length
1,143 words
Files
10
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Take a judge live on production traffic — create an Opik online evaluation rule (LLM-as-judge or Python metric) on a project with sampling, filters, variable mapping, and a cost cap, then confirm…

  • Works in 6 steps: Resolve project, judge, and scope → Look at the traces before mapping… → Check what already runs → …
  • Score production traces
  • SKILL.md covers Inputs, Activation — the only in-scope…, Blockers and Output, plus 3 more sections
  • Runs Python scripts from its folder; needs OPIK_API_KEY

What it does

Opik Online Eval is an agent skill from comet-ml/opik-mcp. Take a judge live on production traffic — create an Opik online evaluation rule (LLM-as-judge or Python metric) on a project with sampling, filters, variable mapping, and a cost cap, then confirm new traces are being scored. Works over the SDK's REST client; reads rules and score names via the MCP when connected. Returns the rule, the score name it emits, and how to watch it. Use for "score production traces", "monitor hallucinations in prod", "take this judge live", "set up an online evaluation rule", "alert me…

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files (for example `evals/HARNESS.md`, `evals/cases.yaml` and `evals/fixtures/traffic/seed.py`). Compatibility notes: Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a project that receives…

It sits in AI & LLM Engineering, covering MCP servers, Third-party API integration and LLM evaluation. It works with Model Context Protocol and Python. The repository describes itself as: Model Context Protocol (MCP) server for Opik, the open-source LLM observability and evaluation platform, built by Comet. Read traces, log scores, and manage prompts from Claude… The licence is Apache-2.0.

When your agent uses it

  • Score production traces
  • Monitor hallucinations in prod
  • Take this judge live
  • Set up an online evaluation rule

Example prompts

  • “score production traces”
  • “monitor hallucinations in prod”
  • “take this judge live”
  • “/opik-online-eval”

Requirements

  • Python 3
  • A credential in OPIK_API_KEY
  • Compatibility (from SKILL.md): Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a project that receives traces. Install the `opik` skill alongside this one — it holds the shared production and observability references; without it, this skill falls back to the public docs.
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Resolve project, judge, and scope
  2. Look at the traces before mapping variables
  3. Check what already runs
  4. Create the rule
  5. Verify on a real trace
  6. Report

What it can do on your machine

Read from SKILL.md and the folder at commit e737581. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • comet.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPIK_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a project that receives traces. Install the `opik` skill alongside this one — it holds the shared production and observability references; without it, this skill falls back to the public docs.

    From compatibility in the SKILL.md frontmatter.

Context cost

Opik Online Eval loads about 3k tokens when it runs. Until then it costs about 166 tokens; SKILL.md has 1,143 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~166
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from comet-ml/opik-mcp at commit e737581, republished under its Apache-2.0 licence (© comet-ml). 1,143 words, ~2,983 tokens.

Download SKILL.mdSave it as .claude/skills/opik-online-eval/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
opik-online-eval
description
Take a judge live on production traffic — create an Opik online evaluation rule (LLM-as-judge or Python metric) on a project with sampling, filters, variable mapping, and a cost cap, then confirm new traces are being scored. Works over the SDK's REST client; reads rules and score names via the MCP when connected. Returns the rule, the score name it emits, and how to watch it. Use for "score production traces", "monitor hallucinations in prod", "take this judge live", "set up an online evaluation rule", "alert me when quality drops". Not for offline experiments (use evaluate or compare) or for finding what is already broken (use diagnose).
allowed-tools
Read, Grep, Glob, Bash
compatibility
Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a project that receives traces. Install the `opik` skill alongside this one — it holds the shared production and observability references; without it, this skill falls back to the public docs.
metadata.last_updated
2026-09-15
metadata.source_commit
2.0.0
metadata.argument-hint
[what to score, or a judge from the evaluate skill; optional project]

Online Eval — Take a Judge Live

Definition of done: an online evaluation rule exists on the project, is enabled, and has scored at least one new trace — confirmed by reading a fresh trace's feedback score, not by the create call returning. The rule scores one failure mode, samples at a rate the project's volume can afford, carries a cost cap, and maps its variables to the fields the traces actually have. If traffic hasn't arrived yet, the done state is "rule live, unverified — watch this filter". If the rule can't be created, stop at the first genuine blocker and return exactly one next step.

Operate: one failure mode per rule, sample before you scale, cap the spend, verify on a real trace — and change no application code. This skill writes to Opik only.

Inputs

The entry point is /opik-online-eval <what to score> ("hallucinations", "the refund-window answer", "refusals"), /opik-online-eval right after /opik-evaluate validated a judge (take that judge live), or /opik-online-eval <project>. Infer the rest; treat these as optional overrides:

  • project (default: configured) · scope (default: trace; span for one model call, thread for whole conversations) · sampling rate (default: 1.0 under ~1k traces/day, else 0.1–0.2) · filters (default: none; typical: an environment or tag filter) · judge model · cost cap (default: set one) · variable mapping (default: input → input, output → output).

Ask only at a genuine, non-inferable blocker (see Blockers).

Activation — the only in-scope work

1. Resolve project, judge, and scope

Confirm Opik is reachable (~/.opik.config or OPIK_API_KEY; otherwise → Blocker). Resolve the project id from a trace or by name. Decide the judge:

  • A judge /opik-evaluate already validated → reuse its prompt and output schema verbatim.
  • A named metric (hallucination, answer relevance, moderation) → a minimal binary judge for that one failure mode (../opik-evaluate/references/write-judge-prompt.md).
  • A mechanical check (JSON valid, contains a disclaimer, latency budget) → a Python metric rule, not a judge.

Scope: trace by default; span when the check is about one LLM/tool call; thread when the check needs the whole conversation (thread rules wait for the thread to go inactive — 15 min by default).

2. Look at the traces before mapping variables

Read three recent production traces and note the real shape of input and output. Experiment runs land in the same project with a different input shape (their metadata carries test_suite_experiment_id), so skip those — filtering on the app's entrypoint is the reliable way: client.search_traces(project_name=…, max_results=3, filter_string='name = "<entrypoint>"'). (OQL has no is_empty for metadata.* keys.) Variables are plain field paths, dot-notation for nested keys (output.answer, input.messages) — never {{ }} templates. A wrong path is the most common reason a rule silently scores nothing.

3. Check what already runs
python
import opik

client = opik.Opik()
existing = client.rest_client.automation_rule_evaluators.find_evaluators(project_id="<project_id>")

Same name or same failure mode already there → don't create a second one; report exists (offer to adjust sampling/enable). When the hosted MCP is connected, list('online_rule', project_id=…) and list('score_name', project_id=…) show the same, with each rule's type, enabled flag, and sampling rate.

4. Create the rule

There is no high-level SDK wrapper; use the REST client. LLM-as-judge, trace scope:

python
from opik.rest_api.types import (
    AutomationRuleEvaluatorWrite_LlmAsJudge,
    LlmAsJudgeCodeWrite,
    LlmAsJudgeModelParametersWrite,
    LlmAsJudgeMessageWrite,
    LlmAsJudgeOutputSchemaWrite,
)

rule = AutomationRuleEvaluatorWrite_LlmAsJudge(
    action="evaluator",  # required literal; the model rejects the payload without it
    name="refund_window_correct",  # becomes the feedback-score name on every scored trace — use underscores, not hyphens: OQL parses `feedback_scores.a-b` as an operator
    project_ids=["<project_id>"],
    sampling_rate=0.2,  # fraction of SDK-logged traces scored
    enabled=True,
    filters=[],  # e.g. [{"field": "tags", "operator": "contains", "value": "production"}]
    code=LlmAsJudgeCodeWrite(
        model=LlmAsJudgeModelParametersWrite(name="<judge model>", temperature=0.0),
        messages=[
            LlmAsJudgeMessageWrite(
                role="USER", content="<the validated judge prompt using {{input}} and {{output}}>"
            )
        ],
        variables={"input": "input", "output": "output"},  # field paths from step 2
        schema_=[
            LlmAsJudgeOutputSchemaWrite(
                name="refund_window_correct",
                type="BOOLEAN",
                description="True if the response states 5-7 business days",
            )
        ],
        max_cost_usd=5.0,  # per-rule spend cap — set it
    ),
)
created = client.rest_client.automation_rule_evaluators.create_automation_rule_evaluator(
    request=rule
)

Notes: on Opik Cloud without your own provider key, model.name="opik-free-model" uses the workspace's built-in free provider; the Python attribute is schema_ (wire name schema); sampling_rate applies to production traces only (experiment traces are always scored in full); trigger_scope defaults to production. Span and thread variants: AutomationRuleEvaluatorWrite_SpanLlmAsJudge, AutomationRuleEvaluatorWrite_TraceThreadLlmAsJudge. Python metric: AutomationRuleEvaluatorWrite_UserDefinedMetricPython with code={"metric": "<python source defining a BaseMetric>", "arguments": {"output": "output"}}.

Endpoint, if scripting outside Python: POST /v1/private/automations/evaluators/ with the same body.

5. Verify on a real trace
python
import time

for _ in range(12):  # ~2 min
    scored = client.search_traces(
        project_name="<project>",
        max_results=1,
        filter_string="feedback_scores.refund_window_correct is_not_empty",
    )
    if scored:
        break
    time.sleep(10)

If the score name already contains a hyphen (an existing rule), double-quote the key or the OQL parser reads the hyphen as an operator: filter_string='feedback_scores."refund-window-correct" is_not_empty'. A scored trace → live. None, and the project had no new traces in the window → live_unverified with the filter to watch. None, but traces did arrive → read the rule's logs (get_evaluator_logs_by_id(id)) — a variable-path error or model failure shows there; fix and re-verify.

6. Report

Rule name/id, scope, sampling, cost cap, the score name, the verification trace link, and one next step (see Output). Natural next steps: /opik-diagnose will now surface low scores on this name; an alert on the score threshold; or, if the judge wasn't validated first, "validate it against 20 human labels (/opik-evaluate) before anyone acts on it".

Show full SKILL.md (451 more words)Show less

Blockers

Stop at the earliest blocker and return exactly one next step:

  • "Run opik configure, then rerun /opik-online-eval."
  • "Which project should this score? Pass /opik-online-eval <what> <project>."
  • "The traces' output is {answer, sources} — should the judge read output.answer? (I'll map it that way unless you say otherwise.)" — ask only when the mapping is genuinely ambiguous.
  • "Which failure mode should the rule catch? Name one (e.g. hallucination, wrong refund window, unsafe content)."

Output

User-facing: a short human message — the rule (name, scope, sampling, cap), the score name, the verification trace as a clickable Opik UI link (or the watch filter if unverified), and the single next step. Not JSON.

Underneath (for composition / evals), one shape:

  • status: live | live_unverified | exists | blocked
  • rule: id, name, type (llm_as_judge | user_defined_metric_python | span/thread variants), scope, sampling_rate, filters, max_cost_usd, enabled
  • score_name
  • variables: the field-path mapping used
  • verification: trace_id, trace_url, value (when live); watch_filter (when unverified)
  • source: sdk | mcp
  • next_step: exactly one

Invariants: live carries a verification.trace_id; every created rule has a max_cost_usd and a sampling_rate; one failure mode per rule; exists created nothing; blocked carries exactly one next_step; every path leaves the codebase unchanged.

Examples

Validated judge goes live. /opik-evaluate calibrated "states the refund window as 5–7 business days" (TPR 0.95). /opik-online-eval: project support-bot, ~5k traces/day → sampling 0.1, filter tags contains "production", cap $5, mapping output → output.answer (traces nest it). Created; 40 s later a trace carries refund-window-correct = 1. → live; next step = "alert when the 1h mean drops below 0.8".

Mechanical check. "Make sure every prod answer is valid JSON." → Python metric rule (IsJson on output), sampling 1.0 (cheap), no judge. → live.

No traffic yet. Rule created on a project that receives traces only in business hours. → live_unverified: "watch feedback_scores.<name> is_not_empty — or send one traced request".

Already there. A rule named hallucination exists, enabled, at 0.5. → exists; next step = "lower sampling to 0.1 if cost is the concern".

Anti-patterns

A rule at sampling_rate: 1.0 on a high-volume project without saying what it costs; no max_cost_usd; {{input}}-style template syntax in variables (they are field paths); a holistic "quality" judge as a rule; a judge for a mechanical check; a second rule for the same failure mode; declaring success from the create call without a scored trace; taking an unvalidated judge live and calling its scores truth; editing application code.

References

Production and observability detail live in the opik skill, installed beside this one — paths relative to this file: ../opik/references/production.md (online evaluation, variable mapping, feedback scores, alerts), ../opik/references/observability.md (trace/span/thread model), ../opik/references/evaluation-datasets.md (OQL filters, feedback_scores.<name> operators). Judge design: ../opik-evaluate/references/write-judge-prompt.md, ../opik-evaluate/references/validate-evaluator.md. If your host lays skills out differently, locate the opik skill's references/ directory.

If the opik skill isn't installed, say so in the report and use https://www.comet.com/docs/opik/ rather than working from memory.

© comet-ml, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files in src/opik_mcp/skills/opik-online-eval of comet-ml/opik-mcp.

  • SKILL.md
  • evals/.gitignore
  • evals/HARNESS.md
  • evals/cases.yaml
  • evals/fixtures/traffic/pyproject.toml
  • evals/fixtures/traffic/seed.py
  • evals/fixtures/traffic/traffic.py
  • evals/grader.py
  • evals/metrics.py
  • evals/run_evals.py

Open the folder on GitHubat commit e737581

Compare with similar skills

Opik Online Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Opik Online Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Opik Online Eval this skillcomet-ml/opik-mcp220—~3kAutomated safety check: NotesApache-2.0
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
MCP Server BuildershareAI-lab/learn-claude-code78k5 repos~1.2kAutomated safety check: PassMIT
LexGuard MCP Developer GuideSeoNaRu/lexguard-mcp131—~1.1kAutomated safety check: PassCustom licence
Agent Observability Eval Bootstrapdatadog-labs/agent-skills177—~25kAutomated safety check: PassMIT
Tiger Brokers OpenAPI SDKqusong0627/QuantMind1.7k—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • MCP Server Builder

    shareAI-lab/learn-claude-code

    Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.

    78k GitHub starsUsed in 5 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • LexGuard MCP Developer Guide

    SeoNaRu/lexguard-mcp

    Developer guide for the LexGuard Korean law MCP server: layer rules, adding tools and repositories, JSON-RPC responses, law API handling, answer rules and tests.

    131 GitHub stars~1.1k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Agent Observability Eval Bootstrap

    datadog-labs/agent-skills

    Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit…

    177 GitHub stars~25k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Tiger Brokers OpenAPI SDK

    qusong0627/QuantMind

    Covers the Tiger Brokers OpenAPI Python SDK for market data, stock, futures and options trading, push subscriptions, a CLI and an MCP server, defaulting to paper trading.

    1.7k GitHub stars~1.4k tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Strands

    strands-agents/harness-sdk

    Build, extend, evaluate, or migrate applications with Strands Agents in Python or TypeScript.

    8.7k GitHub stars~1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from comet-ml/opik-mcp

All 10 skills in this repo
  • Opik

    comet-ml/opik-mcp

    Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST).

    220 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Opik Compare

    comet-ml/opik-mcp

    Run a candidate against the baseline over an Opik test suite and read the numbers back — which cases broke, which got fixed, the per-metric deltas, worst rows, and whether the two runs are…

    220 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Opik Diagnose

    comet-ml/opik-mcp

    Surface the Opik traces worth a developer's attention, ranked by signal — Diagnostics issues first, then errors, failed tool calls, latency, regressions, and low online-eval scores.

    220 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes
  • Opik Evaluate

    comet-ml/opik-mcp

    Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link.

    220 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check: notes
  • Opik Instrument

    comet-ml/opik-mcp

    Add Opik tracing to an existing app and verify a real trace lands.

    220 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes
  • Opik Optimize

    comet-ml/opik-mcp

    Improve a prompt with the Opik Agent Optimizer — resolve the prompt, a dataset, and a metric, pick the algorithm, run a bounded optimization, check the gain on held-out data, and save the winner as…

    220 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes

Questions about Opik Online Eval

What does Opik Online Eval do?

Take a judge live on production traffic — create an Opik online evaluation rule (LLM-as-judge or Python metric) on a project with sampling, filters, variable mapping, and a cost cap, then confirm…. Opik Online Eval is an agent skill from comet-ml/opik-mcp. Take a judge live on production traffic — create an Opik online evaluation rule (LLM-as-judge or Python metric) on a project with sampling, filters, variable mapping, and a cost cap, then confirm new traces are being scored.

When should I use Opik Online Eval?

Opik Online Eval fits situations like: score production traces; monitor hallucinations in prod; take this judge live; set up an online evaluation rule.

How do I install Opik Online Eval in Claude Code?

Run `npx skills add comet-ml/opik-mcp --skill opik-online-eval -a claude-code`. Or copy the skill folder (src/opik_mcp/skills/opik-online-eval in comet-ml/opik-mcp) into .claude/skills/opik-online-eval in your project. Claude Code loads it when a task matches its description.

How do I install Opik Online Eval in Codex?

Run `npx skills add comet-ml/opik-mcp --skill opik-online-eval -a codex`. Or copy the skill folder (src/opik_mcp/skills/opik-online-eval in comet-ml/opik-mcp) into .agents/skills/opik-online-eval in your project. Codex loads it when a task matches its description.

Can I use Opik Online Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add comet-ml/opik-mcp --skill opik-online-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/opik-online-eval, .gemini/skills/opik-online-eval, .github/skills/opik-online-eval and .opencode/skills/opik-online-eval in your project.

What does Opik Online Eval need to run?

Going by SKILL.md and its folder, Opik Online Eval needs Python for the scripts in its folder and credentials named OPIK_API_KEY. Our summary lists: Python 3; A credential in OPIK_API_KEY. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash. Compatibility (from SKILL.md): Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a project that receives traces. Install the `opik` skill alongside this one — it holds the shared production and observability references; without it, this skill falls back to the public docs..

Does Opik Online Eval access the network?

SKILL.md names 1 domain. As links in the text: comet.com. This is read from the text; nothing was executed.

Is Opik Online Eval safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Opik Online Eval use?

Opik Online Eval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Opik Online Eval use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Opik Online Eval?

Skills that share tags, products or a category with Opik Online Eval: MCP Server Builder (anthropics/skills, 180k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars), LexGuard MCP Developer Guide (SeoNaRu/lexguard-mcp, 131 stars) and Agent Observability Eval Bootstrap (datadog-labs/agent-skills, 177 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Opik Online Eval?

comet-ml (a GitHub organization) maintains it in comet-ml/opik-mcp, which has 220 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 8, 2026.

Source: comet-ml/opik-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.