Agent skill

Opik

by comet-ml in comet-ml/opik-mcp

Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST).

Apache-2.0Auto-check passedAI & LLM Engineering

Install Opik

skills CLI
$ npx skills add comet-ml/opik-mcp --skill opik -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install comet-ml/opik-mcp opik --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/opik_mcp/skills/opik .claude/skills/opik && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
opik
GitHub stars
219
Token cost
~2.1k tokens
SKILL.md length
635 words
Files
11 (incl. references)
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST).

  • What span types exist
  • SKILL.md covers Core concepts, Python — tracing, TypeScript — tracing and Framework integrations, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Version a prompt

What it does

Opik is an agent skill from comet-ml/opik-mcp. Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST). Use for "what span types exist", "how do I flush", "trackopenai", "add OpikTracer", "version a prompt". To instrument a repo end to end, use the opik-instrument skill.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `references/agent-patterns.md`, `references/best-practices.md` and `references/evaluation-datasets.md`). Compatibility notes: Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). A reference — needs no Opik connection to read; the…

It sits in AI & LLM Engineering, covering Prompt engineering and MCP servers. It works with Python, TypeScript, OpenAI and Model Context Protocol. The repository describes itself as: Model Context Protocol (MCP) server for Opik, the open-source LLM observability and evaluation platform, built by Comet. Read traces, log scores, and manage prompts from Claude… The licence is Apache-2.0.

When your agent uses it

  • What span types exist
  • Version a prompt

Example prompts

  • “what span types exist”
  • “how do I flush”
  • “trackopenai”
  • “/opik”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). A reference — needs no Opik connection to read; the snippets assume the `opik` Python or TypeScript SDK 2.x. The task-shaped skills (opik-instrument, opik-diagnose, opik-explain, opik-test, opik-compare, opik-evaluate, opik-online-eval, opik-optimize, opik-verify) read this skill's references and expect it installed beside them.
  • Pre-approved tools (allowed-tools): Read, Grep, Glob

What it can do on your machine

Read from SKILL.md and the folder at commit f1dd464. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). A reference — needs no Opik connection to read; the snippets assume the `opik` Python or TypeScript SDK 2.x. The task-shaped skills (opik-instrument, opik-diagnose, opik-explain, opik-test, opik-compare, opik-evaluate, opik-online-eval, opik-optimize, opik-verify) read this skill's references and expect it installed beside them.

    From compatibility in the SKILL.md frontmatter.

Context cost

Opik loads about 2.1k tokens when it runs, and up to ~34k if it reads all its reference files. Until then it costs about 77 tokens; SKILL.md has 635 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~34k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from comet-ml/opik-mcp at commit f1dd464, republished under its Apache-2.0 licence (© comet-ml). 635 words, ~2,105 tokens.

Download SKILL.mdSave it as .claude/skills/opik/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
opik
description
Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST). Use for "what span types exist", "how do I flush", "track_openai", "add OpikTracer", "version a prompt". To instrument a repo end to end, use the `opik-instrument` skill.
allowed-tools
Read, Grep, Glob
compatibility
Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). A reference — needs no Opik connection to read; the snippets assume the `opik` Python or TypeScript SDK 2.x. The task-shaped skills (opik-instrument, opik-diagnose, opik-explain, opik-test, opik-compare, opik-evaluate, opik-online-eval, opik-optimize, opik-verify) read this skill's references and expect it installed beside them.
metadata.last_updated
2026-09-17
metadata.source_commit
2.0.0

Opik SDK Reference

Opik is an open-source LLM observability platform. This skill is a reference for the SDK. To instrument a codebase step by step (detect frameworks, add config, emit and verify a trace), use the task-shaped opik-instrument skill.

Core concepts

A trace is one execution path (one request → one response). Spans are the operations inside it and form a hierarchy.

Span types — the ONLY valid values
TypeUse for
generalorchestration, agent entry points
llmmodel calls
tooltools, retrieval, API / DB calls
guardrailsafety / validation checks

Do NOT use retrieval or any other value.

Python — tracing

python
import opik


@opik.track(name="agent", type="general")
def agent(query: str) -> str:
    return generate(retrieve(query))


@opik.track(type="tool")
def retrieve(query): ...


@opik.track(type="llm")
def generate(ctx): ...


opik.flush_tracker()  # required in scripts

TypeScript — tracing

typescript
import { Opik } from "opik";
const client = new Opik({ projectName: "my-project" });

const trace = client.trace({ name: "agent", input: { query } });
const span = trace.span({ name: "llm-call", type: "llm" });
span.end({ output });
trace.end({ output });
await client.flush();

Framework integrations

Prefer an integration over manual @opik.track — integrations capture tokens, model, and cost automatically. Patterns (full list in references/integrations.md):

  • wrap-the-client — track_openai(OpenAI()), track_anthropic(...)
  • global-enable — track_crewai(crew=crew)
  • callback — dspy.configure(callbacks=[OpikCallback()])
  • tracer — OpikTracer() for LangChain / LangGraph / LlamaIndex
  • agent-specific — track_adk_agent_recursive(agent, OpikTracer())
LiteLLM inside @opik.track (common trap)

If code uses litellm and you add @opik.track, pass current_span_data via metadata on every completion call — otherwise OpikLogger emits orphaned top-level traces instead of nesting under your span.

python
from opik.opik_context import get_current_span_data


@opik.track
def call_llm(messages):
    return litellm.completion(
        model="gpt-4o",
        messages=messages,
        metadata={"opik": {"current_span_data": get_current_span_data()}},
    )

Threads (conversations)

Group turns with thread_id — one turn = one trace, shared thread_id = one thread. Use for chat / multi-turn; skip for single-shot.

python
@opik.track(entrypoint=True)
def handle(session_id: str, message: str) -> str:
    opik.update_current_trace(thread_id=session_id)
    return reply(message)

Prompt library

Version prompts with client.get_prompt / create_prompt (chat variants: get_chat_prompt / create_chat_prompt). Store model + temperature in the prompt metadata so they version with the text. Call get_prompt inside a @opik.track function so the version links to the trace.

python
@opik.track(entrypoint=True)
def run(question: str) -> str:
    p = client.get_prompt(name="system") or client.create_prompt(
        name="system",
        prompt="You help with {{product}}.",
        metadata={"model": "gpt-4o", "temperature": 0.7},
    )
    return llm(p.format(product="Opik"), model=p.metadata["model"])

How a project is doing

With the MCP connected, start at the project, not at its traces:

read("project", "<project name or id>")

One call returns the last 7 days against the 7 before — trace count, error rate, average duration, total cost, SDK traffic only, which is what the Logs page's four cards show — plus the score names and usage keys the project actually records, and the freshest experiment, dataset, prompt version and optimization run in it. since/until pick another window; since="30d" is what the UI opens on. A rate or an average over a window with no traces comes back null rather than 0, because a rate over no samples is undefined and "0% errors" is advice someone may act on.

Then attribute the change rather than restating it:

list("project_metric", project_name="<project>", metric_type="trace_cost")
list("project_metric", project_name="<project>", metric_type="span_count",
     breakdown="model", since="30d")

Rows are time buckets, not records — interval is hourly/daily/weekly/ total, and page/size/sort do not apply. schema("list.project_metric") is the metric list, what each is about, and which groupings each accepts; seven of them accept none. The score names the overview returned are the ones worth filtering on, and list("score_name", project_name=…) has the rest.

Show full SKILL.md (232 more words)Show less

Searching traces

One filter grammar, OQL, serves both the hosted MCP's list tool and the SDK's search_traces / search_spans / search_threads:

<field>[.<key>] <op> <value> [AND ...]
ops: = != > >= < <= contains not_contains starts_with ends_with is_empty is_not_empty in not_in

Strings in double quotes, numbers bare, duration in milliseconds, dates as ISO-8601 instants with a timezone ("2026-09-08T10:00:00Z"). Scores and dictionaries take a key: feedback_scores.accuracy < 0.5, metadata.environment = "prod". AND is the only connector.

error_info is_not_empty AND duration > 5000
type = "llm" AND usage.total_tokens > 10000            # spans
feedback_scores.hallucination > 0.5 AND start_time >= "2026-09-08T00:00:00Z"

With the MCP connected, prefer list — it also sorts (sort="duration desc"), windows (since="1h", "7d"), and searches free text (search="order-42"):

list(entity_type="trace", project_name="<project>", since="1h",
     filters="error_info is_not_empty", sort="duration desc")

Trace, span and thread lists add source = "sdk" unless you name source, so evaluator, playground and experiment traces stay out of the way. A rejected filter comes back with what fixes it; schema("list.trace") (or list.span, list.thread, list.experiment) is the full field and operator reference.

Without the MCP, the same string goes to the SDK:

python
client.search_traces(project_name="<project>", filter_string="error_info is_not_empty")

Anti-patterns

Anti-patternFix
span type retrieval / customuse tool (or general)
get_prompt outside @opik.trackfetch inside — else no trace link
deprecated opik.Prompt / opik.Configuse client.get_prompt / config file
litellm without current_span_datapass it — else orphaned traces
no flush in scriptsopik.flush_tracker() / await client.flush()

References

TopicFile
Python SDK (async, distributed, context)references/tracing-python.md
TypeScript SDKreferences/tracing-typescript.md
REST APIreferences/tracing-rest-api.md
All integrationsreferences/integrations.md
Core concepts (traces, spans, threads)references/observability.md
Best practices (lifecycle, monitoring, anti-patterns)references/best-practices.md
Agent architecture, reliability, securityreferences/agent-patterns.md
Production monitoring, alerts, guardrailsreferences/production.md
Evaluation datasets & test suites (reference)references/evaluation-datasets.md, references/evaluation-test-suites.md

To build and run an evaluation, use the opik-evaluate skill. For repo instrumentation and config, use the opik-instrument skill.

© comet-ml, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in src/opik_mcp/skills/opik of comet-ml/opik-mcp.

  • SKILL.md
  • references/agent-patterns.md
  • references/best-practices.md
  • references/evaluation-datasets.md
  • references/evaluation-test-suites.md
  • references/integrations.md
  • references/observability.md
  • references/production.md
  • references/tracing-python.md
  • references/tracing-rest-api.md
  • references/tracing-typescript.md

Open the folder on GitHubat commit f1dd464

Compare with similar skills

Opik next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Opik compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Opik this skillcomet-ml/opik-mcp219—~2.1kAutomated safety check: PassApache-2.0
Tool Designagentailor/fullstack-langgraph-nextjs-agent132—~3.2kAutomated safety check: PassMIT
Ydc Openai Agent SDK IntegrationLeoYeAI/openclaw-master-skills2.2k—~4.4kAutomated safety check: NotesMIT
Neurolink Guidejuspay/neurolink143—~1.4kAutomated safety check: PassMIT
Strandsstrands-agents/harness-sdk8.7k—~1kAutomated safety check: PassApache-2.0
NaturalNPC-Worldwide/npcpy1.5k—~161Automated safety check: PassMIT

Similar skills

  • Tool Design

    agentailor/fullstack-langgraph-nextjs-agent

    Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise).

    132 GitHub stars~3.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Ydc Openai Agent SDK Integration

    LeoYeAI/openclaw-master-skills

    Integrate OpenAI Agents SDK with You.com MCP server - Hosted and Streamable HTTP support for Python and TypeScript.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Neurolink Guide

    juspay/neurolink

    Guide for using the NeuroLink SDK and CLI. An agent skill from juspay/neurolink.

    143 GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Strands

    strands-agents/harness-sdk

    Build, extend, evaluate, or migrate applications with Strands Agents in Python or TypeScript.

    8.7k GitHub stars~1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Natural

    NPC-Worldwide/npcpy

    Render the provided prompt template with Jinja context and send it to the active NPC's LLM.

    1.5k GitHub stars~161 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Copilot SDK

    github/awesome-copilot

    Official

    Build agentic applications with GitHub Copilot SDK. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 5 repos~6.3k tokens
    AI & LLM EngineeringAuto-check passed

More from comet-ml/opik-mcp

All 10 skills in this repo
  • Opik Compare

    comet-ml/opik-mcp

    Run a candidate against the baseline over an Opik test suite and read the numbers back — which cases broke, which got fixed, the per-metric deltas, worst rows, and whether the two runs are…

    219 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Opik Diagnose

    comet-ml/opik-mcp

    Surface the Opik traces worth a developer's attention, ranked by signal — Diagnostics issues first, then errors, failed tool calls, latency, regressions, and low online-eval scores.

    219 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes
  • Opik Evaluate

    comet-ml/opik-mcp

    Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link.

    219 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check: notes
  • Opik Instrument

    comet-ml/opik-mcp

    Add Opik tracing to an existing app and verify a real trace lands.

    219 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes
  • Opik Optimize

    comet-ml/opik-mcp

    Improve a prompt with the Opik Agent Optimizer — resolve the prompt, a dataset, and a metric, pick the algorithm, run a bounded optimization, check the gain on held-out data, and save the winner as…

    219 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Opik Verify

    comet-ml/opik-mcp

    Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost…

    219 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check: notes

Questions about Opik

What does Opik do?

Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST). Opik is an agent skill from comet-ml/opik-mcp. Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST).

When should I use Opik?

Opik fits situations like: what span types exist; version a prompt.

How do I install Opik in Claude Code?

Run `npx skills add comet-ml/opik-mcp --skill opik -a claude-code`. Or copy the skill folder (src/opik_mcp/skills/opik in comet-ml/opik-mcp) into .claude/skills/opik in your project. Claude Code loads it when a task matches its description.

How do I install Opik in Codex?

Run `npx skills add comet-ml/opik-mcp --skill opik -a codex`. Or copy the skill folder (src/opik_mcp/skills/opik in comet-ml/opik-mcp) into .agents/skills/opik in your project. Codex loads it when a task matches its description.

Can I use Opik in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add comet-ml/opik-mcp --skill opik -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/opik, .gemini/skills/opik, .github/skills/opik and .opencode/skills/opik in your project.

What does Opik need to run?

SKILL.md names no scripts, command-line tools or credentials: Opik is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Grep, Glob. Compatibility (from SKILL.md): Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). A reference — needs no Opik connection to read; the snippets assume the `opik` Python or TypeScript SDK 2.x. The task-shaped skills (opik-instrument, opik-diagnose, opik-explain, opik-test, opik-compare, opik-evaluate, opik-online-eval, opik-optimize, opik-verify) read this skill's references and expect it installed beside them..

Does Opik access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Opik safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Opik use?

Opik is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Opik use?

About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 32k tokens, read only when the agent opens those files.

What are the alternatives to Opik?

Skills that share tags, products or a category with Opik: Tool Design (agentailor/fullstack-langgraph-nextjs-agent, 132 stars), Ydc Openai Agent SDK Integration (LeoYeAI/openclaw-master-skills, 2.2k stars), Neurolink Guide (juspay/neurolink, 143 stars) and Strands (strands-agents/harness-sdk, 8.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Opik?

comet-ml (a GitHub organization) maintains it in comet-ml/opik-mcp, which has 219 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 6, 2026.

Source: comet-ml/opik-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.