Agent skill

Opik Instrument

by comet-ml in comet-ml/opik-mcp

Add Opik tracing to an existing app and verify a real trace lands.

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Opik Instrument

skills CLI
$ npx skills add comet-ml/opik-mcp --skill opik-instrument -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install comet-ml/opik-mcp opik-instrument --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/opik_mcp/skills/opik-instrument .claude/skills/opik-instrument && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
opik-instrument
GitHub stars
220
Token cost
~2.7k tokens
SKILL.md length
1,288 words
Files
16 (incl. references)
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Add Opik tracing to an existing app and verify a real trace lands.

  • Works in 7 steps: Configure Opik (one source of truth) → Detect language & framework → Add the minimum tracing → …
  • Instrument my code
  • SKILL.md covers Inputs, Activation — the only in-scope…, Blockers and Expansion — after the trace…, plus 4 more sections
  • Runs Python scripts from its folder; calls npx and uv; needs OPENAI_API_KEY and OPIK_API_KEY

What it does

Opik Instrument is an agent skill from comet-ml/opik-mcp. Add Opik tracing to an existing app and verify a real trace lands. Installs the Opik package, detects the language and LLM framework, adds the minimum tracing, runs a safe representative path, confirms a trace in Opik, and returns the trace link. Use for "instrument my code", "add opik tracing", "add observability", "trace my agent". Not for building a new app from scratch, or a review-only pass with no code changes.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 20 other files, including reference files (for example `evals/HARNESS.md`, `evals/cases.yaml` and `evals/fixtures/clean/app.py`). Compatibility notes: Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python or TypeScript project. Install the…

It sits in AI & LLM Engineering, covering Observability and MCP servers. It works with Model Context Protocol and Python. The repository describes itself as: Model Context Protocol (MCP) server for Opik, the open-source LLM observability and evaluation platform, built by Comet. Read traces, log scores, and manage prompts from Claude… The licence is Apache-2.0.

When your agent uses it

  • Instrument my code
  • Add opik tracing
  • Add observability

Example prompts

  • “instrument my code”
  • “add opik tracing”
  • “add observability”
  • “/opik-instrument”

Requirements

  • Python 3
  • Node.js
  • A credential in OPIK_API_KEY
  • A credential in OPENAI_API_KEY
  • Compatibility (from SKILL.md): Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python or TypeScript project. Install the `opik` skill alongside this one — it holds the shared SDK and integration references; without it, this skill falls back to the public docs.
  • Pre-approved tools (allowed-tools): Read, Edit, Write, Grep, Glob, Bash

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Configure Opik (one source of truth)
  2. Detect language & framework
  3. Add the minimum tracing
  4. Install the Opik package (by default)
  5. Run a safe representative path
  6. Verify ingestion
  7. Report

What it can do on your machine

Read from SKILL.md and the folder at commit e737581. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Edit
    • Write
    • Grep
    • Glob
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • npx
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • comet.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • OPIK_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python or TypeScript project. Install the `opik` skill alongside this one — it holds the shared SDK and integration references; without it, this skill falls back to the public docs.

    From compatibility in the SKILL.md frontmatter.

Context cost

Opik Instrument loads about 2.7k tokens when it runs, and up to ~3.5k if it reads all its reference files. Until then it costs about 109 tokens; SKILL.md has 1,288 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~109
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:38
    - Only add project-local `.env` vars if the project **already** uses that pattern. Never introduce a second config mecha
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Edit, Write, Grep, Glob, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from comet-ml/opik-mcp at commit e737581, republished under its Apache-2.0 licence (© comet-ml). 1,288 words, ~2,661 tokens.

Download SKILL.mdSave it as .claude/skills/opik-instrument/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.
name
opik-instrument
description
Add Opik tracing to an existing app and verify a real trace lands. Installs the Opik package, detects the language and LLM framework, adds the minimum tracing, runs a safe representative path, confirms a trace in Opik, and returns the trace link. Use for "instrument my code", "add opik tracing", "add observability", "trace my agent". Not for building a new app from scratch, or a review-only pass with no code changes.
allowed-tools
Read, Edit, Write, Grep, Glob, Bash
compatibility
Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python or TypeScript project. Install the `opik` skill alongside this one — it holds the shared SDK and integration references; without it, this skill falls back to the public docs.
metadata.last_updated
2026-09-10
metadata.source_commit
2.0.0
metadata.argument-hint
[optional: file or directory path]

Instrument — Add Opik Tracing and Verify a Real Trace

Definition of done: a representative, safely-executed path produces a trace that is confirmed in Opik and a direct trace link is returned. If verification can't be completed safely or autonomously, stop at the first genuine blocker and return exactly one concrete next step. Code edits alone are not success.

Operate: opinionated in execution, conservative in code changes, automatic in routine decisions, uncompromising about verifying value — but never by running something unsafe.

Inputs

The entry point is just /opik-instrument (optionally /opik-instrument <path>). Infer everything else; treat these only as optional overrides the user may pass, never as required setup:

  • target path (default: project root) · project name (default: inferred from the repo) · run command (default: an inferred safe path) · migrate_prompts (default: false).

Never turn inference into a questionnaire. Ask only when you hit a genuine, non-inferable blocker (see Blockers).

Activation — the only in-scope work

1. Configure Opik (one source of truth)
  • If ~/.opik.config exists or OPIK_API_KEY is set, use it as-is.
  • Otherwise run the official flow: opik configure (Python) / npx opik-ts configure (TypeScript).
  • Verify the config before instrumenting: run opik healthcheck — it validates the config, the install, and backend/workspace reachability. If it fails, stop at that Blocker and fix config or connectivity before adding any tracing.
  • Only add project-local .env vars if the project already uses that pattern. Never introduce a second config mechanism; never copy secret values between mechanisms.
2. Detect language & framework

Python (*.py, pyproject.toml) or TypeScript (*.ts, package.json). Identify the LLM framework from imports and pick its integration:

ImportIntegration
openai / anthropictrack_openai / track_anthropic
langchain / langgraphOpikTracer callback
crewai / dspy / google-genai / bedrock / llama_index / litellmtrack_crewai / OpikCallback / track_genai / track_bedrock / LlamaIndexCallbackHandler / OpikLogger
TS: opik-openai / opik-vercel / opik-langchaintrackOpenAI / OpikExporter / OpikCallbackHandler

Full list: read ../opik/references/integrations.md (the opik skill, installed beside this one). If that file isn't there, read https://www.comet.com/docs/opik/integrations/overview — never settle for manual spans on a framework that has a native integration just because the reference was unreachable. If the project is already instrumented, audit and add only what's missing — do not re-instrument.

3. Add the minimum tracing

Decision policy, in order:

  1. Prefer the framework-native integration for provider LLM spans.
  2. Add manual @opik.track spans only for orchestration/tools the integration doesn't cover (type="tool" / "llm" / "guardrail"). A bare @opik.track produces the default span type, general — the right choice for an entrypoint/orchestrator.
  3. Never instrument the same operation twice (no @opik.track(type="llm") on top of track_openai).
  4. Mark one entrypoint per independently-runnable agent/service — not necessarily one per repo.
  5. Decorator order relative to framework decorators (e.g. @app.route) is framework-dependent — verify per framework; do not assume a universal order.
  6. Scripts: flush at the end (opik.flush_tracker() / await client.flush()). LiteLLM inside @opik.track: pass metadata={"opik": {"current_span_data": get_current_span_data()}} or traces orphan.

Make the smallest change that lets one representative path emit a trace.

4. Install the Opik package (by default)

Add only the required Opik package(s) via the repo's detected package manager (pip / uv / poetry / npm / pnpm / yarn), through normal project conventions. Preserve the lockfile; do not run generic upgrades; do not install globally; treat unusual lifecycle scripts cautiously. Surface it as a change (e.g. "added opik to pyproject.toml"). If the environment blocks installation → Blocker with the one exact command.

5. Run a safe representative path

Infer a safe command — prefer an existing test, example, or dev script, then a bounded single-request entrypoint. Never run anything that looks like production or does irreversible/expensive work (writes, emails, purchases, mass API calls). If no safe path is inferable → Blocker ("which dev command safely exercises this agent?"). Print the command, then run it.

If the run needs an LLM provider credential (e.g. OPENAI_API_KEY) and it's absent, that's a Blocker — the app can't produce a trace. Note some SDK clients raise at construction (module load), before any span runs, so there is no partial trace to wait on: return blocked with the one next step, don't wait on a flush that never happened.

Show full SKILL.md (639 more words)Show less
6. Verify ingestion

Confirm a trace actually arrived — don't assume. Prefer the SDK for verification: it's already installed as part of instrumenting (zero extra moving parts), whereas the Opik MCP is optional and may not be connected.

The SDK read of one trace and its spans: references/verify-ingestion.md (SDK check).

To find the newest trace instead of using a known id, use the client's trace search (e.g. search_traces) scoped to the project. Optionally, if the Opik MCP is connected, list recent traces then read the newest.

Traces are asynchronous — allow a few seconds after the run and make sure the flush ran.

Verify coverage, not just arrival. A trace arriving is necessary but not sufficient — batching can silently drop or truncate spans, so a trace can land incomplete and still look fine. Before reporting verified:

  • Count vs. expected. Compare len(spans) (the search_spans call in references/verify-ingestion.md, project_name included) against the call sites you instrumented on the path you ran (entrypoint + each traced tool/LLM). Fewer spans than expected means spans were dropped — do not report verified.
  • Every span is well-formed. Each span has a non-empty name and type; LLM spans carry input/output (and usage where the integration provides it). A span returned with an empty name/type is the batching-race symptom in references/verify-ingestion.md, not a real span.

With the Opik MCP connected, verify there instead. read(entity_type="trace", id=tid) returns {trace, spans, spansTruncated} with the span tree inlined (up to 200 spans), so both checks above run over that one call, no script needed: count the spans against the instrumented call sites, and confirm each has a name/type and the LLM spans carry input/output. If spansTruncated is true, count with the SDK instead. The SDK stays the default because it is already installed; the MCP is the shortcut when it is there.

If the trace is empty, partial, or has unnamed spans, check references/verify-ingestion.md (Common ingestion traps) before changing the instrumentation.

7. Report

Return a short human result + the trace link (see Output), then make the single expansion offer.

Blockers

When you genuinely can't proceed, stop at the earliest blocker and return exactly one next step — never a checklist — and still report the changes already made (blocked carries changes). An unsupported language or shape is unsupported and modifies nothing. Examples:

  • "Run opik configure, then rerun /opik-instrument."
  • "Install dependencies with uv sync, then rerun /opik-instrument."
  • "Which dev command safely exercises this agent?"
  • "Instrumented and installed, but the run needs a provider credential — set OPENAI_API_KEY (or the relevant provider key) and rerun."
  • "Instrumented and ran, but this environment can't query Opik — open the project and confirm trace <id> arrived."

Expansion — after the trace lands (one offer, not a funnel)

Do not migrate prompts, add threading, or broaden spans during activation. After verification, make a single consolidated offer of what you found, e.g.:

Tracing is verified. I also found ways to deepen it: 3 Prompt Library candidates, missing conversation threading, and 2 untraced tools. Expand?

Output

User-facing: a short human message — what was instrumented, the trace link, and the one expansion offer (or, if blocked, the single next step plus what changed). Not raw JSON.

Underneath (for composition / evals), a small state model, with its invariants: references/output-shape.md.

Examples

Worked runs (no LLM framework, OpenAI, already instrumented): references/examples.md.

Anti-patterns

Double-wrapping (integration + manual span on the same call); orphaned LiteLLM traces (missing current_span_data); missing flush in scripts; overwriting or duplicating config; running an unsafe/production path just to force a trace; broad dependency upgrades when only opik is needed; migrating prompts during activation.

References

SDK detail lives in the opik skill, installed beside this one. Read the files directly — paths are relative to this file: ../opik/references/tracing-python.md, ../opik/references/tracing-typescript.md, ../opik/references/integrations.md, ../opik/references/observability.md. If your host lays skills out differently, locate the opik skill's references/ directory.

If the opik skill isn't installed, say so in the report and use https://www.comet.com/docs/opik/ rather than working from memory.

© comet-ml, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 15 other files (references) in src/opik_mcp/skills/opik-instrument of comet-ml/opik-mcp.

  • SKILL.md
  • evals/.gitignore
  • evals/HARNESS.md
  • evals/cases.yaml
  • evals/fixtures/clean/app.py
  • evals/fixtures/clean/expected.json
  • evals/fixtures/clean/pyproject.toml
  • evals/fixtures/missing_flush/app.py
  • evals/fixtures/missing_flush/expected.json
  • evals/fixtures/missing_flush/pyproject.toml
  • evals/grader.py
  • evals/metrics.py
  • evals/run_evals.py
  • references/examples.md
  • references/output-shape.md
  • references/verify-ingestion.md

Open the folder on GitHubat commit e737581

Compare with similar skills

Opik Instrument next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Opik Instrument compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Opik Instrument this skillcomet-ml/opik-mcp220—~2.7kAutomated safety check: NotesApache-2.0
Agent Observability Eval Bootstrapdatadog-labs/agent-skills177—~25kAutomated safety check: PassMIT
Cloudbase Agent PythonTencentCloudBase/CloudBase-AI-Toolkit1.1k2 repos~2.9kAutomated safety check: NotesMIT
Strandsstrands-agents/harness-sdk8.7k—~1kAutomated safety check: PassApache-2.0
Setupguaardvark/guaardvark257—~1.2kAutomated safety check: PassMIT
NaturalNPC-Worldwide/npcpy1.5k—~161Automated safety check: PassMIT

Similar skills

  • Agent Observability Eval Bootstrap

    datadog-labs/agent-skills

    Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit…

    177 GitHub stars~25k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Cloudbase Agent Python

    TencentCloudBase/CloudBase-AI-Toolkit

    Build production-ready AI agent backends using the CloudBase Agent Python SDK — create agents with LangGraph/CrewAI/LlamaIndex, serve them via FastAPI with AG-UI protocol streaming +…

    1.1k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check: notes
  • Strands

    strands-agents/harness-sdk

    Build, extend, evaluate, or migrate applications with Strands Agents in Python or TypeScript.

    8.7k GitHub stars~1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Setup

    guaardvark/guaardvark

    Connect this agent to a running Guaardvark (self-hosted AI studio) and check what it can do right now.

    257 GitHub stars~1.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Natural

    NPC-Worldwide/npcpy

    Render the provided prompt template with Jinja context and send it to the active NPC's LLM.

    1.5k GitHub stars~161 tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Copilot SDK

    github/awesome-copilot

    Official

    Build agentic applications with GitHub Copilot SDK. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 4 repos~6.3k tokens
    AI & LLM EngineeringAuto-check passed

More from comet-ml/opik-mcp

All 10 skills in this repo
  • Opik

    comet-ml/opik-mcp

    Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST).

    220 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Opik Compare

    comet-ml/opik-mcp

    Run a candidate against the baseline over an Opik test suite and read the numbers back — which cases broke, which got fixed, the per-metric deltas, worst rows, and whether the two runs are…

    220 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Opik Diagnose

    comet-ml/opik-mcp

    Surface the Opik traces worth a developer's attention, ranked by signal — Diagnostics issues first, then errors, failed tool calls, latency, regressions, and low online-eval scores.

    220 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes
  • Opik Evaluate

    comet-ml/opik-mcp

    Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link.

    220 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check: notes
  • Opik Optimize

    comet-ml/opik-mcp

    Improve a prompt with the Opik Agent Optimizer — resolve the prompt, a dataset, and a metric, pick the algorithm, run a bounded optimization, check the gain on held-out data, and save the winner as…

    220 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Opik Verify

    comet-ml/opik-mcp

    Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost…

    220 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check: notes

Questions about Opik Instrument

What does Opik Instrument do?

Add Opik tracing to an existing app and verify a real trace lands. Opik Instrument is an agent skill from comet-ml/opik-mcp. Add Opik tracing to an existing app and verify a real trace lands.

When should I use Opik Instrument?

Opik Instrument fits situations like: instrument my code; add opik tracing; add observability.

How do I install Opik Instrument in Claude Code?

Run `npx skills add comet-ml/opik-mcp --skill opik-instrument -a claude-code`. Or copy the skill folder (src/opik_mcp/skills/opik-instrument in comet-ml/opik-mcp) into .claude/skills/opik-instrument in your project. Claude Code loads it when a task matches its description.

How do I install Opik Instrument in Codex?

Run `npx skills add comet-ml/opik-mcp --skill opik-instrument -a codex`. Or copy the skill folder (src/opik_mcp/skills/opik-instrument in comet-ml/opik-mcp) into .agents/skills/opik-instrument in your project. Codex loads it when a task matches its description.

Can I use Opik Instrument in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add comet-ml/opik-mcp --skill opik-instrument -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/opik-instrument, .gemini/skills/opik-instrument, .github/skills/opik-instrument and .opencode/skills/opik-instrument in your project.

What does Opik Instrument need to run?

Going by SKILL.md and its folder, Opik Instrument needs Python for the scripts in its folder, the command-line tools its instructions call (npx and uv) and credentials named OPENAI_API_KEY and OPIK_API_KEY. Our summary lists: Python 3; Node.js; A credential in OPIK_API_KEY; A credential in OPENAI_API_KEY. Its frontmatter pre-approves these tools: Read, Edit, Write, Grep, Glob, Bash. Compatibility (from SKILL.md): Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python or TypeScript project. Install the `opik` skill alongside this one — it holds the shared SDK and integration references; without it, this skill falls back to the public docs..

Does Opik Instrument access the network?

SKILL.md names 1 domain. As links in the text: comet.com. This is read from the text; nothing was executed.

Is Opik Instrument safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Opik Instrument use?

Opik Instrument is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Opik Instrument use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 827 tokens, read only when the agent opens those files.

What are the alternatives to Opik Instrument?

Skills that share tags, products or a category with Opik Instrument: Agent Observability Eval Bootstrap (datadog-labs/agent-skills, 177 stars), Cloudbase Agent Python (TencentCloudBase/CloudBase-AI-Toolkit, 1.1k stars), Strands (strands-agents/harness-sdk, 8.7k stars) and Setup (guaardvark/guaardvark, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Opik Instrument?

comet-ml (a GitHub organization) maintains it in comet-ml/opik-mcp, which has 220 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 8, 2026.

Source: comet-ml/opik-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.