Tool Design
agentailor/fullstack-langgraph-nextjs-agent
Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise).
Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST).
$ npx skills add comet-ml/opik-mcp --skill opik -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install comet-ml/opik-mcp opik --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/opik_mcp/skills/opik .claude/skills/opik && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "opik" agent skill from https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opik into .claude/skills/opik/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opik", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opikType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add comet-ml/opik-mcp --skill opik -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install comet-ml/opik-mcp opik --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .agents/skills && cp -r skills-src/src/opik_mcp/skills/opik .agents/skills/opik && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "opik" agent skill from https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opik into .agents/skills/opik/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opik", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add comet-ml/opik-mcp --skill opik -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install comet-ml/opik-mcp opik --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/src/opik_mcp/skills/opik .cursor/skills/opik && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "opik" agent skill from https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opik into .cursor/skills/opik/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opik", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/comet-ml/opik-mcp.git --path src/opik_mcp/skills/opik--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add comet-ml/opik-mcp --skill opik -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install comet-ml/opik-mcp opik --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/src/opik_mcp/skills/opik .gemini/skills/opik && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "opik" agent skill from https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opik into .gemini/skills/opik/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opik", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install comet-ml/opik-mcp opikInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add comet-ml/opik-mcp --skill opik -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .github/skills && cp -r skills-src/src/opik_mcp/skills/opik .github/skills/opik && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "opik" agent skill from https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opik into .github/skills/opik/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opik", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add comet-ml/opik-mcp --skill opik -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install comet-ml/opik-mcp opik --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/src/opik_mcp/skills/opik .opencode/skills/opik && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "opik" agent skill from https://github.com/comet-ml/opik-mcp/tree/main/src/opik_mcp/skills/opik into .opencode/skills/opik/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "opik", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
opikReference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST).
Opik is an agent skill from comet-ml/opik-mcp. Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST). Use for "what span types exist", "how do I flush", "trackopenai", "add OpikTracer", "version a prompt". To instrument a repo end to end, use the opik-instrument skill.
Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `references/agent-patterns.md`, `references/best-practices.md` and `references/evaluation-datasets.md`). Compatibility notes: Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). A reference — needs no Opik connection to read; the…
It sits in AI & LLM Engineering, covering Prompt engineering and MCP servers. It works with Python, TypeScript, OpenAI and Model Context Protocol. The repository describes itself as: Model Context Protocol (MCP) server for Opik, the open-source LLM observability and evaluation platform, built by Comet. Read traces, log scores, and manage prompts from Claude… The licence is Apache-2.0.
Read from SKILL.md and the folder at commit f1dd464. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGrepGlobFrom allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python and typescript).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). A reference — needs no Opik connection to read; the snippets assume the `opik` Python or TypeScript SDK 2.x. The task-shaped skills (opik-instrument, opik-diagnose, opik-explain, opik-test, opik-compare, opik-evaluate, opik-online-eval, opik-optimize, opik-verify) read this skill's references and expect it installed beside them.
From compatibility in the SKILL.md frontmatter.
Opik loads about 2.1k tokens when it runs, and up to ~34k if it reads all its reference files. Until then it costs about 77 tokens; SKILL.md has 635 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from comet-ml/opik-mcp at commit f1dd464, republished under its Apache-2.0 licence (© comet-ml). 635 words, ~2,105 tokens.
.claude/skills/opik/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.Opik is an open-source LLM observability platform. This skill is a reference
for the SDK. To instrument a codebase step by step (detect frameworks, add
config, emit and verify a trace), use the task-shaped opik-instrument skill.
A trace is one execution path (one request → one response). Spans are the operations inside it and form a hierarchy.
| Type | Use for |
|---|---|
general | orchestration, agent entry points |
llm | model calls |
tool | tools, retrieval, API / DB calls |
guardrail | safety / validation checks |
Do NOT use retrieval or any other value.
import opik
@opik.track(name="agent", type="general")
def agent(query: str) -> str:
return generate(retrieve(query))
@opik.track(type="tool")
def retrieve(query): ...
@opik.track(type="llm")
def generate(ctx): ...
opik.flush_tracker() # required in scriptsimport { Opik } from "opik";
const client = new Opik({ projectName: "my-project" });
const trace = client.trace({ name: "agent", input: { query } });
const span = trace.span({ name: "llm-call", type: "llm" });
span.end({ output });
trace.end({ output });
await client.flush();Prefer an integration over manual @opik.track — integrations capture tokens,
model, and cost automatically. Patterns (full list in
references/integrations.md):
track_openai(OpenAI()), track_anthropic(...)track_crewai(crew=crew)dspy.configure(callbacks=[OpikCallback()])OpikTracer() for LangChain / LangGraph / LlamaIndextrack_adk_agent_recursive(agent, OpikTracer())@opik.track (common trap)If code uses litellm and you add @opik.track, pass current_span_data
via metadata on every completion call — otherwise OpikLogger emits orphaned
top-level traces instead of nesting under your span.
from opik.opik_context import get_current_span_data
@opik.track
def call_llm(messages):
return litellm.completion(
model="gpt-4o",
messages=messages,
metadata={"opik": {"current_span_data": get_current_span_data()}},
)Group turns with thread_id — one turn = one trace, shared thread_id = one
thread. Use for chat / multi-turn; skip for single-shot.
@opik.track(entrypoint=True)
def handle(session_id: str, message: str) -> str:
opik.update_current_trace(thread_id=session_id)
return reply(message)Version prompts with client.get_prompt / create_prompt (chat variants:
get_chat_prompt / create_chat_prompt). Store model + temperature in the
prompt metadata so they version with the text. Call get_prompt inside a
@opik.track function so the version links to the trace.
@opik.track(entrypoint=True)
def run(question: str) -> str:
p = client.get_prompt(name="system") or client.create_prompt(
name="system",
prompt="You help with {{product}}.",
metadata={"model": "gpt-4o", "temperature": 0.7},
)
return llm(p.format(product="Opik"), model=p.metadata["model"])With the MCP connected, start at the project, not at its traces:
read("project", "<project name or id>")One call returns the last 7 days against the 7 before — trace count, error
rate, average duration, total cost, SDK traffic only, which is what the Logs
page's four cards show — plus the score names and usage keys the project
actually records, and the freshest experiment, dataset, prompt version and
optimization run in it. since/until pick another window; since="30d" is
what the UI opens on. A rate or an average over a window with no traces comes
back null rather than 0, because a rate over no samples is undefined and
"0% errors" is advice someone may act on.
Then attribute the change rather than restating it:
list("project_metric", project_name="<project>", metric_type="trace_cost")
list("project_metric", project_name="<project>", metric_type="span_count",
breakdown="model", since="30d")Rows are time buckets, not records — interval is hourly/daily/weekly/
total, and page/size/sort do not apply. schema("list.project_metric")
is the metric list, what each is about, and which groupings each accepts; seven
of them accept none. The score names the overview returned are the ones worth
filtering on, and list("score_name", project_name=…) has the rest.
One filter grammar, OQL, serves both the hosted MCP's list tool and the
SDK's search_traces / search_spans / search_threads:
<field>[.<key>] <op> <value> [AND ...]
ops: = != > >= < <= contains not_contains starts_with ends_with is_empty is_not_empty in not_inStrings in double quotes, numbers bare, duration in milliseconds, dates as
ISO-8601 instants with a timezone ("2026-09-08T10:00:00Z"). Scores and
dictionaries take a key: feedback_scores.accuracy < 0.5,
metadata.environment = "prod". AND is the only connector.
error_info is_not_empty AND duration > 5000
type = "llm" AND usage.total_tokens > 10000 # spans
feedback_scores.hallucination > 0.5 AND start_time >= "2026-09-08T00:00:00Z"With the MCP connected, prefer list — it also sorts (sort="duration desc"),
windows (since="1h", "7d"), and searches free text (search="order-42"):
list(entity_type="trace", project_name="<project>", since="1h",
filters="error_info is_not_empty", sort="duration desc")Trace, span and thread lists add source = "sdk" unless you name source, so
evaluator, playground and experiment traces stay out of the way. A rejected
filter comes back with what fixes it; schema("list.trace") (or list.span,
list.thread, list.experiment) is the full field and operator reference.
Without the MCP, the same string goes to the SDK:
client.search_traces(project_name="<project>", filter_string="error_info is_not_empty")| Anti-pattern | Fix |
|---|---|
span type retrieval / custom | use tool (or general) |
get_prompt outside @opik.track | fetch inside — else no trace link |
deprecated opik.Prompt / opik.Config | use client.get_prompt / config file |
litellm without current_span_data | pass it — else orphaned traces |
| no flush in scripts | opik.flush_tracker() / await client.flush() |
| Topic | File |
|---|---|
| Python SDK (async, distributed, context) | references/tracing-python.md |
| TypeScript SDK | references/tracing-typescript.md |
| REST API | references/tracing-rest-api.md |
| All integrations | references/integrations.md |
| Core concepts (traces, spans, threads) | references/observability.md |
| Best practices (lifecycle, monitoring, anti-patterns) | references/best-practices.md |
| Agent architecture, reliability, security | references/agent-patterns.md |
| Production monitoring, alerts, guardrails | references/production.md |
| Evaluation datasets & test suites (reference) | references/evaluation-datasets.md, references/evaluation-test-suites.md |
To build and run an evaluation, use the opik-evaluate skill. For repo instrumentation and config, use the opik-instrument skill.
© comet-ml, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 10 other files (references) in src/opik_mcp/skills/opik of comet-ml/opik-mcp.
Open the folder on GitHubat commit f1dd464
Opik next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Opik this skillcomet-ml/opik-mcp | 219 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Tool Designagentailor/fullstack-langgraph-nextjs-agent | 132 | — | ~3.2k | Automated safety check: Pass | MIT | |
| Ydc Openai Agent SDK IntegrationLeoYeAI/openclaw-master-skills | 2.2k | — | ~4.4k | Automated safety check: Notes | MIT | |
| Neurolink Guidejuspay/neurolink | 143 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Strandsstrands-agents/harness-sdk | 8.7k | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| NaturalNPC-Worldwide/npcpy | 1.5k | — | ~161 | Automated safety check: Pass | MIT |
agentailor/fullstack-langgraph-nextjs-agent
Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise).
LeoYeAI/openclaw-master-skills
Integrate OpenAI Agents SDK with You.com MCP server - Hosted and Streamable HTTP support for Python and TypeScript.
juspay/neurolink
Guide for using the NeuroLink SDK and CLI. An agent skill from juspay/neurolink.
strands-agents/harness-sdk
Build, extend, evaluate, or migrate applications with Strands Agents in Python or TypeScript.
NPC-Worldwide/npcpy
Render the provided prompt template with Jinja context and send it to the active NPC's LLM.
github/awesome-copilot
Build agentic applications with GitHub Copilot SDK. An agent skill from github/awesome-copilot.
comet-ml/opik-mcp
Run a candidate against the baseline over an Opik test suite and read the numbers back — which cases broke, which got fixed, the per-metric deltas, worst rows, and whether the two runs are…
comet-ml/opik-mcp
Surface the Opik traces worth a developer's attention, ranked by signal — Diagnostics issues first, then errors, failed tool calls, latency, regressions, and low online-eval scores.
comet-ml/opik-mcp
Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link.
comet-ml/opik-mcp
Add Opik tracing to an existing app and verify a real trace lands.
comet-ml/opik-mcp
Improve a prompt with the Opik Agent Optimizer — resolve the prompt, a dataset, and a metric, pick the algorithm, run a bounded optimization, check the gain on held-out data, and save the winner as…
comet-ml/opik-mcp
Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost…
Categories
Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST). Opik is an agent skill from comet-ml/opik-mcp. Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST).
Opik fits situations like: what span types exist; version a prompt.
Run `npx skills add comet-ml/opik-mcp --skill opik -a claude-code`. Or copy the skill folder (src/opik_mcp/skills/opik in comet-ml/opik-mcp) into .claude/skills/opik in your project. Claude Code loads it when a task matches its description.
Run `npx skills add comet-ml/opik-mcp --skill opik -a codex`. Or copy the skill folder (src/opik_mcp/skills/opik in comet-ml/opik-mcp) into .agents/skills/opik in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add comet-ml/opik-mcp --skill opik -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/opik, .gemini/skills/opik, .github/skills/opik and .opencode/skills/opik in your project.
SKILL.md names no scripts, command-line tools or credentials: Opik is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Grep, Glob. Compatibility (from SKILL.md): Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). A reference — needs no Opik connection to read; the snippets assume the `opik` Python or TypeScript SDK 2.x. The task-shaped skills (opik-instrument, opik-diagnose, opik-explain, opik-test, opik-compare, opik-evaluate, opik-online-eval, opik-optimize, opik-verify) read this skill's references and expect it installed beside them..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Opik is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 32k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Opik: Tool Design (agentailor/fullstack-langgraph-nextjs-agent, 132 stars), Ydc Openai Agent SDK Integration (LeoYeAI/openclaw-master-skills, 2.2k stars), Neurolink Guide (juspay/neurolink, 143 stars) and Strands (strands-agents/harness-sdk, 8.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
comet-ml (a GitHub organization) maintains it in comet-ml/opik-mcp, which has 219 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 6, 2026.
Source: comet-ml/opik-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.