Agent skill

Opik Optimize

by comet-ml in comet-ml/opik-mcp

Improve a prompt with the Opik Agent Optimizer — resolve the prompt, a dataset, and a metric, pick the algorithm, run a bounded optimization, check the gain on held-out data, and save the winner as…

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Opik Optimize

skills CLI
$ npx skills add comet-ml/opik-mcp --skill opik-optimize -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install comet-ml/opik-mcp opik-optimize --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/comet-ml/opik-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/opik_mcp/skills/opik-optimize .claude/skills/opik-optimize && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
opik-optimize
GitHub stars
219
Token cost
~2.6k tokens
SKILL.md length
1,225 words
Files
12 (incl. references)
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Improve a prompt with the Opik Agent Optimizer — resolve the prompt, a dataset, and a metric, pick the algorithm, run a bounded optimization, check the gain on held-out data, and save the winner as…

  • Works in 7 steps: Resolve the prompt → Resolve the dataset (and hold some out) → Define the metric → …
  • Optimize this prompt
  • SKILL.md covers Inputs, Activation — the only in-scope…, Blockers and Output, plus 3 more sections
  • Runs Python scripts from its folder; calls uv and pip; needs OPENAI_API_KEY

What it does

Opik Optimize is an agent skill from comet-ml/opik-mcp. Improve a prompt with the Opik Agent Optimizer — resolve the prompt, a dataset, and a metric, pick the algorithm, run a bounded optimization, check the gain on held-out data, and save the winner as a new prompt version with the optimization run link. Runs via the opik-optimizer package; reads prompts and datasets via the MCP when connected. Use for "optimize this prompt", "improve my system prompt", "make the agent answer better", "tune the prompt against my dataset", "run the prompt optimizer". Not for measuring…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including reference files (for example `evals/HARNESS.md`, `evals/cases.yaml` and `evals/fixtures/format/seed.py`). Compatibility notes: Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python project with Opik configured, a…

It sits in AI & LLM Engineering, covering Prompt engineering and MCP servers. It works with Model Context Protocol. The repository describes itself as: Model Context Protocol (MCP) server for Opik, the open-source LLM observability and evaluation platform, built by Comet. Read traces, log scores, and manage prompts from Claude… The licence is Apache-2.0.

When your agent uses it

  • Optimize this prompt
  • Improve my system prompt
  • Make the agent answer better
  • Tune the prompt against my dataset

Example prompts

  • “optimize this prompt”
  • “improve my system prompt”
  • “make the agent answer better”
  • “/opik-optimize”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY
  • Compatibility (from SKILL.md): Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python project with Opik configured, a provider API key, and a dataset (or traces to build one). Install the `opik` skill alongside this one — it holds the shared dataset and prompt-library references; without it, this skill falls back to the public docs.
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, Write

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Resolve the prompt
  2. Resolve the dataset (and hold some out)
  3. Define the metric
  4. Pick the algorithm
  5. State the budget, then run
  6. Read the result honestly
  7. Save the winner (library prompts) and hand off

What it can do on your machine

Read from SKILL.md and the folder at commit f1dd464. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • comet.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python project with Opik configured, a provider API key, and a dataset (or traces to build one). Install the `opik` skill alongside this one — it holds the shared dataset and prompt-library references; without it, this skill falls back to the public docs.

    From compatibility in the SKILL.md frontmatter.

Context cost

Opik Optimize loads about 2.6k tokens when it runs, and up to ~3.4k if it reads all its reference files. Until then it costs about 160 tokens; SKILL.md has 1,225 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~160
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from comet-ml/opik-mcp at commit f1dd464, republished under its Apache-2.0 licence (© comet-ml). 1,225 words, ~2,602 tokens.

Download SKILL.mdSave it as .claude/skills/opik-optimize/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
opik-optimize
description
Improve a prompt with the Opik Agent Optimizer — resolve the prompt, a dataset, and a metric, pick the algorithm, run a bounded optimization, check the gain on held-out data, and save the winner as a new prompt version with the optimization run link. Runs via the opik-optimizer package; reads prompts and datasets via the MCP when connected. Use for "optimize this prompt", "improve my system prompt", "make the agent answer better", "tune the prompt against my dataset", "run the prompt optimizer". Not for measuring quality once (use evaluate), before/after on a suite (use compare), or hand-editing a prompt without data.
allowed-tools
Read, Grep, Glob, Bash, Write
compatibility
Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python project with Opik configured, a provider API key, and a dataset (or traces to build one). Install the `opik` skill alongside this one — it holds the shared dataset and prompt-library references; without it, this skill falls back to the public docs.
metadata.last_updated
2026-09-15
metadata.source_commit
2.0.0
metadata.argument-hint
[prompt name or file, optional dataset and metric]

Optimize — Improve a Prompt Against Data

Definition of done: an optimized prompt whose score on the metric beats the baseline on data it was not tuned on, the optimization run link in Opik, the cost of getting there, and the winner saved as a new prompt version (when the prompt lives in the library) — with the swap into code left as the next step. If the optimization can't run within a stated budget, stop at the first genuine blocker and return exactly one next step. A prompt that scores higher only on its own training items is not an improvement.

Operate: measure the baseline first, state the budget before spending it, hold data out, pick the algorithm for the failure you see, save the winner where it can be versioned — and change no application code. The only file this skill writes is a runner outside the repo; the prompt is saved to Opik, not into the codebase.

Inputs

The entry point is /opik-optimize <prompt-name> (a prompt-library prompt), /opik-optimize <path or function> (a prompt in code), or /opik-optimize (find the system prompt in this repo). Infer the rest; treat these as optional overrides:

  • dataset (default: existing dataset for the project → export the regression suite → build from traces) · metric (default: heuristic on expected_output if present, else one binary judge) · algorithm (default: MetaPromptOptimizer) · budget (default: n_samples=50, max_trials=10) · model · validation split (default: hold out 20%).

Ask only at a genuine, non-inferable blocker (see Blockers).

Activation — the only in-scope work

1. Resolve the prompt
  • Library: client.get_chat_prompt(name) (or get_prompt for a text prompt). Note the current version — that is the baseline.
  • Code: grep for the system prompt / messages=[...]; read it verbatim. Note where it lives; you will not edit it.
  • Trace: the llm span's input.messages on a representative trace.

The optimizer's opik_optimizer.ChatPrompt is a different class from the library's opik.ChatPrompt — build it from the raw messages yourself: references/sdk-snippets.md (Resolve the prompt).

2. Resolve the dataset (and hold some out)

The optimizer needs an opik.Dataset whose item keys match the prompt's {variables}: references/sdk-snippets.md (Resolve the dataset).

  • Only a test suite exists → export its items into a dataset once: suite.get_items() → client.get_or_create_dataset("<suite>-optimize", project_name=…) → insert([{**it["data"]} …]).
  • Nothing exists → build from traces (search_traces → {"question": t.input[...], "expected_output": …}) or run /opik-evaluate first.
  • Hold out: split into train and validation datasets and pass validation_dataset=. Note what it does: the optimizer scores every trial on validation_dataset and uses the train set to show the reasoning model examples — so the validation set is the selection set, and it needs ≥10 items or every candidate ties (a 4-item split logs n_samples … larger than evaluation dataset size and cannot separate prompts). If you need a gain measured on items the optimizer never saw, keep a third split and re-score the winner on it with evaluate(). Fewer than ~20 items in total → say the result will be noisy; below 10 → Blocker.
3. Define the metric

A function (dataset_item, llm_output) -> float, higher is better. Give it a real name (def refund_answer_similarity(...)) — its __name__ becomes the Optimization run's objective name in the UI and result.metric_name; a function called metric shows up as "metric".

  • expected_output present → heuristic (LevenshteinRatio, Equals, or a task-specific check) — deterministic and free. Prefer a graded metric over exact match: when the baseline scores 0.0 on every item (observed with Equals on a strict output format), every candidate also scores 0.0 and the optimizer has nothing to climb — five trials of flat zeros is a metric problem, not a prompt problem.
  • Otherwise → one binary judge for the failure mode being optimized (../opik-evaluate/references/write-judge-prompt.md), wrapped to return its score .value. Multi-objective → MultiMetricObjective. Never optimize against a judge nobody validated: an unvalidated judge is the easiest thing to overfit.
4. Pick the algorithm
Failure you seeOptimizer
Instructions unclear / underspecified (general default)MetaPromptOptimizer
The model needs examples of the right answer; few-shot is acceptableFewShotBayesianOptimizer
Failures cluster into a few root causesHierarchicalReflectiveOptimizer (HRPO)
Larger budget, want broad searchEvolutionaryOptimizer or GepaOptimizer
The prompt is fine, temperature/top_p are notParameterOptimizer.optimize_parameter(...)
Tool descriptions are the problemoptimize_prompt(..., optimize_tools=True) (optimize_mcp is deprecated)
5. State the budget, then run

uv add opik-optimizer (or pip install opik-optimizer) in a scratch environment, not the repo's lockfile unless the user wants it. Each trial evaluates n_samples items with the task model plus the reasoning model — tell the user the rough call count before running. Provider key absent → Blocker. The MetaPromptOptimizer run, with the default budget: references/sdk-snippets.md (Run the optimizer). Write the runner as a temp file outside the repo. The run appears in Opik as an Optimization (result.get_run_link()).

Show full SKILL.md (476 more words)Show less
6. Read the result honestly

result.initial_score → result.score on the metric; result.details["stop_reason"] and ["trials_completed"]; result.llm_calls, result.llm_cost_total (may be None when the provider returns no cost — say "cost unavailable", don't invent one). Report the validation score, not the training score. A gain within run-to-run noise (rerun the baseline once if in doubt) is "no measurable improvement" — say so rather than shipping a lateral move, and do not save a new version for it. The common cause of a flat result: the answers depend on context the prompt can't contain (retrieval, tools, account data) — then the prompt isn't the bottleneck and the next step is /opik-explain on the worst items, not more trials.

7. Save the winner (library prompts) and hand off

For a library prompt, save the winner as a new version: references/sdk-snippets.md (Save the winner). For a prompt that lives in code, do not edit the file — return the optimized text and the diff as the next step. Then one next step (see Output): typically "/opik-compare the new version against the regression suite" or "point the app at version vN".

Blockers

Stop at the earliest blocker and return exactly one next step:

  • "Run opik configure, then rerun /opik-optimize."
  • "Which prompt? Name the library prompt or point me at the file/function holding the system prompt."
  • "No dataset with matching keys — run /opik-evaluate to build one, or name an existing dataset."
  • "The optimizer needs a provider credential — set OPENAI_API_KEY (or the relevant key) and rerun."
  • "Only 6 items — too few to optimize without overfitting. Add cases (or say synthetic) and rerun."

Output

User-facing: a short human message — baseline vs optimized score on validation, the run link, the cost, the algorithm, what changed in the prompt (one or two lines), the new version (or the diff for a code prompt), and the single next step. Not the full trial history.

Underneath (for composition / evals), one shape, with its invariants: references/output-shape.md.

Examples

Worked runs (library prompt with a heuristic metric, no real gain, prompt in code, blocked): references/examples.md.

Anti-patterns

Reporting the training-set score as the gain; optimizing against an unvalidated judge; spending an unbounded budget (no n_samples/max_trials) or not stating it; overwriting the prompt in place instead of a new version; editing the prompt in the codebase; treating opik.ChatPrompt and opik_optimizer.ChatPrompt as interchangeable; optimizing on fewer than ~20 items and calling it a result; choosing the algorithm by novelty rather than by the failure observed; using deprecated optimize_mcp.

References

Dataset and prompt-library detail live in the opik skill, installed beside this one — paths relative to this file: ../opik/references/evaluation-datasets.md (datasets, insert, versions, metrics), ../opik/references/best-practices.md (prompt library, versioning), ../opik/references/tracing-python.md (SDK client). Judge design and validation: ../opik-evaluate/references/write-judge-prompt.md, ../opik-evaluate/references/validate-evaluator.md. If your host lays skills out differently, locate the opik skill's references/ directory.

Optimizer API (opik-optimizer): https://www.comet.com/docs/opik/agent_optimization/overview. If the opik skill isn't installed, say so in the report and use https://www.comet.com/docs/opik/ rather than working from memory.

© comet-ml, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (references) in src/opik_mcp/skills/opik-optimize of comet-ml/opik-mcp.

  • SKILL.md
  • evals/.gitignore
  • evals/HARNESS.md
  • evals/cases.yaml
  • evals/fixtures/format/pyproject.toml
  • evals/fixtures/format/seed.py
  • evals/grader.py
  • evals/metrics.py
  • evals/run_evals.py
  • references/examples.md
  • references/output-shape.md
  • references/sdk-snippets.md

Open the folder on GitHubat commit f1dd464

Compare with similar skills

Opik Optimize next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Opik Optimize compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Opik Optimize this skillcomet-ml/opik-mcp219—~2.6kAutomated safety check: NotesApache-2.0
Tool CreatorhAcKlyc/MyAgents915—~1.5kAutomated safety check: PassAGPL-3.0
NaturalNPC-Worldwide/npcpy1.5k—~161Automated safety check: PassMIT
Pi AgentK-Dense-AI/scientific-agent-skills48k1 repos~2.1kAutomated safety check: PassMIT
Workflow Schema Tuningbreaking-brake/cc-wf-studio5.4k—~1.3kAutomated safety check: PassCustom licence
Hcls Build Agentaws-samples/amazon-bedrock-agents-healthcare-lifesciences274—~885Automated safety check: PassMIT-0

Similar skills

  • Tool Creator

    hAcKlyc/MyAgents

    把用户的可复用需求封装成标准化的 Agent-CLI 工具,并用 myagents tool add 注册进 MyAgents 工具注册表——注册后所有未来会话(builtin / DSH / Claude Code / Codex 全 runtime)的 AI 都会在 system prompt 里自动发现它。触发场景:(1) 用户说「把 XX…

    915 GitHub stars~1.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Natural

    NPC-Worldwide/npcpy

    Render the provided prompt template with Jinja context and send it to the active NPC's LLM.

    1.5k GitHub stars~161 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Pi Agent

    K-Dense-AI/scientific-agent-skills

    Builds with and operates Pi, the minimal terminal coding harness.

    48k GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Workflow Schema Tuning

    breaking-brake/cc-wf-studio

    Guides edits to cc-wf-studio's workflow schema so AI agents generate better workflows, treating schema text as prompt engineering rather than validation.

    5.4k GitHub stars~1.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Hcls Build Agent

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    A skill your agent uses when a developer wants to build a new healthcare or life sciences agent, structure tools and system prompts for an HCLS workflow, or create a Strands agent with…

    274 GitHub stars~885 tokensUpdated 6 days ago
    Research & ScienceAuto-check passed
  • Add Prompt

    cyanheads/pubmed-mcp-server

    Scaffold a new MCP prompt template. An agent skill from cyanheads/pubmed-mcp-server.

    154 GitHub stars~1.6k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed

More from comet-ml/opik-mcp

All 10 skills in this repo
  • Opik

    comet-ml/opik-mcp

    Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST).

    219 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Opik Compare

    comet-ml/opik-mcp

    Run a candidate against the baseline over an Opik test suite and read the numbers back — which cases broke, which got fixed, the per-metric deltas, worst rows, and whether the two runs are…

    219 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Opik Diagnose

    comet-ml/opik-mcp

    Surface the Opik traces worth a developer's attention, ranked by signal — Diagnostics issues first, then errors, failed tool calls, latency, regressions, and low online-eval scores.

    219 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes
  • Opik Evaluate

    comet-ml/opik-mcp

    Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link.

    219 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check: notes
  • Opik Instrument

    comet-ml/opik-mcp

    Add Opik tracing to an existing app and verify a real trace lands.

    219 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes
  • Opik Verify

    comet-ml/opik-mcp

    Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost…

    219 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check: notes

Questions about Opik Optimize

What does Opik Optimize do?

Improve a prompt with the Opik Agent Optimizer — resolve the prompt, a dataset, and a metric, pick the algorithm, run a bounded optimization, check the gain on held-out data, and save the winner as…. Opik Optimize is an agent skill from comet-ml/opik-mcp. Improve a prompt with the Opik Agent Optimizer — resolve the prompt, a dataset, and a metric, pick the algorithm, run a bounded optimization, check the gain on held-out data, and save the winner as a new prompt version with the optimization run link.

When should I use Opik Optimize?

Opik Optimize fits situations like: optimize this prompt; improve my system prompt; make the agent answer better; tune the prompt against my dataset.

How do I install Opik Optimize in Claude Code?

Run `npx skills add comet-ml/opik-mcp --skill opik-optimize -a claude-code`. Or copy the skill folder (src/opik_mcp/skills/opik-optimize in comet-ml/opik-mcp) into .claude/skills/opik-optimize in your project. Claude Code loads it when a task matches its description.

How do I install Opik Optimize in Codex?

Run `npx skills add comet-ml/opik-mcp --skill opik-optimize -a codex`. Or copy the skill folder (src/opik_mcp/skills/opik-optimize in comet-ml/opik-mcp) into .agents/skills/opik-optimize in your project. Codex loads it when a task matches its description.

Can I use Opik Optimize in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add comet-ml/opik-mcp --skill opik-optimize -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/opik-optimize, .gemini/skills/opik-optimize, .github/skills/opik-optimize and .opencode/skills/opik-optimize in your project.

What does Opik Optimize need to run?

Going by SKILL.md and its folder, Opik Optimize needs Python for the scripts in its folder, the command-line tools its instructions call (uv and pip) and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, Write. Compatibility (from SKILL.md): Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python project with Opik configured, a provider API key, and a dataset (or traces to build one). Install the `opik` skill alongside this one — it holds the shared dataset and prompt-library references; without it, this skill falls back to the public docs..

Does Opik Optimize access the network?

SKILL.md names 1 domain. As links in the text: comet.com. This is read from the text; nothing was executed.

Is Opik Optimize safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Opik Optimize use?

Opik Optimize is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Opik Optimize use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 768 tokens, read only when the agent opens those files.

What are the alternatives to Opik Optimize?

Skills that share tags, products or a category with Opik Optimize: Tool Creator (hAcKlyc/MyAgents, 915 stars), Natural (NPC-Worldwide/npcpy, 1.5k stars), Pi Agent (K-Dense-AI/scientific-agent-skills, 48k stars) and Workflow Schema Tuning (breaking-brake/cc-wf-studio, 5.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Opik Optimize?

comet-ml (a GitHub organization) maintains it in comet-ml/opik-mcp, which has 219 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 6, 2026.

Source: comet-ml/opik-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.