Agent skill

OpenTelemetry Pipeline Metrics Spec

by comet-ml in comet-ml/opik

Specifies how to instrument an opik-backend pipeline with per-stage OpenTelemetry metrics for throughput, latency, errors and queue delay by workspace.

Apache-2.0Auto-check passedDevOps & Cloud

Install OpenTelemetry Pipeline Metrics Spec

skills CLI
$ npx skills add comet-ml/opik --skill metrics-instrumentation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install comet-ml/opik metrics-instrumentation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/comet-ml/opik.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/metrics-instrumentation .claude/skills/metrics-instrumentation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
metrics-instrumentation
GitHub stars
22k
Token cost
~3.2k tokens
SKILL.md length
1,434 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Specifies how to instrument an opik-backend pipeline with per-stage OpenTelemetry metrics for throughput, latency, errors and queue delay by workspace.

  • Works in 5 steps: Model → Backend metrics (OpenTelemetry) → Tests → …
  • Adding per-stage metrics to a scoring, ingestion or job pipeline in opik-backend
  • SKILL.md covers 1. Model, 2. Backend metrics…, 3. Tests and 4. Constraints (normative), plus 1 more section
  • Calls mvn; needs TYPE_KEY and WORKSPACE_ID_KEY

What it does

This is a normative spec for the emission half of operational observability in the Opik backend, using online scoring as the worked example. Metrics are designed so a flow-ordered Grafana dashboard reads top to bottom and the failing stage is where the numbers break, with a per-workspace drill-down. Building the dashboard itself is a separate skill, as is product analytics with PostHog.

The patterns include producer metrics counted at the source, such as sampler decisions and enqueue results; consumer metrics for throughput and per-stage timing; RED measurements at the HTTP ingest entrypoint; a dedicated error counter per stage by error type; and success derived as throughput minus errors. Queue delay is kept apart from processing time so a backlog can be told from a slow scorer, backpressure skips are counted as benign, and gauges for in-flight work and JVM heap show saturation before failures start.

When your agent uses it

  • Adding per-stage metrics to a scoring, ingestion or job pipeline in opik-backend
  • Making a pipeline's failing stage visible by customer workspace
  • Separating queue backlog from slow processing in metrics

Example prompts

  • “Instrument the experiment processing pipeline with per-stage OpenTelemetry metrics.”
  • “Add error counters by error type and workspace to the ingestion path.”
  • “Which metrics should the enqueue and consumer stages emit?”

Requirements

  • The opik-backend codebase with OpenTelemetry metrics support

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Model
  2. Backend metrics (OpenTelemetry)
  3. Tests
  4. Constraints (normative)
  5. Delivery

What it can do on your machine

Read from SKILL.md and the folder at commit ddde888. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • mvn

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • TYPE_KEY
    • WORKSPACE_ID_KEY
    • WORKSPACE_NAME_KEY
    • RESULT_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

OpenTelemetry Pipeline Metrics Spec loads about 3.2k tokens when it runs. Until then it costs about 120 tokens; SKILL.md has 1,434 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from comet-ml/opik at commit ddde888, republished under its Apache-2.0 licence (© comet-ml). 1,434 words, ~3,206 tokens.

Download SKILL.mdSave it as .claude/skills/metrics-instrumentation/SKILL.md (or your agent's skills folder).
name
metrics-instrumentation
description
Specification for instrumenting an opik-backend workflow with operational OpenTelemetry metrics — per-stage throughput/latency/error counters and native histograms, dimensioned per-customer (workspace). Use when a pipeline (scoring, ingestion, experiments, jobs) needs per-stage visibility. Covers metric emission only; building the Grafana dashboard from these metrics is a separate skill. Distinct from analytics-instrumentation (PostHog product events).

Metrics Instrumentation

Normative spec for the backend half of operational observability: per-stage OpenTelemetry metrics in apps/opik-backend. The metrics are designed so a flow-ordered Grafana dashboard can read top-to-bottom — the failing stage is the one whose numbers break — with a per-workspace drill. Building that dashboard (layout, query contracts, per-customer name resolution, dependency panels, validation) is a separate concern, specified by the dashboard-authoring skill in comet monitoring tooling; this skill covers only what to emit.

Patterns applied in this implementation (online scoring is the worked example):

  • Producer metrics — every stage that emits work is counted at the source: sampler decisions (sampler_decisions_total{decision}) and enqueue-to-Redis (enqueue_total{result}). A producer that emits nothing makes downstream starvation explainable rather than mysterious.
  • Consumer metrics — throughput and per-stage timing on the side that drains the queue (processing_time, plus per-Redis-op read/claim/ack_and_remove/list_pending_time).
  • Entrypoint RED — the workflow's front door (the HTTP ingest route) is measured Rate / Errors / Duration from http_server_request_duration_seconds, with 5xx broken down by endpoint × error_type × workspace, so an ingestion problem is never mistaken for a scoring problem.
  • Errors — a dedicated error counter per stage, dimensioned by error_type (the exception class) and, for shared async plumbing, by component (listener/subscriber): processing_errors_total, unexpected_errors_total, and enqueue_total{result="error"} (a push failure = real loss).
  • Success — success is derived, never double-counted: throughput − errors, surfaced as one "success rate" headline tile.
  • Queue time & end-to-end latency — queue_delay (enqueue→pickup) is kept separate from processing_time (scorer/LLM work) so a backlog is distinguishable from a slow scorer; end-to-end = queue_delay + processing_time.
  • Backpressure — poll-tick skips are counted but are benign (consumer busy), never lost work.
  • Saturation & resource levels — gauges for in-flight work (max per pod) and JVM heap used-vs-limit per pod expose the pipeline approaching a ceiling before it starts failing (the USE method — utilization / saturation / errors — complementing RED).
  • Volume & payload size — byte/char counters (bandwidth, total bytes) and payload-size distributions, broken down by content type and workspace, for cost and impact attribution.
  • Per-workspace dimensioning — the customer drill (§1.3) is a first-class label, not an afterthought.
  • Infrastructure dependencies — the datastores the flow leans on (Redis streams, ClickHouse, locks, MySQL) are surfaced on the dashboard from their exporters and system.query_log, so a "slow pipeline" resolves to the dependency causing it.

To see these conventions already in practice, grep apps/opik-backend for the existing metric families rather than specific classes (metric names are the stable contract; class locations move): the online-scoring *_sampler_decisions_total, *_enqueue_total{result}, *_processing_time_milliseconds, *_queue_delay_milliseconds, *_processing_errors_total{error_type} / *_unexpected_errors_total, the per-Redis-op *_{read,claim,ack_and_remove,list_pending}_time_milliseconds, and the attachment upload byte/size families.

Apply alongside the opik-backend skill (general conventions, logging rule).


1. Model

1.1 Decompose the workflow into ordered stages, each defined by: input, output, failure mode, existing instrumentation.

1.2 Choose the instrument by the question it answers. Most stages need a counter (did it happen, and how often?) and a histogram (how long did it take?); add a gauge only for a level you cannot derive from those.

  • Counter — a monotonically increasing total, read as a rate. Use for events: throughput, decisions, results, errors, and cumulative volume (messages, bytes). Anything you phrase "per second" or "how many since start". Surfaces with a _total suffix. MUST NOT be used for a value that can decrease.
  • Native histogram — a latency/size distribution, read with histogram_quantile. Use whenever a p95/p99 matters, not just an average: processing time, queue delay, per-dependency op time, end-to-end latency, and payload size (bytes/chars). Prefer native (no explicit buckets, §2.1) over classic le-bucketed histograms; reach for classic buckets only when a downstream consumer (e.g. an existing exporter) forces them. Do NOT use a histogram where a counter suffices — a plain success/error tally needs no distribution.
  • Gauge — an instantaneous level, read as-is. Use for a quantity that rises and falls and can't be reconstructed from a counter: current queue depth / backlog, in-flight / in-progress count (saturation), batch/read/claim size, lock waiters, heap used. If you actually want a trend, count events with a counter rather than sampling a level with a gauge.
  • UpDownCounter (OTel) — when the level is naturally maintained as signed +1/−1 deltas (in-flight = increment on start, decrement on finish) rather than sampled. Exports like a gauge but is cheaper and less racy than reading a size on every observation.

Rule of thumb: "how many happened" → counter; "how long / how big" → histogram; "how much is there right now" → gauge (or UpDownCounter). Cover each stage with RED (Rate, Errors, Duration) for the work flowing through it and USE (Utilization, Saturation, Errors) for the resource it runs on.

1.3 Define identity dimensions, each bounded in cardinality:

  • workspace_id and workspace_name — the customer drill. Always paired; workspace_name falls back to workspace_id when the name is absent (§2.3).
  • a stage/type label — what kind of work this is (evaluator_type, decision, the Redis op, content/mime type). Lets the dashboard sum by(...) per stage (§2.2).
  • result ∈ {success, error} on outcome counters.
  • error_type on every error counter — the exception class / failure category — plus a component label (listener/subscriber/endpoint) where one counter serves many call sites, so the outcome panel breaks errors down by both cause and origin.

Cardinality MUST stay bounded by #workspaces × #types × #error_types; an unbounded value (trace id, user input, raw message, full URL) MUST NOT be placed on a label.


Show full SKILL.md (597 more words)Show less

2. Backend metrics (OpenTelemetry)

2.1 Meters MUST be created via the OTel API (GlobalOpenTelemetry.getMeter(namespace)), one namespace per workflow:

java
private static final String METRIC_NAMESPACE = "<workflow>";
var meter = GlobalOpenTelemetry.getMeter(METRIC_NAMESPACE);
meter.counterBuilder("%s_<stage>_total".formatted(METRIC_NAMESPACE)).setDescription("…").build();
meter.histogramBuilder("%s_<stage>_time".formatted(METRIC_NAMESPACE)).setUnit("ms").ofLongs().build(); // native histogram
meter.gaugeBuilder("%s_<stage>_size".formatted(METRIC_NAMESPACE)).build();
meter.upDownCounterBuilder("%s_<stage>_in_flight".formatted(METRIC_NAMESPACE)).build(); // signed level (inc/dec)

Counters surface in Prometheus with a _total suffix. Native histograms MUST NOT define explicit buckets (no _bucket/_sum/_count/le).

2.2 The stage/type dimension MUST be a label, not part of the metric name (e.g. online_scoring_enqueue_total{evaluator_type, result}), so the dashboard can sum by(evaluator_type)(rate(...)). Name-encoding the dimension is permitted ONLY to extend an existing metric family; it complicates dashboard aggregation (a name-encoded dimension cannot be rate()d across names in one selector), so prefer a label.

2.3 workspace_id and workspace_name MUST be read from the reactive request context (the workspace-id / workspace-name context keys), reusing the shared workspace attribute-key constants rather than redeclaring them per call site:

java
var workspaceId   = ctx.getOrDefault(WORKSPACE_ID, "");
var workspaceName = StringUtils.defaultIfBlank(ctx.getOrDefault(WORKSPACE_NAME, workspaceId), workspaceId);
counter.add(1, Attributes.of(TYPE_KEY, type, WORKSPACE_ID_KEY, workspaceId, WORKSPACE_NAME_KEY, workspaceName, RESULT_KEY, "success"));
  • workspace_name MUST fall back to workspace_id when the name is absent.
  • A name-service lookup MUST NOT be used to resolve the name on a path where the reactive context already carries it.
  • If an event does not yet carry the name, add a nullable workspaceName field and populate it from the workspace-name context key at the publish site (as the entity-created events feeding this workflow already do).

2.4 The instrumented operation MUST be reactive and read the context with deferContextual:

java
public Mono<Void> enqueue(List<?> messages, Type type) {
    return Flux.deferContextual(ctx -> { /* resolve workspace per 2.3 */
        return Flux.fromIterable(messages).flatMap(m -> redisAdd(m)
            .doOnNext(id -> counter.add(1, successAttrs))
            .doOnError(e -> { counter.add(1, errorAttrs); log.error("Error … id='{}'", id, e); }));
    }).then().subscribeOn(Schedulers.boundedElastic());
}
  • The method MUST return Mono<Void> and MUST NOT self-subscribe. Callers MUST subscribe or compose it.
  • Request-scoped reactive callers MUST compose it (.then(...) / flatMap) so it inherits the workspace context.
  • Event-driven / synchronous callers MUST fire-and-forget via a single shared helper that subscribes with an explicit .contextWrite(ctx -> ctx.put(WORKSPACE_ID, id).put(WORKSPACE_NAME, defaultIfBlank(name, id))) and an error-logging consumer (one helper, not duplicated per caller).
  • Blocking lookups (JDBC/findById) MUST run via Mono.fromCallable(...).subscribeOn(Schedulers.boundedElastic()). The enqueue/IO work MUST run on a bounded scheduler, not the caller's (e.g. EventBus) thread. Where a caller already holds a resolved object, provide an overload that skips the lookup.

2.5 Log statements MUST single-quote placeholders (evaluator='{}' workspaceId='{}') per .agents/skills/opik-backend/SKILL.md, and SHOULD include the batch size on enqueue logs.

2.6 mvn -o compile MUST succeed (spotless clean) before delivery.


3. Tests

3.1 Introducing a Mono<Void> (lazy) return breaks tests that called the method for its side effect. Restore green by stubbing the reactive method (lenient().when(pub.enqueue(any(),any())).thenReturn(Mono.empty())) where production composes it, .block()-ing the returned Mono in unit tests that assert downstream effects, and replacing verifyNoInteractions(mock) with verify(mock, never()).method(...) where a lenient stub now exists.


4. Constraints (normative)

4.1 A returned Mono is inert until subscribed; an unsubscribed enqueue is a silent no-op. Every call site MUST compose or subscribe it.

4.2 The workspace MUST be sourced from the reactive context, not a name-service lookup, where the context carries it (§2.3).

4.3 IO/enqueue work MUST run on a bounded scheduler, off the caller's thread (§2.4).

4.4 A backpressure / poll-tick counter (e.g. backpressure_drops_total) counts skipped scheduler ticks while a consumer is busy; it is NOT lost work and MUST NOT be alerted on alone. Emit it, but document it as benign for whoever builds the dashboard.

4.5 Work MUST be done against origin/main (create the worktree from it), not a possibly-stale local checkout.


5. Delivery

5.1 The change is delivered as two PRs on their own branches <user>/<TASK>-<name>: this metrics PR (opik repo) and a companion dashboard PR built per the dashboard-authoring skill and delivered to comet monitoring.

5.2 The metrics PR (this repo) MUST contain no customer, cluster, or infrastructure identifiers, MUST follow .github/pull_request_template.md including a ## Documentation section (the PR linter fails without it), ## Issues (Resolves OPIK-XXXX), and ## AI-WATERMARK: yes with tools/model/scope/human-verification. Title [OPIK-XXXX] [BE] …. Announce per .agents/commands/comet/send-code-review-slack.md.

5.3 The dashboard panel for a new metric stays empty until this backend PR deploys — so the two PRs are independent and can land in either order.

© comet-ml, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/metrics-instrumentation of comet-ml/opik.

Open the folder on GitHubat commit ddde888

Compare with similar skills

OpenTelemetry Pipeline Metrics Spec next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

OpenTelemetry Pipeline Metrics Spec compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
OpenTelemetry Pipeline Metrics Spec this skillcomet-ml/opik22k—~3.2kAutomated safety check: PassApache-2.0
Archestra Dev Observabilityarchestra-ai/archestra4.4k—~1.2kAutomated safety check: PassCustom licence
Frontmcp Observabilityagentfront/frontmcp146—~4.6kAutomated safety check: PassApache-2.0
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Monitoring ExpertJeffallan/claude-skills12k—~1.6kAutomated safety check: PassMIT
Alloygrafana/skills282—~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • Archestra Dev Observability

    archestra-ai/archestra

    A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup.

    4.4k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Frontmcp Observability

    agentfront/frontmcp

    A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server.

    146 GitHub stars~4.6k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 6 mo ago
    DevOps & CloudAuto-check passed
  • Monitoring Expert

    Jeffallan/claude-skills

    Sets up application monitoring: structured logs, Prometheus metrics, OpenTelemetry tracing, Grafana dashboards, alert rules and load tests with k6 or Artillery.

    12k GitHub stars~1.6k tokensUpdated 6 days ago
    DevOps & CloudAuto-check passed
  • Alloy

    grafana/skills

    Official

    Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /…

    282 GitHub stars~1.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Golang Observability

    context-labs/whip

    Go observability — always-on production signals: slog logging, Prometheus metrics, OpenTelemetry tracing, pprof profiling, alerting, Grafana.

    1.1k GitHub starsUsed in 1 repo~3.3k tokens
    DevOps & CloudAuto-check passed

More from comet-ml/opik

All 19 skills in this repo
  • Checklist for wiring a new linter into Opik's Code Quality pipeline: the four files to edit, the silent-failure gotchas and the pass/fail verification loop.

    22k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Shows how to add product analytics events to Opik's frontend, Java backend and Python SDK, all reporting through Segment to PostHog with an opik_ name prefix.

    22k GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests.

    22k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Rules for writing PR descriptions, changelog entries and feature documentation in the Opik repository, including the exact headings that CI requires.

    22k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Turns a code change into one committed, passing Playwright end-to-end spec by resolving the change scope and handing authoring to a companion skill.

    22k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Starts, rebuilds, and troubleshoots the Opik local dev stack, including an optional Comet Platform integration mode for the Opik team.

    22k GitHub stars~734 tokensUpdated today
    Auto-check passed

Categories

Questions about OpenTelemetry Pipeline Metrics Spec

What does OpenTelemetry Pipeline Metrics Spec do?

Specifies how to instrument an opik-backend pipeline with per-stage OpenTelemetry metrics for throughput, latency, errors and queue delay by workspace. This is a normative spec for the emission half of operational observability in the Opik backend, using online scoring as the worked example. Metrics are designed so a flow-ordered Grafana dashboard reads top to bottom and the failing stage is where the numbers break, with a per-workspace drill-down.

When should I use OpenTelemetry Pipeline Metrics Spec?

OpenTelemetry Pipeline Metrics Spec fits situations like: adding per-stage metrics to a scoring, ingestion or job pipeline in opik-backend; making a pipeline's failing stage visible by customer workspace; separating queue backlog from slow processing in metrics.

How do I install OpenTelemetry Pipeline Metrics Spec in Claude Code?

Run `npx skills add comet-ml/opik --skill metrics-instrumentation -a claude-code`. Or copy the skill folder (.agents/skills/metrics-instrumentation in comet-ml/opik) into .claude/skills/metrics-instrumentation in your project. Claude Code loads it when a task matches its description.

How do I install OpenTelemetry Pipeline Metrics Spec in Codex?

Run `npx skills add comet-ml/opik --skill metrics-instrumentation -a codex`. Or copy the skill folder (.agents/skills/metrics-instrumentation in comet-ml/opik) into .agents/skills/metrics-instrumentation in your project. Codex loads it when a task matches its description.

Can I use OpenTelemetry Pipeline Metrics Spec in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add comet-ml/opik --skill metrics-instrumentation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/metrics-instrumentation, .gemini/skills/metrics-instrumentation, .github/skills/metrics-instrumentation and .opencode/skills/metrics-instrumentation in your project.

What does OpenTelemetry Pipeline Metrics Spec need to run?

Going by SKILL.md and its folder, OpenTelemetry Pipeline Metrics Spec needs the command-line tools its instructions call (mvn) and credentials named TYPE_KEY, WORKSPACE_ID_KEY, WORKSPACE_NAME_KEY and RESULT_KEY. Our summary lists: The opik-backend codebase with OpenTelemetry metrics support.

Does OpenTelemetry Pipeline Metrics Spec access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is OpenTelemetry Pipeline Metrics Spec safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does OpenTelemetry Pipeline Metrics Spec use?

OpenTelemetry Pipeline Metrics Spec is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does OpenTelemetry Pipeline Metrics Spec use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to OpenTelemetry Pipeline Metrics Spec?

Skills that share tags, products or a category with OpenTelemetry Pipeline Metrics Spec: Archestra Dev Observability (archestra-ai/archestra, 4.4k stars), Frontmcp Observability (agentfront/frontmcp, 146 stars), Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars) and Monitoring Expert (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains OpenTelemetry Pipeline Metrics Spec?

comet-ml (a GitHub organization) maintains it in comet-ml/opik, which has 22,478 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 10, 2026.

Source: comet-ml/opik on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.