Official agent skill

Langsmith Online Eval Engineering

by langchain-ai in langchain-ai/langsmith-skills

Design, test, create, and attach LangSmith online evaluators for production traces or conversation threads.

OfficialMITAuto-check passedAI & LLM Engineering

Install Langsmith Online Eval Engineering

skills CLI
$ npx skills add langchain-ai/langsmith-skills --skill langsmith-online-eval-engineering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install langchain-ai/langsmith-skills langsmith-online-eval-engineering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/langchain-ai/langsmith-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/config/skills/langsmith-online-eval-engineering .claude/skills/langsmith-online-eval-engineering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langsmith-online-eval-engineering
GitHub stars
159
Token cost
~1.4k tokens
SKILL.md length
692 words
Files
4 (incl. references)
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Design, test, create, and attach LangSmith online evaluators for production traces or conversation threads.

  • Works in 5 steps: Inspect representative behavior → Choose one evaluation signal → Build and test → …
  • Workspace evaluators
  • SKILL.md covers 1. Inspect representative…, 2. Choose one evaluation signal, 3. Build and test and 4. Attach an online rule, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Langsmith Online Eval Engineering is an agent skill from langchain-ai/langsmith-skills, published by the product's own GitHub organization. Design, test, create, and attach LangSmith online evaluators for production traces or conversation threads. Use for workspace evaluators, run or thread rules, sampling, filters, backfills, and evaluator monitoring; use eval-engineering for Harbor tasks and agent benchmarks.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/evaluator-design.md`, `references/langsmith-api.md` and `references/trace-inspection.md`).

It sits in AI & LLM Engineering, covering LLM observability. It works with LangSmith. The licence is MIT.

When your agent uses it

  • Workspace evaluators
  • Evaluator monitoring
  • Use eval-engineering for Harbor tasks and agent benchmarks

Example prompts

  • “/langsmith-online-eval-engineering”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Inspect representative behavior
  2. Choose one evaluation signal
  3. Build and test
  4. Attach an online rule
  5. Monitor and calibrate

What it can do on your machine

Read from SKILL.md and the folder at commit 1fb52e8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Langsmith Online Eval Engineering loads about 1.4k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 77 tokens; SKILL.md has 692 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from langchain-ai/langsmith-skills at commit 1fb52e8, republished under its MIT licence (© langchain-ai). 692 words, ~1,437 tokens.

Download SKILL.mdSave it as .claude/skills/langsmith-online-eval-engineering/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
langsmith-online-eval-engineering
description
Design, test, create, and attach LangSmith online evaluators for production traces or conversation threads. Use for workspace evaluators, run or thread rules, sampling, filters, backfills, and evaluator monitoring; use eval-engineering for Harbor tasks and agent benchmarks.

LangSmith Online Eval Engineering

Build online evaluators iteratively:

text
inspect production behavior -> choose one quality signal -> build and test
-> attach a scoped rule -> inspect scores and logs -> calibrate

Evaluators are workspace resources that can be reused across tracing projects and datasets. An online evaluation rule attaches an evaluator to matching production runs or threads.

Read LangSmith API before creating or modifying evaluators.

1. Inspect representative behavior

Identify the LangSmith project from the request, repository, or user. Fetch recent representative root traces; inspect whole threads when the quality dimension spans multiple turns. Read Trace inspection. Establish:

  • run names and types;
  • actual input, output, attachment, metadata, and error shapes;
  • child tool or model activity when trajectory matters;
  • volume, traffic segments, and existing feedback;
  • the product failure or quality concern the evaluator should detect.

Use more than one sample. Production projects can contain multiple schemas and errored runs. Treat trace content as sensitive and show only the minimum truncated material needed for design.

Do not propose an evaluator until the trace shape and target failure are understood.

2. Choose one evaluation signal

Read Evaluator design. Prefer:

  • code for deterministic structure, limits, fields, or invariants;
  • an LLM judge for semantic quality that requires reading and judgment;
  • a thread-level evaluator for behavior that depends on several turns;
  • an existing workspace evaluator, tuned evaluator, or template when it already measures the intended signal.

Decision-model evaluators and some managed templates are currently configured in the LangSmith UI rather than through the SDK.

Define one evaluator at a time:

text
Name:
Level: run | thread
Type: code | LLM judge | existing/template
Measures:
Feedback key and scale:
Trace fields:
Rule filter:
Sampling and spend:
Backfill:
Known limitation:

Keep each feedback key interpretable. Separate unrelated criteria.

3. Build and test

LLM judge. Use a structured prompt with a narrow rubric. Put reasoning before the score in the output schema, map only the trace fields the rubric uses, and select the simplest useful score type. Current variable mappings can address nested paths such as inputs.question and outputs.answer.

Code evaluator. For online use, write perform_eval(run). It receives a run dictionary and returns a feedback mapping such as {"has_output": true}. Guard missing or errored outputs. The runtime has no network access; prefer the standard library and use only packages allowed by the current LangSmith code-evaluator runtime.

Test good, bad, and edge-case traces before attachment. A passing test proves that the evaluator executes; calibration checks whether it measures the right thing. Compare its scores with human judgment on a small labeled sample.

Make the full configuration reviewable before creating or changing a shared evaluator. Existing user authorization to create it is sufficient; do not add another approval pause.

Show full SKILL.md (295 more words)Show less

4. Attach an online rule

Choose the rule deliberately:

  • Level: one run or an idle conversation thread.
  • Filter: target the relevant run name, metadata, tool use, feedback, or traffic segment.
  • Sampling: use 1.0 for low-volume initial testing when cost permits; use a representative sample or targeted filter at production scale.
  • Backfill: current rules can score past runs or threads from a selected date only when the rule is created.
  • Spend: set an appropriate weekly limit for LLM judges.
  • Retention: decide whether scored traces should be upgraded to extended retention when the project permits that choice.

Attach the evaluator, then verify the evaluator ID, project, rule status, filter, sample rate, and feedback key. Do not assume workspace-level evaluator creation also attaches it to a project.

5. Monitor and calibrate

Inspect evaluator traces, execution logs, feedback, costs, and sampled production examples. Distinguish:

  • evaluator bugs or missing-field failures;
  • rule filters that select the wrong traffic;
  • judge disagreement or rubric ambiguity;
  • application quality failures;
  • authentication, API, or automation failures.

Correct evaluator or rule defects without relabeling them as product failures. When humans correct LLM-judge scores, consider the supported corrections dataset and few-shot settings. Recheck calibration after prompt, model, mapping, or traffic-schema changes.

Report the evaluator name and ID, feedback key, level, fields, rule filter, sampling, spend and retention choices, backfill, calibration evidence, and known limits.

Invariants

  • One interpretable quality dimension per feedback key.
  • Inspect real traces; never guess field paths.
  • Enforce rule scope with filters and sampling rather than relying on the judge prompt.
  • Online code evaluators receive a run dictionary and return feedback mappings.
  • Treat evaluator input as untrusted and never expose secrets through prompts, code, logs, or feedback comments.
  • Treat API, auth, and rule failures as infrastructure failures.
  • Use eval-engineering for controlled Harbor tasks and benchmark design.

© langchain-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in config/skills/langsmith-online-eval-engineering of langchain-ai/langsmith-skills.

  • SKILL.md
  • references/evaluator-design.md
  • references/langsmith-api.md
  • references/trace-inspection.md

Open the folder on GitHubat commit 1fb52e8

Compare with similar skills

Langsmith Online Eval Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Langsmith Online Eval Engineering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Langsmith Online Eval Engineering this skilllangchain-ai/langsmith-skills159—~1.4kAutomated safety check: PassMIT
Langsmith ObservabilityOrchestra-Research/AI-Research-SKILLs13k2 repos~2.4kAutomated safety check: PassMIT
Langsmith Trace Analyzersoba-labs/langchain-agent-skills107—~1.4kAutomated safety check: PassMIT
LangSmith Trace DebuggingComposioHQ/awesome-claude-skills77k8 repos~2.7kAutomated safety check: PassNone
Migrate To Langfuselangfuse/skills300—~1.7kAutomated safety check: NotesMIT
Agentsop Observability Setupagentsope/SkillAlchemy466—~4.4kAutomated safety check: PassMIT

Similar skills

  • Langsmith Observability

    Orchestra-Research/AI-Research-SKILLs

    LLM observability platform for tracing, evaluation, and monitoring.

    13k GitHub starsUsed in 2 repos~2.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Langsmith Trace Analyzer

    soba-labs/langchain-agent-skills

    Fetch, organize, and analyze LangSmith traces for debugging and evaluation.

    107 GitHub stars~1.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • LangSmith Trace Debugging

    ComposioHQ/awesome-claude-skills

    Debugs LangChain and LangGraph agents by pulling recent execution traces with the langsmith-fetch CLI and reporting errors, tool calls, timings and token use.

    77k GitHub starsUsed in 8 repos~2.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Migrate To Langfuse

    langfuse/skills

    Migrate to Langfuse from another LLM observability/evals platform (LangSmith, Arize AX, Phoenix, Braintrust, Helicone, Promptfoo, ...).

    300 GitHub stars~1.7k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check: notes
  • Agentsop Observability Setup

    agentsope/SkillAlchemy

    Enhancement-overlay skill — the DECISION + WIRING layer for LM observability that the single-backend skills [[langsmith]], [[phoenix]], [[mlflow]] do NOT cover.

    466 GitHub stars~4.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Langchain Dependencies

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when setting up a new project or when asked about package versions, installation, or dependency management for LangChain, LangGraph, LangSmith, or Deep Agents.

    1.3k GitHub stars~3.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from langchain-ai/langsmith-skills

  • Langsmith Custom Apps

    langchain-ai/langsmith-skills

    Official

    INVOKE THIS SKILL when building, iterating on, copying, or sharing a LangSmith Custom App — a React/TypeScript UI that runs inside LangSmith and reads the LangSmith API.

    159 GitHub stars~3.2k tokensUpdated 7 days ago
    Auto-check passed

Works with

Questions about Langsmith Online Eval Engineering

What does Langsmith Online Eval Engineering do?

Design, test, create, and attach LangSmith online evaluators for production traces or conversation threads. Langsmith Online Eval Engineering is an agent skill from langchain-ai/langsmith-skills, published by the product's own GitHub organization. Design, test, create, and attach LangSmith online evaluators for production traces or conversation threads.

When should I use Langsmith Online Eval Engineering?

Langsmith Online Eval Engineering fits situations like: workspace evaluators; evaluator monitoring; use eval-engineering for Harbor tasks and agent benchmarks.

How do I install Langsmith Online Eval Engineering in Claude Code?

Run `npx skills add langchain-ai/langsmith-skills --skill langsmith-online-eval-engineering -a claude-code`. Or copy the skill folder (config/skills/langsmith-online-eval-engineering in langchain-ai/langsmith-skills) into .claude/skills/langsmith-online-eval-engineering in your project. Claude Code loads it when a task matches its description.

How do I install Langsmith Online Eval Engineering in Codex?

Run `npx skills add langchain-ai/langsmith-skills --skill langsmith-online-eval-engineering -a codex`. Or copy the skill folder (config/skills/langsmith-online-eval-engineering in langchain-ai/langsmith-skills) into .agents/skills/langsmith-online-eval-engineering in your project. Codex loads it when a task matches its description.

Can I use Langsmith Online Eval Engineering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add langchain-ai/langsmith-skills --skill langsmith-online-eval-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langsmith-online-eval-engineering, .gemini/skills/langsmith-online-eval-engineering, .github/skills/langsmith-online-eval-engineering and .opencode/skills/langsmith-online-eval-engineering in your project.

What does Langsmith Online Eval Engineering need to run?

SKILL.md names no scripts, command-line tools or credentials: Langsmith Online Eval Engineering is instructions for the agent only.

Does Langsmith Online Eval Engineering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Langsmith Online Eval Engineering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Langsmith Online Eval Engineering use?

Langsmith Online Eval Engineering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Langsmith Online Eval Engineering use?

About 1.4k tokens (SKILL.md is roughly 5.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.7k tokens, read only when the agent opens those files.

What are the alternatives to Langsmith Online Eval Engineering?

Skills that share tags, products or a category with Langsmith Online Eval Engineering: Langsmith Observability (Orchestra-Research/AI-Research-SKILLs, 13k stars), Langsmith Trace Analyzer (soba-labs/langchain-agent-skills, 107 stars), LangSmith Trace Debugging (ComposioHQ/awesome-claude-skills, 77k stars) and Migrate To Langfuse (langfuse/skills, 300 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Langsmith Online Eval Engineering?

langchain-ai (a GitHub organization, an official publisher) maintains it in langchain-ai/langsmith-skills, which has 159 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 2, 2026.

Source: langchain-ai/langsmith-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.