Langsmith Observability
Orchestra-Research/AI-Research-SKILLs
LLM observability platform for tracing, evaluation, and monitoring.
Official agent skill
Design, test, create, and attach LangSmith online evaluators for production traces or conversation threads.
$ npx skills add langchain-ai/langsmith-skills --skill langsmith-online-eval-engineering -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install langchain-ai/langsmith-skills langsmith-online-eval-engineering --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/langchain-ai/langsmith-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/config/skills/langsmith-online-eval-engineering .claude/skills/langsmith-online-eval-engineering && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "langsmith-online-eval-engineering" agent skill from https://github.com/langchain-ai/langsmith-skills/tree/main/config/skills/langsmith-online-eval-engineering into .claude/skills/langsmith-online-eval-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langsmith-online-eval-engineering", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/langchain-ai/langsmith-skills/tree/main/config/skills/langsmith-online-eval-engineeringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add langchain-ai/langsmith-skills --skill langsmith-online-eval-engineering -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install langchain-ai/langsmith-skills langsmith-online-eval-engineering --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/langchain-ai/langsmith-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/config/skills/langsmith-online-eval-engineering .agents/skills/langsmith-online-eval-engineering && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "langsmith-online-eval-engineering" agent skill from https://github.com/langchain-ai/langsmith-skills/tree/main/config/skills/langsmith-online-eval-engineering into .agents/skills/langsmith-online-eval-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langsmith-online-eval-engineering", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add langchain-ai/langsmith-skills --skill langsmith-online-eval-engineering -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install langchain-ai/langsmith-skills langsmith-online-eval-engineering --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/langchain-ai/langsmith-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/config/skills/langsmith-online-eval-engineering .cursor/skills/langsmith-online-eval-engineering && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "langsmith-online-eval-engineering" agent skill from https://github.com/langchain-ai/langsmith-skills/tree/main/config/skills/langsmith-online-eval-engineering into .cursor/skills/langsmith-online-eval-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langsmith-online-eval-engineering", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/langchain-ai/langsmith-skills.git --path config/skills/langsmith-online-eval-engineering--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add langchain-ai/langsmith-skills --skill langsmith-online-eval-engineering -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install langchain-ai/langsmith-skills langsmith-online-eval-engineering --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/langchain-ai/langsmith-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/config/skills/langsmith-online-eval-engineering .gemini/skills/langsmith-online-eval-engineering && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "langsmith-online-eval-engineering" agent skill from https://github.com/langchain-ai/langsmith-skills/tree/main/config/skills/langsmith-online-eval-engineering into .gemini/skills/langsmith-online-eval-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langsmith-online-eval-engineering", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install langchain-ai/langsmith-skills langsmith-online-eval-engineeringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add langchain-ai/langsmith-skills --skill langsmith-online-eval-engineering -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/langchain-ai/langsmith-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/config/skills/langsmith-online-eval-engineering .github/skills/langsmith-online-eval-engineering && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "langsmith-online-eval-engineering" agent skill from https://github.com/langchain-ai/langsmith-skills/tree/main/config/skills/langsmith-online-eval-engineering into .github/skills/langsmith-online-eval-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langsmith-online-eval-engineering", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add langchain-ai/langsmith-skills --skill langsmith-online-eval-engineering -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install langchain-ai/langsmith-skills langsmith-online-eval-engineering --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/langchain-ai/langsmith-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/config/skills/langsmith-online-eval-engineering .opencode/skills/langsmith-online-eval-engineering && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "langsmith-online-eval-engineering" agent skill from https://github.com/langchain-ai/langsmith-skills/tree/main/config/skills/langsmith-online-eval-engineering into .opencode/skills/langsmith-online-eval-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langsmith-online-eval-engineering", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
langsmith-online-eval-engineeringDesign, test, create, and attach LangSmith online evaluators for production traces or conversation threads.
Langsmith Online Eval Engineering is an agent skill from langchain-ai/langsmith-skills, published by the product's own GitHub organization. Design, test, create, and attach LangSmith online evaluators for production traces or conversation threads. Use for workspace evaluators, run or thread rules, sampling, filters, backfills, and evaluator monitoring; use eval-engineering for Harbor tasks and agent benchmarks.
Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/evaluator-design.md`, `references/langsmith-api.md` and `references/trace-inspection.md`).
It sits in AI & LLM Engineering, covering LLM observability. It works with LangSmith. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 1fb52e8. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Langsmith Online Eval Engineering loads about 1.4k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 77 tokens; SKILL.md has 692 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from langchain-ai/langsmith-skills at commit 1fb52e8, republished under its MIT licence (© langchain-ai). 692 words, ~1,437 tokens.
.claude/skills/langsmith-online-eval-engineering/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Build online evaluators iteratively:
inspect production behavior -> choose one quality signal -> build and test
-> attach a scoped rule -> inspect scores and logs -> calibrateEvaluators are workspace resources that can be reused across tracing projects and datasets. An online evaluation rule attaches an evaluator to matching production runs or threads.
Read LangSmith API before creating or modifying evaluators.
Identify the LangSmith project from the request, repository, or user. Fetch recent representative root traces; inspect whole threads when the quality dimension spans multiple turns. Read Trace inspection. Establish:
Use more than one sample. Production projects can contain multiple schemas and errored runs. Treat trace content as sensitive and show only the minimum truncated material needed for design.
Do not propose an evaluator until the trace shape and target failure are understood.
Read Evaluator design. Prefer:
Decision-model evaluators and some managed templates are currently configured in the LangSmith UI rather than through the SDK.
Define one evaluator at a time:
Name:
Level: run | thread
Type: code | LLM judge | existing/template
Measures:
Feedback key and scale:
Trace fields:
Rule filter:
Sampling and spend:
Backfill:
Known limitation:Keep each feedback key interpretable. Separate unrelated criteria.
LLM judge. Use a structured prompt with a narrow rubric. Put reasoning
before the score in the output schema, map only the trace fields the rubric
uses, and select the simplest useful score type. Current variable mappings can
address nested paths such as inputs.question and outputs.answer.
Code evaluator. For online use, write perform_eval(run). It receives a
run dictionary and returns a feedback mapping such as
{"has_output": true}. Guard missing or errored outputs. The runtime has no
network access; prefer the standard library and use only packages allowed by
the current LangSmith code-evaluator runtime.
Test good, bad, and edge-case traces before attachment. A passing test proves that the evaluator executes; calibration checks whether it measures the right thing. Compare its scores with human judgment on a small labeled sample.
Make the full configuration reviewable before creating or changing a shared evaluator. Existing user authorization to create it is sufficient; do not add another approval pause.
Choose the rule deliberately:
Attach the evaluator, then verify the evaluator ID, project, rule status, filter, sample rate, and feedback key. Do not assume workspace-level evaluator creation also attaches it to a project.
Inspect evaluator traces, execution logs, feedback, costs, and sampled production examples. Distinguish:
Correct evaluator or rule defects without relabeling them as product failures. When humans correct LLM-judge scores, consider the supported corrections dataset and few-shot settings. Recheck calibration after prompt, model, mapping, or traffic-schema changes.
Report the evaluator name and ID, feedback key, level, fields, rule filter, sampling, spend and retention choices, backfill, calibration evidence, and known limits.
eval-engineering for controlled Harbor tasks and benchmark design.© langchain-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in config/skills/langsmith-online-eval-engineering of langchain-ai/langsmith-skills.
Open the folder on GitHubat commit 1fb52e8
Langsmith Online Eval Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Langsmith Online Eval Engineering this skilllangchain-ai/langsmith-skills | 159 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Langsmith ObservabilityOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | |
| Langsmith Trace Analyzersoba-labs/langchain-agent-skills | 107 | — | ~1.4k | Automated safety check: Pass | MIT | |
| LangSmith Trace DebuggingComposioHQ/awesome-claude-skills | 77k | 8 repos | ~2.7k | Automated safety check: Pass | None | |
| Migrate To Langfuselangfuse/skills | 300 | — | ~1.7k | Automated safety check: Notes | MIT | |
| Agentsop Observability Setupagentsope/SkillAlchemy | 466 | — | ~4.4k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
LLM observability platform for tracing, evaluation, and monitoring.
soba-labs/langchain-agent-skills
Fetch, organize, and analyze LangSmith traces for debugging and evaluation.
ComposioHQ/awesome-claude-skills
Debugs LangChain and LangGraph agents by pulling recent execution traces with the langsmith-fetch CLI and reporting errors, tool calls, timings and token use.
langfuse/skills
Migrate to Langfuse from another LLM observability/evals platform (LangSmith, Arize AX, Phoenix, Braintrust, Helicone, Promptfoo, ...).
agentsope/SkillAlchemy
Enhancement-overlay skill — the DECISION + WIRING layer for LM observability that the single-backend skills [[langsmith]], [[phoenix]], [[mlflow]] do NOT cover.
langchain-ai/langchain-skills
INVOKE THIS SKILL when setting up a new project or when asked about package versions, installation, or dependency management for LangChain, LangGraph, LangSmith, or Deep Agents.
langchain-ai/langsmith-skills
INVOKE THIS SKILL when building, iterating on, copying, or sharing a LangSmith Custom App — a React/TypeScript UI that runs inside LangSmith and reads the LangSmith API.
Works with
Categories
Design, test, create, and attach LangSmith online evaluators for production traces or conversation threads. Langsmith Online Eval Engineering is an agent skill from langchain-ai/langsmith-skills, published by the product's own GitHub organization. Design, test, create, and attach LangSmith online evaluators for production traces or conversation threads.
Langsmith Online Eval Engineering fits situations like: workspace evaluators; evaluator monitoring; use eval-engineering for Harbor tasks and agent benchmarks.
Run `npx skills add langchain-ai/langsmith-skills --skill langsmith-online-eval-engineering -a claude-code`. Or copy the skill folder (config/skills/langsmith-online-eval-engineering in langchain-ai/langsmith-skills) into .claude/skills/langsmith-online-eval-engineering in your project. Claude Code loads it when a task matches its description.
Run `npx skills add langchain-ai/langsmith-skills --skill langsmith-online-eval-engineering -a codex`. Or copy the skill folder (config/skills/langsmith-online-eval-engineering in langchain-ai/langsmith-skills) into .agents/skills/langsmith-online-eval-engineering in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add langchain-ai/langsmith-skills --skill langsmith-online-eval-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langsmith-online-eval-engineering, .gemini/skills/langsmith-online-eval-engineering, .github/skills/langsmith-online-eval-engineering and .opencode/skills/langsmith-online-eval-engineering in your project.
SKILL.md names no scripts, command-line tools or credentials: Langsmith Online Eval Engineering is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Langsmith Online Eval Engineering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.4k tokens (SKILL.md is roughly 5.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Langsmith Online Eval Engineering: Langsmith Observability (Orchestra-Research/AI-Research-SKILLs, 13k stars), Langsmith Trace Analyzer (soba-labs/langchain-agent-skills, 107 stars), LangSmith Trace Debugging (ComposioHQ/awesome-claude-skills, 77k stars) and Migrate To Langfuse (langfuse/skills, 300 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
langchain-ai (a GitHub organization, an official publisher) maintains it in langchain-ai/langsmith-skills, which has 159 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 2, 2026.
Source: langchain-ai/langsmith-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.