LangSmith Trace Debugging
ComposioHQ/awesome-claude-skills
Debugs LangChain and LangGraph agents by pulling recent execution traces with the langsmith-fetch CLI and reporting errors, tool calls, timings and token use.
Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory…
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-eval-harness -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-eval-harness --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/langchain-eval-harness .claude/skills/langchain-eval-harness && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "langchain-eval-harness" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-eval-harness into .claude/skills/langchain-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-eval-harness", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-eval-harnessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-eval-harness -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-eval-harness --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/.curated/langchain-eval-harness .agents/skills/langchain-eval-harness && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "langchain-eval-harness" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-eval-harness into .agents/skills/langchain-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-eval-harness", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-eval-harness -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-eval-harness --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/.curated/langchain-eval-harness .cursor/skills/langchain-eval-harness && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "langchain-eval-harness" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-eval-harness into .cursor/skills/langchain-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-eval-harness", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jeremylongshore/tons-of-skills-marketplace.git --path skills/.curated/langchain-eval-harness--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-eval-harness -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-eval-harness --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/.curated/langchain-eval-harness .gemini/skills/langchain-eval-harness && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "langchain-eval-harness" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-eval-harness into .gemini/skills/langchain-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-eval-harness", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-eval-harnessInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-eval-harness -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/.curated/langchain-eval-harness .github/skills/langchain-eval-harness && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "langchain-eval-harness" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-eval-harness into .github/skills/langchain-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-eval-harness", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-eval-harness -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-eval-harness --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/.curated/langchain-eval-harness .opencode/skills/langchain-eval-harness && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "langchain-eval-harness" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-eval-harness into .opencode/skills/langchain-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-eval-harness", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
langchain-eval-harnessBuild reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory…
Langchain Eval Harness is an agent skill from jeremylongshore/tons-of-skills-marketplace. Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory analysis, and CI gating on quality regressions. Use when setting up quality measurement for a new chain, diagnosing regression after a model switch, or building an evaluation gate for a pull request. Trigger with "langchain eval", "langsmith evaluate", "ragas", "llm-as-judge", "agent trajectory eval", "eval regression gate".
Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/agent-trajectory-eval.md`, `references/ci-integration.md` and `references/framework-comparison.md`). Compatibility notes: Designed for Claude Code
It sits in AI & LLM Engineering, covering Building AI agents, LLM evaluation and LLM observability. It works with LangChain, LangSmith and LangGraph. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditBash(python:*)Bash(pip:*)Bash(pytest:*)From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.smith.langchain.comdocs.ragas.iodocs.confident-ai.comdocs.scipy.orgarxiv.orgFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
LANGSMITH_API_KEYOPENAI_API_KEYANTHROPIC_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Designed for Claude Code
From compatibility in the SKILL.md frontmatter.
Langchain Eval Harness loads about 3.7k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 135 tokens; SKILL.md has 1,075 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,075 words, ~3,712 tokens.
.claude/skills/langchain-eval-harness/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.A team swapped gpt-4o for claude-sonnet-4-6 to save money and a week later CS
noticed answer quality dropped on 15% of refund tickets — the regression was
invisible in code review and invisible in CI because no golden set existed.
Fix: a versioned golden set, a stacked eval pipeline (LangSmith + ragas + deepeval + custom trajectory), and a PR-blocking regression gate with paired Wilcoxon significance. The tooling exists; the patterns for wiring it into a statistically honest loop are scattered across five doc sites.
Build a 100-example JSONL golden set, wire LangSmith evaluate() with a
custom correctness evaluator, add a ragas quartet (faithfulness, answer
relevance, context precision/recall) for RAG, add deepeval LLM-as-judge
with N=3 judge quorum, score LangGraph trajectories on coverage/precision/
order, and gate PRs on a 2% aggregate drop or 5% per-example drop. Pin:
langchain-core 1.0.x, langgraph 1.0.x, langsmith>=0.2, ragas>=0.2,
deepeval>=2.0. Pain-catalog anchors: P01, P11, P12, P22, P33.
langchain-core >= 1.0, < 2.0, langgraph >= 1.0, < 2.0 for the system under evalpip install langsmith>=0.2 ragas>=0.2 deepeval>=2.0 scipyLANGSMITH_API_KEY (free tier is sufficient for dataset versioning)OPENAI_API_KEY and/or ANTHROPIC_API_KEYFormat: JSONL, one example per line, with a dataset_version tag. Minimum 20
examples to start; grow to 100 for PR gating, 200+ for absolute-metric claims.
# evals/golden_set/v2026.04.jsonl
{"id": "gs-0001", "input": "Refund policy for SKU ABC-42?", "expected": "30 days with receipt", "contexts": ["policy_v3.md"], "tags": ["refund"], "difficulty": "easy", "dataset_version": "2026.04"}
{"id": "gs-0002", "input": "Return policy for opened software?", "expected": "No, opened software is final sale", "contexts": ["policy_v3.md#returns"], "tags": ["refund"], "difficulty": "medium", "dataset_version": "2026.04"}Sample from real traffic (redacted), not imagination. Stratify by tag and difficulty (aim for 30% hard). Two annotators per example, disagreements reconciled — reconciliation rate under 90% means your task definition is ambiguous. Treat the file as immutable within a version; bump the version to refresh. See Golden Set Curation for sourcing strategy, annotation tool options, and the refresh cadence.
evaluate() with a custom evaluatorfrom langsmith import Client
from langsmith.evaluation import evaluate, EvaluationResult
from langchain_anthropic import ChatAnthropic
client = Client()
DATASET_VERSION = "2026.04"
# One-time: upload golden set as a versioned dataset
def upload_golden_set(jsonl_path, dataset_name):
examples = [json.loads(line) for line in open(jsonl_path)]
client.create_dataset(dataset_name)
client.create_examples(
inputs=[{"input": e["input"]} for e in examples],
outputs=[{"expected": e["expected"]} for e in examples],
metadata=[{"id": e["id"], "tags": e["tags"]} for e in examples],
dataset_name=dataset_name,
)
chain = ChatAnthropic(model="claude-sonnet-4-6", temperature=0, timeout=30)
def target(inputs):
return {"answer": chain.invoke(inputs["input"]).content}
def correctness(outputs, reference_outputs):
"""Deterministic exact-match floor — baseline, not ceiling."""
match = outputs["answer"].strip().lower() == reference_outputs["expected"].strip().lower()
return EvaluationResult(key="exact_match", score=float(match))
results = evaluate(
target,
data=f"golden-set-v{DATASET_VERSION}",
evaluators=[correctness],
experiment_prefix="refund-bot-v3",
max_concurrency=10, # Avoid 429s on judge LLM (P22)
)Free-form outputs need semantic scoring (ragas, deepeval, or LLM-as-judge — Step 4).
For a RAG chain returning {answer, contexts}, ragas scores four standard
dimensions. The default judge is gpt-4o-mini; override to pin model +
cost:
from ragas import evaluate as ragas_evaluate
from ragas.metrics import faithfulness, answer_relevancy, context_precision, context_recall
from langchain_openai import ChatOpenAI
from langchain_openai import OpenAIEmbeddings
from datasets import Dataset
judge = ChatOpenAI(model="gpt-4o-mini", temperature=0)
embed = OpenAIEmbeddings(model="text-embedding-3-small")
# Prepare rows — ragas wants HuggingFace Dataset shape
rows = []
for ex in golden_examples:
result = rag_chain.invoke({"question": ex["input"]})
rows.append({
"question": ex["input"],
"answer": result["answer"],
"contexts": [d.page_content for d in result["source_documents"]],
"ground_truth": ex["expected"],
})
ragas_results = ragas_evaluate(
Dataset.from_list(rows),
metrics=[faithfulness, answer_relevancy, context_precision, context_recall],
llm=judge,
embeddings=embed,
)
# ragas_results is a dict of per-metric means; call .to_pandas() for per-rowDo not use ragas on non-RAG chains — context_precision against an empty
context list returns 0 and looks like a regression. See
Framework Comparison for when each
tool fits.
deepeval is pytest-shaped — each example is an LLMTestCase asserting against
metrics. Run N=3 judge invocations per example and take the median to tame
LLM-as-judge variance (±5-15% across runs; single-run scores are not CI-ready):
import statistics
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, LLMTestCaseParams
def eval_with_quorum(test_case, metric, n=3):
scores = []
for _ in range(n):
metric.measure(test_case)
scores.append(metric.score)
return statistics.median(scores), statistics.stdev(scores) if n > 1 else 0.0
correctness = GEval(
name="Correctness",
criteria="Does the actual output match the expected output in meaning?",
evaluation_params=[LLMTestCaseParams.INPUT, LLMTestCaseParams.ACTUAL_OUTPUT, LLMTestCaseParams.EXPECTED_OUTPUT],
model="gpt-4o-mini",
)
for ex in golden_examples:
result = chain.invoke({"input": ex["input"]})
case = LLMTestCase(input=ex["input"], actual_output=result, expected_output=ex["expected"])
median, sd = eval_with_quorum(case, correctness, n=3)
if sd > 0.2: # judge disagreeing with itself — flag, don't gate
flag_for_review(ex["id"], median, sd)For agents, final-answer correctness misses the process. Score the tool-call sequence on three axes — coverage (did required tools run?), precision (were extra tools used?), and order (Kendall's tau on shared tools):
from langchain_core.messages import AIMessage
def extract_trajectory(final_state: dict) -> list[dict]:
return [
{"tool": tc["name"], "args": tc["args"]}
for msg in final_state["messages"] if isinstance(msg, AIMessage)
for tc in (msg.tool_calls or [])
]
def trajectory_score(expected: list[str], actual: list[str]) -> dict:
e_set, a_set = set(expected), set(actual)
coverage = len(e_set & a_set) / len(e_set) if e_set else 1.0
precision = len(e_set & a_set) / len(a_set) if a_set else 0.0
shared = [t for t in actual if t in e_set]
order = _kendall_tau(expected, shared) if len(shared) >= 2 else 1.0
return {"coverage": coverage, "precision": precision, "order": order}
# Composite: 0.5 * coverage + 0.3 * precision + 0.2 * orderSet temperature=0 for the agent during eval — temperature > 0 produces
different trajectories across runs (P11) and makes paired comparison
statistically invalid. See Agent Trajectory Eval
for args-level matching, efficiency/safety scoring, and the LLM-as-judge
fallback for non-deterministic trajectories.
A PR touching prompts, chain code, or model config runs the eval suite on
PR branch and main, then blocks merge on any of: aggregate mean drop > 2.0%,
any single-example drop > 5.0%, or paired Wilcoxon signed-rank p < 0.05
with negative mean delta.
from scipy.stats import wilcoxon
def paired_regression_check(baseline, candidate, alpha=0.05):
"""Wilcoxon — right test when metric distribution is non-normal (most LLM metrics)."""
n = len(baseline)
if n < 50:
return {"verdict": "too_small_n", "n": n}
diffs = [c - b for b, c in zip(baseline, candidate)]
_, p = wilcoxon(diffs, alternative="less")
return {"n": n, "mean_delta": sum(diffs) / n, "p_value": float(p),
"regression": p < alpha and sum(diffs) < 0}At n=100 and α=0.05 this detects a ~3-5% true regression at ~80% power. See CI Integration for the GitHub Actions workflow, PR-comment delta table, bootstrap CI, and spend/rate-limit safety rails.
evals/golden_set/v2026.04.jsonl with an immutable version tagLLMTestCase assertions in pytest, with median-of-3 judge quorum| Use case | LangSmith | ragas | deepeval | Custom |
|---|---|---|---|---|
| RAG metrics (faithfulness, context recall) | — | Primary | Fallback | — |
| Pytest-style assertion in CI | Secondary | — | Primary | — |
| Trace capture + dataset versioning | Primary | Complementary | Complementary | — |
| Agent trajectory (tool-call sequence) | Secondary (traces) | — | — | Primary |
| Exact match / JSON schema / structured output | — | — | — | Primary |
| Free-form paraphrase scoring | Via custom evaluator | — | Primary (G-Eval) | — |
Most real pipelines stack two or three. The anti-pattern is running all four on every example — you pay $10-30 per run for signal you are not using. See Framework Comparison for the full decision tree and dependency weight comparison.
| Error / Failure mode | Cause | Fix |
|---|---|---|
TimeoutError on eval runs > 20 min | Long agent trajectories on slow models; 100 examples × 30s each exceeds default GH Actions job timeout | Cap max_concurrency=10, use asyncio.gather with asyncio.Semaphore, split eval into sharded jobs |
| Judge disagreement (stdev > 0.2 on [0,1] scale across N=3 runs) | LLM-as-judge variance on ambiguous examples | Flag example for manual review; do not use that row's score for gating |
ValidationError: missing 'contexts' in ragas | Chain does not return retrieved docs | Modify chain to surface source_documents, or switch to non-RAG evaluator |
| Wilcoxon p-value is NaN | All paired diffs are 0 (identical outputs) | Expected when the PR did not change behavior — no regression, skip the stat test |
| LangSmith 429 rate limit during upload | > 50 examples/sec to create_examples | Batch with client.create_examples(..., batch_size=20) and sleep between batches |
| Spend overrun ($50+ per run) | Judge calls scaling with N_examples × N_metrics × N_judge_runs | Use gpt-4o-mini not gpt-4o for judge; cache per (dataset_version, chain_version) |
AttributeError: 'list' has no attribute 'lower' in custom evaluator | Claude AIMessage.content is list[dict] not str (P02 — see langchain-model-inference) | Use msg.text() or iterate content blocks |
| Trajectory comparison drifts week-over-week on unchanged agent | temperature > 0 non-determinism (P11) | Set temperature=0 for all eval runs; pin seed where supported |
Start with 20 production-sampled golden examples, wire up ragas_evaluate
with four metrics, record scores to evals/baselines/ as the reference,
and promote to LangSmith dataset versioning once two engineers annotate in
parallel. See Golden Set Curation.
Run the main-branch chain on the golden set, then swap the model and rerun. Diff per-example scores sorted by delta — the top-10 regressions usually cluster by tag (long contexts, one-shot lookups). Report paired Wilcoxon and per-tag breakdown before deciding to ship. See CI Integration.
Record expected tool-call sequences for 50 tasks, capture actual trajectories
via extract_trajectory, and score on coverage/precision/order. Composite
drops indicate a policy change — diff sequences to find the drift. See
Agent Trajectory Eval.
evaluate() referencedocs/pain-catalog.md (entries P01, P11, P12, P22, P33)© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/.curated/langchain-eval-harness of jeremylongshore/tons-of-skills-marketplace.
Open the folder on GitHubat commit cfae287
Langchain Eval Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Langchain Eval Harness this skilljeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~3.7k | Automated safety check: Pass | MIT | |
| LangSmith Trace DebuggingComposioHQ/awesome-claude-skills | 77k | 8 repos | ~2.7k | Automated safety check: Pass | None | |
| Failproof AI SDK IntegrationFailproofAI/failproofai | 5.3k | — | ~6k | Automated safety check: Pass | Custom licence | |
| Agentsop Observability Setupagentsope/SkillAlchemy | 436 | — | ~4.4k | Automated safety check: Pass | MIT | |
| Langchain Dependencieslangchain-ai/langchain-skills | 1.3k | — | ~3.6k | Automated safety check: Pass | MIT | |
| Langgraph Testing Evaluationsoba-labs/langchain-agent-skills | 107 | — | ~2.3k | Automated safety check: Pass | MIT |
ComposioHQ/awesome-claude-skills
Debugs LangChain and LangGraph agents by pulling recent execution traces with the langsmith-fetch CLI and reporting errors, tool calls, timings and token use.
FailproofAI/failproofai
Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.
agentsope/SkillAlchemy
Enhancement-overlay skill — the DECISION + WIRING layer for LM observability that the single-backend skills [[langsmith]], [[phoenix]], [[mlflow]] do NOT cover.
langchain-ai/langchain-skills
INVOKE THIS SKILL when setting up a new project or when asked about package versions, installation, or dependency management for LangChain, LangGraph, LangSmith, or Deep Agents.
soba-labs/langchain-agent-skills
A skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…
langchain-ai/docs
Trace, evaluate, and deploy AI agents and LLM applications with LangSmith.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.
jeremylongshore/tons-of-skills-marketplace
Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.
jeremylongshore/tons-of-skills-marketplace
Execute proactive auto-loading: automatically detects and loads agents.md files.
jeremylongshore/tons-of-skills-marketplace
Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.
jeremylongshore/tons-of-skills-marketplace
Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.
Categories
Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory…. Langchain Eval Harness is an agent skill from jeremylongshore/tons-of-skills-marketplace.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory analysis, and CI gating on quality regressions.
Langchain Eval Harness fits situations like: setting up quality measurement for a new chain; diagnosing regression after a model switch; building an evaluation gate for a pull request; with langchain eval.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-eval-harness -a claude-code`. Or copy the skill folder (skills/.curated/langchain-eval-harness in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langchain-eval-harness in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-eval-harness -a codex`. Or copy the skill folder (skills/.curated/langchain-eval-harness in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langchain-eval-harness in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-eval-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langchain-eval-harness, .gemini/skills/langchain-eval-harness, .github/skills/langchain-eval-harness and .opencode/skills/langchain-eval-harness in your project.
Going by SKILL.md and its folder, Langchain Eval Harness needs the command-line tools its instructions call (pip) and credentials named LANGSMITH_API_KEY, OPENAI_API_KEY and ANTHROPIC_API_KEY. Our summary lists: Python 3; A credential in LANGSMITH_API_KEY; A credential in OPENAI_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(python:*), Bash(pip:*), Bash(pytest:*). Compatibility (from SKILL.md): Designed for Claude Code.
SKILL.md names 5 domains. As links in the text: docs.smith.langchain.com, docs.ragas.io, docs.confident-ai.com, docs.scipy.org and arxiv.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Langchain Eval Harness is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Langchain Eval Harness: LangSmith Trace Debugging (ComposioHQ/awesome-claude-skills, 77k stars), Failproof AI SDK Integration (FailproofAI/failproofai, 5.3k stars), Agentsop Observability Setup (agentsope/SkillAlchemy, 436 stars) and Langchain Dependencies (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.
Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.