Agent skill

Langchain Eval Harness

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory…

MITAuto-check passedAI & LLM Engineering

Install Langchain Eval Harness

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-eval-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-eval-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/langchain-eval-harness .claude/skills/langchain-eval-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langchain-eval-harness
GitHub stars
2.8k
Token cost
~3.7k tokens
SKILL.md length
1,075 words
Files
6 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory…

  • Works in 6 steps: Build a versioned golden set → Wire LangSmith evaluate() with a custom… → Add ragas metrics for RAG pipelines → …
  • Setting up quality measurement for a new chain
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 4 more sections
  • Calls pip; needs LANGSMITH_API_KEY and OPENAI_API_KEY

What it does

Langchain Eval Harness is an agent skill from jeremylongshore/tons-of-skills-marketplace. Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory analysis, and CI gating on quality regressions. Use when setting up quality measurement for a new chain, diagnosing regression after a model switch, or building an evaluation gate for a pull request. Trigger with "langchain eval", "langsmith evaluate", "ragas", "llm-as-judge", "agent trajectory eval", "eval regression gate".

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/agent-trajectory-eval.md`, `references/ci-integration.md` and `references/framework-comparison.md`). Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering Building AI agents, LLM evaluation and LLM observability. It works with LangChain, LangSmith and LangGraph. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Setting up quality measurement for a new chain
  • Diagnosing regression after a model switch
  • Building an evaluation gate for a pull request
  • With langchain eval

Example prompts

  • “langchain eval”
  • “langsmith evaluate”
  • “llm-as-judge”
  • “/langchain-eval-harness”

Requirements

  • Python 3
  • A credential in LANGSMITH_API_KEY
  • A credential in OPENAI_API_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(python:*), Bash(pip:*), Bash(pytest:*)

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Build a versioned golden set
  2. Wire LangSmith evaluate() with a custom evaluator
  3. Add ragas metrics for RAG pipelines
  4. Add deepeval LLM-as-judge for free-form outputs
  5. LangGraph agent trajectory eval
  6. Gate PRs on regression

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(python:*)
    • Bash(pip:*)
    • Bash(pytest:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.smith.langchain.com
    • docs.ragas.io
    • docs.confident-ai.com
    • docs.scipy.org
    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • LANGSMITH_API_KEY
    • OPENAI_API_KEY
    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Langchain Eval Harness loads about 3.7k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 135 tokens; SKILL.md has 1,075 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~135
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,075 words, ~3,712 tokens.

Download SKILL.mdSave it as .claude/skills/langchain-eval-harness/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
langchain-eval-harness
description
Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory analysis, and CI gating on quality regressions. Use when setting up quality measurement for a new chain, diagnosing regression after a model switch, or building an evaluation gate for a pull request. Trigger with "langchain eval", "langsmith evaluate", "ragas", "llm-as-judge", "agent trajectory eval", "eval regression gate".
allowed-tools
Read, Write, Edit, Bash(python:*), Bash(pip:*), Bash(pytest:*)
compatibility
Designed for Claude Code
version
2.7.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, langchain, langgraph, python, langchain-1.0, evaluation, langsmith, ragas, deepeval, research

LangChain Eval Harness (Python)

Overview

A team swapped gpt-4o for claude-sonnet-4-6 to save money and a week later CS noticed answer quality dropped on 15% of refund tickets — the regression was invisible in code review and invisible in CI because no golden set existed.

Fix: a versioned golden set, a stacked eval pipeline (LangSmith + ragas + deepeval + custom trajectory), and a PR-blocking regression gate with paired Wilcoxon significance. The tooling exists; the patterns for wiring it into a statistically honest loop are scattered across five doc sites.

Build a 100-example JSONL golden set, wire LangSmith evaluate() with a custom correctness evaluator, add a ragas quartet (faithfulness, answer relevance, context precision/recall) for RAG, add deepeval LLM-as-judge with N=3 judge quorum, score LangGraph trajectories on coverage/precision/ order, and gate PRs on a 2% aggregate drop or 5% per-example drop. Pin: langchain-core 1.0.x, langgraph 1.0.x, langsmith>=0.2, ragas>=0.2, deepeval>=2.0. Pain-catalog anchors: P01, P11, P12, P22, P33.

Prerequisites

  • Python 3.10+
  • langchain-core >= 1.0, < 2.0, langgraph >= 1.0, < 2.0 for the system under eval
  • pip install langsmith>=0.2 ragas>=0.2 deepeval>=2.0 scipy
  • LangSmith account + LANGSMITH_API_KEY (free tier is sufficient for dataset versioning)
  • Provider API keys for the judge LLM: OPENAI_API_KEY and/or ANTHROPIC_API_KEY

Instructions

Step 1 — Build a versioned golden set

Format: JSONL, one example per line, with a dataset_version tag. Minimum 20 examples to start; grow to 100 for PR gating, 200+ for absolute-metric claims.

python
# evals/golden_set/v2026.04.jsonl
{"id": "gs-0001", "input": "Refund policy for SKU ABC-42?", "expected": "30 days with receipt", "contexts": ["policy_v3.md"], "tags": ["refund"], "difficulty": "easy", "dataset_version": "2026.04"}
{"id": "gs-0002", "input": "Return policy for opened software?", "expected": "No, opened software is final sale", "contexts": ["policy_v3.md#returns"], "tags": ["refund"], "difficulty": "medium", "dataset_version": "2026.04"}

Sample from real traffic (redacted), not imagination. Stratify by tag and difficulty (aim for 30% hard). Two annotators per example, disagreements reconciled — reconciliation rate under 90% means your task definition is ambiguous. Treat the file as immutable within a version; bump the version to refresh. See Golden Set Curation for sourcing strategy, annotation tool options, and the refresh cadence.

Step 2 — Wire LangSmith evaluate() with a custom evaluator
python
from langsmith import Client
from langsmith.evaluation import evaluate, EvaluationResult
from langchain_anthropic import ChatAnthropic

client = Client()
DATASET_VERSION = "2026.04"

# One-time: upload golden set as a versioned dataset
def upload_golden_set(jsonl_path, dataset_name):
    examples = [json.loads(line) for line in open(jsonl_path)]
    client.create_dataset(dataset_name)
    client.create_examples(
        inputs=[{"input": e["input"]} for e in examples],
        outputs=[{"expected": e["expected"]} for e in examples],
        metadata=[{"id": e["id"], "tags": e["tags"]} for e in examples],
        dataset_name=dataset_name,
    )

chain = ChatAnthropic(model="claude-sonnet-4-6", temperature=0, timeout=30)

def target(inputs):
    return {"answer": chain.invoke(inputs["input"]).content}

def correctness(outputs, reference_outputs):
    """Deterministic exact-match floor — baseline, not ceiling."""
    match = outputs["answer"].strip().lower() == reference_outputs["expected"].strip().lower()
    return EvaluationResult(key="exact_match", score=float(match))

results = evaluate(
    target,
    data=f"golden-set-v{DATASET_VERSION}",
    evaluators=[correctness],
    experiment_prefix="refund-bot-v3",
    max_concurrency=10,   # Avoid 429s on judge LLM (P22)
)

Free-form outputs need semantic scoring (ragas, deepeval, or LLM-as-judge — Step 4).

Step 3 — Add ragas metrics for RAG pipelines

For a RAG chain returning {answer, contexts}, ragas scores four standard dimensions. The default judge is gpt-4o-mini; override to pin model + cost:

python
from ragas import evaluate as ragas_evaluate
from ragas.metrics import faithfulness, answer_relevancy, context_precision, context_recall
from langchain_openai import ChatOpenAI
from langchain_openai import OpenAIEmbeddings
from datasets import Dataset

judge = ChatOpenAI(model="gpt-4o-mini", temperature=0)
embed = OpenAIEmbeddings(model="text-embedding-3-small")

# Prepare rows — ragas wants HuggingFace Dataset shape
rows = []
for ex in golden_examples:
    result = rag_chain.invoke({"question": ex["input"]})
    rows.append({
        "question": ex["input"],
        "answer": result["answer"],
        "contexts": [d.page_content for d in result["source_documents"]],
        "ground_truth": ex["expected"],
    })

ragas_results = ragas_evaluate(
    Dataset.from_list(rows),
    metrics=[faithfulness, answer_relevancy, context_precision, context_recall],
    llm=judge,
    embeddings=embed,
)
# ragas_results is a dict of per-metric means; call .to_pandas() for per-row

Do not use ragas on non-RAG chains — context_precision against an empty context list returns 0 and looks like a regression. See Framework Comparison for when each tool fits.

Step 4 — Add deepeval LLM-as-judge for free-form outputs

deepeval is pytest-shaped — each example is an LLMTestCase asserting against metrics. Run N=3 judge invocations per example and take the median to tame LLM-as-judge variance (±5-15% across runs; single-run scores are not CI-ready):

python
import statistics
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, LLMTestCaseParams

def eval_with_quorum(test_case, metric, n=3):
    scores = []
    for _ in range(n):
        metric.measure(test_case)
        scores.append(metric.score)
    return statistics.median(scores), statistics.stdev(scores) if n > 1 else 0.0

correctness = GEval(
    name="Correctness",
    criteria="Does the actual output match the expected output in meaning?",
    evaluation_params=[LLMTestCaseParams.INPUT, LLMTestCaseParams.ACTUAL_OUTPUT, LLMTestCaseParams.EXPECTED_OUTPUT],
    model="gpt-4o-mini",
)

for ex in golden_examples:
    result = chain.invoke({"input": ex["input"]})
    case = LLMTestCase(input=ex["input"], actual_output=result, expected_output=ex["expected"])
    median, sd = eval_with_quorum(case, correctness, n=3)
    if sd > 0.2:  # judge disagreeing with itself — flag, don't gate
        flag_for_review(ex["id"], median, sd)
Step 5 — LangGraph agent trajectory eval

For agents, final-answer correctness misses the process. Score the tool-call sequence on three axes — coverage (did required tools run?), precision (were extra tools used?), and order (Kendall's tau on shared tools):

python
from langchain_core.messages import AIMessage

def extract_trajectory(final_state: dict) -> list[dict]:
    return [
        {"tool": tc["name"], "args": tc["args"]}
        for msg in final_state["messages"] if isinstance(msg, AIMessage)
        for tc in (msg.tool_calls or [])
    ]

def trajectory_score(expected: list[str], actual: list[str]) -> dict:
    e_set, a_set = set(expected), set(actual)
    coverage = len(e_set & a_set) / len(e_set) if e_set else 1.0
    precision = len(e_set & a_set) / len(a_set) if a_set else 0.0
    shared = [t for t in actual if t in e_set]
    order = _kendall_tau(expected, shared) if len(shared) >= 2 else 1.0
    return {"coverage": coverage, "precision": precision, "order": order}

# Composite: 0.5 * coverage + 0.3 * precision + 0.2 * order

Set temperature=0 for the agent during eval — temperature > 0 produces different trajectories across runs (P11) and makes paired comparison statistically invalid. See Agent Trajectory Eval for args-level matching, efficiency/safety scoring, and the LLM-as-judge fallback for non-deterministic trajectories.

Step 6 — Gate PRs on regression

A PR touching prompts, chain code, or model config runs the eval suite on PR branch and main, then blocks merge on any of: aggregate mean drop > 2.0%, any single-example drop > 5.0%, or paired Wilcoxon signed-rank p < 0.05 with negative mean delta.

python
from scipy.stats import wilcoxon

def paired_regression_check(baseline, candidate, alpha=0.05):
    """Wilcoxon — right test when metric distribution is non-normal (most LLM metrics)."""
    n = len(baseline)
    if n < 50:
        return {"verdict": "too_small_n", "n": n}
    diffs = [c - b for b, c in zip(baseline, candidate)]
    _, p = wilcoxon(diffs, alternative="less")
    return {"n": n, "mean_delta": sum(diffs) / n, "p_value": float(p),
            "regression": p < alpha and sum(diffs) < 0}

At n=100 and α=0.05 this detects a ~3-5% true regression at ~80% power. See CI Integration for the GitHub Actions workflow, PR-comment delta table, bootstrap CI, and spend/rate-limit safety rails.

Output

  • JSONL golden set at evals/golden_set/v2026.04.jsonl with an immutable version tag
  • LangSmith dataset uploaded and versioned; experiment runs linked to traces
  • Ragas scores (faithfulness, answer relevance, context precision/recall) on RAG chains
  • Deepeval LLMTestCase assertions in pytest, with median-of-3 judge quorum
  • LangGraph trajectory scores (coverage, precision, order) with composite summary
  • GitHub Actions workflow gating PRs on 2% aggregate / 5% per-example / Wilcoxon p < 0.05
  • PR-comment delta table posted on every eval run
Show full SKILL.md (451 more words)Show less

Framework selection at a glance

Use caseLangSmithragasdeepevalCustom
RAG metrics (faithfulness, context recall)—PrimaryFallback—
Pytest-style assertion in CISecondary—Primary—
Trace capture + dataset versioningPrimaryComplementaryComplementary—
Agent trajectory (tool-call sequence)Secondary (traces)——Primary
Exact match / JSON schema / structured output———Primary
Free-form paraphrase scoringVia custom evaluator—Primary (G-Eval)—

Most real pipelines stack two or three. The anti-pattern is running all four on every example — you pay $10-30 per run for signal you are not using. See Framework Comparison for the full decision tree and dependency weight comparison.

Error Handling

Error / Failure modeCauseFix
TimeoutError on eval runs > 20 minLong agent trajectories on slow models; 100 examples × 30s each exceeds default GH Actions job timeoutCap max_concurrency=10, use asyncio.gather with asyncio.Semaphore, split eval into sharded jobs
Judge disagreement (stdev > 0.2 on [0,1] scale across N=3 runs)LLM-as-judge variance on ambiguous examplesFlag example for manual review; do not use that row's score for gating
ValidationError: missing 'contexts' in ragasChain does not return retrieved docsModify chain to surface source_documents, or switch to non-RAG evaluator
Wilcoxon p-value is NaNAll paired diffs are 0 (identical outputs)Expected when the PR did not change behavior — no regression, skip the stat test
LangSmith 429 rate limit during upload> 50 examples/sec to create_examplesBatch with client.create_examples(..., batch_size=20) and sleep between batches
Spend overrun ($50+ per run)Judge calls scaling with N_examples × N_metrics × N_judge_runsUse gpt-4o-mini not gpt-4o for judge; cache per (dataset_version, chain_version)
AttributeError: 'list' has no attribute 'lower' in custom evaluatorClaude AIMessage.content is list[dict] not str (P02 — see langchain-model-inference)Use msg.text() or iterate content blocks
Trajectory comparison drifts week-over-week on unchanged agenttemperature > 0 non-determinism (P11)Set temperature=0 for all eval runs; pin seed where supported

Examples

Setting up eval for a new RAG chain

Start with 20 production-sampled golden examples, wire up ragas_evaluate with four metrics, record scores to evals/baselines/ as the reference, and promote to LangSmith dataset versioning once two engineers annotate in parallel. See Golden Set Curation.

Diagnosing regression after a model swap

Run the main-branch chain on the golden set, then swap the model and rerun. Diff per-example scores sorted by delta — the top-10 regressions usually cluster by tag (long contexts, one-shot lookups). Report paired Wilcoxon and per-tag breakdown before deciding to ship. See CI Integration.

Evaluating a LangGraph tool-calling agent

Record expected tool-call sequences for 50 tasks, capture actual trajectories via extract_trajectory, and score on coverage/precision/order. Composite drops indicate a policy change — diff sequences to find the drift. See Agent Trajectory Eval.

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/.curated/langchain-eval-harness of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/agent-trajectory-eval.md
  • references/ci-integration.md
  • references/framework-comparison.md
  • references/golden-set-curation.md
  • references/one-pager.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Langchain Eval Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Langchain Eval Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Langchain Eval Harness this skilljeremylongshore/tons-of-skills-marketplace2.8k—~3.7kAutomated safety check: PassMIT
LangSmith Trace DebuggingComposioHQ/awesome-claude-skills77k8 repos~2.7kAutomated safety check: PassNone
Failproof AI SDK IntegrationFailproofAI/failproofai5.3k—~6kAutomated safety check: PassCustom licence
Agentsop Observability Setupagentsope/SkillAlchemy436—~4.4kAutomated safety check: PassMIT
Langchain Dependencieslangchain-ai/langchain-skills1.3k—~3.6kAutomated safety check: PassMIT
Langgraph Testing Evaluationsoba-labs/langchain-agent-skills107—~2.3kAutomated safety check: PassMIT

Similar skills

  • LangSmith Trace Debugging

    ComposioHQ/awesome-claude-skills

    Debugs LangChain and LangGraph agents by pulling recent execution traces with the langsmith-fetch CLI and reporting errors, tool calls, timings and token use.

    77k GitHub starsUsed in 8 repos~2.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Agentsop Observability Setup

    agentsope/SkillAlchemy

    Enhancement-overlay skill — the DECISION + WIRING layer for LM observability that the single-backend skills [[langsmith]], [[phoenix]], [[mlflow]] do NOT cover.

    436 GitHub stars~4.4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Langchain Dependencies

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when setting up a new project or when asked about package versions, installation, or dependency management for LangChain, LangGraph, LangSmith, or Deep Agents.

    1.3k GitHub stars~3.6k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Langgraph Testing Evaluation

    soba-labs/langchain-agent-skills

    A skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…

    107 GitHub stars~2.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Langsmith

    langchain-ai/docs

    Official

    Trace, evaluate, and deploy AI agents and LLM applications with LangSmith.

    426 GitHub stars~935 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Langchain Eval Harness

What does Langchain Eval Harness do?

Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory…. Langchain Eval Harness is an agent skill from jeremylongshore/tons-of-skills-marketplace.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory analysis, and CI gating on quality regressions.

When should I use Langchain Eval Harness?

Langchain Eval Harness fits situations like: setting up quality measurement for a new chain; diagnosing regression after a model switch; building an evaluation gate for a pull request; with langchain eval.

How do I install Langchain Eval Harness in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-eval-harness -a claude-code`. Or copy the skill folder (skills/.curated/langchain-eval-harness in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langchain-eval-harness in your project. Claude Code loads it when a task matches its description.

How do I install Langchain Eval Harness in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-eval-harness -a codex`. Or copy the skill folder (skills/.curated/langchain-eval-harness in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langchain-eval-harness in your project. Codex loads it when a task matches its description.

Can I use Langchain Eval Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-eval-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langchain-eval-harness, .gemini/skills/langchain-eval-harness, .github/skills/langchain-eval-harness and .opencode/skills/langchain-eval-harness in your project.

What does Langchain Eval Harness need to run?

Going by SKILL.md and its folder, Langchain Eval Harness needs the command-line tools its instructions call (pip) and credentials named LANGSMITH_API_KEY, OPENAI_API_KEY and ANTHROPIC_API_KEY. Our summary lists: Python 3; A credential in LANGSMITH_API_KEY; A credential in OPENAI_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(python:*), Bash(pip:*), Bash(pytest:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Langchain Eval Harness access the network?

SKILL.md names 5 domains. As links in the text: docs.smith.langchain.com, docs.ragas.io, docs.confident-ai.com, docs.scipy.org and arxiv.org. This is read from the text; nothing was executed.

Is Langchain Eval Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Langchain Eval Harness use?

Langchain Eval Harness is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Langchain Eval Harness use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.8k tokens, read only when the agent opens those files.

What are the alternatives to Langchain Eval Harness?

Skills that share tags, products or a category with Langchain Eval Harness: LangSmith Trace Debugging (ComposioHQ/awesome-claude-skills, 77k stars), Failproof AI SDK Integration (FailproofAI/failproofai, 5.3k stars), Agentsop Observability Setup (agentsope/SkillAlchemy, 436 stars) and Langchain Dependencies (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Langchain Eval Harness?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.