Agent skill

Langgraph Testing Evaluation

by soba-labs in soba-labs/langchain-agent-skills

A skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…

MITAuto-check passedAI & LLM Engineering

Install Langgraph Testing Evaluation

skills CLI
$ npx skills add soba-labs/langchain-agent-skills --skill langgraph-testing-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install soba-labs/langchain-agent-skills langgraph-testing-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/soba-labs/langchain-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/langgraph-testing-evaluation .claude/skills/langgraph-testing-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langgraph-testing-evaluation
GitHub stars
107
Token cost
~2.3k tokens
SKILL.md length
730 words
Files
18 (incl. scripts, references, assets)
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…

  • Works in 4 steps: Unit tests → Integration/trajectory checks → Dataset evaluation in LangSmith → …
  • You need to test
  • SKILL.md covers Start Here, Quick Commands, Core Workflow and Current References (Load On…, plus 4 more sections
  • Runs Python and JavaScript scripts from its folder; calls uv and node

What it does

Langgraph Testing Evaluation is an agent skill from soba-labs/langchain-agent-skills. Use this skill when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory evaluation (match or LLM-as-judge), running LangSmith dataset evaluations, and comparing two agent versions with A/B-style offline analysis. Use it for Python and JavaScript/TypeScript workflows, evaluator design, experiment setup, regression gates, and debugging flaky/incorrect evaluation results.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 23 other files, including scripts, reference files and assets (for example `assets/datasets/sample_dataset.json`, `assets/examples/README.md` and `assets/templates/test_template.py`).

It sits in AI & LLM Engineering, covering Building AI agents, LLM observability and Unit testing. It works with LangGraph, LangSmith, LangChain and JavaScript. The repository describes itself as: A collection of agent-optimized LangChain, LangGraph and LangSmith skills for AI coding assistants. The licence is MIT.

When your agent uses it

  • You need to test
  • Evaluate LangGraph/LangChain agents: writing unit
  • Integration tests
  • Generating test scaffolds

Example prompts

  • “/langgraph-testing-evaluation”

Requirements

  • Python 3
  • Node.js

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Unit tests
  2. Integration/trajectory checks
  3. Dataset evaluation in LangSmith
  4. A/B comparison before deployment

What it can do on your machine

Read from SKILL.md and the folder at commit a2d4a10. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (Python and JavaScript, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Langgraph Testing Evaluation loads about 2.3k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 128 tokens; SKILL.md has 730 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~128
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~14k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from soba-labs/langchain-agent-skills at commit a2d4a10, republished under its MIT licence (© soba-labs). 730 words, ~2,316 tokens.

Download SKILL.mdSave it as .claude/skills/langgraph-testing-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
langgraph-testing-evaluation
description
Use this skill when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory evaluation (match or LLM-as-judge), running LangSmith dataset evaluations, and comparing two agent versions with A/B-style offline analysis. Use it for Python and JavaScript/TypeScript workflows, evaluator design, experiment setup, regression gates, and debugging flaky/incorrect evaluation results.

LangGraph Testing & Evaluation

Practical workflows for validating agent quality with:

  • Unit/integration tests
  • Trajectory evaluation
  • LangSmith dataset evaluations
  • A/B-style comparisons between versions

Use this file for high-level flow. Load references/* for detailed implementation.

Start Here

Choose the smallest approach that answers your question:

GoalPrimary methodLoad first
Validate node logic quicklyUnit tests with mocksreferences/unit-testing-patterns.md
Validate multi-step agent behaviorTrajectory evaluationreferences/trajectory-evaluation.md
Track quality over datasets over timeLangSmith evaluationreferences/langsmith-evaluation.md
Compare old vs new agent versionsA/B comparisonreferences/ab-testing.md

Recommended order:

  1. Unit tests
  2. Integration/trajectory checks
  3. Dataset evaluation in LangSmith
  4. A/B comparison before deployment

Quick Commands

Run from repo root.

Generate test scaffolding
bash
# Python (preferred)
uv run skills/langgraph-testing-evaluation/scripts/generate_test_cases.py my_agent:graph --output tests/ --framework pytest

# JavaScript/TypeScript
node skills/langgraph-testing-evaluation/scripts/generate_test_cases.js ./my-agent.ts:graph --output tests/ --framework vitest
Run trajectory evaluation
bash
# Python: LLM-as-judge
uv run skills/langgraph-testing-evaluation/scripts/run_trajectory_eval.py my_agent:run_agent my_dataset --method llm-judge --model openai:o3-mini

# Python: trajectory match
uv run skills/langgraph-testing-evaluation/scripts/run_trajectory_eval.py my_agent:run_agent dataset.json --method match --trajectory-match-mode strict --reference-trajectory reference.json

# JavaScript/TypeScript
node skills/langgraph-testing-evaluation/scripts/run_trajectory_eval.js ./agent.ts:runAgent my_dataset --method llm-judge --model openai:o3-mini --max-concurrency 4
Run LangSmith dataset evaluation
bash
# Python
uv run skills/langgraph-testing-evaluation/scripts/evaluate_with_langsmith.py my_agent:run_agent my_dataset --evaluators accuracy,latency --max-concurrency 4

# Python (do not upload experiment results)
uv run skills/langgraph-testing-evaluation/scripts/evaluate_with_langsmith.py my_agent:run_agent my_dataset --evaluators accuracy --no-upload

# JavaScript/TypeScript
node skills/langgraph-testing-evaluation/scripts/evaluate_with_langsmith.js ./agent.ts:runAgent my_dataset --evaluators accuracy,latency --max-concurrency 4
Compare two agent versions
bash
# Python
uv run skills/langgraph-testing-evaluation/scripts/compare_agents.py my_agent:v1 my_agent:v2 dataset.json --output comparison_report.json

# JavaScript/TypeScript
node skills/langgraph-testing-evaluation/scripts/compare_agents.js ./v1.ts:run ./v2.ts:run dataset.json --output comparison_report.json

# JavaScript/TypeScript (force local dataset file only)
node skills/langgraph-testing-evaluation/scripts/compare_agents.js ./v1.ts:run ./v2.ts:run dataset.json --no-langsmith
Create mock response configs
bash
# Python
uv run skills/langgraph-testing-evaluation/scripts/mock_llm_responses.py create --type sequence --output mock_config.json

# JavaScript/TypeScript
node skills/langgraph-testing-evaluation/scripts/mock_llm_responses.js create --type sequence --output mock_config.json

Core Workflow

  1. Define test scope.
  • Unit: deterministic logic in one node/function.
  • Integration: node interactions and routing.
  • End-to-end: complete response quality on realistic inputs.
  1. Start from deterministic checks.
  • Mock LLM/tool IO for speed and repeatability.
  • Keep real-model tests as a smaller, explicit suite.
  1. Build/curate dataset examples.
  • Use stable inputs and expected outputs.
  • Keep schema simple: inputs and outputs objects (optional metadata).
  • Compatibility note: scripts also accept singular keys (input, output) for legacy datasets.
  1. Run evaluation with explicit gates.
  • Use evaluator keys that map to deployment decisions.
  • Set thresholds in CI for regression prevention.
  1. Compare versions before rollout.
  • Run same dataset on both versions.
  • Check both quality and latency.
  1. Diagnose failures from traces/experiments.
  • Inspect low-scoring examples.
  • Split failures by pattern (routing, tool usage, hallucination, latency spikes).

Current References (Load On Demand)

references/unit-testing-patterns.md

Load when:

  • You need node-level and routing test patterns.
  • You need pytest/vitest/Jest integration patterns.
  • You need robust mocking and flaky-test reduction.
references/trajectory-evaluation.md

Load when:

  • You need trajectory match evaluation (strict, unordered, subset, superset).
  • You need LLM-as-judge trajectory scoring.
  • You need LangSmith experiment comparison for trajectory results.
references/langsmith-evaluation.md

Load when:

  • You need dataset creation/management in LangSmith.
  • You need evaluator signatures and experiment runs in Python/TS.
  • You need CI-friendly workflows with quality thresholds.
references/ab-testing.md

Load when:

  • You need offline A/B comparison methodology.
  • You need significance testing and interpretation.
  • You need production traffic split strategy and guardrails.

Assets

assets/templates/test_template.py
  • Runnable Python pytest template aligned with current LangGraph testing patterns.
  • Includes:
    • Compiled-graph invocation with thread_id
    • Single-node testing via compiled_graph.nodes[...]
    • Integration-test placeholder
assets/datasets/sample_dataset.json
  • Deterministic seed dataset for LangSmith ingestion.
  • Uses examples: [{ inputs, outputs, metadata }] format.
assets/examples/README.md
  • Documentation-only index for current asset usage.
  • Notes where runnable assets live today.

Script Interface Summary

scripts/generate_test_cases.py / .js

Use for fast test scaffolding.

Inputs:

  • Graph module path
    • Python: my_module:graph or my_module.graph
    • JS/TS: ./file.ts:graph

Outputs:

  • Framework-specific starter tests in target directory.
Show full SKILL.md (301 more words)Show less
scripts/run_trajectory_eval.py / .js

Use for trajectory scoring with either:

  • --method match
  • --method llm-judge

Supports:

  • Local dataset files (.json)
  • LangSmith dataset names
  • Optional reference trajectory file with --reference-trajectory
  • Match modes: strict, unordered, subset, superset

Local-only mode:

  • --no-langsmith in both Python and JavaScript scripts (requires local JSON dataset file)
scripts/evaluate_with_langsmith.py / .js

Use for dataset-based evaluation runs and experiment tracking.

Supports:

  • Existing dataset by name
  • Dataset creation from JSON examples file
  • Multiple evaluators (--evaluators accuracy,latency,...)
  • Concurrency control (--max-concurrency)

Python-only:

  • --no-upload to run without uploading experiment results
scripts/compare_agents.py / .js

Use for offline version comparisons:

  • Shared dataset input
  • Success/latency summaries
  • JSON report output for CI artifacts
  • Local JSON datasets or LangSmith datasets (JS supports --no-langsmith to disable remote loading)
scripts/mock_llm_responses.py / .js

Use for deterministic test doubles:

  • single
  • sequence
  • conditional

Decision Rules

If behavior is deterministic and local:

  • Use unit tests first.

If behavior depends on tool sequence/routing:

  • Add trajectory evaluation.

If behavior depends on realistic distribution quality:

  • Run LangSmith dataset evaluation.

If approving a replacement model/prompt/graph:

  • Run A/B comparison and check both quality and latency.

Common Failure Patterns

Flaky tests
  • Cause: real-model nondeterminism in unit scope.
  • Fix: mock LLM/tool calls for unit tests; reserve real-model tests for separate integration marks.
High trajectory variance
  • Cause: overly strict matching for workflows with equivalent paths.
  • Fix: switch match mode (unordered, subset, or superset) where appropriate.
Regressions hidden by averages
  • Cause: only aggregate score monitored.
  • Fix: inspect per-example failures and segment by category metadata.
Latency regressions with same quality
  • Cause: no explicit latency gate.
  • Fix: include latency evaluator and CI threshold.

Minimal Best Practices

  1. Keep fast deterministic tests as the largest share.
  2. Version datasets and keep them stable.
  3. Track both correctness and latency.
  4. Add explicit go/no-go thresholds in CI.
  5. Compare candidate vs baseline before production rollout.
  6. Investigate failures with trace-level evidence, not only aggregate scores.

© soba-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (scripts, references, assets) in skills/langgraph-testing-evaluation of soba-labs/langchain-agent-skills.

  • SKILL.md
  • assets/datasets/sample_dataset.json
  • assets/examples/README.md
  • assets/templates/test_template.py
  • references/ab-testing.md
  • references/langsmith-evaluation.md
  • references/trajectory-evaluation.md
  • references/unit-testing-patterns.md
  • scripts/compare_agents.js
  • scripts/compare_agents.py
  • scripts/evaluate_with_langsmith.js
  • scripts/evaluate_with_langsmith.py
  • scripts/generate_test_cases.js
  • scripts/generate_test_cases.py
  • scripts/mock_llm_responses.js
  • … and 3 more

Open the folder on GitHubat commit a2d4a10

Compare with similar skills

Langgraph Testing Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Langgraph Testing Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Langgraph Testing Evaluation this skillsoba-labs/langchain-agent-skills107—~2.3kAutomated safety check: PassMIT
Langchain Dependencieslangchain-ai/langchain-skills1.3k1 repos~3.6kAutomated safety check: PassMIT
Failproof AI SDK IntegrationFailproofAI/failproofai5.3k—~6kAutomated safety check: PassCustom licence
LangSmith Trace DebuggingComposioHQ/awesome-claude-skills77k9 repos~2.7kAutomated safety check: PassNone
Add Example AgentGetBindu/Bindu10k—~1.1kAutomated safety check: NotesCustom licence
Add Docs Pagelangchain-ai/docs424—~2kAutomated safety check: PassMIT

Similar skills

  • Langchain Dependencies

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when setting up a new project or when asked about package versions, installation, or dependency management for LangChain, LangGraph, LangSmith, or Deep Agents.

    1.3k GitHub starsUsed in 1 repo~3.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LangSmith Trace Debugging

    ComposioHQ/awesome-claude-skills

    Debugs LangChain and LangGraph agents by pulling recent execution traces with the langsmith-fetch CLI and reporting errors, tool calls, timings and token use.

    77k GitHub starsUsed in 9 repos~2.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Example Agent

    GetBindu/Bindu

    Add a new self-contained example agent under examples/. An agent skill from GetBindu/Bindu.

    10k GitHub stars~1.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Add Docs Page

    langchain-ai/docs

    Official

    Add, move, rename, or delete a page on the LangChain docs site.

    424 GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • LangGraph Decision Models

    langchain-ai/langchain-skills

    Official

    Routes LangGraph agents with typed decision models that return probabilities, and finds LLM calls that only exist to produce a routing decision.

    1.3k GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from soba-labs/langchain-agent-skills

All 9 skills in this repo
  • Deepagents Planning Todos

    soba-labs/langchain-agent-skills

    Use the writetodos tool effectively for task planning and decomposition in Deep Agents.

    107 GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Deepagents Setup Configuration

    soba-labs/langchain-agent-skills

    Initialize, validate, and troubleshoot Deep Agents projects in Python or JavaScript using the deepagents package.

    107 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Langgraph Agent Patterns

    soba-labs/langchain-agent-skills

    Implement multi-agent coordination patterns (supervisor-subagent, router, orchestrator-worker, handoffs) for LangGraph applications.

    107 GitHub stars~3.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Langgraph Error Handling

    soba-labs/langchain-agent-skills

    Implement LangGraph error handling with current v1 patterns.

    107 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Langgraph Project Setup

    soba-labs/langchain-agent-skills

    Initialize and configure LangGraph projects with proper structure, langgraph.json configuration, environment variables, and dependency management.

    107 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check: notes
  • Langgraph State Management

    soba-labs/langchain-agent-skills

    Design state schemas, implement reducers, configure persistence, and debug state issues for LangGraph applications.

    107 GitHub stars~3.4k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Langgraph Testing Evaluation

What does Langgraph Testing Evaluation do?

A skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…. Langgraph Testing Evaluation is an agent skill from soba-labs/langchain-agent-skills. Use this skill when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory evaluation (match or LLM-as-judge), running LangSmith dataset evaluations, and comparing two agent versions with A/B-style offline analysis.

When should I use Langgraph Testing Evaluation?

Langgraph Testing Evaluation fits situations like: you need to test; evaluate LangGraph/LangChain agents: writing unit; integration tests; generating test scaffolds.

How do I install Langgraph Testing Evaluation in Claude Code?

Run `npx skills add soba-labs/langchain-agent-skills --skill langgraph-testing-evaluation -a claude-code`. Or copy the skill folder (skills/langgraph-testing-evaluation in soba-labs/langchain-agent-skills) into .claude/skills/langgraph-testing-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Langgraph Testing Evaluation in Codex?

Run `npx skills add soba-labs/langchain-agent-skills --skill langgraph-testing-evaluation -a codex`. Or copy the skill folder (skills/langgraph-testing-evaluation in soba-labs/langchain-agent-skills) into .agents/skills/langgraph-testing-evaluation in your project. Codex loads it when a task matches its description.

Can I use Langgraph Testing Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add soba-labs/langchain-agent-skills --skill langgraph-testing-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langgraph-testing-evaluation, .gemini/skills/langgraph-testing-evaluation, .github/skills/langgraph-testing-evaluation and .opencode/skills/langgraph-testing-evaluation in your project.

What does Langgraph Testing Evaluation need to run?

Going by SKILL.md and its folder, Langgraph Testing Evaluation needs Python and JavaScript for the scripts in its folder and the command-line tools its instructions call (uv and node). Our summary lists: Python 3; Node.js.

Does Langgraph Testing Evaluation access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Langgraph Testing Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Langgraph Testing Evaluation use?

Langgraph Testing Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Langgraph Testing Evaluation use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.

What are the alternatives to Langgraph Testing Evaluation?

Skills that share tags, products or a category with Langgraph Testing Evaluation: Langchain Dependencies (langchain-ai/langchain-skills, 1.3k stars), Failproof AI SDK Integration (FailproofAI/failproofai, 5.3k stars), LangSmith Trace Debugging (ComposioHQ/awesome-claude-skills, 77k stars) and Add Example Agent (GetBindu/Bindu, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Langgraph Testing Evaluation?

soba-labs (a GitHub organization) maintains it in soba-labs/langchain-agent-skills, which has 107 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on August 17, 2026.

Source: soba-labs/langchain-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.