Langchain Dependencies
langchain-ai/langchain-skills
INVOKE THIS SKILL when setting up a new project or when asked about package versions, installation, or dependency management for LangChain, LangGraph, LangSmith, or Deep Agents.
A skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…
$ npx skills add soba-labs/langchain-agent-skills --skill langgraph-testing-evaluation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install soba-labs/langchain-agent-skills langgraph-testing-evaluation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/soba-labs/langchain-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/langgraph-testing-evaluation .claude/skills/langgraph-testing-evaluation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "langgraph-testing-evaluation" agent skill from https://github.com/soba-labs/langchain-agent-skills/tree/main/skills/langgraph-testing-evaluation into .claude/skills/langgraph-testing-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langgraph-testing-evaluation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/soba-labs/langchain-agent-skills/tree/main/skills/langgraph-testing-evaluationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add soba-labs/langchain-agent-skills --skill langgraph-testing-evaluation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install soba-labs/langchain-agent-skills langgraph-testing-evaluation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/soba-labs/langchain-agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/langgraph-testing-evaluation .agents/skills/langgraph-testing-evaluation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "langgraph-testing-evaluation" agent skill from https://github.com/soba-labs/langchain-agent-skills/tree/main/skills/langgraph-testing-evaluation into .agents/skills/langgraph-testing-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langgraph-testing-evaluation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add soba-labs/langchain-agent-skills --skill langgraph-testing-evaluation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install soba-labs/langchain-agent-skills langgraph-testing-evaluation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/soba-labs/langchain-agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/langgraph-testing-evaluation .cursor/skills/langgraph-testing-evaluation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "langgraph-testing-evaluation" agent skill from https://github.com/soba-labs/langchain-agent-skills/tree/main/skills/langgraph-testing-evaluation into .cursor/skills/langgraph-testing-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langgraph-testing-evaluation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/soba-labs/langchain-agent-skills.git --path skills/langgraph-testing-evaluation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add soba-labs/langchain-agent-skills --skill langgraph-testing-evaluation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install soba-labs/langchain-agent-skills langgraph-testing-evaluation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/soba-labs/langchain-agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/langgraph-testing-evaluation .gemini/skills/langgraph-testing-evaluation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "langgraph-testing-evaluation" agent skill from https://github.com/soba-labs/langchain-agent-skills/tree/main/skills/langgraph-testing-evaluation into .gemini/skills/langgraph-testing-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langgraph-testing-evaluation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install soba-labs/langchain-agent-skills langgraph-testing-evaluationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add soba-labs/langchain-agent-skills --skill langgraph-testing-evaluation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/soba-labs/langchain-agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/langgraph-testing-evaluation .github/skills/langgraph-testing-evaluation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "langgraph-testing-evaluation" agent skill from https://github.com/soba-labs/langchain-agent-skills/tree/main/skills/langgraph-testing-evaluation into .github/skills/langgraph-testing-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langgraph-testing-evaluation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add soba-labs/langchain-agent-skills --skill langgraph-testing-evaluation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install soba-labs/langchain-agent-skills langgraph-testing-evaluation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/soba-labs/langchain-agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/langgraph-testing-evaluation .opencode/skills/langgraph-testing-evaluation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "langgraph-testing-evaluation" agent skill from https://github.com/soba-labs/langchain-agent-skills/tree/main/skills/langgraph-testing-evaluation into .opencode/skills/langgraph-testing-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langgraph-testing-evaluation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
langgraph-testing-evaluationA skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…
Langgraph Testing Evaluation is an agent skill from soba-labs/langchain-agent-skills. Use this skill when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory evaluation (match or LLM-as-judge), running LangSmith dataset evaluations, and comparing two agent versions with A/B-style offline analysis. Use it for Python and JavaScript/TypeScript workflows, evaluator design, experiment setup, regression gates, and debugging flaky/incorrect evaluation results.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 23 other files, including scripts, reference files and assets (for example `assets/datasets/sample_dataset.json`, `assets/examples/README.md` and `assets/templates/test_template.py`).
It sits in AI & LLM Engineering, covering Building AI agents, LLM observability and Unit testing. It works with LangGraph, LangSmith, LangChain and JavaScript. The repository describes itself as: A collection of agent-optimized LangChain, LangGraph and LangSmith skills for AI coding assistants. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit a2d4a10. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 7 files in scripts/ (Python and JavaScript, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
uvnodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Langgraph Testing Evaluation loads about 2.3k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 128 tokens; SKILL.md has 730 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from soba-labs/langchain-agent-skills at commit a2d4a10, republished under its MIT licence (© soba-labs). 730 words, ~2,316 tokens.
.claude/skills/langgraph-testing-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.Practical workflows for validating agent quality with:
Use this file for high-level flow. Load references/* for detailed implementation.
Choose the smallest approach that answers your question:
| Goal | Primary method | Load first |
|---|---|---|
| Validate node logic quickly | Unit tests with mocks | references/unit-testing-patterns.md |
| Validate multi-step agent behavior | Trajectory evaluation | references/trajectory-evaluation.md |
| Track quality over datasets over time | LangSmith evaluation | references/langsmith-evaluation.md |
| Compare old vs new agent versions | A/B comparison | references/ab-testing.md |
Recommended order:
Run from repo root.
# Python (preferred)
uv run skills/langgraph-testing-evaluation/scripts/generate_test_cases.py my_agent:graph --output tests/ --framework pytest
# JavaScript/TypeScript
node skills/langgraph-testing-evaluation/scripts/generate_test_cases.js ./my-agent.ts:graph --output tests/ --framework vitest# Python: LLM-as-judge
uv run skills/langgraph-testing-evaluation/scripts/run_trajectory_eval.py my_agent:run_agent my_dataset --method llm-judge --model openai:o3-mini
# Python: trajectory match
uv run skills/langgraph-testing-evaluation/scripts/run_trajectory_eval.py my_agent:run_agent dataset.json --method match --trajectory-match-mode strict --reference-trajectory reference.json
# JavaScript/TypeScript
node skills/langgraph-testing-evaluation/scripts/run_trajectory_eval.js ./agent.ts:runAgent my_dataset --method llm-judge --model openai:o3-mini --max-concurrency 4# Python
uv run skills/langgraph-testing-evaluation/scripts/evaluate_with_langsmith.py my_agent:run_agent my_dataset --evaluators accuracy,latency --max-concurrency 4
# Python (do not upload experiment results)
uv run skills/langgraph-testing-evaluation/scripts/evaluate_with_langsmith.py my_agent:run_agent my_dataset --evaluators accuracy --no-upload
# JavaScript/TypeScript
node skills/langgraph-testing-evaluation/scripts/evaluate_with_langsmith.js ./agent.ts:runAgent my_dataset --evaluators accuracy,latency --max-concurrency 4# Python
uv run skills/langgraph-testing-evaluation/scripts/compare_agents.py my_agent:v1 my_agent:v2 dataset.json --output comparison_report.json
# JavaScript/TypeScript
node skills/langgraph-testing-evaluation/scripts/compare_agents.js ./v1.ts:run ./v2.ts:run dataset.json --output comparison_report.json
# JavaScript/TypeScript (force local dataset file only)
node skills/langgraph-testing-evaluation/scripts/compare_agents.js ./v1.ts:run ./v2.ts:run dataset.json --no-langsmith# Python
uv run skills/langgraph-testing-evaluation/scripts/mock_llm_responses.py create --type sequence --output mock_config.json
# JavaScript/TypeScript
node skills/langgraph-testing-evaluation/scripts/mock_llm_responses.js create --type sequence --output mock_config.jsoninputs and outputs objects (optional metadata).input, output) for legacy datasets.references/unit-testing-patterns.mdLoad when:
references/trajectory-evaluation.mdLoad when:
strict, unordered, subset, superset).references/langsmith-evaluation.mdLoad when:
references/ab-testing.mdLoad when:
assets/templates/test_template.pythread_idcompiled_graph.nodes[...]assets/datasets/sample_dataset.jsonexamples: [{ inputs, outputs, metadata }] format.assets/examples/README.mdscripts/generate_test_cases.py / .jsUse for fast test scaffolding.
Inputs:
my_module:graph or my_module.graph./file.ts:graphOutputs:
scripts/run_trajectory_eval.py / .jsUse for trajectory scoring with either:
--method match--method llm-judgeSupports:
.json)--reference-trajectorystrict, unordered, subset, supersetLocal-only mode:
--no-langsmith in both Python and JavaScript scripts (requires local JSON dataset file)scripts/evaluate_with_langsmith.py / .jsUse for dataset-based evaluation runs and experiment tracking.
Supports:
--evaluators accuracy,latency,...)--max-concurrency)Python-only:
--no-upload to run without uploading experiment resultsscripts/compare_agents.py / .jsUse for offline version comparisons:
--no-langsmith to disable remote loading)scripts/mock_llm_responses.py / .jsUse for deterministic test doubles:
If behavior is deterministic and local:
If behavior depends on tool sequence/routing:
If behavior depends on realistic distribution quality:
If approving a replacement model/prompt/graph:
unordered, subset, or superset) where appropriate.© soba-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 17 other files (scripts, references, assets) in skills/langgraph-testing-evaluation of soba-labs/langchain-agent-skills.
Open the folder on GitHubat commit a2d4a10
Langgraph Testing Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Langgraph Testing Evaluation this skillsoba-labs/langchain-agent-skills | 107 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Langchain Dependencieslangchain-ai/langchain-skills | 1.3k | 1 repos | ~3.6k | Automated safety check: Pass | MIT | |
| Failproof AI SDK IntegrationFailproofAI/failproofai | 5.3k | — | ~6k | Automated safety check: Pass | Custom licence | |
| LangSmith Trace DebuggingComposioHQ/awesome-claude-skills | 77k | 9 repos | ~2.7k | Automated safety check: Pass | None | |
| Add Example AgentGetBindu/Bindu | 10k | — | ~1.1k | Automated safety check: Notes | Custom licence | |
| Add Docs Pagelangchain-ai/docs | 424 | — | ~2k | Automated safety check: Pass | MIT |
langchain-ai/langchain-skills
INVOKE THIS SKILL when setting up a new project or when asked about package versions, installation, or dependency management for LangChain, LangGraph, LangSmith, or Deep Agents.
FailproofAI/failproofai
Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.
ComposioHQ/awesome-claude-skills
Debugs LangChain and LangGraph agents by pulling recent execution traces with the langsmith-fetch CLI and reporting errors, tool calls, timings and token use.
GetBindu/Bindu
Add a new self-contained example agent under examples/. An agent skill from GetBindu/Bindu.
langchain-ai/docs
Add, move, rename, or delete a page on the LangChain docs site.
langchain-ai/langchain-skills
Routes LangGraph agents with typed decision models that return probabilities, and finds LLM calls that only exist to produce a routing decision.
soba-labs/langchain-agent-skills
Use the writetodos tool effectively for task planning and decomposition in Deep Agents.
soba-labs/langchain-agent-skills
Initialize, validate, and troubleshoot Deep Agents projects in Python or JavaScript using the deepagents package.
soba-labs/langchain-agent-skills
Implement multi-agent coordination patterns (supervisor-subagent, router, orchestrator-worker, handoffs) for LangGraph applications.
soba-labs/langchain-agent-skills
Implement LangGraph error handling with current v1 patterns.
soba-labs/langchain-agent-skills
Initialize and configure LangGraph projects with proper structure, langgraph.json configuration, environment variables, and dependency management.
soba-labs/langchain-agent-skills
Design state schemas, implement reducers, configure persistence, and debug state issues for LangGraph applications.
Categories
A skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…. Langgraph Testing Evaluation is an agent skill from soba-labs/langchain-agent-skills. Use this skill when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory evaluation (match or LLM-as-judge), running LangSmith dataset evaluations, and comparing two agent versions with A/B-style offline analysis.
Langgraph Testing Evaluation fits situations like: you need to test; evaluate LangGraph/LangChain agents: writing unit; integration tests; generating test scaffolds.
Run `npx skills add soba-labs/langchain-agent-skills --skill langgraph-testing-evaluation -a claude-code`. Or copy the skill folder (skills/langgraph-testing-evaluation in soba-labs/langchain-agent-skills) into .claude/skills/langgraph-testing-evaluation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add soba-labs/langchain-agent-skills --skill langgraph-testing-evaluation -a codex`. Or copy the skill folder (skills/langgraph-testing-evaluation in soba-labs/langchain-agent-skills) into .agents/skills/langgraph-testing-evaluation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add soba-labs/langchain-agent-skills --skill langgraph-testing-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langgraph-testing-evaluation, .gemini/skills/langgraph-testing-evaluation, .github/skills/langgraph-testing-evaluation and .opencode/skills/langgraph-testing-evaluation in your project.
Going by SKILL.md and its folder, Langgraph Testing Evaluation needs Python and JavaScript for the scripts in its folder and the command-line tools its instructions call (uv and node). Our summary lists: Python 3; Node.js.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Langgraph Testing Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Langgraph Testing Evaluation: Langchain Dependencies (langchain-ai/langchain-skills, 1.3k stars), Failproof AI SDK Integration (FailproofAI/failproofai, 5.3k stars), LangSmith Trace Debugging (ComposioHQ/awesome-claude-skills, 77k stars) and Add Example Agent (GetBindu/Bindu, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
soba-labs (a GitHub organization) maintains it in soba-labs/langchain-agent-skills, which has 107 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on August 17, 2026.
Source: soba-labs/langchain-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.