Build a fast, deterministic local test loop for LangChain 1.0 / LangGraph 1.0 — FakeListChatModel fixtures, pytest config, VCR cassettes with key redaction, warning-filter policy.
Install the "langchain-local-dev-loop" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-local-dev-loop into .claude/skills/langchain-local-dev-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-local-dev-loop", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-local-dev-loop -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "langchain-local-dev-loop" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-local-dev-loop into .agents/skills/langchain-local-dev-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-local-dev-loop", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-local-dev-loop -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "langchain-local-dev-loop" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-local-dev-loop into .cursor/skills/langchain-local-dev-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-local-dev-loop", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-local-dev-loop -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "langchain-local-dev-loop" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-local-dev-loop into .gemini/skills/langchain-local-dev-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-local-dev-loop", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-local-dev-loop -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "langchain-local-dev-loop" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-local-dev-loop into .github/skills/langchain-local-dev-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-local-dev-loop", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-local-dev-loop -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "langchain-local-dev-loop" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-local-dev-loop into .opencode/skills/langchain-local-dev-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-local-dev-loop", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
langchain-local-dev-loop
GitHub stars
2.8k
Token cost
~4.1k tokens
SKILL.md length
1,104 words
Files
6 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT
At a glance
Build a fast, deterministic local test loop for LangChain 1.0 / LangGraph 1.0 — FakeListChatModel fixtures, pytest config, VCR cassettes with key redaction, warning-filter policy.
Works in 7 steps: Deterministic unit tests with… → Subclass FakeListChatModel to emit… → pytest fixtures that wire the fake into… → …
Adding tests to a new chain
SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 4 more sections
Calls pytest, git and pip; needs ANTHROPIC_API_KEY
What it does
Langchain Local Dev Loop is an agent skill from jeremylongshore/tons-of-skills-marketplace. Build a fast, deterministic local test loop for LangChain 1.0 / LangGraph 1.0 — FakeListChatModel fixtures, pytest config, VCR cassettes with key redaction, warning-filter policy. Use when adding tests to a new chain, fixing a flaky test, or making integration tests reproducible. Trigger with "langchain pytest", "FakeListChatModel", "VCR langchain", "langchain test fixtures", "langchain integration test".
Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/fake-model-fixtures.md`, `references/langgraph-test-patterns.md` and `references/one-pager.md`). Compatibility notes: Designed for Claude Code
It sits in Testing & QA, covering Building AI agents, Integration testing and Unit testing. It works with LangChain, pytest and LangGraph. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.
When your agent uses it
Adding tests to a new chain
Fixing a flaky test
Making integration tests reproducible
With langchain pytest
Example prompts
“langchain pytest”
“FakeListChatModel”
“VCR langchain”
“/langchain-local-dev-loop”
Requirements
Python 3
A credential in ANTHROPIC_API_KEY
Compatibility (from SKILL.md): Designed for Claude Code
Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves these tools, so the agent can use them without asking each time:
Read
Write
Edit
Bash(pytest:*)
Bash(python:*)
Bash(pip:*)
From allowed-tools in the SKILL.md frontmatter.
Runs code
Shell commands in SKILL.md call:
pytest
git
pip
From the folder's file list and the shell code blocks in SKILL.md.
Network
Links to these hosts (documentation or services it may open):
python.langchain.com
vcrpy.readthedocs.io
pytest-vcr.readthedocs.io
docs.pytest.org
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names these keys or tokens, usually read from environment variables:
ANTHROPIC_API_KEY
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Compatibility
Designed for Claude Code
From compatibility in the SKILL.md frontmatter.
Context cost
Langchain Local Dev Loop loads about 4.1k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 108 tokens; SKILL.md has 1,104 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~108
When it runs· the whole SKILL.md, loaded when a task matches
~4.1k
With references· SKILL.md plus every file in references/, read only if the agent opens them
~11k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
Download SKILL.mdSave it as .claude/skills/langchain-local-dev-loop/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
langchain-local-dev-loop
description
Build a fast, deterministic local test loop for LangChain 1.0 / LangGraph 1.0
— FakeListChatModel fixtures, pytest config, VCR cassettes with key redaction,
warning-filter policy. Use when adding tests to a new chain, fixing a flaky
test, or making integration tests reproducible.
Trigger with "langchain pytest", "FakeListChatModel", "VCR langchain",
"langchain test fixtures", "langchain integration test".
It passes locally against Claude at temperature=0. It fails in CI on the third
run with a one-token delta in the output. That is P05: Anthropic's temperature=0
is not greedy — it still samples. Tests against live Claude are not deterministic,
period.
So the engineer swaps in FakeListChatModel(responses=["expected summary"]) and
the assertion passes. Then the downstream callback that logs cost blows up in CI
with KeyError: 'token_usage' — because FakeListChatModel does not emit
response_metadata["token_usage"] (P43). Production code reads that key, so
either the fake has to synthesize it or the test has to skip the callback.
Meanwhile, the first integration test under VCR records a cassette that ships
Authorization: Bearer sk-ant-api03-... in the repo (P44). PR review catches it;
the reviewer revokes the key; the dev loop is hosed for an afternoon.
And none of this matters if pytest cannot even collect the suite because
import langchain_community emits a DeprecationWarning that -W error promotes
to failure (P45).
This skill installs the four layers that make the whole loop fast and safe:
FakeListChatModel / FakeListLLM with a metadata-emitting subclass (fixes P43);
VCR with filter_headers plus a pre-commit hook (fixes P44); pytest
filterwarnings policy in pyproject.toml (fixes P45); and an env-var-gated
integration marker so the default pytest run never touches live APIs.
Speed targets: unit tests with FakeListChatModel run in < 100ms per
test; VCR-replayed integration tests run in 500ms – 2s per test; live
integration tests (the RUN_INTEGRATION=1 gate) run only in nightly or
manual workflows.
For integration tests: at least one provider key (ANTHROPIC_API_KEY, etc.)
Project uses pyproject.toml (PEP 621) for pytest config
Instructions
Step 1 — Deterministic unit tests with FakeListChatModel
Use FakeListChatModel from langchain_core.language_models.fake for chat
chains and FakeListLLM for legacy completion LLMs. Responses cycle through
the list.
python
from langchain_core.language_models.fake import FakeListChatModel
from langchain_core.prompts import ChatPromptTemplate
def test_classifier_picks_positive():
fake = FakeListChatModel(responses=["positive"])
prompt = ChatPromptTemplate.from_messages([("user", "Classify: {text}")])
chain = prompt | fake
out = chain.invoke({"text": "I love it"})
assert out.content == "positive"
This is deterministic, runs in single-digit milliseconds, and has zero provider
dependency. Use it for every chain assertion that does not specifically require
real model behavior.
Step 2 — Subclass FakeListChatModel to emit response_metadata (P43 fix)
The stock fake emits no response_metadata["token_usage"]. If your chain has a
callback that records cost, the callback crashes under the fake. Subclass and
synthesize the metadata instead of mocking around the callback:
python
from langchain_core.language_models.fake import FakeListChatModel
from langchain_core.outputs import ChatGeneration, ChatResult
from langchain_core.messages import AIMessage
class FakeChatWithUsage(FakeListChatModel):
"""FakeListChatModel that emits response_metadata['token_usage'] so
downstream callbacks reading token usage do not crash under test."""
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
response = self.responses[self.i % len(self.responses)]
self.i += 1
message = AIMessage(
content=response,
response_metadata={
"token_usage": {
"input_tokens": 10,
"output_tokens": len(response.split()),
"total_tokens": 10 + len(response.split()),
},
"model_name": "fake-chat",
},
usage_metadata={
"input_tokens": 10,
"output_tokens": len(response.split()),
"total_tokens": 10 + len(response.split()),
},
)
return ChatResult(generations=[ChatGeneration(message=message)])
Use FakeChatWithUsage whenever a chain's observability / cost path is in the
assertion surface. See Fake Model Fixtures
for agent, retriever, and embedder fakes.
Step 3 — pytest fixtures that wire the fake into chains
Put fixtures in tests/conftest.py so they are shared across the suite:
python
# tests/conftest.py
import pytest
from langchain_core.prompts import ChatPromptTemplate
from tests.fakes import FakeChatWithUsage
@pytest.fixture
def fake_chat():
"""Reusable fake chat model. Override responses per-test via
monkeypatch.setattr(fake_chat, 'responses', [...])."""
return FakeChatWithUsage(responses=["ok"])
@pytest.fixture
def summarize_chain(fake_chat):
prompt = ChatPromptTemplate.from_messages([
("system", "Summarize the user's text in one line."),
("user", "{text}"),
])
return prompt | fake_chat
Step 4 — VCR cassettes for integration tests with key redaction (P44 fix)
Unit tests should never touch the network. Integration tests do, exactly once —
to record a cassette — and every subsequent run replays from the cassette file.
vcrpy records headers by default, which means Authorization: Bearer sk-...
lands in the fixture unless you filter it.
import pytest
@pytest.mark.vcr # cassette at tests/cassettes/<test_name>.yaml
@pytest.mark.integration
def test_live_claude_short_answer():
from langchain_anthropic import ChatAnthropic
chat = ChatAnthropic(model="claude-sonnet-4-6", temperature=0, timeout=30)
out = chat.invoke("Say 'ok' and nothing else.")
assert "ok" in out.content.lower()
To record (once, locally, with a real key): pytest --record-mode=once tests/.
Every other run replays — cassettes are committed, real API is never hit again.
Pre-commit hook to block key leaks:
bash
# .git/hooks/pre-commit or .pre-commit-config.yaml entry
#!/usr/bin/env bash
set -e
if git diff --cached --name-only | grep -q '^tests/cassettes/'; then
if git diff --cached -U0 -- 'tests/cassettes/' | \
grep -E '(sk-ant-[a-zA-Z0-9_-]+|sk-[a-zA-Z0-9]{20,}|Bearer\s+[a-zA-Z0-9_-]{20,})'; then
echo "ERROR: API key pattern found in staged cassette." >&2
exit 1
fi
fi
See VCR Cassette Hygiene for the full
pre-commit config, record-new-episodes flow, shared-cassette patterns, and the
PR review checklist.
langchain_community and some provider SDKs emit DeprecationWarning at import
time. If the suite runs -W error, collection fails before any test does. Set
the policy once in pyproject.toml:
toml
[tool.pytest.ini_options]
minversion = "8.0"
testpaths = ["tests"]
addopts = [
"-ra",
"--strict-markers",
"--strict-config",
"-W", "error",
]
markers = [
"integration: hits real APIs or replays VCR cassettes (set RUN_INTEGRATION=1)",
"slow: takes > 1s per test",
"smoke: minimal healthcheck run in CI",
]
filterwarnings = [
"error",
"ignore::DeprecationWarning:langchain_community.*",
"ignore::DeprecationWarning:pydantic.*",
"ignore::PendingDeprecationWarning:langchain_core.*",
]
See Pytest Config for the full skeleton
including coverage config and parallel execution notes.
Step 6 — Integration-test gating via env var
Default pytest must never hit real APIs. Gate on RUN_INTEGRATION=1:
python
# tests/conftest.py (continued)
import os
import pytest
def pytest_collection_modifyitems(config, items):
if os.getenv("RUN_INTEGRATION") == "1":
return
skip_integration = pytest.mark.skip(reason="set RUN_INTEGRATION=1 to run")
for item in items:
if "integration" in item.keywords:
item.add_marker(skip_integration)
Step 7 — LangGraph tests: per-test thread_id + state assertions
LangGraph state is scoped to a thread_id. Tests that share a thread_id leak
state between each other. Give every test a fresh thread_id and a fresh
MemorySaver:
python
from langgraph.checkpoint.memory import MemorySaver
import uuid, pytest
@pytest.fixture
def graph_config():
return {"configurable": {"thread_id": str(uuid.uuid4())}}
@pytest.fixture
def checkpointed_graph(fake_chat):
from my_app.graphs import build_graph
return build_graph(fake_chat).compile(checkpointer=MemorySaver())
def test_node_emits_plan(checkpointed_graph, graph_config, fake_chat):
fake_chat.responses = ["step 1\nstep 2\nstep 3"]
result = checkpointed_graph.invoke({"goal": "deploy"}, graph_config)
# Assert state shape per node, not just the final output:
assert result["plan"] == ["step 1", "step 2", "step 3"]
# Time-travel: inspect every checkpoint for debugging
history = list(checkpointed_graph.get_state_history(graph_config))
assert history[-1].values == {"goal": "deploy"} # initial state
Subgraph isolation testing cross-references langchain-langgraph-subgraphs
(pain P21 — parent cannot read child state unless the key is in the parent
schema). See LangGraph Test Patterns
for the subgraph-shared-state test recipe.
Show full SKILL.md (415 more words)Show less
Output
tests/fakes.py with FakeChatWithUsage subclass that emits response_metadata
tests/conftest.py with fake-model fixtures, VCR config, and RUN_INTEGRATION gate
pyproject.toml[tool.pytest.ini_options] block with markers and filterwarnings
tests/cassettes/ committed with filtered headers (no Authorization / x-api-key)
Commit 1 — failing test uses real ChatAnthropic, passes locally, fails
1-in-5 in CI at temperature=0 (P05).
Commit 2 — swap to fake model uses FakeListChatModel, passes
deterministically, but the cost-logging callback crashes (P43).
Commit 3 — fake with metadata uses FakeChatWithUsage, the callback
reads response_metadata["token_usage"] cleanly, the test is green and
runs in 40ms.
See Fake Model Fixtures for the full
worked example including agent and retriever fakes.
Recording a cassette without leaking a key
bash
# 1. Ensure conftest.py has filter_headers configured FIRST
# 2. Record with real key present in the environment
ANTHROPIC_API_KEY=sk-ant-... pytest --record-mode=once tests/integration/test_summarize.py
# 3. Verify no leak
grep -E 'sk-|Bearer' tests/cassettes/*.yaml && echo "LEAK" || echo "clean"
# 4. Commit cassettes/ — pre-commit hook runs the same grep as a hard gate
git add tests/cassettes/ && git commit -m "test: record summarize cassette"
See VCR Cassette Hygiene for
record-new-episodes mode, rerecord-on-mismatch, and the PR review checklist.
LangGraph time-travel debugging on a failing test
When a graph test fails mid-graph, get_state_history(config) returns every
checkpoint — you can replay from any point by passing its config.checkpoint_id
back into graph.invoke. See
LangGraph Test Patterns for the full
time-travel debugging recipe and the subgraph-shared-state test pattern
(cross-ref langchain-langgraph-subgraphs / pain L30).
Langchain Local Dev Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
Langchain Local Dev Loop compared with similar skills
Skill
Stars
Used in
Tokens
Auto-check
Licence
Repo updated
Langchain Local Dev Loop this skilljeremylongshore/tons-of-skills-marketplace
A skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…
Tests JavaScript embedded in an HTML file in two layers: pytest checks of the logic ported to Python, and Playwright runs in a real browser for the DOM.
Build a fast, deterministic local test loop for LangChain 1.0 / LangGraph 1.0 — FakeListChatModel fixtures, pytest config, VCR cassettes with key redaction, warning-filter policy. Langchain Local Dev Loop is an agent skill from jeremylongshore/tons-of-skills-marketplace.0 — FakeListChatModel fixtures, pytest config, VCR cassettes with key redaction, warning-filter policy.
When should I use Langchain Local Dev Loop?
Langchain Local Dev Loop fits situations like: adding tests to a new chain; fixing a flaky test; making integration tests reproducible; with langchain pytest.
How do I install Langchain Local Dev Loop in Claude Code?
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-local-dev-loop -a claude-code`. Or copy the skill folder (skills/.curated/langchain-local-dev-loop in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langchain-local-dev-loop in your project. Claude Code loads it when a task matches its description.
How do I install Langchain Local Dev Loop in Codex?
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-local-dev-loop -a codex`. Or copy the skill folder (skills/.curated/langchain-local-dev-loop in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langchain-local-dev-loop in your project. Codex loads it when a task matches its description.
Can I use Langchain Local Dev Loop in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-local-dev-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langchain-local-dev-loop, .gemini/skills/langchain-local-dev-loop, .github/skills/langchain-local-dev-loop and .opencode/skills/langchain-local-dev-loop in your project.
What does Langchain Local Dev Loop need to run?
Going by SKILL.md and its folder, Langchain Local Dev Loop needs the command-line tools its instructions call (pytest, git and pip) and credentials named ANTHROPIC_API_KEY. Our summary lists: Python 3; A credential in ANTHROPIC_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(pytest:*), Bash(python:*), Bash(pip:*). Compatibility (from SKILL.md): Designed for Claude Code.
Does Langchain Local Dev Loop access the network?
SKILL.md names 4 domains. As links in the text: python.langchain.com, vcrpy.readthedocs.io, pytest-vcr.readthedocs.io and docs.pytest.org. This is read from the text; nothing was executed.
Is Langchain Local Dev Loop safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
What licence does Langchain Local Dev Loop use?
Langchain Local Dev Loop is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Langchain Local Dev Loop use?
About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.7k tokens, read only when the agent opens those files.
What are the alternatives to Langchain Local Dev Loop?
Skills that share tags, products or a category with Langchain Local Dev Loop: Langgraph Testing Evaluation (soba-labs/langchain-agent-skills, 107 stars), Agent Eval Cases (agentailor/fullstack-langgraph-nextjs-agent, 132 stars), Designing Tests (CloudAI-X/claude-workflow-v2, 1.4k stars) and NIC Testing Patterns (nginx/kubernetes-ingress, 5.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Langchain Local Dev Loop?
jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.