Install the "langchain-ci-integration" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-ci-integration into .claude/skills/langchain-ci-integration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-ci-integration", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-ci-integration -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "langchain-ci-integration" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-ci-integration into .agents/skills/langchain-ci-integration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-ci-integration", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-ci-integration -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "langchain-ci-integration" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-ci-integration into .cursor/skills/langchain-ci-integration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-ci-integration", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-ci-integration -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "langchain-ci-integration" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-ci-integration into .gemini/skills/langchain-ci-integration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-ci-integration", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-ci-integration -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "langchain-ci-integration" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-ci-integration into .github/skills/langchain-ci-integration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-ci-integration", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-ci-integration -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "langchain-ci-integration" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-ci-integration into .opencode/skills/langchain-ci-integration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-ci-integration", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
langchain-ci-integration
GitHub stars
2.8k
Token cost
~4.4k tokens
SKILL.md length
1,482 words
Files
6 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT
At a glance
Wire LangChain 1.0 / LangGraph 1.0 tests into a GitHub Actions pipeline — unit tests with FakeListChatModel, VCR-gated integration tests, warning-filter policy, and eval-regression merge gates.
Works in 6 steps: GHA workflow skeleton with four jobs → Unit job: -W error + filterwarnings to… → Integration job: VCR replay +… → …
Setting up GHA for a new LLM service
SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 3 more sections
Calls pytest and git; reaches github.com
What it does
Langchain CI Integration is an agent skill from jeremylongshore/tons-of-skills-marketplace. Wire LangChain 1.0 / LangGraph 1.0 tests into a GitHub Actions pipeline — unit tests with FakeListChatModel, VCR-gated integration tests, warning-filter policy, and eval-regression merge gates. Complements langchain-local-dev-loop (F23) which covers the inner loop; THIS covers the CI wire-up. Use when setting up GHA for a new LLM service, after a VCR cassette leak incident, or hardening an existing pipeline. Trigger with "langchain ci", "langchain github actions", "langchain test pipeline", "vcr ci", "langchain…
Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/eval-regression-gate.md`, `references/github-actions-workflow.md` and `references/integration-gating.md`). Compatibility notes: Designed for Claude Code
It sits in AI & LLM Engineering, covering Building AI agents, Unit testing and CI/CD. It works with LangChain, GitHub Actions, pytest and LangGraph. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.
When your agent uses it
Setting up GHA for a new LLM service
After a VCR cassette leak incident
Hardening an existing pipeline
With langchain ci
Example prompts
“langchain ci”
“langchain github actions”
“langchain test pipeline”
“/langchain-ci-integration”
Requirements
Python 3
Compatibility (from SKILL.md): Designed for Claude Code
Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves these tools, so the agent can use them without asking each time:
Read
Write
Edit
Bash(python:*)
Bash(pytest:*)
From allowed-tools in the SKILL.md frontmatter.
Runs code
Shell commands in SKILL.md call:
pytest
git
From the folder's file list and the shell code blocks in SKILL.md.
Network
Hosts in commands or code, which the agent is likely to contact:
github.com
Also links to:
python.langchain.com
vcrpy.readthedocs.io
docs.github.com
docs.pytest.org
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Compatibility
Designed for Claude Code
From compatibility in the SKILL.md frontmatter.
Context cost
Langchain CI Integration loads about 4.4k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 146 tokens; SKILL.md has 1,482 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~146
When it runs· the whole SKILL.md, loaded when a task matches
~4.4k
With references· SKILL.md plus every file in references/, read only if the agent opens them
~12k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
Download SKILL.mdSave it as .claude/skills/langchain-ci-integration/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
langchain-ci-integration
description
Wire LangChain 1.0 / LangGraph 1.0 tests into a GitHub Actions pipeline —
unit tests with FakeListChatModel, VCR-gated integration tests, warning-filter
policy, and eval-regression merge gates. Complements langchain-local-dev-loop
(F23) which covers the inner loop; THIS covers the CI wire-up. Use when setting
up GHA for a new LLM service, after a VCR cassette leak incident, or hardening
an existing pipeline.
Trigger with "langchain ci", "langchain github actions", "langchain test pipeline",
"vcr ci", "langchain eval gate", "pytest -W error langchain".
The org runs pytest -W error and a provider SDK emitted a DeprecationWarning
at import time, which the warning filter promoted to an exception while pytest
was still walking the test tree. This is P45 and it blocks every PR for the
team until someone pins a filterwarnings config.
Meanwhile the integration suite has its own failure mode: a VCR cassette
recorded three months ago at temperature=0 against Anthropic is now flaking
against a snapshot. temperature=0 is not deterministic on Claude — it still
nucleus-samples (P05) — so the cassette captured one valid completion, not
the valid completion. And yesterday a reviewer caught
Authorization: Bearer sk-ant-... in a cassette file that had been committed
six weeks ago (P44) because vcrpy records all request headers by default.
This skill covers the outer loop: the GitHub Actions workflow, the unit /
integration / eval gate separation, VCR cassette hygiene, pytest warning
policy, and a merge-blocking eval regression gate. The inner loop — fake
model fixtures, VCR recording workflow, local determinism tricks — lives in
langchain-local-dev-loop (F23); cross-reference it, do not duplicate it.
Pin: langchain-core 1.0.x, langgraph 1.0.x, actions/checkout@v4,
actions/setup-python@v5, vcrpy 6.x. Pain-catalog anchors: P05, P43, P44, P45.
See GHA Workflow Reference for the full
job definitions including the secret-injection pattern, the matrix caching
nuance, and the softprops/action-gh-release-style PR comment action used by
the eval job.
Step 2 — Unit job: -W error + filterwarnings to neutralize P45
Root cause of the collection abort: pytest collects tests by importing them.
Some provider SDKs emit DeprecationWarning on import. With -W error those
become exceptions during collection. Fix at the filter level, not by dropping
-W error (which would mask real warnings).
In pyproject.toml:
toml
[tool.pytest.ini_options]
filterwarnings = [
"error",
# P45 — neutralize known import-time noise; scoped per module so new
# warnings from YOUR code still fail the build.
"ignore::DeprecationWarning:langchain_community.*",
"ignore::DeprecationWarning:pydantic.*",
"ignore:Pydantic serializer warnings:UserWarning",
]
asyncio_mode = "auto"
testpaths = ["tests"]
The ordering matters — "error" first, specific "ignore" entries after, so
the filters override the global promote-to-error. Keep the list narrow: a
blanket ignore::DeprecationWarning hides regressions you need to see.
Unit tests use FakeListChatModel fixtures from F23 (do not redefine them
here). One CI-specific gotcha (P43): FakeListChatModel does not emit
response_metadata["token_usage"], so any callback that asserts on token counts
will break. Either subclass the fake and inject generation_info, or gate the
assertion:
python
def test_chain_uses_tokens(patched_chat_model):
result = chain.invoke({"input": "hi"})
if patched_chat_model.__class__.__name__ == "FakeListChatModel":
pytest.skip("fake model doesn't emit token_usage (P43)")
assert result.response_metadata["token_usage"]["total_tokens"] > 0
Budget: unit job should finish in <2 minutes across the 3-version matrix.
If it doesn't, something is calling out to a real provider — check with
pytest --collect-only -q | wc -l and audit which tests lack fake-model
fixtures.
Integration tests replay pre-recorded VCR cassettes. Three rules:
Gate the job. if: contains(github.event.pull_request.labels.*.name, 'run-integration') or env.RUN_INTEGRATION == "1", plus a nightly cron that flips to VCR_MODE=once and re-records against live APIs. PRs default to pure replay.
Enforce filter_headers at the fixture level — not per-test. A single conftest.py prevents any contributor from recording a cassette with raw credentials.
Pre-commit + CI both scan cassettes for leaked keys. Belt and suspenders.
Fixture (lives in tests/integration/conftest.py, owned by this skill's
pipeline concern — F23 owns the recording workflow):
Integration suite must finish in <5 minutes wall-clock on the runner, or
you will start getting cancellation flakes from the concurrency block. If
you exceed 5 minutes, split into a nightly-only long tier.
See Integration Gating for the full
live-vs-replay decision tree, cost-per-run budget worksheet, and the
VCR_MODE flip pattern.
The eval job runs the langchain-eval-harness harness (see that skill for the
harness itself — this skill only covers the CI wire-up) against both the PR
branch and the merge base. Post a comment; block merge on regression.
scripts/run_eval.py is a thin CI wrapper: check out baseline and head via
git worktree, run the harness at each ref, diff the results, post a PR
comment, exit nonzero on regression. Full implementation in
Eval Regression Gate.
Thresholds:
Gate
Threshold
Rationale
Aggregate score
drop >2%
One-sigma noise on n=100 with well-behaved evals
Per-example score
drop >5% on any single case
Catches quiet regressions masked by aggregate averaging
Sample size floor
n ≥ 100
Below this, aggregate delta is dominated by noise
The PR comment is a Markdown table with before / after / Δ per metric plus a
bold red line if the gate failed. Required-status-check on the eval job
completes the enforcement. See Eval Regression Gate
for the comment template and the noise-budget calculation.
Two layers: local (pre-commit) and CI (re-runs the same hooks as a final
catch). Local alone is not sufficient — contributors can skip with -n. CI
alone is slow. Run both.
scan_cassettes.py greps for sk-[A-Za-z0-9]{20,}, sk-ant-[A-Za-z0-9_-]{20,},
AIza[A-Za-z0-9_-]{35} (Google), xoxb-, and Bearer [A-Za-z0-9._-]{20,}.
Fail on any match. This is your last line of defense before P44 ships to
main. See Pre-Commit Hooks for the full
pattern list, the prompt-convention lint rules (aligned with
claude-prompt-conventions), and the detect-secrets baseline-rotation policy.
LangChain 0.x → 1.0 moved integrations into provider packages. A chain that
imports from langchain.chat_models import ChatOpenAI works in local dev if
you still have the old compat shim installed, and explodes in CI. Dry-run-load
every chain module at lint time:
python
# scripts/dryrun_load_chains.py
import importlib, pathlib, sys, traceback
failures = []
for py in pathlib.Path("src/chains").rglob("*.py"):
mod = str(py.with_suffix("")).replace("/", ".")
try:
importlib.import_module(mod)
except Exception:
failures.append((mod, traceback.format_exc()))
if failures:
for mod, tb in failures:
print(f"::error::chain {mod} failed to import\n{tb}")
sys.exit(1)
Runs in the lint job. Costs ~5 seconds. Catches every ImportError and
every top-level NameError from a bad rename before a single unit test fires.
Output
GHA workflow with four isolated jobs (unit / integration / eval / lint)
pyproject.tomlfilterwarnings config that survives -W error (P45)
VCR conftest.py fixture with enforced filter_headers (P44)
run_eval.py CI wrapper that posts PR comments and blocks merge on regression
.pre-commit-config.yaml with cassette secret scan + prompt lint + ruff
Dry-run chain loader that catches migration ImportErrors
Gate policy
Gate
Required?
Target speed
On failure
unit (3 Python versions)
yes, every PR
<2 min
block PR
lint + dryrun-load
yes, every PR
<30 s
block PR
integration (VCR replay)
on run-integration label or nightly
<5 min
block merge when run
integration (live, nightly cron)
no
<15 min
open issue on fail
eval regression (n≥100)
yes, every PR
<10 min
block merge if agg >2% or per-example >5%
pre-commit (local)
yes
<10 s
reject commit
Error Handling
Error
Cause
Fix
PytestUnraisableExceptionWarning during collection
Add scoped filterwarnings = ["ignore::DeprecationWarning:langchain_community.*"] to pyproject.toml
VCR replay mismatch after weeks of passing
Cassette recorded at temp=0 on Anthropic (P05); model drift
Re-record on nightly cron with VCR_MODE=once; treat replay mismatches as eval-gate concerns, not unit failures
sk-ant-... in cassette flagged by reviewer
vcrpy records all headers by default (P44)
Enforce filter_headers in conftest.py; add scan_cassettes.py to pre-commit AND CI
Callback AssertionError: 'token_usage' not in response_metadata
FakeListChatModel doesn't emit metadata (P43)
Subclass the fake to inject generation_info, or pytest.skip on fake-model detection
ImportError: cannot import name 'ChatOpenAI' from 'langchain.chat_models' in CI only
Legacy compat shim installed locally, not in CI
Add dryrun_load_chains.py to lint job; fail at lint, not at test
Eval job times out at 10 min
n too large or harness not using asyncio concurrency
Cap at n=100 for PRs; run n=500 nightly; see F23 for async harness pattern
Concurrency block cancels integration run
Long job + rapid pushes
Do not disable; keep integration <5 min or split long tier to nightly
Examples
Wiring a new repo from scratch
Copy the Step 1 workflow, the Step 2 pyproject.toml block, and the Step 5
pre-commit config. Create tests/unit/, tests/integration/cassettes/,
scripts/run_eval.py, scripts/dryrun_load_chains.py,
scripts/scan_cassettes.py. Apply langchain-local-dev-loop (F23) first so
fake-model fixtures exist before the unit job runs. Enable required status
checks: unit (3.10), unit (3.11), unit (3.12), lint, eval.
Integration stays optional (label-gated).
Rotate every leaked key first (not a CI concern — incident response).
Then: add scan_cassettes.py to pre-commit, re-scan the full history with
git log -p -- tests/integration/cassettes/, rewrite history with
git-filter-repo if keys hit main, enforce the filter_headers fixture
going forward. See Pre-Commit Hooks for the
full pattern list and the detect-secrets baseline-rotation playbook.
Wiring the eval harness into an existing repo
The harness itself lives in langchain-eval-harness. THIS skill only supplies
run_eval.py (the CI wrapper that reads the harness output, computes deltas,
and posts PR comments) plus the gate thresholds. Drop in the Step 4 script,
add the eval job to .github/workflows/tests.yml, make eval a required
status check. See Eval Regression Gate
for the PR-comment Markdown template and the n≥100 noise-budget derivation.
Langchain CI Integration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
Langchain CI Integration compared with similar skills
Skill
Stars
Used in
Tokens
Auto-check
Licence
Repo updated
Langchain CI Integration this skilljeremylongshore/tons-of-skills-marketplace
A skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…
Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.
Scans Python agent code for framework imports and recommends the matching Omnigent executor type, or says when the framework is not natively supported yet.
11k GitHub stars~610 tokensUpdated today
AI & LLM EngineeringAuto-check passed
More from jeremylongshore/tons-of-skills-marketplace
Wire LangChain 1.0 / LangGraph 1.0 tests into a GitHub Actions pipeline — unit tests with FakeListChatModel, VCR-gated integration tests, warning-filter policy, and eval-regression merge gates. Langchain CI Integration is an agent skill from jeremylongshore/tons-of-skills-marketplace.0 tests into a GitHub Actions pipeline — unit tests with FakeListChatModel, VCR-gated integration tests, warning-filter policy, and eval-regression merge gates.
When should I use Langchain CI Integration?
Langchain CI Integration fits situations like: setting up GHA for a new LLM service; after a VCR cassette leak incident; hardening an existing pipeline; with langchain ci.
How do I install Langchain CI Integration in Claude Code?
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-ci-integration -a claude-code`. Or copy the skill folder (skills/.curated/langchain-ci-integration in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langchain-ci-integration in your project. Claude Code loads it when a task matches its description.
How do I install Langchain CI Integration in Codex?
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-ci-integration -a codex`. Or copy the skill folder (skills/.curated/langchain-ci-integration in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langchain-ci-integration in your project. Codex loads it when a task matches its description.
Can I use Langchain CI Integration in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-ci-integration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langchain-ci-integration, .gemini/skills/langchain-ci-integration, .github/skills/langchain-ci-integration and .opencode/skills/langchain-ci-integration in your project.
What does Langchain CI Integration need to run?
Going by SKILL.md and its folder, Langchain CI Integration needs the command-line tools its instructions call (pytest and git). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(python:*), Bash(pytest:*). Compatibility (from SKILL.md): Designed for Claude Code.
Does Langchain CI Integration access the network?
SKILL.md names 5 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: python.langchain.com, vcrpy.readthedocs.io, docs.github.com and docs.pytest.org. This is read from the text; nothing was executed.
Is Langchain CI Integration safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
What licence does Langchain CI Integration use?
Langchain CI Integration is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Langchain CI Integration use?
About 4.4k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.6k tokens, read only when the agent opens those files.
What are the alternatives to Langchain CI Integration?
Skills that share tags, products or a category with Langchain CI Integration: Langgraph Testing Evaluation (soba-labs/langchain-agent-skills, 107 stars), Simple Modern Uv (jlevy/simple-modern-uv, 301 stars), Code Patterns (Aedelon/claude-code-blueprint, 120 stars) and Add Example Agent (GetBindu/Bindu, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Langchain CI Integration?
jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.