Agent skill

Langchain Performance Tuning

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Tune LangChain 1.0 / LangGraph 1.0 Python chains and agents for throughput, latency, and cost — streaming modes, explicit batch concurrency, semantic plus exact caches, persistent message history…

MITAuto-check passedAI & LLM Engineering

Install Langchain Performance Tuning

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-performance-tuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-performance-tuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/langchain-performance-tuning .claude/skills/langchain-performance-tuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langchain-performance-tuning
GitHub stars
2.8k
Token cost
~3.5k tokens
SKILL.md length
1,242 words
Files
6 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Tune LangChain 1.0 / LangGraph 1.0 Python chains and agents for throughput, latency, and cost — streaming modes, explicit batch concurrency, semantic plus exact caches, persistent message history…

  • Works in 10 steps: Establish a latency budget and baseline.… → Convert every hot path to async (P48).… → Fix .abatch() concurrency (P08). Every… → …
  • P95 latency exceeds target
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Langchain Performance Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Tune LangChain 1.0 / LangGraph 1.0 Python chains and agents for throughput, latency, and cost — streaming modes, explicit batch concurrency, semantic plus exact caches, persistent message history, and async-safe retriever patterns. Use when p95 latency exceeds target, batching "does not work", cost grows linearly with traffic, or a process restart wipes chat history. Trigger with "langchain performance", "langchain slow batch", "langchain throughput", "langchain p95 latency", "semantic cache hit rate".

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/async-safety-checklist.md`, `references/batch-concurrency-per-provider.md` and `references/cache-tuning.md`). Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering Building AI agents. It works with LangChain, LangGraph, Python and Redis. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • P95 latency exceeds target
  • Batching does not work
  • Cost grows linearly with traffic
  • A process restart wipes chat history

Example prompts

  • “does not work”
  • “langchain performance”
  • “langchain slow batch”
  • “/langchain-performance-tuning”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(python:*), Bash(redis-cli:*)

Workflow steps

10 steps, taken from the first numbered list in SKILL.md.

  1. Establish a latency budget and baseline. Pick explicit targets before changing code: TTFT under 1s, p95 total under 5s, throughput over 20…
  2. Convert every hot path to async (P48). Inside async def handlers, replace invoke, stream, batch, get_relevant_documents, and tool.run with…
  3. Fix .abatch() concurrency (P08). Every .abatch / .batch call must pass config={"max_concurrency": N} where N is chosen from the provider…
  4. Instrument TTFT with astream_events(version="v2") (P01). Measure time to first token separately from total latency — user-perceived…
  5. Enable an exact LLM cache. For deterministic (temperature=0) prompts, set RedisCache or SQLiteCache globally. LangChain 1.0 keys include…
  6. Add a semantic cache with a tuned threshold (P62). The RedisSemanticCache default score_threshold=0.95 produces < 5% hit rate on real…
  7. Replace InMemoryChatMessageHistory (P22). Every production chat path must use RedisChatMessageHistory (with ttl) or a LangGraph…
  8. Close retriever connection pools in FastAPI lifespan (P59). Build the vector store once at startup, expose it via app.state, close it in…
  9. Stream tokens with SSE, not BackgroundTasks (P60). BackgroundTasks runs after the response body is flushed; per-token dispatch via it…
  10. Re-run the load test and diff the four metrics. TTFT, p95, throughput, cost per 1k. If any regressed, revert that step and investigate…

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(python:*)
    • Bash(redis-cli:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • python.langchain.com
    • langchain-ai.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Langchain Performance Tuning loads about 3.5k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 134 tokens; SKILL.md has 1,242 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~134
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,242 words, ~3,544 tokens.

Download SKILL.mdSave it as .claude/skills/langchain-performance-tuning/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
langchain-performance-tuning
description
Tune LangChain 1.0 / LangGraph 1.0 Python chains and agents for throughput, latency, and cost — streaming modes, explicit batch concurrency, semantic plus exact caches, persistent message history, and async-safe retriever patterns. Use when p95 latency exceeds target, batching "does not work", cost grows linearly with traffic, or a process restart wipes chat history. Trigger with "langchain performance", "langchain slow batch", "langchain throughput", "langchain p95 latency", "semantic cache hit rate".
allowed-tools
Read, Write, Edit, Bash(python:*), Bash(redis-cli:*)
compatibility
Designed for Claude Code
version
2.7.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, langchain, langgraph, python, langchain-1.0, performance, caching, async

LangChain Performance Tuning

Overview

An engineer calls chain.batch(inputs_1000) expecting 1000 parallel LLM calls. Actual behavior: Runnable.batch and Runnable.abatch in LangChain 1.0 default to max_concurrency=1, so the 1000 inputs run sequentially with bookkeeping overhead — sometimes slower than a plain for loop. This is pain-catalog entry P08. The fix is one line:

python
# Before: serial, ~1000 * per_call_latency
await chain.abatch(inputs)

# After: 10x throughput at 10 providers' worth of concurrency
await chain.abatch(inputs, config={"max_concurrency": 10})

Other silent regressions in the same pain catalog: P48 (invoke inside async def blocks the FastAPI event loop), P22 (InMemoryChatMessageHistory loses every user's chat on restart), P62 (RedisSemanticCache at the default score_threshold=0.95 returns under 5% hit rate), P59 (async retrievers leak connections on cancellation), P60 (BackgroundTasks fires after the response — wrong for per-token SSE), P01 (streaming token counts are only reliable on the on_chat_model_end event).

This skill wires a production performance baseline: explicit batch concurrency, async-only code paths, Redis-backed caches tuned on a golden set, persistent chat history with TTL, and TTFT instrumentation from astream_events(version="v2").

Prerequisites

  • Python 3.11+ with langchain>=1.0,<2, langgraph>=1.0,<2, langchain-openai or langchain-anthropic, langchain-community, langchain-redis or redis>=5.
  • A working LangChain 1.0 chain or LangGraph 1.0 graph that already passes functional tests.
  • Redis 7+ reachable from the app for cache and history (local Docker is fine for dev).
  • A FastAPI / Starlette async endpoint, or an equivalent async entrypoint.
  • Observability: a place to emit metrics (Prometheus, OpenTelemetry, or LangSmith) — needed to measure TTFT, p95, and cache hit rate.

Instructions

  1. Establish a latency budget and baseline. Pick explicit targets before changing code: TTFT under 1s, p95 total under 5s, throughput over 20 req/s per worker, cost under $X per 1k interactions. Run a 5-minute load test with locust or wrk against the current chain and record p50 / p95 / p99 / TTFT / total cost. Without these numbers every downstream change is theater.

  2. Convert every hot path to async (P48). Inside async def handlers, replace invoke, stream, batch, get_relevant_documents, and tool.run with ainvoke, astream / astream_events(version="v2"), abatch, aget_relevant_documents, and tool.arun. See references/async-safety-checklist.md for a grep pattern and a CI linter. Target: zero sync LangChain calls inside any async function.

  3. Fix .abatch() concurrency (P08). Every .abatch / .batch call must pass config={"max_concurrency": N} where N is chosen from the provider table in references/batch-concurrency-per-provider.md (Anthropic 10-20, OpenAI 20-50, local vLLM 100+). For multi-worker deploys, cap account-wide calls with a LiteLLM / Portkey proxy or a Redis semaphore — max_concurrency only governs one process.

  4. Instrument TTFT with astream_events(version="v2") (P01). Measure time to first token separately from total latency — user-perceived performance hinges on TTFT. Read usage metadata only on the on_chat_model_end event; per-chunk usage fields lag and are not reliable mid-stream.

    python
    from time import perf_counter
    async def run(chain, query: str):
        t0 = perf_counter(); ttft = None; tokens = 0
        async for ev in chain.astream_events({"input": query}, version="v2"):
            if ev["event"] == "on_chat_model_stream" and ttft is None:
                ttft = perf_counter() - t0
            if ev["event"] == "on_chat_model_end":
                tokens = ev["data"]["output"].usage_metadata["total_tokens"]
        return {"ttft_s": ttft, "total_s": perf_counter() - t0, "tokens": tokens}
  5. Enable an exact LLM cache. For deterministic (temperature=0) prompts, set RedisCache or SQLiteCache globally. LangChain 1.0 keys include the bound tools signature (P61 fix), which prevents cache poisoning when an agent's tool list changes. Always set an explicit TTL on Redis keys — default Redis keys are immortal.

    python
    from langchain_core.globals import set_llm_cache
    from langchain_community.cache import RedisCache
    import redis
    set_llm_cache(RedisCache(redis.Redis.from_url("redis://cache:6379/0")))
  6. Add a semantic cache with a tuned threshold (P62). The RedisSemanticCache default score_threshold=0.95 produces < 5% hit rate on real traffic. Collect a 200-500 prompt golden set with labeled near-duplicates, measure cosine similarity with your embedding model, and pick the F1-maximizing threshold — typically 0.85-0.90 for text-embedding-3-small. Full procedure in references/cache-tuning.md. Do not run semantic cache behind temperature > 0; users will see prior random draws.

  7. Replace InMemoryChatMessageHistory (P22). Every production chat path must use RedisChatMessageHistory (with ttl) or a LangGraph checkpointer (AsyncPostgresSaver / AsyncSqliteSaver). Add a restart test: mid-conversation, kill and restart the worker, assert the next user turn still sees prior messages. See references/persistent-history.md for migration steps and trim policies.

  8. Close retriever connection pools in FastAPI lifespan (P59). Build the vector store once at startup, expose it via app.state, close it in the finally block. Never construct a retriever per request — cancellations leak pg connections.

  9. Stream tokens with SSE, not BackgroundTasks (P60). BackgroundTasks runs after the response body is flushed; per-token dispatch via it delivers tokens the client will never read. Use EventSourceResponse (sse-starlette) or a WebSocket and pipe events from astream_events.

  10. Re-run the load test and diff the four metrics. TTFT, p95, throughput, cost per 1k. If any regressed, revert that step and investigate — do not stack changes without verification. Execute in this order to isolate effects:

    1. Run the baseline load test and save results.
    2. Set max_concurrency on every .abatch call and re-run.
    3. Add exact cache, re-run, check cache hit rate.
    4. Configure semantic cache with tuned threshold, re-run, check hit rate again.
    5. Verify persistent history survives a worker restart.
Show full SKILL.md (514 more words)Show less
Throughput Tuning Table (starting values)
ProviderSafe max_concurrencyCeiling signal
Anthropic (sonnet-4.5/4.6)10-20429 rate_limit_error
OpenAI (gpt-4o / 4o-mini)20-50429 + TPM exhaustion header
OpenAI o1 / reasoning2-5Cost + latency, not rate
Google Gemini 1.5/2.510-30429
Cohere20-40429
Local vLLM / TGI100-500 (batch N≈32-64)GPU KV-cache OOM
Ollama on consumer GPU1-4Process queue backpressure
Latency Breakdown Template

Record these for every change, not just total:

MetricTargetSource
TTFT p50 / p95500ms / 1sfirst on_chat_model_stream event
Total p50 / p952s / 5send-to-end handler
Tool-call p95< 1s per toolon_tool_end - on_tool_start
Retriever p95< 300mson_retriever_end - on_retriever_start
Provider p95measure per modelsplit by LLM node
Batch Sweet-Spot Numbers
  • Anthropic tier 2 chat: max_concurrency=10 saturates at roughly 8 req/s, p95 doubles past 20.
  • OpenAI gpt-4o-mini tier 3: knee of the curve around max_concurrency=30-40; ~40 req/s throughput.
  • Local vLLM A100: server-side batch sweet spot N=32-64, client max_concurrency=100+.

Verify on your own account — these are starting points, not promises.

Output

Deliverables from running this skill end-to-end:

  • A perf/ directory with baseline.json and tuned.json load-test results.
  • All async handlers use ainvoke / astream_events / abatch with explicit max_concurrency.
  • set_llm_cache wired to RedisCache (exact) and optionally RedisSemanticCache (tuned threshold).
  • RunnableWithMessageHistory or LangGraph checkpointer backed by Redis or Postgres, with TTL.
  • FastAPI lifespan closing vector store pools on shutdown.
  • SSE endpoint streaming from astream_events(version="v2").
  • A tests/test_no_sync_in_async.py CI guard (see async-safety reference).
  • Metrics exported: ttft_seconds, total_latency_seconds, cache_hit_total, cache_miss_total, batch_concurrency_current.
  • Runbook entry with the tuned max_concurrency per provider and the semantic-cache threshold, versioned in git.

Error Handling

SymptomRoot causeFix
.abatch(inputs) no faster than a for loopmax_concurrency=1 default (P08)Pass config={"max_concurrency": N}
FastAPI TTFT collapses under loadSync invoke inside async def (P48)Switch to ainvoke / astream_events
Chat forgets prior turns after deployInMemoryChatMessageHistory (P22)Move to RedisChatMessageHistory with TTL
Semantic cache hit rate < 5%score_threshold=0.95 default (P62)Tune on golden set to 0.85-0.90
pg pool exhausted hours into load testRetriever not closed on cancel (P59)Close vector store in FastAPI lifespan
SSE client sees zero tokensDispatching via BackgroundTasks (P60)Use EventSourceResponse and astream_events
Per-chunk token counts fluctuateUsage metadata lags during stream (P01)Read only on on_chat_model_end
429 storm after tuning concurrencyPer-worker limit * N workers > account RPMAdd LiteLLM/Portkey proxy or Redis semaphore
Semantic cache returns off-brand outputCache hit on temperature > 0 routeDisable semantic cache or force temperature=0
Cache poisoning after tool changeMissing tools in cache keyUpgrade LangChain to 1.0.x post-P61 fix

Examples

Example 1 — Fix a sequential batch job.

python
# Before — 1000 items, 18 minutes end-to-end
results = await chain.abatch(inputs)

# After — 1000 items, ~2 minutes; Anthropic tier-2 account, N=10
results = await chain.abatch(inputs, config={"max_concurrency": 10})

Example 2 — Wire persistent history and an exact cache on a FastAPI app.

python
from contextlib import asynccontextmanager
from fastapi import FastAPI
from langchain_core.globals import set_llm_cache
from langchain_core.runnables.history import RunnableWithMessageHistory
from langchain_community.cache import RedisCache
from langchain_community.chat_message_histories import RedisChatMessageHistory
import redis

@asynccontextmanager
async def lifespan(app: FastAPI):
    r = redis.Redis.from_url("redis://cache:6379/0")
    set_llm_cache(RedisCache(r))
    app.state.r = r
    yield
    r.close()

app = FastAPI(lifespan=lifespan)

def history_for(session_id: str) -> RedisChatMessageHistory:
    return RedisChatMessageHistory(
        session_id=session_id,
        url="redis://history:6379/2",
        ttl=60 * 60 * 24 * 14,
    )

chain_with_history = RunnableWithMessageHistory(
    base_chain, history_for,
    input_messages_key="input",
    history_messages_key="history",
)

Example 3 — Stream tokens with measured TTFT.

python
from sse_starlette.sse import EventSourceResponse
from time import perf_counter

@app.post("/chat")
async def chat(req: ChatReq):
    async def gen():
        t0 = perf_counter()
        async for ev in chain_with_history.astream_events(
            {"input": req.text},
            config={"configurable": {"session_id": req.session_id}},
            version="v2",
        ):
            if ev["event"] == "on_chat_model_stream":
                yield {"data": ev["data"]["chunk"].content}
        app.state.r.incrbyfloat("ttft_sum_s", perf_counter() - t0)
    return EventSourceResponse(gen())

Resources

  • One-pager — problem / solution / key features snapshot.
  • batch-concurrency-per-provider — per-provider max_concurrency table, sweep procedure, semaphore patterns.
  • cache-tuning — exact vs semantic, Redis key design, golden-set threshold procedure, TTL strategy.
  • persistent-history — Redis / Postgres / LangGraph checkpointer migration off InMemoryChatMessageHistory.
  • async-safety-checklist — sync-in-async grep + linter, lifespan pool cleanup, SSE vs BackgroundTasks.
  • LangChain streaming / batching — official docs for Runnable.batch and streaming modes.
  • LangChain caching — set_llm_cache, Redis and SQLite backends.
  • LangGraph checkpointers — persistence for graph state.
  • Companion skills in langchain-py-pack: langchain-model-inference (token accounting), langchain-embeddings-search (retrieval tuning), langchain-middleware-patterns (tool-signature cache keying, P61).

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/.curated/langchain-performance-tuning of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/async-safety-checklist.md
  • references/batch-concurrency-per-provider.md
  • references/cache-tuning.md
  • references/one-pager.md
  • references/persistent-history.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Langchain Performance Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Langchain Performance Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Langchain Performance Tuning this skilljeremylongshore/tons-of-skills-marketplace2.8k—~3.5kAutomated safety check: PassMIT
Add Example AgentGetBindu/Bindu10k—~1.1kAutomated safety check: NotesCustom licence
Failproof AI SDK IntegrationFailproofAI/failproofai5.3k—~6kAutomated safety check: PassCustom licence
Omnigent Framework Detectionomnigent-ai/omnigent11k—~610Automated safety check: PassApache-2.0
Deep Agents to Pydantic AI Migrationpydantic/pydantic-ai21k—~1.7kAutomated safety check: PassMIT
LangGraph Decision Modelslangchain-ai/langchain-skills1.3k—~2.3kAutomated safety check: PassMIT

Similar skills

  • Add Example Agent

    GetBindu/Bindu

    Add a new self-contained example agent under examples/. An agent skill from GetBindu/Bindu.

    10k GitHub stars~1.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Omnigent Framework Detection

    omnigent-ai/omnigent

    Scans Python agent code for framework imports and recommends the matching Omnigent executor type, or says when the framework is not natively supported yet.

    11k GitHub stars~610 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Migrates Python LangChain Deep Agents applications to Pydantic AI and Pydantic AI Harness while preserving the application's observed behavior.

    21k GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • LangGraph Decision Models

    langchain-ai/langchain-skills

    Official

    Routes LangGraph agents with typed decision models that return probabilities, and finds LLM calls that only exist to produce a routing decision.

    1.3k GitHub stars~2.3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Migrate Python LangChain, LangGraph, or Deep Agents applications to Pydantic AI and, when the source uses harness features, Pydantic AI Harness.

    21k GitHub stars~2.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Langchain Performance Tuning

What does Langchain Performance Tuning do?

Tune LangChain 1.0 / LangGraph 1.0 Python chains and agents for throughput, latency, and cost — streaming modes, explicit batch concurrency, semantic plus exact caches, persistent message history…. Langchain Performance Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace.0 Python chains and agents for throughput, latency, and cost — streaming modes, explicit batch concurrency, semantic plus exact caches, persistent message history, and async-safe retriever patterns.

When should I use Langchain Performance Tuning?

Langchain Performance Tuning fits situations like: P95 latency exceeds target; batching does not work; cost grows linearly with traffic; A process restart wipes chat history.

How do I install Langchain Performance Tuning in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-performance-tuning -a claude-code`. Or copy the skill folder (skills/.curated/langchain-performance-tuning in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langchain-performance-tuning in your project. Claude Code loads it when a task matches its description.

How do I install Langchain Performance Tuning in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-performance-tuning -a codex`. Or copy the skill folder (skills/.curated/langchain-performance-tuning in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langchain-performance-tuning in your project. Codex loads it when a task matches its description.

Can I use Langchain Performance Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-performance-tuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langchain-performance-tuning, .gemini/skills/langchain-performance-tuning, .github/skills/langchain-performance-tuning and .opencode/skills/langchain-performance-tuning in your project.

What does Langchain Performance Tuning need to run?

SKILL.md names no scripts, command-line tools or credentials: Langchain Performance Tuning is instructions for the agent only. Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(python:*), Bash(redis-cli:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Langchain Performance Tuning access the network?

SKILL.md names 2 domains. As links in the text: python.langchain.com and langchain-ai.github.io. This is read from the text; nothing was executed.

Is Langchain Performance Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Langchain Performance Tuning use?

Langchain Performance Tuning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Langchain Performance Tuning use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.4k tokens, read only when the agent opens those files.

What are the alternatives to Langchain Performance Tuning?

Skills that share tags, products or a category with Langchain Performance Tuning: Add Example Agent (GetBindu/Bindu, 10k stars), Failproof AI SDK Integration (FailproofAI/failproofai, 5.3k stars), Omnigent Framework Detection (omnigent-ai/omnigent, 11k stars) and Deep Agents to Pydantic AI Migration (pydantic/pydantic-ai, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Langchain Performance Tuning?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.