Agent skill

Langchain Cost Tuning

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Control LangChain 1.0 AI spend with accurate streaming token accounting, model tiering, provider-specific cache hit tuning, per-tenant budgets, and retry dedup.

MITAuto-check passedAI & LLM Engineering

Install Langchain Cost Tuning

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-cost-tuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-cost-tuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/langchain-cost-tuning .claude/skills/langchain-cost-tuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langchain-cost-tuning
GitHub stars
2.8k
Token cost
~4.9k tokens
SKILL.md length
1,774 words
Files
6 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Control LangChain 1.0 AI spend with accurate streaming token accounting, model tiering, provider-specific cache hit tuning, per-tenant budgets, and retry dedup.

  • Works in 9 steps: Read usage_metadata, never… → Stream-accurate aggregation via… → Dedup retries on run_id, not prompt hash → …
  • AI spend grows faster than traffic
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 3 more sections
  • Calls pip

What it does

Langchain Cost Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Control LangChain 1.0 AI spend with accurate streaming token accounting, model tiering, provider-specific cache hit tuning, per-tenant budgets, and retry dedup. Use when AI spend grows faster than traffic, a cost regression lands, or you need per-tenant budget enforcement. Trigger with "langchain cost", "langchain token accounting", "langchain per-tenant budget", "langchain model tiering", "prompt cache savings".

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/cache-economics.md`, `references/model-tiering.md` and `references/one-pager.md`). Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering Building AI agents and Accounting and bookkeeping. It works with LangChain. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • AI spend grows faster than traffic
  • A cost regression lands
  • You need per-tenant budget enforcement
  • With langchain cost

Example prompts

  • “langchain cost”
  • “langchain token accounting”
  • “langchain per-tenant budget”
  • “/langchain-cost-tuning”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(python:*), Bash(redis-cli:*)

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Read usage_metadata, never response_metadata["token_usage"]
  2. Stream-accurate aggregation via astream_events(version="v2")
  3. Dedup retries on run_id, not prompt hash
  4. Model tiering: draft cheap, finalize expensive
  5. Aggregate Anthropic cache usage per session and tenant (P04)
  6. Cache keys must include bound tools (P61)
  7. Tune semantic-cache threshold; ship a calibrated value, not the default (P62)
  8. Per-tenant budget middleware: soft warn, hard refuse
  9. Cap agent recursion (P10, P23)

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(python:*)
    • Bash(redis-cli:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • anthropic.com
    • openai.com
    • python.langchain.com
    • platform.claude.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Langchain Cost Tuning loads about 4.9k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 110 tokens; SKILL.md has 1,774 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,774 words, ~4,887 tokens.

Download SKILL.mdSave it as .claude/skills/langchain-cost-tuning/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
langchain-cost-tuning
description
Control LangChain 1.0 AI spend with accurate streaming token accounting, model tiering, provider-specific cache hit tuning, per-tenant budgets, and retry dedup. Use when AI spend grows faster than traffic, a cost regression lands, or you need per-tenant budget enforcement. Trigger with "langchain cost", "langchain token accounting", "langchain per-tenant budget", "langchain model tiering", "prompt cache savings".
allowed-tools
Read, Write, Edit, Bash(python:*), Bash(redis-cli:*)
compatibility
Designed for Claude Code
version
2.7.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, langchain, langgraph, python, langchain-1.0, cost, tokens, budget

LangChain Cost Tuning (Python)

Overview

An engineer shipped a new research agent Tuesday. By Friday the Anthropic bill had grown 6x while traffic grew 1.4x. The cost dashboard — wired to on_llm_end — showed spend up maybe 2x. Reconciling against the provider console on Monday surfaced two compounding bugs: (1) the agent's ChatOpenAI fallback kept the default max_retries=6, so each logical call billed as up to 7 requests (P30); (2) retry middleware was registered below token accounting, so every retry fired on_llm_end twice — the aggregator summed both emissions while LangSmith deduped them by generation ID, undercounting the dashboard by ~50% against actual billed rate (P25).

The fix took an afternoon: cap retries at 2, tag retries with a stable request_id, and migrate token accounting to AIMessage.usage_metadata read from astream_events(version="v2"). Finding the bug took a week. This skill is that week compressed into a runbook.

Cost tuning for a LangChain 1.0 production app has five levers, each with a sharp failure mode:

  • Token accounting — on_llm_end lags streams by 5-30s (P01); retries double-count (P25); Anthropic cache savings aggregate per-call, never per-session (P04).
  • Retry discipline — max_retries=6 default on ChatOpenAI (P30); Anthropic 50 RPM tier throttles cached and uncached calls against the same budget (P31).
  • Agent loop caps — create_react_agent defaults to recursion_limit=25; vague prompts burn a session's budget before GraphRecursionError surfaces (P10).
  • Caching — InMemoryCache ignores bound tools in the cache key and returns wrong answers (P61); RedisSemanticCache ships with a 0.95 threshold that hits <5% of the time (P62).
  • Model tiering — Running claude-opus-4-5 on intent classification is 30-60x more expensive than claude-haiku-4-5 for a task the cheaper model solves at equal quality.

Pin: langchain-core 1.0.x, langchain-anthropic 1.0.x, langchain-openai 1.0.x. Pain-catalog anchors: P01, P04, P10, P23, P25, P30, P31, P61, P62.

Prerequisites

  • Python 3.10+
  • langchain-core >= 1.0, < 2.0
  • At least one provider package: pip install langchain-anthropic langchain-openai
  • redis-py >= 5.0 for budget middleware (optional; in-process dict works for dev)
  • Provider console access (Anthropic, OpenAI) to reconcile usage_metadata against billed spend — you will need this to verify any instrumentation fix

Instructions

Step 1 — Read usage_metadata, never response_metadata["token_usage"]

LangChain 1.0 standardizes all provider usage into AIMessage.usage_metadata. response_metadata["token_usage"] still exists as a compatibility shim but its shape is provider-specific (Anthropic nests under usage, OpenAI flat, Gemini uses different keys). Code that reads it directly will break when you switch providers or when a provider SDK upgrades.

python
from langchain_core.messages import AIMessage

def read_usage(msg: AIMessage) -> dict:
    """Canonical shape: input_tokens, output_tokens, input_token_details,
    output_token_details. Safe across Anthropic, OpenAI, Gemini."""
    meta = msg.usage_metadata or {}
    details_in = meta.get("input_token_details", {}) or {}
    details_out = meta.get("output_token_details", {}) or {}
    return {
        "input": meta.get("input_tokens", 0),
        "output": meta.get("output_tokens", 0),
        "cache_read": details_in.get("cache_read", 0),       # Anthropic
        "cache_creation": details_in.get("cache_creation", 0),
        "reasoning": details_out.get("reasoning", 0),        # OpenAI o1/o3
    }

Include reasoning in your output-billable total for o1/o3. A call with output_tokens=500 and reasoning=2000 actually bills 2500 output tokens.

Step 2 — Stream-accurate aggregation via astream_events(version="v2")

on_llm_end fires once after the stream closes, so dashboards lag by stream duration (P01). Anthropic populates usage_metadata on the message_start and message_delta events; OpenAI populates only the final chunk. Both show up as on_chat_model_stream events in astream_events.

python
async def metered_invoke(chain, inputs, meter):
    async for event in chain.astream_events(inputs, version="v2"):
        if event["event"] == "on_chat_model_stream":
            chunk = event["data"]["chunk"]
            if getattr(chunk, "usage_metadata", None):
                meter.record(
                    run_id=event["run_id"],
                    usage=chunk.usage_metadata,
                )

See Token Accounting Pitfalls for the full streaming-delta behavior across providers and reconciliation against provider dashboards.

Step 3 — Dedup retries on run_id, not prompt hash

Retry middleware runs the model twice on transient errors. Both emit usage events. If the aggregator keys on prompt hash, it looks like one call cost twice as much. If it keys on run_id (LangChain assigns one per generation attempt), you can attach a stable request_id at the chain level and dedupe on that (P25).

python
from uuid import uuid4

class RetryAwareMeter:
    def __init__(self):
        self._seen: set[str] = set()
        self.totals = {"input": 0, "output": 0, "cache_read": 0}

    def record(self, run_id: str, usage: dict, request_id: str | None = None):
        # Keep only the last emission per logical request.
        # On retry: same request_id, different run_id -> overwrite.
        key = request_id or run_id
        if key in self._seen:
            # Retry emission — subtract prior, add new (last wins).
            prior = self._prior_by_key.get(key, {})
            for k in self.totals:
                self.totals[k] -= prior.get(k, 0)
        self._seen.add(key)
        self._prior_by_key[key] = usage
        self.totals["input"]  += usage.get("input_tokens", 0)
        self.totals["output"] += usage.get("output_tokens", 0)
        details = usage.get("input_token_details", {}) or {}
        self.totals["cache_read"] += details.get("cache_read", 0)

Inject request_id via config={"metadata": {"request_id": str(uuid4())}} on each invoke. The meter reads event["metadata"]["request_id"] alongside run_id.

Alternative: place token accounting above retry middleware in the chain — retries happen inside, so only the successful attempt emits. This is simpler but makes retries invisible to observability, which you usually want to see.

Step 4 — Model tiering: draft cheap, finalize expensive

Most chains have a structural split: a cheap "understand the request" call and an expensive "produce the final artifact" call. Running the expensive model on both roughly triples cost for no quality gain.

Per-1M pricing snapshot, 2026-04 (verify current prices before shipping at https://www.anthropic.com/pricing and https://openai.com/api/pricing/):

ModelInput $/1MOutput $/1MCache read $/1MRole
claude-haiku-4-5$1.00$5.00$0.10Draft, classify, route
claude-sonnet-4-6$3.00$15.00$0.30Finalize, reason, extract
claude-opus-4-5$15.00$75.00$1.50High-stakes, long-horizon
gpt-4o-mini$0.15$0.60n/a (prefix cache only)Draft, classify
gpt-4o$2.50$10.00n/aFinalize
gpt-o3-mini$1.10$4.40n/aReasoning, planning

Anthropic cache reads cost 10% of input. Cache creation costs 125% of input. Break-even is ~4 uses of a cached prefix. See Cache Economics.

Decision tree:

input
  └── intent classification / routing
      └── gpt-4o-mini OR claude-haiku-4-5        (~$0.15-$1 per 1M in)
  └── generation / reasoning
      ├── single-pass, low-stakes
      │   └── gpt-4o-mini                         (draft)
      ├── single-pass, high-stakes (extraction, contracts)
      │   └── claude-sonnet-4-6                   (finalize)
      ├── multi-step reasoning
      │   └── gpt-o3-mini OR claude-sonnet-4-6    (plan)
      └── mission-critical long-horizon
          └── claude-opus-4-5                     (expensive, used sparingly)

Tiering is wrong when quality degrades silently — high-stakes extraction on Haiku misses entities the Sonnet would catch. Always evaluate both tiers on a gold set before committing. See Model Tiering for the evaluation harness and a worked draft-then-finalize chain.

Step 5 — Aggregate Anthropic cache usage per session and tenant (P04)

usage_metadata["input_token_details"]["cache_read"] reports per-call. To see whether caching is paying for itself you need to aggregate per-session or per-tenant and compare against cache-creation cost.

python
class CacheLedger:
    def __init__(self, tenant_id: str):
        self.tenant_id = tenant_id
        self.read = 0           # billed at 0.1x input rate
        self.creation = 0       # billed at 1.25x input rate
        self.uncached_input = 0 # billed at 1.0x input rate

    def ingest(self, usage: dict):
        details = usage.get("input_token_details", {}) or {}
        self.read     += details.get("cache_read", 0)
        self.creation += details.get("cache_creation", 0)
        total_input = usage.get("input_tokens", 0)
        self.uncached_input += total_input - self.read - self.creation

    def savings_vs_no_cache(self, price_per_1m_input: float) -> float:
        # What we paid with cache vs. paying full price on all input.
        actual = (self.uncached_input * 1.00
                  + self.creation      * 1.25
                  + self.read          * 0.10) * price_per_1m_input / 1_000_000
        naive  = (self.uncached_input + self.creation + self.read) * price_per_1m_input / 1_000_000
        return naive - actual

Persist CacheLedger to Redis or Postgres keyed by (tenant_id, day). If savings is negative over a 24h window, caching is costing more than it saves — your cached prefix is either too short or hit too rarely. See Cache Economics.

Step 6 — Cache keys must include bound tools (P61)

set_llm_cache(InMemoryCache()) hashes the prompt string only. A chain that binds different tool sets will return wrong answers from the cache. This is the most dangerous cache failure mode — it silently returns semantically incorrect responses rather than missing.

Do not use InMemoryCache on any chain that calls bind_tools(). Use SQLiteCache or RedisSemanticCache with a composite key:

python
import hashlib, json

def tool_aware_key(prompt: str, tools: list) -> str:
    tools_fingerprint = hashlib.sha256(
        json.dumps([t.args_schema.model_json_schema() for t in tools],
                   sort_keys=True).encode()
    ).hexdigest()[:16]
    return f"{tools_fingerprint}:{hashlib.sha256(prompt.encode()).hexdigest()}"

Cross-reference: langchain-middleware-patterns covers cache-key layering in middleware order (redact → cache → model) to avoid cross-tenant PII leaks.

Step 7 — Tune semantic-cache threshold; ship a calibrated value, not the default (P62)

RedisSemanticCache defaults to score_threshold=0.95. On real workloads this hits under 5% of the time. Production-tuned values land between 0.85 and 0.90. Ship with a calibrated threshold, not the default:

  1. Collect 200 real query pairs labeled "should return same answer" (positive) and "should return different answer" (negative).
  2. Embed both sides of each pair with your embedding model.
  3. For thresholds in [0.80, 0.82, 0.84, …, 0.95], compute:
    • hit rate (% positive pairs above threshold)
    • false positive rate (% negative pairs above threshold)
  4. Pick the lowest threshold where FPR < 2%.
  5. Ship with a daily audit: sample 1% of cache hits, log to review queue.

If the curve flattens above 0.92 on positives, your embeddings are too weak for semantic caching — consider exact-match SQLiteCache instead. See Cache Economics for the calibration worksheet.

Show full SKILL.md (735 more words)Show less
Step 8 — Per-tenant budget middleware: soft warn, hard refuse

A single runaway tenant will consume the pack if left uncapped. LangChain 1.0 middleware slots cleanly in front of the model call; back it with a Redis counter keyed per tenant per day.

python
# Sketch — see references/per-tenant-budgets.md for the full middleware class.
async def budget_check(tenant_id: str, estimated_tokens: int) -> str:
    day_key = f"budget:{tenant_id}:{date.today().isoformat()}"
    used = int(await redis.get(day_key) or 0)
    soft = TENANT_SOFT_CAPS[tenant_id]  # alert only
    hard = TENANT_HARD_CAPS[tenant_id]  # refuse
    projected = used + estimated_tokens
    if projected > hard:
        raise BudgetExceeded(tenant_id, used, hard)
    if projected > soft:
        emit_alert(tenant_id, used, soft)
    return day_key  # caller increments on completion

Grace period: on hard-cap hit, allow in-flight calls (they already billed) but reject new ones. Reset counter on UTC day boundary. See Per-Tenant Budgets for full middleware class, Redis schema, alert wiring, and grace-period semantics.

Step 9 — Cap agent recursion (P10, P23)

create_react_agent defaults to recursion_limit=25. Agents on vague prompts loop until limit, then raise GraphRecursionError — but every loop billed. Cap at 5-10 for interactive, 10-15 for batch. If you use trim_messages on the loop to control context, pass include_system=True and start_on="human" so the trimmer does not drop the system prompt under pressure (P23).

Pair this with Step 8's budget middleware — the budget is the hard stop even if the recursion cap is generous. Cross-reference: langchain-langgraph-agents for routing patterns that terminate early on repeated tool calls.

Output

  • Canonical token read via usage_metadata, including reasoning and cache fields
  • Streaming-accurate aggregation via astream_events(version="v2") on on_chat_model_stream
  • Retry-aware meter that dedupes on request_id
  • Model-tier decision tree with 2026-04 per-1M price snapshot
  • CacheLedger that reports Anthropic cache savings per session/tenant
  • Tool-aware cache keys for tool-binding chains
  • Semantic-cache threshold calibrated (0.85-0.90) against a gold set
  • Per-tenant budget middleware with soft and hard caps, Redis-backed
  • recursion_limit set to match the workload, not the default 25

Error Handling

Error / symptomCauseFix
Dashboard lags provider console by minutes on streaming callson_llm_end fires at stream close (P01)Migrate to astream_events(version="v2") + on_chat_model_stream (Step 2)
Dashboard shows ~50% of billed spendRetry middleware double-emits on_llm_end, aggregator sums both (P25)Dedupe on request_id (Step 3), or move accounting above retry middleware
Cache hits return wrong answer for tool-binding chainInMemoryCache hashes prompt only, ignores tools (P61)Switch to SQLiteCache / RedisSemanticCache with tool-aware key (Step 6)
RedisSemanticCache hit rate < 5% on similar queriesDefault score_threshold=0.95 too strict (P62)Calibrate to 0.85-0.90 against a gold pair set (Step 7)
Cost spike then GraphRecursionError: Recursion limit of 25 reachedAgent loops on vague prompt (P10)Set recursion_limit=5-10; add budget middleware (Step 8)
Agent loses persona mid-conversation after many turnstrim_messages dropped system prompt (P23)Pass include_system=True, start_on="human"
max_retries=6 billing 7 requests per logical callChatOpenAI default (P30)Set max_retries=2; log every retry via callback to verify
429 on cache reads while input-token budget shows headroomAnthropic 50 RPM throttles cached and uncached together (P31)Budget RPM at client level (semaphore); separate monitors for read vs uncached
"Cache savings" metric always zeroinput_token_details.cache_read reset per call (P04)Aggregate via CacheLedger keyed per session/tenant (Step 5)

Examples

Reconciling a 6x cost spike

A team saw Anthropic spend 6x over a week with traffic up 1.4x. The dashboard showed only 2x. Root cause: max_retries=6 on a flaky downstream API (P30) plus retry middleware double-emit (P25) undercounting on the dashboard.

Fix sequence: (1) set max_retries=2, (2) attach request_id metadata at chain entry, (3) migrate meter to astream_events with run_id dedup. After the fix, dashboard matched billed spend within 1%.

See Token Accounting Pitfalls for the full reconciliation procedure against provider console CSVs.

Draft-then-finalize chain, measured savings

A document-extraction chain runs claude-haiku-4-5 to extract a rough outline, then claude-sonnet-4-6 to validate and fill missing fields. On a 10K-doc batch:

  • Sonnet-only: ~$14/1K docs
  • Haiku draft + Sonnet finalize: ~$4.20/1K docs
  • Quality on the gold set: equivalent (F1 within 0.01)

The draft step burned 80% of the input tokens on the cheaper model.

See Model Tiering for the full chain, the gold set, and the evaluation harness.

Per-tenant runaway

One tenant's prompt template had an accidental double-interpolation of the message history that grew context unboundedly each turn. Spend 400x'd overnight. The per-tenant budget middleware (Step 8) hit hard cap at 10x normal, alerted on soft cap at 5x, refused new requests, allowed in-flight calls to complete.

See Per-Tenant Budgets for the full middleware, alert wiring, and grace-period semantics.

Resources

  • Pair skill: langchain-performance-tuning (latency, throughput, cache hit rate) — this skill focuses on spend, not speed; reference each other on cache tuning
  • Related: langchain-middleware-patterns (cache-key order, retry telemetry), langchain-langgraph-agents (recursion caps, early termination), langchain-rate-limits (RPM/ITPM budgeting, companion to Step 4)
  • LangChain Python: usage_metadata
  • LangChain Python: astream_events v2
  • Anthropic pricing — verify current rates before shipping
  • OpenAI pricing — verify current rates before shipping
  • Anthropic prompt caching
  • Pack pain catalog: docs/pain-catalog.md (P01, P04, P10, P23, P25, P30, P31, P61, P62)

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/.curated/langchain-cost-tuning of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/cache-economics.md
  • references/model-tiering.md
  • references/one-pager.md
  • references/per-tenant-budgets.md
  • references/token-accounting-pitfalls.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Langchain Cost Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Langchain Cost Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Langchain Cost Tuning this skilljeremylongshore/tons-of-skills-marketplace2.8k—~4.9kAutomated safety check: PassMIT
Swapper Depositswapperfinance/swapper-toolkit852—~1.8kAutomated safety check: PassMIT
LLM Developmentmeleantonio/ChernyCode516—~499Automated safety check: PassNone
Tool CreatorAgentTeam-TaichuAI/ScienceClaw671—~4.7kAutomated safety check: PassNone
Langchain ArchitectureHermeticOrmus/LibreUIUX-Claude-Code11210 repos~2.5kAutomated safety check: PassMIT
AI EngineerDokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI508—~1.1kAutomated safety check: PassCustom licence

Similar skills

  • Swapper Deposit

    swapperfinance/swapper-toolkit

    Deposit and bridge funds into a wallet or protocol using Swapper Finance.

    852 GitHub stars~1.8k tokensUpdated 6 mo ago
    Business, Finance & HRAuto-check passed
  • LLM Development

    meleantonio/ChernyCode

    LLM and ML development best practices with LangChain and transformers.

    516 GitHub stars~499 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Tool Creator

    AgentTeam-TaichuAI/ScienceClaw

    Create new tools or upgrade existing tools for the agent. An agent skill from AgentTeam-TaichuAI/ScienceClaw.

    671 GitHub stars~4.7k tokensUpdated 5 mo ago
    AI & LLM EngineeringAuto-check passed
  • Langchain Architecture

    HermeticOrmus/LibreUIUX-Claude-Code

    Design LLM applications using the LangChain framework with agents, memory, and tool integration patterns.

    112 GitHub starsUsed in 10 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • AI Engineer

    Dokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI

    Principal AI Architect and Machine Learning Engineer. An agent skill from Dokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI.

    508 GitHub stars~1.1k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Solana Agent Kit

    internet-court/internet-court-skill

    Walks through building AI agents that run Solana operations such as token deploys, NFT minting, swaps and staking with SendAI's toolkit, in chat or fully autonomous mode.

    6.6k GitHub starsUsed in 2 repos~3.8k tokens
    AI & LLM EngineeringAuto-check: notes

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Langchain Cost Tuning

What does Langchain Cost Tuning do?

Control LangChain 1.0 AI spend with accurate streaming token accounting, model tiering, provider-specific cache hit tuning, per-tenant budgets, and retry dedup. Langchain Cost Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace.0 AI spend with accurate streaming token accounting, model tiering, provider-specific cache hit tuning, per-tenant budgets, and retry dedup.

When should I use Langchain Cost Tuning?

Langchain Cost Tuning fits situations like: AI spend grows faster than traffic; A cost regression lands; you need per-tenant budget enforcement; with langchain cost.

How do I install Langchain Cost Tuning in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-cost-tuning -a claude-code`. Or copy the skill folder (skills/.curated/langchain-cost-tuning in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langchain-cost-tuning in your project. Claude Code loads it when a task matches its description.

How do I install Langchain Cost Tuning in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-cost-tuning -a codex`. Or copy the skill folder (skills/.curated/langchain-cost-tuning in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langchain-cost-tuning in your project. Codex loads it when a task matches its description.

Can I use Langchain Cost Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-cost-tuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langchain-cost-tuning, .gemini/skills/langchain-cost-tuning, .github/skills/langchain-cost-tuning and .opencode/skills/langchain-cost-tuning in your project.

What does Langchain Cost Tuning need to run?

Going by SKILL.md and its folder, Langchain Cost Tuning needs the command-line tools its instructions call (pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(python:*), Bash(redis-cli:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Langchain Cost Tuning access the network?

SKILL.md names 4 domains. As links in the text: anthropic.com, openai.com, python.langchain.com and platform.claude.com. This is read from the text; nothing was executed.

Is Langchain Cost Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Langchain Cost Tuning use?

Langchain Cost Tuning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Langchain Cost Tuning use?

About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.2k tokens, read only when the agent opens those files.

What are the alternatives to Langchain Cost Tuning?

Skills that share tags, products or a category with Langchain Cost Tuning: Swapper Deposit (swapperfinance/swapper-toolkit, 852 stars), LLM Development (meleantonio/ChernyCode, 516 stars), Tool Creator (AgentTeam-TaichuAI/ScienceClaw, 671 stars) and Langchain Architecture (HermeticOrmus/LibreUIUX-Claude-Code, 112 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Langchain Cost Tuning?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.