Agent skill

Langchain Rate Limits

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Rate-limit LangChain 1.0 calls correctly across multi-worker deployments — Redis-backed limiters, asyncio.Semaphore, narrow exception whitelists, and provider-specific throttle handling.

MITAuto-check passedAI & LLM Engineering

Install Langchain Rate Limits

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-rate-limits -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-rate-limits --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/langchain-rate-limits .claude/skills/langchain-rate-limits && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langchain-rate-limits
GitHub stars
2.8k
Token cost
~4.4k tokens
SKILL.md length
1,471 words
Files
6 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Rate-limit LangChain 1.0 calls correctly across multi-worker deployments — Redis-backed limiters, asyncio.Semaphore, narrow exception whitelists, and provider-specific throttle handling.

  • Works in 9 steps: Measure actual demand before picking a… → InMemoryRateLimiter for single-process… → Redis-backed limiter for cluster-wide… → …
  • Hitting 429s in production
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 3 more sections
  • Calls pip, python and uvicorn

What it does

Langchain Rate Limits is an agent skill from jeremylongshore/tons-of-skills-marketplace. Rate-limit LangChain 1.0 calls correctly across multi-worker deployments — Redis-backed limiters, asyncio.Semaphore, narrow exception whitelists, and provider-specific throttle handling. Use when hitting 429s in production, scaling workers horizontally, or tuning throughput against Anthropic, OpenAI, or Gemini tier limits. Trigger with "langchain rate limit", "langchain 429", "langchain semaphore", "langchain token bucket", "anthropic rpm", "openai rpm throttling", "InMemoryRateLimiter", "redis rate limiter".

Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/backoff-and-retry.md`, `references/measuring-demand.md` and `references/one-pager.md`). Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering Rate limiting and Building AI agents. It works with LangChain, Redis, OpenAI and Python. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Hitting 429s in production
  • Scaling workers horizontally
  • Tuning throughput against Anthropic
  • Gemini tier limits

Example prompts

  • “langchain rate limit”
  • “langchain 429”
  • “langchain semaphore”
  • “/langchain-rate-limits”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(python:*), Bash(redis-cli:*)

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Measure actual demand before picking a number
  2. InMemoryRateLimiter for single-process dev only; never multi-worker prod
  3. Redis-backed limiter for cluster-wide enforcement
  4. asyncio.Semaphore for per-worker in-flight concurrency cap
  5. Narrow with_fallbacks(exceptions_to_handle=...) — never (Exception,)
  6. max_retries=2, never the default max_retries=6
  7. Understand the provider limit taxonomy
  8. Decision tree: which limiter to use
  9. Provider tier snapshot (verify before shipping)

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(python:*)
    • Bash(redis-cli:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • python
    • uvicorn

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.claude.com
    • platform.openai.com
    • ai.google.dev
    • python.langchain.com
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Langchain Rate Limits loads about 4.4k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 134 tokens; SKILL.md has 1,471 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~134
When it runs · the whole SKILL.md, loaded when a task matches
~4.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,471 words, ~4,365 tokens.

Download SKILL.mdSave it as .claude/skills/langchain-rate-limits/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
langchain-rate-limits
description
Rate-limit LangChain 1.0 calls correctly across multi-worker deployments — Redis-backed limiters, asyncio.Semaphore, narrow exception whitelists, and provider-specific throttle handling. Use when hitting 429s in production, scaling workers horizontally, or tuning throughput against Anthropic, OpenAI, or Gemini tier limits. Trigger with "langchain rate limit", "langchain 429", "langchain semaphore", "langchain token bucket", "anthropic rpm", "openai rpm throttling", "InMemoryRateLimiter", "redis rate limiter".
allowed-tools
Read, Write, Edit, Bash(python:*), Bash(redis-cli:*)
compatibility
Designed for Claude Code
version
2.7.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, langchain, langgraph, python, langchain-1.0, rate-limits, throttling, concurrency

LangChain Rate Limits (Python)

Overview

A team deploys 10 Cloud Run workers. Each worker initializes its ChatAnthropic with InMemoryRateLimiter(requests_per_second=10) — they read the docs, they picked a safe-looking number, they shipped. Thirty seconds later the dashboard lights up with 429s: the cluster is pushing 100 RPS to Anthropic's 50 RPM tier-1 ceiling, not the 10 RPS they configured. The name is the fix — InMemoryRateLimiter is in-process. Each worker has its own counter. Ten workers × 10 RPS = 100 RPS to the provider. This is pain-catalog entry P29 and it lands on every team that scales past one pod.

Three more traps wait on the same code path:

  • P07 — .with_fallbacks([backup]) defaults exceptions_to_handle=(Exception,), which on Python <3.12 swallows KeyboardInterrupt. Ctrl+C during a 429 retry storm silently falls through to the backup chain and keeps billing.
  • P30 — ChatOpenAI and ChatAnthropic default max_retries=6. That is retries, not attempts: 7 total requests per logical call on flaky networks. One .invoke() can bill 7x.
  • P31 — Anthropic's RPM counts cache reads, cache writes, and uncached calls uniformly. Cache-heavy workloads at 50 RPM can 429 on cache writes while the ITPM dashboard shows headroom.

This skill covers measuring demand before picking a limit; the InMemoryRateLimiter vs Redis-backed limiter vs asyncio.Semaphore decision tree; the narrow exceptions_to_handle whitelist; max_retries=2 math; and the provider-specific limit taxonomy (RPM, ITPM, OTPM, concurrent, cached-vs-uncached). Pin: langchain-core 1.0.x, langchain-anthropic 1.0.x, langchain-openai 1.0.x. Pain-catalog anchors: P07, P08, P29, P30, P31. For .batch(max_concurrency=...) tuning, see the sibling skill langchain-performance-tuning — this skill is about provider-facing rate caps.

Prerequisites

  • Python 3.10+ (3.12+ fixes the KeyboardInterrupt half of P07)
  • langchain-core >= 1.0, < 2.0
  • At least one provider: pip install langchain-anthropic langchain-openai
  • For multi-worker prod: redis >= 4.5 client and a Redis server reachable from every worker
  • Completed langchain-model-inference — the chat-model factory from that skill is where rate_limiter= gets attached

Instructions

Step 1 — Measure actual demand before picking a number

Do not guess at requests_per_second. Instrument first, size second. Attach a BaseCallbackHandler that logs per-call input_tokens, output_tokens, and cache_read_input_tokens from response.generations[].message.usage_metadata:

python
chain.with_config({"callbacks": [DemandLogger()]})

Collect 24-48 hours of representative traffic. Roll up: p50 and p95 RPM, p95 ITPM, p95 OTPM, cache hit rate. Size the limiter at 70% of the binding constraint's tier ceiling on your p95.

See Measuring Demand for the full DemandLogger implementation, pandas roll-up, OTEL integration, load-test harness, and multi-tenant sizing strategies.

Step 2 — InMemoryRateLimiter for single-process dev only; never multi-worker prod

LangChain 1.0 ships InMemoryRateLimiter as a first-class BaseChatModel parameter:

python
from langchain_anthropic import ChatAnthropic
from langchain_core.rate_limiters import InMemoryRateLimiter

limiter = InMemoryRateLimiter(
    requests_per_second=0.58,    # 35 RPM = 70% of Anthropic tier-1 50 RPM
    check_every_n_seconds=0.1,
    max_bucket_size=5,           # burst capacity
)

llm = ChatAnthropic(
    model="claude-sonnet-4-6",
    rate_limiter=limiter,
    max_retries=2,
    timeout=30,
)

InMemoryRateLimiter is per-process. Safe for:

  • Single-process local dev (python script.py)
  • Single-worker uvicorn (uvicorn --workers 1)
  • Jupyter notebooks, batch scripts

Unsafe for (this is P29):

  • Multi-worker uvicorn / gunicorn (--workers 4)
  • Any container orchestrator with replica count > 1 (Cloud Run min-instances > 1, K8s, ECS)
  • Distributed job runners (Celery, Temporal, Cloud Tasks fanout)
Step 3 — Redis-backed limiter for cluster-wide enforcement

For multi-worker deployments, cluster-wide rate limiting requires shared state. Redis is the default answer — atomic Lua script for sliding-window, or Redis 6.2+ CL.THROTTLE for GCRA.

python
import redis
from langchain_anthropic import ChatAnthropic
# RedisRateLimiter class defined in references/redis-limiter-pattern.md
from your_app.limiters import RedisRateLimiter

client = redis.Redis.from_url("redis://redis.internal:6379/0")

limiter = RedisRateLimiter(
    client,
    key="anthropic:prod",
    requests_per_second=35 / 60,  # 35 RPM cluster-wide, not per-worker
)

llm = ChatAnthropic(
    model="claude-sonnet-4-6",
    rate_limiter=limiter,
    max_retries=2,
    timeout=30,
)

Key scoping decisions:

  • key="anthropic:prod" — all tenants share one global budget (simplest)
  • key=f"anthropic:tenant:{tenant_id}" — per-tenant quota (requires cleanup for dead tenants)
  • Two-level: per-tenant + global, acquire both (best for multi-tenant SaaS)

See Redis Limiter Pattern for the full RedisRateLimiter implementation (atomic Lua sliding window), the GCRA alternative via CL.THROTTLE, failure modes (Redis down, clock skew), and per-tenant cleanup strategy.

Step 4 — asyncio.Semaphore for per-worker in-flight concurrency cap

The rate limiter throttles request rate. A semaphore throttles in-flight count. Use both:

python
import asyncio

# Cluster: 35 RPM (Redis enforces)
# Worker: 20 in-flight at once (semaphore enforces)
worker_sem = asyncio.Semaphore(20)

async def bounded_invoke(inp):
    async with worker_sem:
        return await llm.ainvoke(inp)

# Fanout
results = await asyncio.gather(*[bounded_invoke(x) for x in inputs])

Why both: a semaphore prevents a single worker from queueing hundreds of pending limiter acquires against Redis (head-of-line blocking on the event loop). The limiter prevents the cluster from exceeding the provider tier. They solve different problems.

Semaphore sizing: target latency-bandwidth-product. If p95 request latency is 2s and the worker's RPS cap is 10, in-flight count ≈ 2 × 10 = 20. Overshoot is wasted memory; undershoot leaves throughput on the table.

Step 5 — Narrow with_fallbacks(exceptions_to_handle=...) — never (Exception,)

.with_fallbacks([backup]) defaults to catching Exception. This is P07 — on Python <3.12, Exception edge-cases include KeyboardInterrupt propagation. Ctrl+C during a retry storm silently hands off to the backup and keeps running. Always narrow the tuple:

python
from anthropic import (
    RateLimitError, APITimeoutError, APIConnectionError, InternalServerError,
)

resilient = (prompt | claude | parser).with_fallbacks(
    [prompt | gpt4o | parser],
    exceptions_to_handle=(
        RateLimitError, APITimeoutError,
        APIConnectionError, InternalServerError,
    ),
    # NEVER: Exception, BaseException, AuthenticationError,
    # BadRequestError, ValidationError
)

The whitelist is only transient provider errors. AuthenticationError, BadRequestError, and ValidationError are bugs in your code/credentials — fallback produces the same crash. See the sibling skill's reference langchain-sdk-patterns/references/fallback-exception-list.md for the full per-provider whitelist (Anthropic, OpenAI, Gemini).

Step 6 — max_retries=2, never the default max_retries=6

max_retries is retries, not attempts. Default max_retries=6 on ChatOpenAI / ChatAnthropic means initial + 6 retries = 7 billed requests per logical call (P30). On a flaky network, one .invoke() costs 7x what you budgeted.

python
# BAD — default
llm = ChatOpenAI(model="gpt-4o")  # max_retries=6

# GOOD — production default
llm = ChatOpenAI(
    model="gpt-4o",
    max_retries=2,      # initial + 2 retries = 3 total billed requests max
    timeout=30,
    rate_limiter=redis_limiter,
)

Trade resilience off to the fallback layer — with_fallbacks is strictly cheaper than retry amplification when the primary is genuinely unhealthy. Instrument retry count via callback and alert if retry rate exceeds ~5%.

See Backoff and Retry for the full math, Retry-After header handling, and circuit-breaker pattern for sustained overload.

Step 7 — Understand the provider limit taxonomy

Different providers expose different limit types. Know which one binds your workload before you size:

LimitMeaningWho enforcesBinds for
RPMRequests/minute (counts every call)All three providersShort chat replies
ITPMInput tokens/minuteAnthropic, OpenAI (as TPM combined)Long document Q&A
OTPMOutput tokens/minuteAnthropic separately; OpenAI as combined TPMLong completions
ConcurrentIn-flight request capMainly OpenAI higher tiersBurst traffic
Cached readsCache-read input tokens (Anthropic)Anthropic separate budget lineCache-heavy workloads (but still counts toward RPM — P31)

Critical for Anthropic cache workloads (P31): RPM counts uniformly across cached reads, cache writes, and uncached calls. A workload at 90% cache hit rate still trips the 50 RPM ceiling at 51 requests/min. Separate monitors for cache_read_input_tokens vs input_tokens (minus cache read/write) give early warning.

Show full SKILL.md (550 more words)Show less
Step 8 — Decision tree: which limiter to use
┌─ Single process (dev, notebooks, sync CLI, --workers 1)?
│  └─ InMemoryRateLimiter
│
├─ Multi-process but single host (same-machine pool, local gunicorn)?
│  └─ Redis-backed limiter (even localhost Redis beats InMemoryRateLimiter —
│     which still has per-process counters)
│
├─ Multi-host cluster (Cloud Run --min-instances>1, K8s, ECS)?
│  └─ Redis-backed limiter (mandatory)
│
├─ Multi-region or cross-cloud?
│  └─ Regional Redis per zone + provider-side account quota
│     (cross-region Redis latency adds 30-200ms per acquire)
│
└─ Any of the above + multi-tenant SaaS?
   └─ Two-level Redis limiter: per-tenant + global, acquire both

Always pair with asyncio.Semaphore(N) per-worker for in-flight concurrency.

Step 9 — Provider tier snapshot (verify before shipping)

2026-04-21 snapshot — re-verify against the official console before shipping.

ProviderFree tier RPMTier-1 RPMHigh tier RPMSource
Anthropic550 (Build 1)4000 (Build 4)https://platform.claude.com/docs/en/api/rate-limits
OpenAI350010000 (Tier 5)https://platform.openai.com/docs/guides/rate-limits
Google Gemini152000 (Paid 1)30000 (Paid 3)https://ai.google.dev/gemini-api/docs/rate-limits

Tiers change quarterly. A limiter sized six months ago on a different tier is a liability. See Provider Tier Matrix for the full matrix including ITPM / OTPM / cached-read separation, binding-limit math, and the pre-ship verification checklist.

Output

  • Instrumented DemandLogger callback attached to your chains for 24-48h before sizing
  • InMemoryRateLimiter in dev / notebooks / single-worker only
  • RedisRateLimiter (sliding-window Lua or CL.THROTTLE GCRA) for any multi-worker deployment, keyed per-tenant or global
  • asyncio.Semaphore(N) per-worker in-flight cap paired with the cluster-wide limiter
  • max_retries=2 on every ChatAnthropic / ChatOpenAI / ChatGoogleGenerativeAI
  • .with_fallbacks(exceptions_to_handle=(RateLimitError, APITimeoutError, APIConnectionError, InternalServerError)) — never (Exception,)
  • Per-provider tier re-verified from the official console, sized at 70% of the binding constraint

Error Handling

ErrorCauseFix
anthropic.RateLimitError: 429 THROTTLED at cluster RPM = N × InMemoryRateLimiter ceilingInMemoryRateLimiter is per-process; N workers each send at their limit (P29)Switch to Redis-backed limiter (Step 3)
429 on cache writes while ITPM dashboard shows headroomAnthropic RPM counts cache writes uniformly (P31)Budget at RPM level with limiter; separate cached vs uncached metrics
One .invoke() bills as 7 requests on flaky networksDefault max_retries=6 (P30)max_retries=2 + fallback layer for resilience
Ctrl+C during retry storm silently falls through to backup chainexceptions_to_handle=(Exception,) catches KeyboardInterrupt on Python <3.12 (P07)Narrow tuple to (RateLimitError, APITimeoutError, APIConnectionError, InternalServerError)
Limiter queue p95 wait > 500msLimiter is oversubscribed for real trafficRe-measure demand (Step 1); upgrade provider tier OR shed load
redis.exceptions.ConnectionError blocks all LLM callsRedis unavailable and limiter is fail-closedInstrument Redis health; decide fail-open (log loudly) vs fail-closed (shed load) — for provider safety, prefer fail-closed
retry-after header climbing 2→4→8→16Pushing past tier; backoff amplifying, not absorbingLower limiter target RPS by 20%; upgrade tier if sustained
google.api_core.exceptions.ResourceExhausted on GeminiGemini free tier 15 RPM is brutalUpgrade to paid Gemini tier 1 (2000 RPM) or use Redis limiter at 10 RPM

Examples

Multi-worker Cloud Run deployment with Anthropic tier-1 50 RPM

Ten workers, single region, Redis in same VPC. Target: 35 RPM cluster-wide (70% of 50 RPM ceiling), 20 in-flight per worker.

python
import asyncio, os, redis
from langchain_anthropic import ChatAnthropic
from anthropic import (
    RateLimitError, APITimeoutError, APIConnectionError, InternalServerError,
)
from your_app.redis_limiter import RedisRateLimiter  # see references

_client = redis.Redis.from_url(os.environ["REDIS_URL"])
anthropic_limiter = RedisRateLimiter(
    _client, key="anthropic:prod",
    requests_per_second=35 / 60,    # 35 RPM cluster-wide
)

llm = ChatAnthropic(
    model="claude-sonnet-4-6",
    rate_limiter=anthropic_limiter, # cluster gate
    max_retries=2,                  # not 6 (P30)
    timeout=30,
)

chain = (prompt | llm | parser).with_fallbacks(
    [prompt | gpt4o_backup | parser],
    exceptions_to_handle=(          # narrow tuple (P07)
        RateLimitError, APITimeoutError,
        APIConnectionError, InternalServerError,
    ),
)

worker_sem = asyncio.Semaphore(20)  # per-worker in-flight cap
async def invoke_bounded(inp):
    async with worker_sem:
        return await chain.ainvoke(inp)

Cluster behavior: every worker's limiter call hits the same Redis key. At 35 RPM cluster-wide, individual workers see fair-share throughput. max_retries=2

  • narrow fallback tuple means transient 429s surface quickly and hand off to GPT-4o instead of amplifying cost.
Multi-tenant SaaS with per-tenant isolation

Two-level Redis limiter. Per-tenant limit prevents noisy neighbors; global limit protects the provider tier.

See Redis Limiter Pattern for the two-level acquire implementation (acquire tenant key first, then global key; release tenant if global fails) and the per-tenant cleanup cron.

Single-process dev — InMemoryRateLimiter is fine

For local debugging, notebook work, or a sync CLI tool:

python
from langchain_core.rate_limiters import InMemoryRateLimiter

limiter = InMemoryRateLimiter(requests_per_second=0.5, max_bucket_size=3)
llm = ChatAnthropic(model="claude-sonnet-4-6", rate_limiter=limiter, max_retries=2)

Do not carry this into production without re-reading Step 2.

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/.curated/langchain-rate-limits of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/backoff-and-retry.md
  • references/measuring-demand.md
  • references/one-pager.md
  • references/provider-tier-matrix.md
  • references/redis-limiter-pattern.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Langchain Rate Limits next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Langchain Rate Limits compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Langchain Rate Limits this skilljeremylongshore/tons-of-skills-marketplace2.8k—~4.4kAutomated safety check: PassMIT
Add Example AgentGetBindu/Bindu10k—~1.1kAutomated safety check: NotesCustom licence
Upgrade Stripekanchengw/cnllm1733 repos~1.4kAutomated safety check: PassApache-2.0
Tool Designagentailor/fullstack-langgraph-nextjs-agent132—~3.2kAutomated safety check: PassMIT
Sentry Setup AI MonitoringLiorVainer/data-israel130—~1.8kAutomated safety check: PassApache-2.0
Phoenix Integration SnippetsArize-ai/phoenix12k—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Add Example Agent

    GetBindu/Bindu

    Add a new self-contained example agent under examples/. An agent skill from GetBindu/Bindu.

    10k GitHub stars~1.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Upgrade Stripe

    kanchengw/cnllm

    Guide for upgrading Stripe API versions and SDKs. An agent skill from kanchengw/cnllm.

    173 GitHub starsUsed in 3 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Tool Design

    agentailor/fullstack-langgraph-nextjs-agent

    Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise).

    132 GitHub stars~3.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Sentry Setup AI Monitoring

    LiorVainer/data-israel

    Setup Sentry AI Agent Monitoring in any project. An agent skill from LiorVainer/data-israel.

    130 GitHub stars~1.8k tokensUpdated 5 mo ago
    AI & LLM EngineeringAuto-check passed
  • Generates onboarding code snippets for Phoenix tracing integrations and wires them into the project onboarding UI.

    12k GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Langchain RAG

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.

    1.3k GitHub stars~3.9k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Langchain Rate Limits

What does Langchain Rate Limits do?

Rate-limit LangChain 1.0 calls correctly across multi-worker deployments — Redis-backed limiters, asyncio.Semaphore, narrow exception whitelists, and provider-specific throttle handling. Langchain Rate Limits is an agent skill from jeremylongshore/tons-of-skills-marketplace.Semaphore, narrow exception whitelists, and provider-specific throttle handling.

When should I use Langchain Rate Limits?

Langchain Rate Limits fits situations like: hitting 429s in production; scaling workers horizontally; tuning throughput against Anthropic; gemini tier limits.

How do I install Langchain Rate Limits in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-rate-limits -a claude-code`. Or copy the skill folder (skills/.curated/langchain-rate-limits in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langchain-rate-limits in your project. Claude Code loads it when a task matches its description.

How do I install Langchain Rate Limits in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-rate-limits -a codex`. Or copy the skill folder (skills/.curated/langchain-rate-limits in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langchain-rate-limits in your project. Codex loads it when a task matches its description.

Can I use Langchain Rate Limits in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-rate-limits -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langchain-rate-limits, .gemini/skills/langchain-rate-limits, .github/skills/langchain-rate-limits and .opencode/skills/langchain-rate-limits in your project.

What does Langchain Rate Limits need to run?

Going by SKILL.md and its folder, Langchain Rate Limits needs the command-line tools its instructions call (pip, python and uvicorn). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(python:*), Bash(redis-cli:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Langchain Rate Limits access the network?

SKILL.md names 5 domains. As links in the text: platform.claude.com, platform.openai.com, ai.google.dev, python.langchain.com and github.com. This is read from the text; nothing was executed.

Is Langchain Rate Limits safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Langchain Rate Limits use?

Langchain Rate Limits is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Langchain Rate Limits use?

About 4.4k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.3k tokens, read only when the agent opens those files.

What are the alternatives to Langchain Rate Limits?

Skills that share tags, products or a category with Langchain Rate Limits: Add Example Agent (GetBindu/Bindu, 10k stars), Upgrade Stripe (kanchengw/cnllm, 173 stars), Tool Design (agentailor/fullstack-langgraph-nextjs-agent, 132 stars) and Sentry Setup AI Monitoring (LiorVainer/data-israel, 130 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Langchain Rate Limits?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.