Implement Anthropic Claude API rate limiting, backoff, and quota management.

MITAuto-check passedBackend & APIs

Install Anth Rate Limits

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-rate-limits -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace anth-rate-limits --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/anth-rate-limits .claude/skills/anth-rate-limits && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
anth-rate-limits
GitHub stars
2.8k
Token cost
~1.9k tokens
SKILL.md length
453 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Implement Anthropic Claude API rate limiting, backoff, and quota management.

  • Works in 5 steps: Read the response headers after each… → Honor retry-after when present, apply… → Coordinate RPM, input-token, and… → …
  • Handling 429 errors
  • SKILL.md covers Overview, Rate Limit Dimensions, Usage Tiers (Auto-Upgrade) and SDK Built-In Retry, plus 10 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Anth Rate Limits is an agent skill from jeremylongshore/tons-of-skills-marketplace. Implement Anthropic Claude API rate limiting, backoff, and quota management. Use when handling 429 errors, optimizing request throughput, or managing RPM/TPM limits across usage tiers. Trigger with phrases like "anthropic rate limit", "claude 429", "anthropic throttling", "claude retry", "anthropic backoff".

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Claude Code

It sits in Backend & APIs, covering Rate limiting. It works with Anthropic API. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Handling 429 errors
  • Optimizing request throughput
  • Managing RPM/TPM limits across usage tiers
  • With phrases like anthropic rate limit

Example prompts

  • “anthropic rate limit”
  • “claude 429”
  • “anthropic throttling”
  • “/anth-rate-limits”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Read the response headers after each permitted call and update a shared limiter using the provider's remaining/reset values. Reserve…
  2. Honor retry-after when present, apply jitter and a maximum delay, and stop after a bounded number of attempts. Never let every worker…
  3. Coordinate RPM, input-token, and output-token budgets across instances. Queue or batch offline work and apply backpressure when the shared…
  4. Canary limiter changes with synthetic traffic and compare 429 rate, latency, queue age, token totals, and side_effects=0. Roll back the…
  5. Expire queued test items and temporary counters according to the retention policy; retain a redacted receipt for the decision.

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.claude.com
    • console.anthropic.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Anth Rate Limits loads about 1.9k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 453 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 453 words, ~1,879 tokens.

Download SKILL.mdSave it as .claude/skills/anth-rate-limits/SKILL.md (or your agent's skills folder).
name
anth-rate-limits
description
Implement Anthropic Claude API rate limiting, backoff, and quota management. Use when handling 429 errors, optimizing request throughput, or managing RPM/TPM limits across usage tiers. Trigger with phrases like "anthropic rate limit", "claude 429", "anthropic throttling", "claude retry", "anthropic backoff".
allowed-tools
Read, Write, Edit
compatibility
Designed for Claude Code
version
1.7.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, ai, anthropic

Anthropic Rate Limits

Overview

The Claude API uses token-bucket rate limiting measured in three dimensions: requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM). Limits increase automatically as you move through usage tiers.

Rate Limit Dimensions

DimensionHeaderDescription
RPManthropic-ratelimit-requests-limitRequests per minute
ITPManthropic-ratelimit-tokens-limitInput tokens per minute
OTPManthropic-ratelimit-tokens-limitOutput tokens per minute

Limits are per-organization and per-model-class. Cached input tokens do NOT count toward ITPM limits.

Usage Tiers (Auto-Upgrade)

TierMonthly SpendKey Benefit
Tier 1 (Free)$0Evaluation access
Tier 2$40+Higher RPM
Tier 3$200+Production-grade limits
Tier 4$2,000+High-throughput access
ScaleCustomCustom limits via sales

Check your current tier and limits at console.anthropic.com.

SDK Built-In Retry

python
import anthropic

# The SDK retries 429 and 5xx errors automatically (2 retries by default)
client = anthropic.Anthropic(max_retries=5)  # Increase for high-traffic apps

# Disable auto-retry for manual control
client = anthropic.Anthropic(max_retries=0)
typescript
const client = new Anthropic({ maxRetries: 5 });

Custom Rate Limiter with Header Awareness

python
import time
import anthropic

class RateLimitedClient:
    def __init__(self):
        self.client = anthropic.Anthropic(max_retries=0)  # We handle retries
        self.remaining_requests = 100
        self.remaining_tokens = 100000
        self.reset_at = 0.0

    def create_message(self, **kwargs):
        # Pre-check: wait if near limit
        if self.remaining_requests < 3 and time.time() < self.reset_at:
            wait = self.reset_at - time.time()
            print(f"Pre-throttle: waiting {wait:.1f}s")
            time.sleep(wait)

        for attempt in range(5):
            try:
                response = self.client.messages.create(**kwargs)

                # Update from response headers (via _response)
                headers = response._response.headers
                self.remaining_requests = int(headers.get("anthropic-ratelimit-requests-remaining", 100))
                self.remaining_tokens = int(headers.get("anthropic-ratelimit-tokens-remaining", 100000))
                reset = headers.get("anthropic-ratelimit-requests-reset")
                if reset:
                    from datetime import datetime
                    self.reset_at = datetime.fromisoformat(reset.replace("Z", "+00:00")).timestamp()

                return response
            except anthropic.RateLimitError as e:
                retry_after = float(e.response.headers.get("retry-after", 2 ** attempt))
                print(f"429 — retry in {retry_after}s (attempt {attempt + 1})")
                time.sleep(retry_after)

        raise Exception("Exhausted rate limit retries")

Queue-Based Throughput Control

typescript
import PQueue from 'p-queue';
import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

// Enforce 50 RPM with concurrency limit
const queue = new PQueue({
  concurrency: 10,
  interval: 60_000,
  intervalCap: 50,
});

async function rateLimitedCall(prompt: string) {
  return queue.add(() =>
    client.messages.create({
      model: 'claude-sonnet-4-20250514',
      max_tokens: 1024,
      messages: [{ role: 'user', content: prompt }],
    })
  );
}

// Process 200 prompts without hitting limits
const results = await Promise.all(
  prompts.map(p => rateLimitedCall(p))
);

Cost-Saving: Use Batches for Bulk Work

python
# Message Batches API: 50% cheaper, no rate limit pressure on real-time quota
batch = client.messages.batches.create(
    requests=[
        {"custom_id": f"req-{i}", "params": {
            "model": "claude-sonnet-4-20250514",
            "max_tokens": 1024,
            "messages": [{"role": "user", "content": prompt}]
        }}
        for i, prompt in enumerate(prompts)
    ]
)

Error Handling

HeaderDescriptionAction
retry-afterSeconds until next request allowedSleep this duration exactly
anthropic-ratelimit-requests-remainingRequests left in windowThrottle if < 5
anthropic-ratelimit-tokens-remainingTokens left in windowReduce max_tokens if low
anthropic-ratelimit-requests-resetISO timestamp of window resetSchedule retry after this time

Prerequisites

  • Record the authorized organization/model limits, budget ceiling, retry cap, and shared limiter policy. Do not infer a production limit from a local load test.
  • Use synthetic prompts and a sandbox workspace for experiments, with no-op downstream effects and aggregate-only telemetry.
  • Ensure logs exclude API keys, prompts, completions, tool arguments, and user identifiers; retain only headers needed to explain throttling, request IDs, and counts.
Show full SKILL.md (209 more words)Show less

Instructions

  1. Read the response headers after each permitted call and update a shared limiter using the provider's remaining/reset values. Reserve headroom for interactive traffic.
  2. Honor retry-after when present, apply jitter and a maximum delay, and stop after a bounded number of attempts. Never let every worker retry at the same instant.
  3. Coordinate RPM, input-token, and output-token budgets across instances. Queue or batch offline work and apply backpressure when the shared budget is exhausted.
  4. Canary limiter changes with synthetic traffic and compare 429 rate, latency, queue age, token totals, and side_effects=0. Roll back the limiter/configuration if thresholds or scope checks fail.
  5. Expire queued test items and temporary counters according to the retention policy; retain a redacted receipt for the decision.

Output

Produce a rate-limit receipt with model class, configured and observed aggregate limits, limiter version, request/token counts, retry-after handling, 429 count, queue/batch disposition, canary result, rollback reference, and cleanup status. Do not include payloads or secrets.

Examples

Queue 20 synthetic OK prompts behind a shared 10-RPM limiter, permit only the configured window, and assert 429_retries_bounded=true; side_effects=0. The receipt may contain submitted=20; completed=<aggregate>; deferred=<aggregate>; headers_captured=true; canary=pass; cleanup=verified without user content.

Resources

Next Steps

For security configuration, see anth-security-basics.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/.curated/anth-rate-limits of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit cfae287

Compare with similar skills

Anth Rate Limits next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Anth Rate Limits compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Anth Rate Limits this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.9kAutomated safety check: PassMIT
Anthropic Product Knowledgesyahiidkamil/Software-Engineer-AI-Agent-Atlas4014 repos~651Automated safety check: PassNone
Neon AI Gatewayneondatabase/agent-skills100—~5.1kAutomated safety check: NotesApache-2.0
Add Hosted Keysimstudioai/sim30k—~3.4kAutomated safety check: PassApache-2.0
Upstash Ratelimit TSupstash/ratelimit-js2k—~313Automated safety check: PassMIT
Repo2skillzhangyanxs/repo2skill246—~3.6kAutomated safety check: PassNone

Similar skills

  • Anthropic Product Knowledge

    syahiidkamil/Software-Engineer-AI-Agent-Atlas

    Stop and consult this skill whenever your response would include specific facts about Anthropic's products.

    401 GitHub starsUsed in 4 repos~651 tokens
    AI & LLM EngineeringAuto-check passed
  • Neon AI Gateway

    neondatabase/agent-skills

    Official

    One API and one credential for frontier and open-source LLMs, built into your Neon branch and powered by Databricks.

    100 GitHub stars~5.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Add Hosted Key

    simstudioai/sim

    Add hosted API key support to a tool so Sim provides the key (metered and billed to the workspace) when a user has not brought their own.

    30k GitHub stars~3.4k tokensUpdated today
    Backend & APIsAuto-check passed
  • Upstash Ratelimit TS

    upstash/ratelimit-js

    Official

    Lightweight guidance for using the Redis Rate Limit TypeScript SDK, including setup steps, basic usage, and pointers to advanced algorithm, features, pricing, and traffic‑protection docs.

    2k GitHub stars~313 tokensUpdated 15 days ago
    Backend & APIsAuto-check passed
  • Repo2skill

    zhangyanxs/repo2skill

    Convert GitHub/GitLab/Gitee repositories into comprehensive OpenCode Skills using embedded LLM calls with multiple mirrors and rate limit handling

    246 GitHub stars~3.6k tokensUpdated 7 mo ago
    Backend & APIsAuto-check passed
  • Better Auth security hardening: rate limits, secrets, CSRF, trusted origins, cookies, sessions, OAuth tokens, and audit logging.

    4.8k GitHub stars~896 tokensUpdated 2 days ago
    Backend & APIsAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Anth Rate Limits

What does Anth Rate Limits do?

Implement Anthropic Claude API rate limiting, backoff, and quota management. Anth Rate Limits is an agent skill from jeremylongshore/tons-of-skills-marketplace. Implement Anthropic Claude API rate limiting, backoff, and quota management.

When should I use Anth Rate Limits?

Anth Rate Limits fits situations like: handling 429 errors; optimizing request throughput; managing RPM/TPM limits across usage tiers; with phrases like anthropic rate limit.

How do I install Anth Rate Limits in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-rate-limits -a claude-code`. Or copy the skill folder (skills/.curated/anth-rate-limits in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/anth-rate-limits in your project. Claude Code loads it when a task matches its description.

How do I install Anth Rate Limits in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-rate-limits -a codex`. Or copy the skill folder (skills/.curated/anth-rate-limits in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/anth-rate-limits in your project. Codex loads it when a task matches its description.

Can I use Anth Rate Limits in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-rate-limits -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/anth-rate-limits, .gemini/skills/anth-rate-limits, .github/skills/anth-rate-limits and .opencode/skills/anth-rate-limits in your project.

What does Anth Rate Limits need to run?

SKILL.md names no scripts, command-line tools or credentials: Anth Rate Limits is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit. Compatibility (from SKILL.md): Designed for Claude Code.

Does Anth Rate Limits access the network?

SKILL.md names 2 domains. As links in the text: platform.claude.com and console.anthropic.com. This is read from the text; nothing was executed.

Is Anth Rate Limits safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Anth Rate Limits use?

Anth Rate Limits is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Anth Rate Limits use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Anth Rate Limits?

Skills that share tags, products or a category with Anth Rate Limits: Anthropic Product Knowledge (syahiidkamil/Software-Engineer-AI-Agent-Atlas, 401 stars), Neon AI Gateway (neondatabase/agent-skills, 100 stars), Add Hosted Key (simstudioai/sim, 30k stars) and Upstash Ratelimit TS (upstash/ratelimit-js, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Anth Rate Limits?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.