Agent skill

Openrouter Rate Limits

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Understand and handle OpenRouter rate limits. An agent skill from jeremylongshore/tons-of-skills-marketplace.

MITAuto-check passedBackend & APIs

Install Openrouter Rate Limits

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill openrouter-rate-limits -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace openrouter-rate-limits --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/openrouter-rate-limits .claude/skills/openrouter-rate-limits && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
openrouter-rate-limits
GitHub stars
2.8k
Token cost
~2.4k tokens
SKILL.md length
539 words
Files
7 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Understand and handle OpenRouter rate limits. An agent skill from jeremylongshore/tons-of-skills-marketplace.

  • Works in 6 steps: Query your key's limits via GET… → Place yourself in the Rate Limit Tiers… → Inspect live headroom with… → …
  • Hitting 429 errors
  • SKILL.md covers Overview, Prerequisites, Instructions and Check Your Rate Limits, plus 10 more sections
  • Calls curl and jq; reaches openrouter.ai; needs OPENROUTER_API_KEY

What it does

Openrouter Rate Limits is an agent skill from jeremylongshore/tons-of-skills-marketplace. Understand and handle OpenRouter rate limits. Use when hitting 429 errors, building high-throughput systems, or implementing retry logic. Triggers: 'openrouter rate limit', 'openrouter 429', 'openrouter throttle', 'rate limiting openrouter'.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `references/batch-processing-with-rate-limits.md`, `references/best-practices.md` and `references/errors.md`). Compatibility notes: Designed for Claude Code

It sits in Backend & APIs, covering Rate limiting and Model routing and gateways. It works with OpenRouter and OpenAI. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Hitting 429 errors
  • Building high-throughput systems
  • Implementing retry logic

Example prompts

  • “openrouter rate limit”
  • “openrouter 429”
  • “openrouter throttle”
  • “/openrouter-rate-limits”

Requirements

  • Python 3
  • A credential in OPENROUTER_API_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Grep, Bash(python3:*), Bash(curl:*), Bash(jq:*)

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Query your key's limits via GET /api/v1/auth/key per Check Your Rate Limits — note rate_limit.requests and rate_limit.interval.
  2. Place yourself in the Rate Limit Tiers table, remembering free models carry separate daily caps (50 req/day free, 1000 req/day with $10+…
  3. Inspect live headroom with check_rate_headers() per Read Rate Limit Headers — watch x-ratelimit-remaining and retry-after.
  4. Configure SDK retries per Retry Strategy with OpenAI SDK: max_retries=5, timeout=60.0; the SDK catches 429s and backs off with jitter…
  5. Add the client-side TokenBucket limiter from Custom Rate Limiter, set below the server limit (e.g. 150 per 10s under a 200/10s cap) so you…
  6. For bulk jobs, use batch_with_rate_limit() per Batch Processing with Rate Awareness — staggered starts plus semaphore-capped concurrency…

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Grep
    • Bash(python3:*)
    • Bash(curl:*)
    • Bash(jq:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openrouter.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Openrouter Rate Limits loads about 2.4k tokens when it runs, and up to ~5.3k if it reads all its reference files. Until then it costs about 66 tokens; SKILL.md has 539 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 539 words, ~2,423 tokens.

Download SKILL.mdSave it as .claude/skills/openrouter-rate-limits/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
openrouter-rate-limits
description
Understand and handle OpenRouter rate limits. Use when hitting 429 errors, building high-throughput systems, or implementing retry logic. Triggers: 'openrouter rate limit', 'openrouter 429', 'openrouter throttle', 'rate limiting openrouter'.
allowed-tools
Read, Write, Edit, Grep, Bash(python3:*), Bash(curl:*), Bash(jq:*)
compatibility
Designed for Claude Code
version
1.20.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, openrouter, rate-limits, throttling

OpenRouter Rate Limits

Overview

OpenRouter rate limits are per-key, not per-account. Free tier keys get lower limits; paid keys get higher limits that scale with credit balance. The OpenAI SDK has built-in retry with exponential backoff for 429 responses. Check your current limits via GET /api/v1/auth/key. Rate limit headers are returned on every response.

Prerequisites

  • An OpenRouter API key (sk-or-v1-...) exported as OPENROUTER_API_KEY — see the openrouter-install-auth skill for setup
  • curl and jq for querying your key's limits from GET /api/v1/auth/key
  • Python 3.8+ with the OpenAI SDK (sync OpenAI and AsyncOpenAI) plus the requests package for reading rate-limit headers directly
  • Awareness of your tier: free keys get 20 req/10s, keys with any credits 200 req/10s (see Rate Limit Tiers)

Instructions

  1. Query your key's limits via GET /api/v1/auth/key per Check Your Rate Limits — note rate_limit.requests and rate_limit.interval.
  2. Place yourself in the Rate Limit Tiers table, remembering free models carry separate daily caps (50 req/day free, 1000 req/day with $10+ credits).
  3. Inspect live headroom with check_rate_headers() per Read Rate Limit Headers — watch x-ratelimit-remaining and retry-after.
  4. Configure SDK retries per Retry Strategy with OpenAI SDK: max_retries=5, timeout=60.0; the SDK catches 429s and backs off with jitter automatically.
  5. Add the client-side TokenBucket limiter from Custom Rate Limiter, set below the server limit (e.g. 150 per 10s under a 200/10s cap) so you rarely hit 429 at all.
  6. For bulk jobs, use batch_with_rate_limit() per Batch Processing with Rate Awareness — staggered starts plus semaphore-capped concurrency instead of bursts.

Check Your Rate Limits

bash
# Query current rate limit configuration for your key
curl -s https://openrouter.ai/api/v1/auth/key \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" | jq '{
    label: .data.label,
    rate_limit: .data.rate_limit,
    is_free_tier: .data.is_free_tier,
    credits_used: .data.usage,
    credit_limit: .data.limit
  }'
# Example output:
# {
#   "label": "my-app-prod",
#   "rate_limit": {"requests": 200, "interval": "10s"},
#   "is_free_tier": false,
#   "credits_used": 12.34,
#   "credit_limit": 100
# }

Rate Limit Tiers

TierRequestsIntervalWho
Free (no credits)2010sNew accounts
Free (with credits)20010sAccounts with any credits
PaidHigherVariesBased on credit balance

Free models have separate limits: 50 req/day (free users), 1000 req/day (with $10+ credits).

Read Rate Limit Headers

python
import os
from openai import OpenAI
import requests as http_requests

# The OpenAI SDK abstracts headers, so use requests for direct access
def check_rate_headers():
    """Make a request and inspect rate limit headers."""
    resp = http_requests.post(
        "https://openrouter.ai/api/v1/chat/completions",
        headers={
            "Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}",
            "Content-Type": "application/json",
            "HTTP-Referer": "https://my-app.com",
        },
        json={
            "model": "openai/gpt-4o-mini",
            "messages": [{"role": "user", "content": "hi"}],
            "max_tokens": 1,
        },
    )
    return {
        "status": resp.status_code,
        "x-ratelimit-limit": resp.headers.get("x-ratelimit-limit"),
        "x-ratelimit-remaining": resp.headers.get("x-ratelimit-remaining"),
        "x-ratelimit-reset": resp.headers.get("x-ratelimit-reset"),
        "retry-after": resp.headers.get("retry-after"),
    }

Retry Strategy with OpenAI SDK

python
from openai import OpenAI

# The SDK handles 429 retries automatically with exponential backoff
client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
    max_retries=5,           # Default is 2; increase for high-throughput
    timeout=60.0,            # Per-request timeout
    default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
)

# The SDK will:
# 1. Catch 429 responses
# 2. Read Retry-After header
# 3. Wait with exponential backoff (+ jitter)
# 4. Retry up to max_retries times
response = client.chat.completions.create(
    model="anthropic/claude-3.5-sonnet",
    messages=[{"role": "user", "content": "Hello"}],
    max_tokens=200,
)

Custom Rate Limiter (Client-Side)

python
import time, threading
from collections import deque

class TokenBucket:
    """Client-side rate limiter to prevent hitting server limits."""

    def __init__(self, rate: int = 200, interval: float = 10.0):
        self.rate = rate           # Max requests per interval
        self.interval = interval
        self._timestamps = deque()
        self._lock = threading.Lock()

    def acquire(self, timeout: float = 30.0) -> bool:
        """Block until a request slot is available."""
        deadline = time.monotonic() + timeout
        while time.monotonic() < deadline:
            with self._lock:
                now = time.monotonic()
                # Remove timestamps outside the window
                while self._timestamps and now - self._timestamps[0] > self.interval:
                    self._timestamps.popleft()

                if len(self._timestamps) < self.rate:
                    self._timestamps.append(now)
                    return True

            time.sleep(0.1)  # Wait and retry
        return False  # Timed out

limiter = TokenBucket(rate=150, interval=10.0)  # Stay under 200 limit

def rate_limited_completion(messages, **kwargs):
    """Completion with client-side rate limiting."""
    if not limiter.acquire(timeout=30):
        raise TimeoutError("Rate limiter timeout")
    return client.chat.completions.create(messages=messages, **kwargs)

Batch Processing with Rate Awareness

python
import asyncio
from openai import AsyncOpenAI

async def batch_with_rate_limit(prompts: list[str], model="openai/gpt-4o-mini",
                                 max_concurrent=10, delay_between=0.05):
    """Process a batch of prompts with rate-aware concurrency."""
    semaphore = asyncio.Semaphore(max_concurrent)
    aclient = AsyncOpenAI(
        base_url="https://openrouter.ai/api/v1",
        api_key=os.environ["OPENROUTER_API_KEY"],
        max_retries=5,
        default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
    )

    async def process(prompt, idx):
        await asyncio.sleep(idx * delay_between)  # Stagger requests
        async with semaphore:
            response = await aclient.chat.completions.create(
                model=model,
                messages=[{"role": "user", "content": prompt}],
                max_tokens=200,
            )
            return response.choices[0].message.content

    return await asyncio.gather(*[process(p, i) for i, p in enumerate(prompts)])
Show full SKILL.md (226 more words)Show less

Output

  • A key-limit snapshot from /api/v1/auth/key: label, rate_limit (requests + interval), is_free_tier, and credit usage
  • Per-request header readings from check_rate_headers(): x-ratelimit-limit, x-ratelimit-remaining, x-ratelimit-reset, retry-after
  • A rate-limited client: SDK auto-retry on 429 plus a TokenBucket that blocks (up to a timeout) instead of erroring
  • Ordered batch results from batch_with_rate_limit() produced without triggering a retry storm

Examples

Read your server-side limit, then size the client-side limiter under it:

bash
curl -s https://openrouter.ai/api/v1/auth/key \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" | jq '.data.rate_limit'
# {"requests": 200, "interval": "10s"}

With that 200/10s ceiling, configure TokenBucket(rate=150, interval=10.0) so steady-state traffic stays ~25% below the limit, and let the SDK's max_retries=5 absorb whatever bursts through. More worked examples: references/examples.md.

Error Handling

ErrorCauseFix
429 Too Many RequestsExceeded requests per intervalSDK auto-retries; increase max_retries
Retry stormMultiple clients retrying simultaneouslyAdd random jitter (0-1s) to retry delay
Silent throttlingResponses slow down before 429Monitor latency; proactively reduce rate
Free tier limit hit50 req/day on free modelsAdd credits ($10+) for 1000 req/day limit

Enterprise Considerations

  • Rate limits are per-key: use multiple keys to multiply effective throughput
  • The OpenAI SDK handles 429 retries automatically -- configure max_retries (default 2)
  • Implement client-side rate limiting to stay under limits proactively (cheaper than retries)
  • Free models have daily limits separate from the per-key rate limit
  • Monitor x-ratelimit-remaining headers to detect approaching limits before hitting 429
  • For batch workloads, use staggered concurrent requests rather than burst patterns

References

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in skills/.curated/openrouter-rate-limits of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/batch-processing-with-rate-limits.md
  • references/best-practices.md
  • references/errors.md
  • references/examples.md
  • references/rate-limiter-implementation.md
  • references/retry-strategies.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Openrouter Rate Limits next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Openrouter Rate Limits compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Openrouter Rate Limits this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2.4kAutomated safety check: PassMIT
LLM GatewayBagelHole/DevOps-Security-Agent-Skills1.2k—~2kAutomated safety check: PassMIT
LLM Gatewaysickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT
Embeddings via 9Routerdecolua/9router31k—~604Automated safety check: PassMIT
Using Ccproxy APIstarbaser/ccproxy350—~4kAutomated safety check: PassCustom licence
Auth Architecturemajiayu000/litellm-rs118—~1.9kAutomated safety check: PassMIT

Similar skills

  • LLM Gateway

    BagelHole/DevOps-Security-Agent-Skills

    Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

    1.2k GitHub stars~2k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLM Gateway

    sickn33/agentic-awesome-skills

    Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Backend & APIsAuto-check passed
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    31k GitHub stars~604 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Using Ccproxy API

    starbaser/ccproxy

    Guides users through ccproxy as an OpenAI-compatible and Anthropic-compatible LLM API server with SDK integration, OAuth authentication, sentinel key substitution, model routing, and troubleshooting.

    350 GitHub stars~4k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Auth Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Authentication Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~1.9k tokensUpdated today
    Backend & APIsAuto-check passed
  • Quota Axi

    kunchenguid/quota-axi

    Report local Claude, Codex, Cursor, GitHub Copilot, Grok, Kimi, Z.AI, Alibaba, OpenCode Go, Antigravity, Command Code, MiniMax, MiMo, DeepSeek, OpenRouter, ElevenLabs, Devin, Muse, and Higgsfield…

    147 GitHub stars~547 tokensUpdated yesterday
    Backend & APIsAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Openrouter Rate Limits

What does Openrouter Rate Limits do?

Understand and handle OpenRouter rate limits. An agent skill from jeremylongshore/tons-of-skills-marketplace. Openrouter Rate Limits is an agent skill from jeremylongshore/tons-of-skills-marketplace. Understand and handle OpenRouter rate limits.

When should I use Openrouter Rate Limits?

Openrouter Rate Limits fits situations like: hitting 429 errors; building high-throughput systems; implementing retry logic.

How do I install Openrouter Rate Limits in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill openrouter-rate-limits -a claude-code`. Or copy the skill folder (skills/.curated/openrouter-rate-limits in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/openrouter-rate-limits in your project. Claude Code loads it when a task matches its description.

How do I install Openrouter Rate Limits in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill openrouter-rate-limits -a codex`. Or copy the skill folder (skills/.curated/openrouter-rate-limits in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/openrouter-rate-limits in your project. Codex loads it when a task matches its description.

Can I use Openrouter Rate Limits in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill openrouter-rate-limits -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/openrouter-rate-limits, .gemini/skills/openrouter-rate-limits, .github/skills/openrouter-rate-limits and .opencode/skills/openrouter-rate-limits in your project.

What does Openrouter Rate Limits need to run?

Going by SKILL.md and its folder, Openrouter Rate Limits needs the command-line tools its instructions call (curl and jq) and credentials named OPENROUTER_API_KEY. Our summary lists: Python 3; A credential in OPENROUTER_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit, Grep, Bash(python3:*), Bash(curl:*), Bash(jq:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Openrouter Rate Limits access the network?

SKILL.md names 1 domain. In commands or code: openrouter.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Openrouter Rate Limits safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Openrouter Rate Limits use?

Openrouter Rate Limits is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Openrouter Rate Limits use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.

What are the alternatives to Openrouter Rate Limits?

Skills that share tags, products or a category with Openrouter Rate Limits: LLM Gateway (BagelHole/DevOps-Security-Agent-Skills, 1.2k stars), LLM Gateway (sickn33/agentic-awesome-skills, 47k stars), Embeddings via 9Router (decolua/9router, 31k stars) and Using Ccproxy API (starbaser/ccproxy, 350 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Openrouter Rate Limits?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.