Implement Groq rate limit handling with backoff, queuing, and header parsing.

MITAuto-check passedBackend & APIs

Install Groq Rate Limits

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-rate-limits -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace groq-rate-limits --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/groq-rate-limits .claude/skills/groq-rate-limits && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
groq-rate-limits
GitHub stars
2.8k
Token cost
~1.5k tokens
SKILL.md length
617 words
Files
3 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Implement Groq rate limit handling with backoff, queuing, and header parsing.

  • Works in 5 steps: Parse Rate Limit Headers → Exponential Backoff with Retry-After → Request Queue with Concurrency Control → …
  • Handling rate limit errors
  • SKILL.md covers Overview, Prerequisites, Rate Limits at a Glance and Instructions, plus 5 more sections
  • Calls npm; needs GROQ_API_KEY

What it does

Groq Rate Limits is an agent skill from jeremylongshore/tons-of-skills-marketplace. Implement Groq rate limit handling with backoff, queuing, and header parsing. Use when handling rate limit errors, implementing retry logic, or optimizing API request throughput for Groq. Trigger with phrases like "groq rate limit", "groq throttling", "groq 429", "groq retry", "groq backoff".

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/implementation.md` and `references/reference.md`). Compatibility notes: Designed for Claude Code

It sits in Backend & APIs, covering Rate limiting. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Handling rate limit errors
  • Implementing retry logic
  • Optimizing API request throughput for Groq
  • With phrases like groq rate limit

Example prompts

  • “groq rate limit”
  • “groq throttling”
  • “groq 429”
  • “/groq-rate-limits”

Requirements

  • Node.js
  • A credential in GROQ_API_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Parse Rate Limit Headers
  2. Exponential Backoff with Retry-After
  3. Request Queue with Concurrency Control
  4. Proactive Rate Limit Monitor
  5. Model-Aware Rate Limit Strategy

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • console.groq.com
    • groq.com
    • npmjs.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GROQ_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Groq Rate Limits loads about 1.5k tokens when it runs, and up to ~3.7k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 617 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 617 words, ~1,508 tokens.

Download SKILL.mdSave it as .claude/skills/groq-rate-limits/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
groq-rate-limits
description
Implement Groq rate limit handling with backoff, queuing, and header parsing. Use when handling rate limit errors, implementing retry logic, or optimizing API request throughput for Groq. Trigger with phrases like "groq rate limit", "groq throttling", "groq 429", "groq retry", "groq backoff".
allowed-tools
Read, Write, Edit
compatibility
Designed for Claude Code
version
1.11.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, groq, api

Groq Rate Limits

Overview

Handle Groq rate limits using the retry-after header, exponential backoff, and request queuing. Groq enforces limits at the organization level with both RPM (requests/minute) and TPM (tokens/minute) constraints -- hitting either one triggers a 429.

The workflow builds up in five composable layers: parse the rate-limit headers, wrap calls in retry-with-backoff, gate concurrency through a queue, monitor remaining capacity proactively, and fall back across models when one pool is exhausted. Read SKILL.md for the high-level flow, then drill into the full implementation for every code block and the reference tables + worked examples for header definitions and composed clients.

Prerequisites

  • A Groq API key (GROQ_API_KEY) — get one at console.groq.com.
  • groq-sdk installed: npm install groq-sdk.
  • For queuing (Step 3): p-queue installed: npm install p-queue.
  • Node.js 18+ (for native fetch and the SDK).
  • Know your plan's limits — check console.groq.com/settings/limits.

Rate Limits at a Glance

Groq applies RPM, RPD, TPM, and TPD limits simultaneously — you must stay under every one, and either RPM or TPM can trip a 429. Every response (even a success) carries x-ratelimit-* headers describing remaining capacity and reset timing; 429 responses add a retry-after header. Full header and constraint tables: reference.md.

Instructions

Compose these five steps into one client wrapper (queue → monitor → retry). Each step's complete, copy-pasteable code is in implementation.md.

Step 1: Parse Rate Limit Headers

Read the x-ratelimit-* headers off every response into a typed RateLimitInfo so downstream logic can reason about remaining capacity. Groq reports reset times as strings like "1.2s" or "120ms" — normalize them to milliseconds.

Step 2: Exponential Backoff with Retry-After

Wrap each API call in a retry loop. Prefer Groq's retry-after header when present; otherwise back off exponentially with jitter, capped at maxDelayMs. Retry only 429 and 5xx — other 4xx errors are not retryable.

Step 3: Request Queue with Concurrency Control

Gate all requests through a p-queue sized to your plan's RPM (intervalCap over a 60s interval) so bursts never exceed the limit in the first place.

Step 4: Proactive Rate Limit Monitor

Track remaining requests/tokens from response headers and pause before hitting zero (shouldThrottle() → waitIfNeeded()), instead of reacting to 429s after the fact.

Step 5: Model-Aware Rate Limit Strategy

Different models draw from different limit pools. When the preferred model is throttled, fall back to another model to keep making progress without waiting for a reset.

Show full SKILL.md (231 more words)Show less

Output

Applying this skill produces:

  • A withRateLimitRetry() wrapper that transparently retries 429/5xx with retry-after-aware backoff.
  • A RateLimitMonitor that surfaces live status (getStatus() → "Requests: N remaining | Tokens: M remaining") and throttles proactively.
  • A p-queue-backed client that caps throughput to your RPM so 100 fan-out calls complete without tripping the limit.
  • Console diagnostics on every backoff/throttle event (e.g. Rate limited (attempt 2/5). Waiting 1.4s...).

The observable end state: sustained request volume that stays under RPM/TPM with zero unhandled 429s.

Error Handling

ScenarioSymptomSolution
Burst of requestsMany 429s in quick successionUse queue with p-queue interval limiting (Step 3)
Large prompts burn TPM429 on tokens, not requestsReduce max_tokens, compress prompts
Free tier too restrictiveConstant 429sUpgrade to Developer plan at console.groq.com
Multiple services sharing keyCascading 429sUse separate API keys per service
retry-after absent on 429Retries hammer too fastFall back to exponential backoff + jitter (Step 2)

Examples

Start from this minimal 429 handler, then graduate to the composed client (queue + monitor + retry) in reference.md:

typescript
try {
  await groq.chat.completions.create({ model, messages });
} catch (err) {
  if (err instanceof Groq.APIError && err.status === 429) {
    const retryAfter = parseInt(err.headers?.["retry-after"] || "0");
    console.log(`Rate limited. retry-after says wait ${retryAfter}s.`);
    // -> feed retryAfter into withRateLimitRetry (Step 2)
  }
}
  • Full five-step implementation (every code block, verbatim): implementation.md
  • Composed client + header/limit tables + a 100-request fan-out: reference.md

Resources

Next Steps

For security configuration, see the groq-security-basics skill in this pack, which covers API key storage, rotation, and request signing to complement the throughput handling above.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/.curated/groq-rate-limits of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/implementation.md
  • references/reference.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Groq Rate Limits next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Groq Rate Limits compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Groq Rate Limits this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.5kAutomated safety check: PassMIT
Add Hosted Keysimstudioai/sim30k—~3.6kAutomated safety check: PassApache-2.0
Upstash Ratelimit TSupstash/ratelimit-js2k—~313Automated safety check: PassMIT
Repo2skillzhangyanxs/repo2skill246—~3.6kAutomated safety check: PassNone
Better Auth Security Best PracticesEpicenterHQ/epicenter4.8k—~896Automated safety check: PassCustom licence
Dload Fetch Toolphp-internal/dload105—~1.1kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Add Hosted Key

    simstudioai/sim

    Add hosted API key support to a tool so Sim provides the key (metered and billed to the workspace) when a user has not brought their own.

    30k GitHub stars~3.6k tokensUpdated today
    Backend & APIsAuto-check passed
  • Upstash Ratelimit TS

    upstash/ratelimit-js

    Official

    Lightweight guidance for using the Redis Rate Limit TypeScript SDK, including setup steps, basic usage, and pointers to advanced algorithm, features, pricing, and traffic‑protection docs.

    2k GitHub stars~313 tokensUpdated 16 days ago
    Backend & APIsAuto-check passed
  • Repo2skill

    zhangyanxs/repo2skill

    Convert GitHub/GitLab/Gitee repositories into comprehensive OpenCode Skills using embedded LLM calls with multiple mirrors and rate limit handling

    246 GitHub stars~3.6k tokensUpdated 7 mo ago
    Backend & APIsAuto-check passed
  • Better Auth security hardening: rate limits, secrets, CSRF, trusted origins, cookies, sessions, OAuth tokens, and audit logging.

    4.8k GitHub stars~896 tokensUpdated 3 days ago
    Backend & APIsAuto-check passed
  • Dload Fetch Tool

    php-internal/dload

    Get a CLI tool — native binary or PHAR — from a GitHub release into a project folder with dload (vendor/bin/dload).

    105 GitHub stars~1.1k tokensUpdated today
    Backend & APIsAuto-check passed
  • API Gateway

    itsmostafa/aws-agent-skills

    AWS API Gateway for REST and HTTP API management. An agent skill from itsmostafa/aws-agent-skills.

    1.2k GitHub stars~2.2k tokensUpdated 5 days ago
    Backend & APIsAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Categories

Questions about Groq Rate Limits

What does Groq Rate Limits do?

Implement Groq rate limit handling with backoff, queuing, and header parsing. Groq Rate Limits is an agent skill from jeremylongshore/tons-of-skills-marketplace. Implement Groq rate limit handling with backoff, queuing, and header parsing.

When should I use Groq Rate Limits?

Groq Rate Limits fits situations like: handling rate limit errors; implementing retry logic; optimizing API request throughput for Groq; with phrases like groq rate limit.

How do I install Groq Rate Limits in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-rate-limits -a claude-code`. Or copy the skill folder (skills/.curated/groq-rate-limits in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/groq-rate-limits in your project. Claude Code loads it when a task matches its description.

How do I install Groq Rate Limits in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-rate-limits -a codex`. Or copy the skill folder (skills/.curated/groq-rate-limits in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/groq-rate-limits in your project. Codex loads it when a task matches its description.

Can I use Groq Rate Limits in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-rate-limits -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/groq-rate-limits, .gemini/skills/groq-rate-limits, .github/skills/groq-rate-limits and .opencode/skills/groq-rate-limits in your project.

What does Groq Rate Limits need to run?

Going by SKILL.md and its folder, Groq Rate Limits needs the command-line tools its instructions call (npm) and credentials named GROQ_API_KEY. Our summary lists: Node.js; A credential in GROQ_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit. Compatibility (from SKILL.md): Designed for Claude Code.

Does Groq Rate Limits access the network?

SKILL.md names 3 domains. As links in the text: console.groq.com, groq.com and npmjs.com. This is read from the text; nothing was executed.

Is Groq Rate Limits safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Groq Rate Limits use?

Groq Rate Limits is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Groq Rate Limits use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.

What are the alternatives to Groq Rate Limits?

Skills that share tags, products or a category with Groq Rate Limits: Add Hosted Key (simstudioai/sim, 30k stars), Upstash Ratelimit TS (upstash/ratelimit-js, 2k stars), Repo2skill (zhangyanxs/repo2skill, 246 stars) and Better Auth Security Best Practices (EpicenterHQ/epicenter, 4.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Groq Rate Limits?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.