Implement load testing, auto-scaling, and capacity planning for Claude API.

MITAuto-check passedTesting & QA

Install Anth Load Scale

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-load-scale -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace anth-load-scale --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/anth-load-scale .claude/skills/anth-load-scale && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
anth-load-scale
GitHub stars
2.8k
Token cost
~1.8k tokens
SKILL.md length
396 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Implement load testing, auto-scaling, and capacity planning for Claude API.

  • Works in 5 steps: Calculate RPM, input tokens per minute,… → Start with a small canary, then increase… → Separate real-time traffic from batch… → …
  • Running performance benchmarks
  • SKILL.md covers Overview, Capacity Planning, Load Testing Script and Scaling Strategies, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Anth Load Scale is an agent skill from jeremylongshore/tons-of-skills-marketplace. Implement load testing, auto-scaling, and capacity planning for Claude API. Use when running performance benchmarks, planning for traffic spikes, or configuring horizontal scaling for Claude-powered services. Trigger with phrases like "anthropic load test", "claude scaling", "anthropic capacity planning", "scale claude api".

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Claude Code

It sits in Testing & QA, covering Load testing, Site reliability engineering and LLM API integration. It works with Anthropic API. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Running performance benchmarks
  • Planning for traffic spikes
  • Configuring horizontal scaling for Claude-powered services
  • With phrases like anthropic load test

Example prompts

  • “anthropic load test”
  • “claude scaling”
  • “anthropic capacity planning”
  • “/anth-load-scale”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(npm:*), Grep

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Calculate RPM, input tokens per minute, output tokens per minute, concurrency, and expected cost from the measured workload. Reserve…
  2. Start with a small canary, then increase concurrency in bounded steps while a shared limiter coordinates all workers. Stop immediately at…
  3. Separate real-time traffic from batch work, and use queue backpressure rather than unbounded task creation. Honor provider retry metadata…
  4. Compare baseline and candidate metrics, including aggregate token/cost usage and side_effects=0. Promote only after an owner approves the…
  5. Expire synthetic fixtures, queues, and temporary metrics according to the test retention policy, and keep a redacted capacity receipt.

What it can do on your machine

Read from SKILL.md and the folder at commit 80f86df. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(npm:*)
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.claude.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Anth Load Scale loads about 1.8k tokens when it runs. Until then it costs about 86 tokens; SKILL.md has 396 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~86
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit 80f86df, republished under its MIT licence (© jeremylongshore). 396 words, ~1,806 tokens.

Download SKILL.mdSave it as .claude/skills/anth-load-scale/SKILL.md (or your agent's skills folder).
name
anth-load-scale
description
Implement load testing, auto-scaling, and capacity planning for Claude API. Use when running performance benchmarks, planning for traffic spikes, or configuring horizontal scaling for Claude-powered services. Trigger with phrases like "anthropic load test", "claude scaling", "anthropic capacity planning", "scale claude api".
allowed-tools
Read, Write, Edit, Bash(npm:*), Grep
compatibility
Designed for Claude Code
version
1.7.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, ai, anthropic

Anthropic Load & Scale

Overview

Capacity planning and load testing for Claude API integrations. Key constraint: your rate limits (RPM/ITPM/OTPM) are the ceiling, not your infrastructure.

Capacity Planning

python
# Calculate required tier based on traffic
def plan_capacity(
    requests_per_minute: int,
    avg_input_tokens: int,
    avg_output_tokens: int,
    model: str = "claude-sonnet-4-20250514"
) -> dict:
    itpm = requests_per_minute * avg_input_tokens
    otpm = requests_per_minute * avg_output_tokens

    # Estimate monthly cost
    pricing = {
        "claude-haiku-4-20250514": (0.80, 4.00),
        "claude-sonnet-4-20250514": (3.00, 15.00),
        "claude-opus-4-20250514": (15.00, 75.00),
    }
    rates = pricing[model]
    cost_per_request = (avg_input_tokens * rates[0] + avg_output_tokens * rates[1]) / 1_000_000
    monthly_cost = cost_per_request * requests_per_minute * 60 * 24 * 30

    return {
        "rpm_needed": requests_per_minute,
        "itpm_needed": itpm,
        "otpm_needed": otpm,
        "cost_per_request": f"${cost_per_request:.4f}",
        "monthly_estimate": f"${monthly_cost:,.0f}",
        "recommendation": "Contact Anthropic sales for Scale tier" if requests_per_minute > 500 else "Self-serve tiers sufficient",
    }

print(plan_capacity(100, 500, 200))

Load Testing Script

python
import anthropic
import asyncio
import time
from dataclasses import dataclass

@dataclass
class LoadTestResult:
    total_requests: int = 0
    successful: int = 0
    failed: int = 0
    rate_limited: int = 0
    avg_latency_ms: float = 0
    p99_latency_ms: float = 0
    total_input_tokens: int = 0
    total_output_tokens: int = 0

async def load_test(
    concurrency: int = 10,
    total_requests: int = 100,
    model: str = "claude-haiku-4-20250514"
) -> LoadTestResult:
    client = anthropic.Anthropic()
    result = LoadTestResult()
    latencies = []
    semaphore = asyncio.Semaphore(concurrency)

    async def single_request():
        async with semaphore:
            start = time.monotonic()
            try:
                msg = client.messages.create(
                    model=model,
                    max_tokens=64,
                    messages=[{"role": "user", "content": "Respond with exactly: OK"}]
                )
                duration = (time.monotonic() - start) * 1000
                latencies.append(duration)
                result.successful += 1
                result.total_input_tokens += msg.usage.input_tokens
                result.total_output_tokens += msg.usage.output_tokens
            except anthropic.RateLimitError:
                result.rate_limited += 1
            except Exception:
                result.failed += 1
            result.total_requests += 1

    tasks = [single_request() for _ in range(total_requests)]
    await asyncio.gather(*tasks)

    if latencies:
        latencies.sort()
        result.avg_latency_ms = sum(latencies) / len(latencies)
        result.p99_latency_ms = latencies[int(len(latencies) * 0.99)]

    return result

# Run: asyncio.run(load_test(concurrency=10, total_requests=50))

Scaling Strategies

StrategyWhenImplementation
Queue-based processing> 50 RPM sustainedRedis/SQS queue + worker pool
Model routingMixed workloadsHaiku for simple, Sonnet for complex
Message BatchesOffline processing100K requests, 50% cheaper, no RPM impact
Prompt cachingRepeated system prompts90% input token savings
Request coalescingDuplicate promptsCache identical request hashes

Horizontal Scaling Pattern

python
# Multiple application instances sharing the same API key
# Rate limits are per-organization, NOT per-instance
# Use a shared rate limiter (Redis) to coordinate

import redis

r = redis.Redis()

def check_rate_limit(key: str = "claude:rpm", limit: int = 100, window: int = 60) -> bool:
    current = r.incr(key)
    if current == 1:
        r.expire(key, window)
    return current <= limit

Error Handling

IssueCauseFix
429 during load testExceeded tier limitsReduce concurrency or upgrade tier
Increasing latency under loadOutput queue saturationReduce max_tokens
Uneven request distributionNo load balancingUse queue for fair distribution

Prerequisites

  • Confirm the organization/model rate limits, budget ceiling, test environment, concurrency cap, and success/latency/error thresholds before measuring capacity.
  • Run only against an approved sandbox using synthetic prompts and a no-op result sink. Never stress production or use real customer content for load tests.
  • Configure aggregate metrics and redaction: request counts, status classes, latency, token totals, queue depth, and 429 counts are sufficient; prompts, completions, keys, and tool arguments are not.
Show full SKILL.md (204 more words)Show less

Instructions

  1. Calculate RPM, input tokens per minute, output tokens per minute, concurrency, and expected cost from the measured workload. Reserve headroom below provider and application limits.
  2. Start with a small canary, then increase concurrency in bounded steps while a shared limiter coordinates all workers. Stop immediately at error, budget, data-scope, or latency thresholds.
  3. Separate real-time traffic from batch work, and use queue backpressure rather than unbounded task creation. Honor provider retry metadata and avoid synchronized retries.
  4. Compare baseline and candidate metrics, including aggregate token/cost usage and side_effects=0. Promote only after an owner approves the result; revert autoscaling/limiter changes on regression.
  5. Expire synthetic fixtures, queues, and temporary metrics according to the test retention policy, and keep a redacted capacity receipt.

Output

Return a capacity receipt with workload class, model, concurrency steps, aggregate request/token counts, p50/p95/p99 latency, status/429 counts, queue depth, cost estimate, threshold decision, canary result, rollback reference, and cleanup status. Do not include payloads or secret material.

Examples

Run 50 requests using Respond with exactly: OK in the sandbox, cap concurrency at 10, and assert side_effects=0. A useful receipt is requests=50; successes=50; rate_limited=0; p99_ms=<redacted>; tokens=<aggregate>; canary=pass; cleanup=verified.

Resources

Next Steps

For reliability patterns, see anth-reliability-patterns.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/.curated/anth-load-scale of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit 80f86df

Compare with similar skills

Anth Load Scale next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Anth Load Scale compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Anth Load Scale this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.8kAutomated safety check: PassMIT
E2E Livereceptron/mulmoclaude368—~580Automated safety check: PassMIT
E2E Live Mediareceptron/mulmoclaude368—~220Automated safety check: PassMIT
Performance Testingkid-sid/claude-spellbook190—~2.5kAutomated safety check: PassMIT
Afrexai Performance EngineeringLeoYeAI/openclaw-master-skills2.2k—~7.1kAutomated safety check: PassMIT
E2E Live Dockerreceptron/mulmoclaude368—~544Automated safety check: PassMIT

Similar skills

  • E2E Live

    receptron/mulmoclaude

    実 Claude API を叩く総合テスト(全カテゴリ)を実行する。リリース前ではなく、定期的に手動で回して回帰を検出するための skill。yarn dev が起動済みであることが前提。

    368 GitHub stars~580 tokensUpdated today
    Testing & QAAuto-check passed
  • E2E Live Media

    receptron/mulmoclaude

    実 Claude API を叩く media カテゴリ(画像 / PDF / 動画)の総合テストを実行する。yarn dev が起動済みであることが前提。

    368 GitHub stars~220 tokensUpdated today
    Testing & QAAuto-check passed
  • Performance Testing

    kid-sid/claude-spellbook

    A skill your agent uses when load testing a service before launch or after a significant traffic change — writing k6 or Locust scripts, setting SLO-based pass/fail thresholds, diagnosing bottlenecks…

    190 GitHub stars~2.5k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Afrexai Performance Engineering

    LeoYeAI/openclaw-master-skills

    Complete performance engineering system — profiling, optimization, load testing, capacity planning, and performance culture.

    2.2k GitHub stars~7.1k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • E2E Live Docker

    receptron/mulmoclaude

    実 Claude API + Docker サンドボックスを叩く docker カテゴリの総合テストを実行する。DISABLESANDBOX を unset した状態で yarn dev が起動済みであることが前提(サンドボックス off だと全 test が test.skip で抜ける)。

    368 GitHub stars~544 tokensUpdated today
    DevOps & CloudAuto-check passed
  • HTTP Load Profiler

    zebbern/claude-code-guide

    Run stepped HTTP load tests with ab/wrk, ramping concurrency levels to collect p50/p90/p99 latency, detect performance inflection points, and recommend optimal concurrency.

    4.7k GitHub stars~1.6k tokensUpdated yesterday
    Testing & QAAuto-check: notes

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated yesterday
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Anth Load Scale

What does Anth Load Scale do?

Implement load testing, auto-scaling, and capacity planning for Claude API. Anth Load Scale is an agent skill from jeremylongshore/tons-of-skills-marketplace. Implement load testing, auto-scaling, and capacity planning for Claude API.

When should I use Anth Load Scale?

Anth Load Scale fits situations like: running performance benchmarks; planning for traffic spikes; configuring horizontal scaling for Claude-powered services; with phrases like anthropic load test.

How do I install Anth Load Scale in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-load-scale -a claude-code`. Or copy the skill folder (skills/.curated/anth-load-scale in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/anth-load-scale in your project. Claude Code loads it when a task matches its description.

How do I install Anth Load Scale in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-load-scale -a codex`. Or copy the skill folder (skills/.curated/anth-load-scale in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/anth-load-scale in your project. Codex loads it when a task matches its description.

Can I use Anth Load Scale in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-load-scale -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/anth-load-scale, .gemini/skills/anth-load-scale, .github/skills/anth-load-scale and .opencode/skills/anth-load-scale in your project.

What does Anth Load Scale need to run?

SKILL.md names no scripts, command-line tools or credentials: Anth Load Scale is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(npm:*), Grep. Compatibility (from SKILL.md): Designed for Claude Code.

Does Anth Load Scale access the network?

SKILL.md names 1 domain. As links in the text: platform.claude.com. This is read from the text; nothing was executed.

Is Anth Load Scale safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Anth Load Scale use?

Anth Load Scale is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Anth Load Scale use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Anth Load Scale?

Skills that share tags, products or a category with Anth Load Scale: E2E Live (receptron/mulmoclaude, 368 stars), E2E Live Media (receptron/mulmoclaude, 368 stars), Performance Testing (kid-sid/claude-spellbook, 190 stars) and Afrexai Performance Engineering (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Anth Load Scale?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,825 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 9, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.