Agent skill

Anth Performance Tuning

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Optimize Claude API performance with prompt caching, model selection, streaming, and latency reduction techniques.

MITAuto-check passedAI & LLM Engineering

Install Anth Performance Tuning

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-performance-tuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace anth-performance-tuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/anth-performance-tuning .claude/skills/anth-performance-tuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
anth-performance-tuning
GitHub stars
2.8k
Token cost
~1.9k tokens
SKILL.md length
492 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Optimize Claude API performance with prompt caching, model selection, streaming, and latency reduction techniques.

  • Works in 5 steps: Establish a baseline for… → Change one lever at a time: model,… → Enforce request scope, max_tokens,… → …
  • Experiencing slow responses
  • SKILL.md covers Overview, Prompt Caching (Biggest Win), Model Selection for Speed and Streaming for Perceived Speed, plus 10 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Anth Performance Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize Claude API performance with prompt caching, model selection, streaming, and latency reduction techniques. Use when experiencing slow responses, optimizing token usage, or reducing time-to-first-token in production. Trigger with phrases like "anthropic performance", "claude speed", "optimize claude latency", "anthropic caching", "faster claude responses".

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering LLM cost and token optimization, LLM API integration and Caching. It works with Anthropic API. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Experiencing slow responses
  • Optimizing token usage
  • Reducing time-to-first-token in production
  • With phrases like anthropic performance

Example prompts

  • “anthropic performance”
  • “claude speed”
  • “optimize claude latency”
  • “/anth-performance-tuning”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Grep

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Establish a baseline for time-to-first-token, completion latency, tokens, cache hit rate, throughput, quality, and errors using repeated…
  2. Change one lever at a time: model, prompt/cache layout, token budget, streaming, batching, or concurrency. Keep prompt content out of logs…
  3. Enforce request scope, max_tokens, timeout, retry, and concurrency limits. Stop the run when rate limits, quality, or data-policy checks…
  4. Canary the selected configuration in a sandbox or internal workspace, compare against baseline, and obtain approval before production…
  5. Restore the prior configuration on regression, invalidate temporary cache/test artifacts according to retention policy, and retain a…

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.claude.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Anth Performance Tuning loads about 1.9k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 492 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 492 words, ~1,948 tokens.

Download SKILL.mdSave it as .claude/skills/anth-performance-tuning/SKILL.md (or your agent's skills folder).
name
anth-performance-tuning
description
Optimize Claude API performance with prompt caching, model selection, streaming, and latency reduction techniques. Use when experiencing slow responses, optimizing token usage, or reducing time-to-first-token in production. Trigger with phrases like "anthropic performance", "claude speed", "optimize claude latency", "anthropic caching", "faster claude responses".
allowed-tools
Read, Write, Edit, Grep
compatibility
Designed for Claude Code
version
1.7.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, ai, anthropic

Anthropic Performance Tuning

Overview

Optimize Claude API latency and throughput via prompt caching, model selection, streaming, and request optimization. The biggest wins come from prompt caching (90% input cost reduction) and model selection (Haiku is 4x faster than Sonnet).

Prompt Caching (Biggest Win)

python
import anthropic

client = anthropic.Anthropic()

# Mark long, reusable content with cache_control
# Cached content: 90% cheaper on subsequent requests, near-zero latency for cached portion
message = client.messages.create(
    model="claude-sonnet-4-20250514",
    max_tokens=1024,
    system=[
        {
            "type": "text",
            "text": "You are an expert on the following 50-page document: ...<long document>...",
            "cache_control": {"type": "ephemeral"}  # Cache this block
        }
    ],
    messages=[{"role": "user", "content": "What does section 3.2 say?"}]
)

# Check cache performance
print(f"Cache read tokens: {message.usage.cache_read_input_tokens}")   # Free/cheap
print(f"Cache creation tokens: {message.usage.cache_creation_input_tokens}")  # First call only
print(f"Uncached input tokens: {message.usage.input_tokens}")

Cache requirements: Minimum 1,024 tokens for Sonnet/Opus, 2,048 for Haiku. Cache lives for 5 minutes (refreshed on each hit).

Model Selection for Speed

ModelSpeedCost (per MTok in/out)Best For
Claude HaikuFastest$0.80 / $4.00Classification, extraction, routing
Claude SonnetBalanced$3.00 / $15.00General tasks, tool use, code
Claude OpusDeepest$15.00 / $75.00Complex reasoning, research
python
# Route by task complexity
def select_model(task_type: str) -> str:
    routing = {
        "classify": "claude-haiku-4-20250514",
        "extract": "claude-haiku-4-20250514",
        "summarize": "claude-sonnet-4-20250514",
        "code": "claude-sonnet-4-20250514",
        "research": "claude-opus-4-20250514",
    }
    return routing.get(task_type, "claude-sonnet-4-20250514")

Streaming for Perceived Speed

python
# Streaming reduces time-to-first-token from seconds to ~200ms
with client.messages.stream(
    model="claude-sonnet-4-20250514",
    max_tokens=2048,
    messages=[{"role": "user", "content": prompt}]
) as stream:
    for text in stream.text_stream:
        yield text  # User sees response immediately

Reduce Token Count

python
# 1. Set max_tokens to what you actually need (not max)
msg = client.messages.create(
    model="claude-haiku-4-20250514",
    max_tokens=128,  # Not 4096 — smaller = faster generation
    messages=[{"role": "user", "content": "Classify as positive/negative: 'Great product!'"}]
)

# 2. Use prefill to skip preamble
msg = client.messages.create(
    model="claude-sonnet-4-20250514",
    max_tokens=64,
    messages=[
        {"role": "user", "content": "Classify sentiment: 'Great product!'"},
        {"role": "assistant", "content": "Sentiment:"}  # Skip "Sure, I'd be happy to..."
    ]
)

# 3. Pre-check token count for large inputs
count = client.messages.count_tokens(
    model="claude-sonnet-4-20250514",
    messages=[{"role": "user", "content": large_document}]
)
if count.input_tokens > 100_000:
    # Chunk or summarize first
    pass

Parallel Requests

typescript
import Anthropic from '@anthropic-ai/sdk';
import PQueue from 'p-queue';

const client = new Anthropic();
const queue = new PQueue({ concurrency: 10 });

// Process multiple prompts in parallel (within rate limits)
const results = await Promise.all(
  prompts.map(p => queue.add(() =>
    client.messages.create({
      model: 'claude-haiku-4-20250514',
      max_tokens: 256,
      messages: [{ role: 'user', content: p }],
    })
  ))
);

Performance Benchmarks

OptimizationLatency ImpactCost Impact
Prompt caching-50% (cached portion)-90% input cost
Haiku over Sonnet-75% TTFT-73% cost
Streaming-80% TTFT (perceived)Same cost
Lower max_tokens-10-30% total timeSame cost
Prefill technique-20% output tokensProportional savings

Prerequisites

  • Define latency, throughput, quality, token, and error SLOs plus the owner-approved model, cache, concurrency, and retry policy.
  • Use pinned model IDs, synthetic prompts, an isolated workspace, and representative non-sensitive fixtures; do not benchmark with customer content or production credentials.
  • Configure aggregate-only telemetry, bounded concurrency, rate-limit awareness, and a tested rollback configuration.

Instructions

  1. Establish a baseline for time-to-first-token, completion latency, tokens, cache hit rate, throughput, quality, and errors using repeated synthetic runs.
  2. Change one lever at a time: model, prompt/cache layout, token budget, streaming, batching, or concurrency. Keep prompt content out of logs and verify cache eligibility for sensitive data before enabling it.
  3. Enforce request scope, max_tokens, timeout, retry, and concurrency limits. Stop the run when rate limits, quality, or data-policy checks fail rather than increasing access or disabling controls.
  4. Canary the selected configuration in a sandbox or internal workspace, compare against baseline, and obtain approval before production rollout. Monitor p95/p99 latency, error rate, token use, and spend.
  5. Restore the prior configuration on regression, invalidate temporary cache/test artifacts according to retention policy, and retain a redacted benchmark receipt.
Show full SKILL.md (158 more words)Show less

Output

Produce a performance receipt containing configuration and model IDs, benchmark fixture class, sample size, latency/throughput/token/cache aggregates, quality and error outcomes, workspace/canary scope, approval, retention, and rollback reference. Exclude prompts, responses, user identifiers, and secrets.

Error Handling

FailureResponse
Rate limit or queue saturationReduce bounded concurrency, honor retry guidance, and stop the canary if the SLO remains breached.
Quality falls after model/token changeRestore the baseline configuration and quarantine the comparison until reviewed.
Cache miss or policy-ineligible contentDisable caching for that path and use the approved uncached flow.
Timeout or streaming disconnectApply bounded retry/idempotency handling, return a safe partial-state result, and investigate without logging content.

Examples

Benchmark 500 synthetic fixture-prompt-* requests in a staging workspace with Haiku and the current route, assert content_logged=0, p99_latency<approved_limit, and quality=pass, then canary the winner to internal traffic. If p99 or error thresholds fail, emit canary=halted; rollback=perf-baseline.

Resources

Next Steps

For cost optimization, see anth-cost-tuning.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/.curated/anth-performance-tuning of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit cfae287

Compare with similar skills

Anth Performance Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Anth Performance Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Anth Performance Tuning this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.9kAutomated safety check: PassMIT
LLM Cost Optimizationsickn33/agentic-awesome-skills47k1 repos~2.5kAutomated safety check: PassMIT
Claude APIkid-sid/claude-spellbook190—~2.7kAutomated safety check: PassMIT
LLM Cost OptimizationBagelHole/DevOps-Security-Agent-Skills1.2k—~2.2kAutomated safety check: PassMIT
LLM Cost Latency Budgetmohitagw15856/pm-claude-skills1.4k—~985Automated safety check: PassMIT
Claude APIloulanyue/awesome-claude-notes2722 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • LLM Cost Optimization

    sickn33/agentic-awesome-skills

    Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Claude API

    kid-sid/claude-spellbook

    A skill your agent uses when building or debugging apps that call the Claude API — implementing tool use, streaming, vision, prompt caching, batch processing, extended thinking, or an agentic loop…

    190 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLM Cost Optimization

    BagelHole/DevOps-Security-Agent-Skills

    Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

    1.2k GitHub stars~2.2k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLM Cost Latency Budget

    mohitagw15856/pm-claude-skills

    Model the cost and latency of an LLM feature before it ships and surprises the bill.

    1.4k GitHub stars~985 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Claude API

    loulanyue/awesome-claude-notes

    Anthropic Claude API patterns for Python and TypeScript. An agent skill from loulanyue/awesome-claude-notes.

    272 GitHub starsUsed in 2 repos~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Claude API Development

    warpdotdev/warp

    Guides building, debugging and tuning apps on the Claude API and Anthropic SDK, including prompt caching, and migrating code between Claude model versions.

    65k GitHub starsUsed in 3 repos~8.2k tokens
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Anth Performance Tuning

What does Anth Performance Tuning do?

Optimize Claude API performance with prompt caching, model selection, streaming, and latency reduction techniques. Anth Performance Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize Claude API performance with prompt caching, model selection, streaming, and latency reduction techniques.

When should I use Anth Performance Tuning?

Anth Performance Tuning fits situations like: experiencing slow responses; optimizing token usage; reducing time-to-first-token in production; with phrases like anthropic performance.

How do I install Anth Performance Tuning in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-performance-tuning -a claude-code`. Or copy the skill folder (skills/.curated/anth-performance-tuning in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/anth-performance-tuning in your project. Claude Code loads it when a task matches its description.

How do I install Anth Performance Tuning in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-performance-tuning -a codex`. Or copy the skill folder (skills/.curated/anth-performance-tuning in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/anth-performance-tuning in your project. Codex loads it when a task matches its description.

Can I use Anth Performance Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-performance-tuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/anth-performance-tuning, .gemini/skills/anth-performance-tuning, .github/skills/anth-performance-tuning and .opencode/skills/anth-performance-tuning in your project.

What does Anth Performance Tuning need to run?

SKILL.md names no scripts, command-line tools or credentials: Anth Performance Tuning is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Grep. Compatibility (from SKILL.md): Designed for Claude Code.

Does Anth Performance Tuning access the network?

SKILL.md names 1 domain. As links in the text: platform.claude.com. This is read from the text; nothing was executed.

Is Anth Performance Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Anth Performance Tuning use?

Anth Performance Tuning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Anth Performance Tuning use?

About 1.9k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Anth Performance Tuning?

Skills that share tags, products or a category with Anth Performance Tuning: LLM Cost Optimization (sickn33/agentic-awesome-skills, 47k stars), Claude API (kid-sid/claude-spellbook, 190 stars), LLM Cost Optimization (BagelHole/DevOps-Security-Agent-Skills, 1.2k stars) and LLM Cost Latency Budget (mohitagw15856/pm-claude-skills, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Anth Performance Tuning?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.