Agent skill

Clade Performance Tuning

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Optimize Anthropic API latency — streaming, prompt caching, model selection, Use when working with performance-tuning patterns.

MITAuto-check passedAI & LLM Engineering

Install Clade Performance Tuning

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill clade-performance-tuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace clade-performance-tuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/clade-performance-tuning .claude/skills/clade-performance-tuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
clade-performance-tuning
GitHub stars
2.8k
Token cost
~1.3k tokens
SKILL.md length
217 words
Files
2 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Optimize Anthropic API latency — streaming, prompt caching, model selection, Use when working with performance-tuning patterns.

  • Works in 6 steps: Always Stream → Prompt Caching — Faster TTFT → Use Haiku for Speed-Critical Paths → …
  • Working with performance-tuning patterns
  • SKILL.md covers Overview, Latency Benchmarks (approximate), Optimization Strategies and Instructions, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Clade Performance Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize Anthropic API latency — streaming, prompt caching, model selection, Use when working with performance-tuning patterns. connection reuse, and parallel requests. Trigger with "anthropic slow", "claude latency", "speed up anthropic", "anthropic performance", "claude response time".

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/one-pager.md`). Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering LLM cost and token optimization. It works with Anthropic API. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Working with performance-tuning patterns
  • With anthropic slow
  • Speed up anthropic
  • Anthropic performance

Example prompts

  • “anthropic slow”
  • “claude latency”
  • “speed up anthropic”
  • “/clade-performance-tuning”

Requirements

  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Always Stream
  2. Prompt Caching — Faster TTFT
  3. Use Haiku for Speed-Critical Paths
  4. Reuse Client Instance
  5. Parallel Requests
  6. Minimize Output Tokens

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.claude.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Clade Performance Tuning loads about 1.3k tokens when it runs, and up to ~1.8k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 217 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 217 words, ~1,332 tokens.

Download SKILL.mdSave it as .claude/skills/clade-performance-tuning/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
clade-performance-tuning
description
Optimize Anthropic API latency — streaming, prompt caching, model selection, Use when working with performance-tuning patterns. connection reuse, and parallel requests. Trigger with "anthropic slow", "claude latency", "speed up anthropic", "anthropic performance", "claude response time".
allowed-tools
Read, Write, Edit
compatibility
Designed for Claude Code
version
1.1.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, anthropic, claude, performance, latency

Anthropic Performance Tuning

Overview

Claude latency has two components: time to first token (TTFT) and tokens per second (TPS). Different strategies target each.

Latency Benchmarks (approximate)

ModelTTFT (p50)TTFT (p95)Output TPS
Claude Haiku 4.5200ms600ms~150
Claude Sonnet 4400ms1.2s~90
Claude Opus 4800ms2.5s~40

Optimization Strategies

Instructions

Step 1: Always Stream
typescript
// Streaming delivers the first token ASAP — user sees response instantly
// instead of waiting for the full response to generate

const stream = client.messages.stream({
  model: 'claude-sonnet-4-20250514',
  max_tokens: 1024,
  messages,
});

// First token arrives in ~400ms (Sonnet)
// Full response may take 5-10s, but user sees progress immediately
for await (const event of stream) {
  if (event.type === 'content_block_delta') {
    yield event.delta.text;
  }
}
Step 2: Prompt Caching — Faster TTFT
typescript
// Cached prompts skip re-processing — dramatically lower TTFT for large system prompts
const message = await client.messages.create({
  model: 'claude-sonnet-4-20250514',
  max_tokens: 1024,
  system: [{
    type: 'text',
    text: largeSystemPrompt, // 10K+ tokens
    cache_control: { type: 'ephemeral' },
  }],
  messages,
}, {
  headers: { 'claude-beta': 'prompt-caching-2024-07-31' },
});
// TTFT drops from ~2s to ~500ms on cache hit with large prompts
Step 3: Use Haiku for Speed-Critical Paths
typescript
// Haiku is 2-4x faster than Sonnet with 80% quality for many tasks
// Use for: classification, extraction, simple Q&A, routing decisions

const route = await client.messages.create({
  model: 'claude-haiku-4-5-20251001', // 200ms TTFT
  max_tokens: 10,
  system: 'Classify the intent. Reply with exactly one word: search, create, update, delete.',
  messages: [{ role: 'user', content: userInput }],
});

// Then use Sonnet/Opus for the actual task
Step 4: Reuse Client Instance
typescript
// BAD — creates new connection pool per request
app.get('/api/chat', async (req, res) => {
  const client = new Anthropic(); // DON'T
  // ...
});

// GOOD — single client shared across requests
const client = new Anthropic(); // Module-level singleton

app.get('/api/chat', async (req, res) => {
  const message = await client.messages.create({ ... });
  // ...
});
Step 5: Parallel Requests
typescript
// When you need multiple independent Claude calls, fire them in parallel
const [summary, sentiment, entities] = await Promise.all([
  client.messages.create({ model: 'claude-haiku-4-5-20251001', max_tokens: 200,
    messages: [{ role: 'user', content: `Summarize: ${text}` }] }),
  client.messages.create({ model: 'claude-haiku-4-5-20251001', max_tokens: 20,
    messages: [{ role: 'user', content: `Sentiment (positive/negative/neutral): ${text}` }] }),
  client.messages.create({ model: 'claude-haiku-4-5-20251001', max_tokens: 200,
    messages: [{ role: 'user', content: `Extract named entities from: ${text}` }] }),
]);
Step 6: Minimize Output Tokens
typescript
// Fewer output tokens = faster response
system: 'Be extremely concise. Use bullet points, not paragraphs.',

// Set tight max_tokens
max_tokens: 256, // Don't use 4096 for short answers

Output

  • Streaming enabled for all user-facing responses (first token in ~400ms with Sonnet)
  • Prompt caching reducing TTFT for large system prompts
  • Model routing to Haiku for speed-critical classification/routing tasks
  • Client instance reused across requests (no per-request connection overhead)
  • Parallel requests firing independent Claude calls concurrently

Error Handling

IssueCauseFix
TTFT > 3sLarge uncached promptEnable prompt caching
Slow outputUsing Opus for simple tasksDowngrade to Haiku/Sonnet
TimeoutsLong generation + default timeoutnew Anthropic({ timeout: 120_000 })
529 overloadedAPI capacitySDK auto-retries; add fallback model

Examples

See Latency Benchmarks table and six numbered strategy sections above, each with complete TypeScript code examples.

Resources

Next Steps

See clade-deploy-integration for production deployment patterns.

Prerequisites

  • Completed clade-install-auth
  • User-facing application where latency matters
  • Understanding of streaming and async patterns

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/.curated/clade-performance-tuning of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/one-pager.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Clade Performance Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Clade Performance Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Clade Performance Tuning this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.3kAutomated safety check: PassMIT
Token Optimizationcwinvestments/memstack423—~1.4kAutomated safety check: PassProprietary
Claude APIkid-sid/claude-spellbook190—~2.7kAutomated safety check: PassMIT
Prompt Caching Patternssoftspark/ai-toolkit179—~1.1kAutomated safety check: PassApache-2.0
Claude APIloulanyue/awesome-claude-notes2722 repos~2.1kAutomated safety check: PassMIT
Claude API Developmentwarpdotdev/warp65k3 repos~8.2kAutomated safety check: PassApache-2.0

Similar skills

  • Token Optimization

    cwinvestments/memstack

    A skill your agent uses when the user says 'token optimization', 'save tokens', 'context window', 'reduce tokens', 'token stack', or 'TokenStack', or asks about extending context window capacity.

    423 GitHub stars~1.4k tokensUpdated 14 days ago
    AI & LLM EngineeringAuto-check passed
  • Claude API

    kid-sid/claude-spellbook

    A skill your agent uses when building or debugging apps that call the Claude API — implementing tool use, streaming, vision, prompt caching, batch processing, extended thinking, or an agentic loop…

    190 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Prompt Caching Patterns

    softspark/ai-toolkit

    Anthropic API prompt caching: TTL, breakpoints, stacking, invalidation, hit rate.

    179 GitHub stars~1.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Claude API

    loulanyue/awesome-claude-notes

    Anthropic Claude API patterns for Python and TypeScript. An agent skill from loulanyue/awesome-claude-notes.

    272 GitHub starsUsed in 2 repos~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Claude API Development

    warpdotdev/warp

    Guides building, debugging and tuning apps on the Claude API and Anthropic SDK, including prompt caching, and migrating code between Claude model versions.

    65k GitHub starsUsed in 3 repos~8.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Context Compression

    guanyang/open-agent-hub

    This skill should be used when long-running agent sessions need context compression, structured summarization, compaction, token-per-task optimization, or durable handoff summaries that preserve…

    977 GitHub starsUsed in 2 repos~4.6k tokens
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Clade Performance Tuning

What does Clade Performance Tuning do?

Optimize Anthropic API latency — streaming, prompt caching, model selection, Use when working with performance-tuning patterns. Clade Performance Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize Anthropic API latency — streaming, prompt caching, model selection, Use when working with performance-tuning patterns.

When should I use Clade Performance Tuning?

Clade Performance Tuning fits situations like: working with performance-tuning patterns; with anthropic slow; speed up anthropic; anthropic performance.

How do I install Clade Performance Tuning in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill clade-performance-tuning -a claude-code`. Or copy the skill folder (skills/.curated/clade-performance-tuning in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/clade-performance-tuning in your project. Claude Code loads it when a task matches its description.

How do I install Clade Performance Tuning in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill clade-performance-tuning -a codex`. Or copy the skill folder (skills/.curated/clade-performance-tuning in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/clade-performance-tuning in your project. Codex loads it when a task matches its description.

Can I use Clade Performance Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill clade-performance-tuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/clade-performance-tuning, .gemini/skills/clade-performance-tuning, .github/skills/clade-performance-tuning and .opencode/skills/clade-performance-tuning in your project.

What does Clade Performance Tuning need to run?

SKILL.md names no scripts, command-line tools or credentials: Clade Performance Tuning is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Write, Edit. Compatibility (from SKILL.md): Designed for Claude Code.

Does Clade Performance Tuning access the network?

SKILL.md names 1 domain. As links in the text: platform.claude.com. This is read from the text; nothing was executed.

Is Clade Performance Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Clade Performance Tuning use?

Clade Performance Tuning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Clade Performance Tuning use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 474 tokens, read only when the agent opens those files.

What are the alternatives to Clade Performance Tuning?

Skills that share tags, products or a category with Clade Performance Tuning: Token Optimization (cwinvestments/memstack, 423 stars), Claude API (kid-sid/claude-spellbook, 190 stars), Prompt Caching Patterns (softspark/ai-toolkit, 179 stars) and Claude API (loulanyue/awesome-claude-notes, 272 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Clade Performance Tuning?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.