Optimize Anthropic Claude API costs with model routing, prompt caching, batching, and spend monitoring.

MITAuto-check passedAI & LLM Engineering

Install Anth Cost Tuning

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-cost-tuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace anth-cost-tuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/anth-cost-tuning .claude/skills/anth-cost-tuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
anth-cost-tuning
GitHub stars
2.8k
Token cost
~2.1k tokens
SKILL.md length
527 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Optimize Anthropic Claude API costs with model routing, prompt caching, batching, and spend monitoring.

  • Works in 5 steps: Baseline request volume,… → Define a quality floor and route only… → Cap max_tokens, concurrency, retries,… → …
  • Analyzing Claude API billing
  • SKILL.md covers Overview, Pricing Reference (per million…, Cost Calculator and Strategy 1: Model Routing, plus 11 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Anth Cost Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize Anthropic Claude API costs with model routing, prompt caching, batching, and spend monitoring. Use when analyzing Claude API billing, reducing costs, or implementing cost controls and budget alerts. Trigger with phrases like "anthropic cost", "claude billing", "reduce claude spend", "anthropic budget", "claude pricing optimize".

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering LLM cost and token optimization, Model routing and gateways and LLM API integration. It works with Anthropic API. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Analyzing Claude API billing
  • Implementing cost controls and budget alerts
  • With phrases like anthropic cost
  • Reduce claude spend

Example prompts

  • “anthropic cost”
  • “claude billing”
  • “reduce claude spend”
  • “/anth-cost-tuning”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Grep

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Baseline request volume, input/output/cache tokens, latency, quality, and spend by feature using a redacted measurement window. Do not…
  2. Define a quality floor and route only eligible workloads to the least expensive model that meets it. Use prompt caching only for approved…
  3. Cap max_tokens, concurrency, retries, and batch size. Enforce per-feature and per-workspace budgets before requests are sent; fail closed…
  4. Test the proposed policy on synthetic fixtures in a sandbox, then canary it with aggregate cost, quality, latency, error, and rate-limit…
  5. If quality, spend, or policy thresholds regress, disable the new route/cache/batch policy, restore the prior configuration, and retain a…

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.claude.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Anth Cost Tuning loads about 2.1k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 527 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 527 words, ~2,065 tokens.

Download SKILL.mdSave it as .claude/skills/anth-cost-tuning/SKILL.md (or your agent's skills folder).
name
anth-cost-tuning
description
Optimize Anthropic Claude API costs with model routing, prompt caching, batching, and spend monitoring. Use when analyzing Claude API billing, reducing costs, or implementing cost controls and budget alerts. Trigger with phrases like "anthropic cost", "claude billing", "reduce claude spend", "anthropic budget", "claude pricing optimize".
allowed-tools
Read, Write, Edit, Grep
compatibility
Designed for Claude Code
version
1.7.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, ai, anthropic

Anthropic Cost Tuning

Overview

Optimize Claude API spend through model routing, prompt caching, the Message Batches API, and real-time cost tracking. The four biggest levers: model selection (4-19x), prompt caching (10x input), batches (2x), and max_tokens discipline.

Pricing Reference (per million tokens)

ModelInputOutputCache ReadCache Write
Claude Haiku$0.80$4.00$0.08$1.00
Claude Sonnet$3.00$15.00$0.30$3.75
Claude Opus$15.00$75.00$1.50$18.75

Message Batches: 50% off all model pricing for async processing.

Cost Calculator

python
def estimate_cost(
    input_tokens: int,
    output_tokens: int,
    model: str = "claude-sonnet-4-20250514",
    cached_input: int = 0,
    use_batch: bool = False
) -> float:
    pricing = {
        "claude-haiku-4-20250514": {"input": 0.80, "output": 4.00, "cache_read": 0.08},
        "claude-sonnet-4-20250514": {"input": 3.00, "output": 15.00, "cache_read": 0.30},
        "claude-opus-4-20250514": {"input": 15.00, "output": 75.00, "cache_read": 1.50},
    }
    rates = pricing[model]
    uncached_input = input_tokens - cached_input

    cost = (
        uncached_input * rates["input"] +
        cached_input * rates["cache_read"] +
        output_tokens * rates["output"]
    ) / 1_000_000

    if use_batch:
        cost *= 0.5

    return cost

# Example: 10K requests/day, 500 input + 200 output tokens each
daily = estimate_cost(500, 200, "claude-sonnet-4-20250514") * 10_000
print(f"Daily: ${daily:.2f}")      # ~$0.045 * 10K = $450/day
print(f"Monthly: ${daily * 30:.2f}")  # ~$13,500/month

# Same with Haiku + batching
daily_optimized = estimate_cost(500, 200, "claude-haiku-4-20250514", use_batch=True) * 10_000
print(f"Optimized: ${daily_optimized:.2f}/day")  # ~$22/day (20x cheaper)

Strategy 1: Model Routing

python
def route_to_model(task: str, complexity: str) -> str:
    """Route tasks to cheapest adequate model."""
    # Haiku: classification, extraction, yes/no, routing ($0.80/$4)
    if task in ("classify", "extract", "route", "validate"):
        return "claude-haiku-4-20250514"

    # Sonnet: general tasks, code, tool use ($3/$15)
    if complexity in ("low", "medium"):
        return "claude-sonnet-4-20250514"

    # Opus: only for complex reasoning, research ($15/$75)
    return "claude-opus-4-20250514"

Strategy 2: Prompt Caching

python
# Cache system prompts and reference documents (90% input savings)
# Break-even: 2 requests with same cached content
message = client.messages.create(
    model="claude-sonnet-4-20250514",
    max_tokens=256,
    system=[{
        "type": "text",
        "text": large_reference_document,  # 10K+ tokens
        "cache_control": {"type": "ephemeral"}
    }],
    messages=[{"role": "user", "content": user_question}]
)

Strategy 3: Batches for Non-Real-Time

python
# 50% cost reduction for anything that doesn't need immediate response
# Ideal for: summarization pipelines, data extraction, content generation
batch = client.messages.batches.create(requests=[...])  # Up to 100K requests

Strategy 4: Spend Tracking

python
import anthropic
from dataclasses import dataclass, field

@dataclass
class SpendTracker:
    budget_usd: float = 100.0
    spent_usd: float = 0.0
    requests: int = 0

    def track(self, response):
        cost = estimate_cost(
            response.usage.input_tokens,
            response.usage.output_tokens,
            response.model,
            getattr(response.usage, "cache_read_input_tokens", 0)
        )
        self.spent_usd += cost
        self.requests += 1

        if self.spent_usd > self.budget_usd * 0.8:
            print(f"WARNING: 80% budget used (${self.spent_usd:.2f}/${self.budget_usd})")
        if self.spent_usd > self.budget_usd:
            raise RuntimeError(f"Budget exceeded: ${self.spent_usd:.2f}")

tracker = SpendTracker(budget_usd=50.0)

Cost Reduction Checklist

  • Use Haiku for classification/extraction/routing tasks
  • Enable prompt caching for repeated system prompts
  • Use Message Batches for non-real-time processing
  • Set max_tokens to realistic values (not maximum)
  • Use prefill to reduce output preamble tokens
  • Implement spend tracking and budget alerts
  • Monitor via Usage API

Prerequisites

  • Establish an approved budget, billing owner, cost allocation dimensions, and alert thresholds before changing model routing or batch behavior.
  • Use a sandbox workspace, synthetic prompts, pinned model IDs, and a versioned pricing snapshot; confirm current rates in the official pricing documentation before making a forecast.
  • Configure least-privileged credentials and ensure logs/metrics contain token counts and aggregate cost only, never prompt or response content.

Instructions

  1. Baseline request volume, input/output/cache tokens, latency, quality, and spend by feature using a redacted measurement window. Do not make routing changes from a single outlier.
  2. Define a quality floor and route only eligible workloads to the least expensive model that meets it. Use prompt caching only for approved non-sensitive content and batches only where asynchronous completion is acceptable.
  3. Cap max_tokens, concurrency, retries, and batch size. Enforce per-feature and per-workspace budgets before requests are sent; fail closed when a budget or scope check cannot be evaluated.
  4. Test the proposed policy on synthetic fixtures in a sandbox, then canary it with aggregate cost, quality, latency, error, and rate-limit monitoring. Require owner approval before broader rollout.
  5. If quality, spend, or policy thresholds regress, disable the new route/cache/batch policy, restore the prior configuration, and retain a redacted comparison receipt.
Show full SKILL.md (181 more words)Show less

Output

Produce a cost-control receipt containing the pricing snapshot date, policy version, model/batch/cache decisions, token aggregates, projected and observed spend, quality and latency results, budget outcome, canary scope, approval, and rollback reference. Exclude prompt/response text, customer identifiers, API keys, and raw billing exports.

Error Handling

FailureResponse
Unknown model price or usage fieldStop forecasting, refresh the official pricing/usage source, and mark the estimate provisional.
Budget or quota exceededReject or queue new work, alert the owner, and do not bypass the guard with another key or workspace.
Quality regression after cheaper routingRestore the prior route, quarantine affected output, and rerun the quality fixture before another canary.
Cache or batch unsuitable for data/latency policyDisable that optimization and use the approved synchronous, non-cached path.

Examples

Evaluate 1,000 synthetic classification prompts in a sandbox with a fixed budget, compare pinned Sonnet against Haiku plus an approved batch policy, assert customer_content_logged=0, and emit budget=within_limit; quality=pass; canary=internal; rollback=route-v1. Do not use live customer prompts to tune pricing.

Resources

Next Steps

For architecture patterns, see anth-reference-architecture.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/.curated/anth-cost-tuning of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit cfae287

Compare with similar skills

Anth Cost Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Anth Cost Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Anth Cost Tuning this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2.1kAutomated safety check: PassMIT
Cost Aware LLM PipelineaAAaqwq/AGI-Super-Team1055 repos~1.4kAutomated safety check: PassMIT
Claude APIkid-sid/claude-spellbook190—~2.7kAutomated safety check: PassMIT
Using Ccproxy Inspectorstarbaser/ccproxy350—~2.7kAutomated safety check: PassCustom licence
Using Ccproxy APIstarbaser/ccproxy350—~4kAutomated safety check: PassCustom licence
Claude APIloulanyue/awesome-claude-notes2722 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • Cost Aware LLM Pipeline

    aAAaqwq/AGI-Super-Team

    Cost optimization patterns for LLM API usage — model routing by task complexity, budget tracking, retry logic, and prompt caching.

    105 GitHub starsUsed in 5 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Claude API

    kid-sid/claude-spellbook

    A skill your agent uses when building or debugging apps that call the Claude API — implementing tool use, streaming, vision, prompt caching, batch processing, extended thinking, or an agentic loop…

    190 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Using Ccproxy Inspector

    starbaser/ccproxy

    Operates the ccproxy inspector MITM system for intercepting, inspecting, and transforming LLM API traffic.

    350 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Using Ccproxy API

    starbaser/ccproxy

    Guides users through ccproxy as an OpenAI-compatible and Anthropic-compatible LLM API server with SDK integration, OAuth authentication, sentinel key substitution, model routing, and troubleshooting.

    350 GitHub stars~4k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Claude API

    loulanyue/awesome-claude-notes

    Anthropic Claude API patterns for Python and TypeScript. An agent skill from loulanyue/awesome-claude-notes.

    272 GitHub starsUsed in 2 repos~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Claude API Development

    warpdotdev/warp

    Guides building, debugging and tuning apps on the Claude API and Anthropic SDK, including prompt caching, and migrating code between Claude model versions.

    65k GitHub starsUsed in 3 repos~8.2k tokens
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Anth Cost Tuning

What does Anth Cost Tuning do?

Optimize Anthropic Claude API costs with model routing, prompt caching, batching, and spend monitoring. Anth Cost Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize Anthropic Claude API costs with model routing, prompt caching, batching, and spend monitoring.

When should I use Anth Cost Tuning?

Anth Cost Tuning fits situations like: analyzing Claude API billing; implementing cost controls and budget alerts; with phrases like anthropic cost; reduce claude spend.

How do I install Anth Cost Tuning in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-cost-tuning -a claude-code`. Or copy the skill folder (skills/.curated/anth-cost-tuning in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/anth-cost-tuning in your project. Claude Code loads it when a task matches its description.

How do I install Anth Cost Tuning in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-cost-tuning -a codex`. Or copy the skill folder (skills/.curated/anth-cost-tuning in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/anth-cost-tuning in your project. Codex loads it when a task matches its description.

Can I use Anth Cost Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-cost-tuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/anth-cost-tuning, .gemini/skills/anth-cost-tuning, .github/skills/anth-cost-tuning and .opencode/skills/anth-cost-tuning in your project.

What does Anth Cost Tuning need to run?

SKILL.md names no scripts, command-line tools or credentials: Anth Cost Tuning is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Grep. Compatibility (from SKILL.md): Designed for Claude Code.

Does Anth Cost Tuning access the network?

SKILL.md names 1 domain. As links in the text: platform.claude.com. This is read from the text; nothing was executed.

Is Anth Cost Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Anth Cost Tuning use?

Anth Cost Tuning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Anth Cost Tuning use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Anth Cost Tuning?

Skills that share tags, products or a category with Anth Cost Tuning: Cost Aware LLM Pipeline (aAAaqwq/AGI-Super-Team, 105 stars), Claude API (kid-sid/claude-spellbook, 190 stars), Using Ccproxy Inspector (starbaser/ccproxy, 350 stars) and Using Ccproxy API (starbaser/ccproxy, 350 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Anth Cost Tuning?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.