Agent skill

Prompt Caching Patterns

by softspark in softspark/ai-toolkit

Anthropic API prompt caching: TTL, breakpoints, stacking, invalidation, hit rate.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Prompt Caching Patterns

skills CLI
$ npx skills add softspark/ai-toolkit --skill prompt-caching-patterns -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install softspark/ai-toolkit prompt-caching-patterns --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/app/skills/prompt-caching-patterns .claude/skills/prompt-caching-patterns && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
prompt-caching-patterns
GitHub stars
179
Token cost
~1.1k tokens
SKILL.md length
414 words
Files
1
Skills in repo
112
Repo updated
First seen
Licence
Apache-2.0

At a glance

Anthropic API prompt caching: TTL, breakpoints, stacking, invalidation, hit rate.

  • Tasks that involve LLM cost and token optimization
  • SKILL.md covers Dated cache reference…, Prefix design, Explicit caching example and Invalidation, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Responsive design

What it does

Prompt Caching Patterns is an agent skill from softspark/ai-toolkit. Anthropic API prompt caching: TTL, breakpoints, stacking, invalidation, hit rate. Triggers: prompt caching, cachecontrol, cache breakpoint, cache TTL, hit rate.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM cost and token optimization and Responsive design. It works with Anthropic API. The repository describes itself as: Professional-grade AI coding toolkit: 94 skills, 44 agents, multi-platform (Claude, Cursor, Windsurf, Copilot, Gemini, Cline, Roo Code, Aider, Augment, Antigravity, Codex CLI… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM cost and token optimization
  • Tasks that involve Responsive design

Example prompts

  • “/prompt-caching-patterns”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read

What it can do on your machine

Read from SKILL.md and the folder at commit d64db2b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.claude.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prompt Caching Patterns loads about 1.1k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 414 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from softspark/ai-toolkit at commit d64db2b, republished under its Apache-2.0 licence (© softspark). 414 words, ~1,133 tokens.

Download SKILL.mdSave it as .claude/skills/prompt-caching-patterns/SKILL.md (or your agent's skills folder).
name
prompt-caching-patterns
description
Anthropic API prompt caching: TTL, breakpoints, stacking, invalidation, hit rate. Triggers: prompt caching, cache_control, cache breakpoint, cache TTL, hit rate.
allowed-tools
Read
effort
medium
user-invocable
false

Prompt Caching Patterns

Cache repeated prefixes when reuse offsets write costs. Cache eligibility and pricing depend on the model and platform; a short system prompt is not cached merely because it repeats.

Dated cache reference (2026-09-23)

ModelMinimum eligible prefixCache read / base input price
Claude Opus 5.5512 tokens5%
Claude Fable 5.1512 tokens2.5%
Claude Sonnet 51024 tokens10%
Claude Haiku 4.54096 tokens10%

For the Claude API, a five-minute write costs 1.25 times base input and a one-hour write costs 2 times base input. Recheck current pricing before budgeting; provider-specific billing and model availability can differ.

Prefix design

Cache order is tools → system → messages, regardless of the order of request keys. An explicit breakpoint includes the marked block and everything before it. Keep dynamic material after the stable prefix.

text
[ tool definitions              ] breakpoint 1
[ reusable system instructions  ] breakpoint 2
[ reference documents           ] breakpoint 3
[ stable conversation prefix    ] breakpoint 4
[ current variable content      ]

There are at most four breakpoints. Top-level automatic cache_control moves a breakpoint to the last eligible block and consumes one slot. Explicit markers give control over a static prefix.

Explicit caching example

The caller supplies the approved model, text and output limit. Marking a prefix below its model's minimum silently produces no cache entry.

python
def cached_answer(client, model, policy, document, question, max_tokens):
    return client.messages.create(
        model=model,
        max_tokens=max_tokens,
        system=[{
            "type": "text",
            "text": policy,
            "cache_control": {"type": "ephemeral"},
        }],
        messages=[{
            "role": "user",
            "content": [
                {"type": "text", "text": document,
                 "cache_control": {"type": "ephemeral"}},
                {"type": "text", "text": question},
            ],
        }],
    )

Use {"type": "ephemeral", "ttl": "1h"} for an approved one-hour write. When mixing TTLs, put longer-lived breakpoints before shorter-lived ones. Do not send paid heartbeat requests simply to keep an unused prefix warm.

Show full SKILL.md (193 more words)Show less

Invalidation

Changing tools invalidates subsequent system and message prefixes; changing the system invalidates subsequent messages. Top-level effort changes invalidate message cache blocks and can affect earlier blocks depending on the model. Supported per-message effort updates preserve earlier prefixes. Changing the model is not a promise of cross-model cache reuse.

A static string passed as system alone does not enable caching: configure cache_control at the request or content-block level. Keep tool definitions, document serialization and stable instructions deterministic.

Measuring reuse

Include writes when calculating the fraction of input served from cache.

python
def cache_read_fraction(usage):
    read = usage.cache_read_input_tokens or 0
    written = usage.cache_creation_input_tokens or 0
    uncached = usage.input_tokens or 0
    total = read + written + uncached
    return read / total if total else 0.0

Record write/read counts and actual costs across cold and warm requests. Choose a target from observed reuse; a single universal hit-rate threshold is misleading. Both cache counters remaining zero can indicate an ineligible prefix.

When not to cache

Skip cache writes when no prefix will be reused before expiry or when measured cost exceeds uncached requests. Do not pad prompts with irrelevant content merely to reach a minimum. A one-hour TTL may fit intermittent reuse better than five minutes, within the approved cost policy.

Reviewed 2026-09-23:

Use model-routing-patterns for route evaluation and llm-ops-engineer for application operations.

© softspark, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in app/skills/prompt-caching-patterns of softspark/ai-toolkit.

Open the folder on GitHubat commit d64db2b

Compare with similar skills

Prompt Caching Patterns next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prompt Caching Patterns compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prompt Caching Patterns this skillsoftspark/ai-toolkit179—~1.1kAutomated safety check: PassApache-2.0
Commerce Prompt Cachinganthropics/commerce-agents3.2k—~2kAutomated safety check: PassApache-2.0
Prompt CachingArchive228/loopkit755—~735Automated safety check: PassMIT
Token Optimizationcwinvestments/memstack423—~1.4kAutomated safety check: PassProprietary
Claude API In Prototypesasgeirtj/system_prompts_leaks69k—~281Automated safety check: PassCC0-1.0
Claude APIkid-sid/claude-spellbook189—~2.7kAutomated safety check: PassMIT

Similar skills

  • Commerce Prompt Caching

    anthropics/commerce-agents

    Official

    The reference agents' cache-stable request assembly, covering the static system and per-request context split, the fixed tool list, the rolling conversation breakpoint, which config fields are…

    3.2k GitHub stars~2k tokensUpdated 5 days ago
    Frontend & DesignAuto-check passed
  • Prompt Caching

    Archive228/loopkit

    Cache the parts of the prompt that don't change so a long-running loop stops paying full price on every turn.

    755 GitHub stars~735 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Token Optimization

    cwinvestments/memstack

    A skill your agent uses when the user says 'token optimization', 'save tokens', 'context window', 'reduce tokens', 'token stack', or 'TokenStack', or asks about extending context window capacity.

    423 GitHub stars~1.4k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • Claude API In Prototypes

    asgeirtj/system_prompts_leaks

    Call Claude from your HTML artifacts via window.claude.complete

    69k GitHub stars~281 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Claude API

    kid-sid/claude-spellbook

    A skill your agent uses when building or debugging apps that call the Claude API — implementing tool use, streaming, vision, prompt caching, batch processing, extended thinking, or an agentic loop…

    189 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Claude API Development

    warpdotdev/warp

    Guides building, debugging and tuning apps on the Claude API and Anthropic SDK, including prompt caching, and migrating code between Claude model versions.

    65k GitHub starsUsed in 3 repos~8.2k tokens
    AI & LLM EngineeringAuto-check passed

More from softspark/ai-toolkit

All 112 skills in this repo
  • Prepare Test Env

    softspark/ai-toolkit

    Prepare or verify a project QA environment with source identity, readiness, browser access, evidence paths and owned cleanup.

    179 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check: notes
  • A11y Validate

    softspark/ai-toolkit

    Accessibility validator: WCAG 2.1 AA, EN 301 549, EAA. An agent skill from softspark/ai-toolkit.

    179 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check: notes
  • Analyze

    softspark/ai-toolkit

    Analyzes code quality, complexity, patterns across codebase.

    179 GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • Autonomous Dev

    softspark/ai-toolkit

    Drives a brief, specification, issue or existing PR through implementation, review, tests and QA to a ready PR.

    179 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Brand Voice

    softspark/ai-toolkit

    Direct technical voice for docs, README, user-facing text. An agent skill from softspark/ai-toolkit.

    179 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • CI

    softspark/ai-toolkit

    Detect/generate/debug CI pipeline config (GitHub Actions, GitLab CI).

    179 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check: notes

Works with

Questions about Prompt Caching Patterns

What does Prompt Caching Patterns do?

Anthropic API prompt caching: TTL, breakpoints, stacking, invalidation, hit rate. Prompt Caching Patterns is an agent skill from softspark/ai-toolkit. Anthropic API prompt caching: TTL, breakpoints, stacking, invalidation, hit rate.

When should I use Prompt Caching Patterns?

Prompt Caching Patterns fits situations like: tasks that involve LLM cost and token optimization; tasks that involve Responsive design.

How do I install Prompt Caching Patterns in Claude Code?

Run `npx skills add softspark/ai-toolkit --skill prompt-caching-patterns -a claude-code`. Or copy the skill folder (app/skills/prompt-caching-patterns in softspark/ai-toolkit) into .claude/skills/prompt-caching-patterns in your project. Claude Code loads it when a task matches its description.

How do I install Prompt Caching Patterns in Codex?

Run `npx skills add softspark/ai-toolkit --skill prompt-caching-patterns -a codex`. Or copy the skill folder (app/skills/prompt-caching-patterns in softspark/ai-toolkit) into .agents/skills/prompt-caching-patterns in your project. Codex loads it when a task matches its description.

Can I use Prompt Caching Patterns in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add softspark/ai-toolkit --skill prompt-caching-patterns -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prompt-caching-patterns, .gemini/skills/prompt-caching-patterns, .github/skills/prompt-caching-patterns and .opencode/skills/prompt-caching-patterns in your project.

What does Prompt Caching Patterns need to run?

SKILL.md names no scripts, command-line tools or credentials: Prompt Caching Patterns is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read.

Does Prompt Caching Patterns access the network?

SKILL.md names 1 domain. As links in the text: platform.claude.com. This is read from the text; nothing was executed.

Is Prompt Caching Patterns safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prompt Caching Patterns use?

Prompt Caching Patterns is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prompt Caching Patterns use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Prompt Caching Patterns?

Skills that share tags, products or a category with Prompt Caching Patterns: Commerce Prompt Caching (anthropics/commerce-agents, 3.2k stars), Prompt Caching (Archive228/loopkit, 755 stars), Token Optimization (cwinvestments/memstack, 423 stars) and Claude API In Prototypes (asgeirtj/system_prompts_leaks, 69k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prompt Caching Patterns?

softspark (a GitHub user) maintains it in softspark/ai-toolkit, which has 179 GitHub stars. The repository holds 112 skills in this directory. The repository was last updated on October 7, 2026.

Source: softspark/ai-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.