Agent skill

Prompt Caching Strategy

by affaan-m in affaan-m/ECC

Optimize LLM prompt caching hit rate to reduce API costs and improve latency.

MITAuto-check passedAI & LLM Engineering

Install Prompt Caching Strategy

skills CLI
$ npx skills add affaan-m/ECC --skill prompt-caching-strategy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC prompt-caching-strategy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/prompt-caching-strategy .claude/skills/prompt-caching-strategy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
prompt-caching-strategy
GitHub stars
276k
Token cost
~2.5k tokens
SKILL.md length
989 words
Files
1
Skills in repo
683
Repo updated
First seen
Licence
MIT

At a glance

Optimize LLM prompt caching hit rate to reduce API costs and improve latency.

  • Works in 5 steps: Capture eligible prefix size and… → Run the offline example and edge cases… → If provider traffic is authorized,… → …
  • : user says prompt cache
  • SKILL.md covers When to Activate, How It Works, Provider Requirements and… and Cost Comparison Calculator, plus 3 more sections
  • Calls node

What it does

Prompt Caching Strategy is an agent skill from affaan-m/ECC. Optimize LLM prompt caching hit rate to reduce API costs and improve latency. Analyzes prompt structure, identifies cacheable vs non-cacheable content, and recommends restructuring to maximize cache reuse. TRIGGER when: user says "prompt cache", "cache hit rate", "reduce LLM cost", "optimize token usage", "cheaper LLM calls", or asks about prompt caching strategies for OpenAI/Anthropic/other providers. DO NOT TRIGGER when: user just wants to optimize a single prompt's quality (use prompt-optimizer instead), or…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM cost and token optimization. It works with OpenAI. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • : user says prompt cache
  • Reduce LLM cost
  • Optimize token usage
  • Cheaper LLM calls

Example prompts

  • “prompt cache”
  • “cache hit rate”
  • “reduce LLM cost”
  • “/prompt-caching-strategy”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Capture eligible prefix size and model/platform requirements from current docs.
  2. Run the offline example and edge cases with
  3. If provider traffic is authorized, compare matched requests and aggregate
  4. Report cached reads divided by total input tokens as a token hit rate; report
  5. If no traffic was measured, label hit rate, latency improvement and realized

What it can do on your machine

Read from SKILL.md and the folder at commit 4eb71d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • developers.openai.com
    • ai.google.dev
    • platform.claude.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prompt Caching Strategy loads about 2.5k tokens when it runs. Until then it costs about 146 tokens; SKILL.md has 989 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit 4eb71d9, republished under its MIT licence (© affaan-m). 989 words, ~2,528 tokens.

Download SKILL.mdSave it as .claude/skills/prompt-caching-strategy/SKILL.md (or your agent's skills folder).
name
prompt-caching-strategy
description
Optimize LLM prompt caching hit rate to reduce API costs and improve latency. Analyzes prompt structure, identifies cacheable vs non-cacheable content, and recommends restructuring to maximize cache reuse. TRIGGER when: user says "prompt cache", "cache hit rate", "reduce LLM cost", "optimize token usage", "cheaper LLM calls", or asks about prompt caching strategies for OpenAI/Anthropic/other providers. DO NOT TRIGGER when: user just wants to optimize a single prompt's quality (use prompt-optimizer instead), or asks about general caching infrastructure.
metadata.origin
community
metadata.author
cyberspace-cs
metadata.version
1.0.0

Prompt Caching Strategy

Arrange reusable prompt content to improve cache reuse, then measure input cost and latency. Savings depend on model, platform, cache eligibility, writes, retention, reuse and output volume. Runtime cache hits and savings for this skill are unmeasured; the examples below validate arithmetic without provider calls.

When to Activate

Use for repeated API requests that share authorized instructions, tools or reference material. First identify the exact provider endpoint, model, service tier and caching mode. Consult its current documentation before selecting rates, minimum prefix length or retention. Caching does not reduce context-window use.

How It Works

Preserve a stable prefix

Keep reusable instructions, tool schemas and examples stable; place changing questions, timestamps and retrieved results after that content where the API allows. Respect the provider's serialization order: Claude uses tools, system, then messages. Matching prefixes need identical content, including tool schemas and relevant request settings. A changed prefix can prevent reuse from that point. A dynamic suffix may itself be cached later if reused unchanged.

text
Reusable prefix: approved tools, instructions, stable examples/reference
Changing suffix: current question, fresh retrieval results, request metadata

Choose only relevant context. Padding prompts or broadening access just to meet an eligibility threshold needs its own cost and quality justification.

Preserve user and workspace boundaries

Authorize content access before cache lookup or reuse. Scope application-managed cache handles and accounting keys to the approved user, workspace, provider account, model and content revision. Share only material explicitly approved for that scope. A cache key is a routing/accounting hint, not an authorization gate or a sandbox. Provider cache isolation does not replace application access checks. Invalidate application references when permissions or content change; follow provider retention/deletion controls. Avoid logging private prompt contents. Context selection, capabilities, sandbox enforcement and evidence remain independent of cache optimization.

Provider Requirements and Usage Evidence

The linked documentation is authoritative; endpoint and SDK versions can expose different usage field names. Record that version alongside counters.

ProviderRequirements to verifyCounters to record
OpenAIModel-specific automatic/explicit modes, minimum length, retention and write/read prices; earlier and newer models differResponses usage.input_tokens_details.cached_tokens and, when present, cache_write_tokens; input and output totals. Chat Completions uses prompt_tokens_details.cached_tokens
ClaudeSupported platform, model minimum, cache_control support and TTL; cache writes may cost more than ordinary inputinput_tokens, cache_creation_input_tokens, cache_read_input_tokens, output_tokens; split writes by TTL when rates differ
GeminiModel minimum and API: Interactions supports implicit caching; explicit cache objects require GenerateContentInteractions usage.total_cached_tokens; GenerateContent usage_metadata cache/input/output counters plus explicit-cache token count and retained duration

Use OpenAI's prompt caching guide and pricing for the selected model. For Chat Completions, check the usage schema. Derive ordinary input from the total minus cached reads and any separately reported writes; do not count the same tokens twice.

Use Claude's prompt caching guide for current model/platform limits and TTL rates. Its input_tokens counts ordinary input separately from cache reads and writes. A short prefix can produce zero cache counters even when marked for caching. Do not assume every Claude model or hosted platform uses the same minimum or read price.

Use Gemini's caching guide, GenerateContent caching guide and pricing. Explicit-cache storage is charged by token count and retained duration, separately from reads, ordinary input and output. Implicit hits are not guaranteed. Verify billing categories for the chosen endpoint rather than copying another provider's rates.

Cost Comparison Calculator

Normalize actual usage into nonoverlapping categories and use prices in the same currency per million tokens. Split write categories when TTL, modality, tier or model prices differ. Storage uses million-token-hours; other applicable charges such as tools and grounding need their own lines. Include all warmup, miss, expiration, refresh and retry requests in the measured period.

text
baseline = (ordinary + written + read) * ordinary_input_rate / 1e6
           + output * output_rate / 1e6
actual   = ordinary * ordinary_input_rate / 1e6
           + written * cache_write_rate / 1e6
           + read * cache_read_rate / 1e6
           + output * output_rate / 1e6
           + storage_token_hours * storage_rate / 1e6
savings  = baseline - actual

This counterfactual holds prompt and output volumes fixed. For explicit-cache creation, include any separately billed creation input in the write category; use the endpoint's documented rate. Set nonexistent categories to zero. Compare actual before/after invoices separately if warmup requests or output differ. Never apply a cache discount to the total bill or to output tokens.

The executable calculator takes normalized counters, not raw provider responses:

javascript
function compareCacheCost(usage, rates) {
  const counts = ['uncached', 'written', 'read', 'output', 'storageTokenHours'];
  const prices = ['input', 'write', 'read', 'output', 'storage'];
  for (const [values, keys] of [[usage, counts], [rates, prices]]) {
    for (const key of keys) {
      if (!Number.isFinite(values[key]) || values[key] < 0) {
        throw new Error(`${key} must be a finite nonnegative number`);
      }
    }
  }
  const outputCost = usage.output * rates.output / 1e6;
  const storageCost = usage.storageTokenHours * rates.storage / 1e6;
  const baseline = (usage.uncached + usage.written + usage.read)
    * rates.input / 1e6 + outputCost;
  const actual = (usage.uncached * rates.input + usage.written * rates.write
    + usage.read * rates.read) / 1e6 + outputCost + storageCost;
  return { baseline, actual, savings: baseline - actual, outputCost, storageCost };
}
Show full SKILL.md (336 more words)Show less

Examples

These are hypothetical prices and counts, not provider quotes or measured hits. Ten calls each contain a 10,000-token stable prefix, 2,000-token changing suffix and 2,000 output tokens. Assume one prefix write and nine full prefix reads.

CategoryTokensExample rate per millionCost
Ordinary input20,000$3.00$0.0600
Cache write10,000$3.75$0.0375
Cache read90,000$0.30$0.0270
Output20,000$15.00$0.3000
Storage0 token-hours$0.00$0.0000

Baseline: 120,000 input tokens at $3 plus the same output = $0.6600. With caching: $0.4245, savings $0.2355. The write premium is included; the $0.3000 output bill remains unchanged. An output-heavy workload has the same absolute input savings under these assumptions, a smaller fraction of total cost.

If an explicit cache retains 10,000 tokens for two hours at a hypothetical $1 per million-token-hour, add $0.0200 storage. Actual cost becomes $0.4445. Use the selected Gemini endpoint's prices for a real estimate.

For a one-off call, that prefix write costs $0.0375 rather than $0.0300 ordinary input: $0.0075 more, with no read savings. For a 500-token prefix below a selected model's documented minimum, treat the request as uncached; claiming future hits without counters would overstate savings. Repeated expiry or changing prefixes can also turn expected savings into higher cost.

Verification Workflow

  1. Capture eligible prefix size and model/platform requirements from current docs.
  2. Run the offline example and edge cases with node tests/skills/prompt-caching-strategy.test.js in the ECC repository.
  3. If provider traffic is authorized, compare matched requests and aggregate actual cached-token counts, all input categories, output, storage, latency and realized cost by approved user/workspace. Record misses and rewrites too.
  4. Report cached reads divided by total input tokens as a token hit rate; report request hit rate separately. Keep denominators and sampling period explicit.
  5. If no traffic was measured, label hit rate, latency improvement and realized savings unmeasured. Offline arithmetic establishes no runtime cache behavior.

Relationship to Other Skills

  • prompt-optimizer improves prompt quality before optimizing reuse.
  • context-budget identifies unnecessary context; caching still consumes it.
  • cost-aware-llm-pipeline compares caching with routing, retries and other costs.

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/prompt-caching-strategy of affaan-m/ECC.

Open the folder on GitHubat commit 4eb71d9

Compare with similar skills

Prompt Caching Strategy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prompt Caching Strategy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prompt Caching Strategy this skillaffaan-m/ECC276k—~2.5kAutomated safety check: PassMIT
Continue Enable DefaultsOnlyTerp/prompt-cache-skills114—~977Automated safety check: PassCustom licence
Venice Chatveniceai/skills144—~5.9kAutomated safety check: PassMIT
LLM Routercuriositech/some_claude_skills244—~1.7kAutomated safety check: PassMIT
Agents Best PracticesDenisSergeevitch/agents-best-practices2.4k—~7.4kAutomated safety check: PassMIT
ClawRouter LLM GatewayBlockRunAI/ClawRouter6.6k—~6.8kAutomated safety check: PassMIT

Similar skills

  • Continue Enable Defaults

    OnlyTerp/prompt-cache-skills

    Continue's prompt caching is opt-in via config and off by default.

    114 GitHub stars~977 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Venice Chat

    veniceai/skills

    Call POST /chat/completions on Venice. An agent skill from veniceai/skills.

    144 GitHub stars~5.9k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Router

    curiositech/some_claude_skills

    Selects the optimal LLM model and provider for each task based on complexity, cost budget, and capability requirements.

    244 GitHub stars~1.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Agents Best Practices

    DenisSergeevitch/agents-best-practices

    A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.

    2.4k GitHub stars~7.4k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • ClawRouter LLM Gateway

    BlockRunAI/ClawRouter

    Describes ClawRouter, a local proxy that forwards each LLM request to the blockrun.ai gateway, which routes to a cheaper capable model, paid by USDC wallet or API key credit.

    6.6k GitHub stars~6.8k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Cline OpenAI Prompt Cache Key Fix

    OnlyTerp/prompt-cache-skills

    Patch guide for adding a stable per-task prompt_cache_key to Cline's OpenAI native provider so cached token counts stop reading as zero.

    114 GitHub stars~868 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from affaan-m/ECC

All 682 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    277k GitHub starsUsed in 5 repos~3.1k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    277k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Docs Governance

    affaan-m/ECC

    Route broad documentation-governance requests to existing ECC skills and run an opt-in, read-only audit of mapped documentation roles, links, ADR indexes, and evidence references.

    277k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    277k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    277k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Set an ECC-specific frontend design direction for production UI work.

    277k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed

Works with

Questions about Prompt Caching Strategy

What does Prompt Caching Strategy do?

Optimize LLM prompt caching hit rate to reduce API costs and improve latency. Prompt Caching Strategy is an agent skill from affaan-m/ECC. Optimize LLM prompt caching hit rate to reduce API costs and improve latency.

When should I use Prompt Caching Strategy?

Prompt Caching Strategy fits situations like: : user says prompt cache; reduce LLM cost; optimize token usage; cheaper LLM calls.

How do I install Prompt Caching Strategy in Claude Code?

Run `npx skills add affaan-m/ECC --skill prompt-caching-strategy -a claude-code`. Or copy the skill folder (skills/prompt-caching-strategy in affaan-m/ECC) into .claude/skills/prompt-caching-strategy in your project. Claude Code loads it when a task matches its description.

How do I install Prompt Caching Strategy in Codex?

Run `npx skills add affaan-m/ECC --skill prompt-caching-strategy -a codex`. Or copy the skill folder (skills/prompt-caching-strategy in affaan-m/ECC) into .agents/skills/prompt-caching-strategy in your project. Codex loads it when a task matches its description.

Can I use Prompt Caching Strategy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill prompt-caching-strategy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prompt-caching-strategy, .gemini/skills/prompt-caching-strategy, .github/skills/prompt-caching-strategy and .opencode/skills/prompt-caching-strategy in your project.

What does Prompt Caching Strategy need to run?

Going by SKILL.md and its folder, Prompt Caching Strategy needs the command-line tools its instructions call (node).

Does Prompt Caching Strategy access the network?

SKILL.md names 3 domains. As links in the text: developers.openai.com, ai.google.dev and platform.claude.com. This is read from the text; nothing was executed.

Is Prompt Caching Strategy safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prompt Caching Strategy use?

Prompt Caching Strategy is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prompt Caching Strategy use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Prompt Caching Strategy?

Skills that share tags, products or a category with Prompt Caching Strategy: Continue Enable Defaults (OnlyTerp/prompt-cache-skills, 114 stars), Venice Chat (veniceai/skills, 144 stars), LLM Router (curiositech/some_claude_skills, 244 stars) and Agents Best Practices (DenisSergeevitch/agents-best-practices, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prompt Caching Strategy?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 276,111 GitHub stars. The repository holds 683 skills in this directory. The repository was last updated on October 10, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.