Continue Enable Defaults
OnlyTerp/prompt-cache-skills
Continue's prompt caching is opt-in via config and off by default.
Optimize LLM prompt caching hit rate to reduce API costs and improve latency.
$ npx skills add affaan-m/ECC --skill prompt-caching-strategy -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install affaan-m/ECC prompt-caching-strategy --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/prompt-caching-strategy .claude/skills/prompt-caching-strategy && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "prompt-caching-strategy" agent skill from https://github.com/affaan-m/ECC/tree/main/skills/prompt-caching-strategy into .claude/skills/prompt-caching-strategy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prompt-caching-strategy", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/affaan-m/ECC/tree/main/skills/prompt-caching-strategyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add affaan-m/ECC --skill prompt-caching-strategy -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install affaan-m/ECC prompt-caching-strategy --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/prompt-caching-strategy .agents/skills/prompt-caching-strategy && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "prompt-caching-strategy" agent skill from https://github.com/affaan-m/ECC/tree/main/skills/prompt-caching-strategy into .agents/skills/prompt-caching-strategy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prompt-caching-strategy", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add affaan-m/ECC --skill prompt-caching-strategy -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install affaan-m/ECC prompt-caching-strategy --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/prompt-caching-strategy .cursor/skills/prompt-caching-strategy && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "prompt-caching-strategy" agent skill from https://github.com/affaan-m/ECC/tree/main/skills/prompt-caching-strategy into .cursor/skills/prompt-caching-strategy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prompt-caching-strategy", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/affaan-m/ECC.git --path skills/prompt-caching-strategy--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add affaan-m/ECC --skill prompt-caching-strategy -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install affaan-m/ECC prompt-caching-strategy --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/prompt-caching-strategy .gemini/skills/prompt-caching-strategy && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "prompt-caching-strategy" agent skill from https://github.com/affaan-m/ECC/tree/main/skills/prompt-caching-strategy into .gemini/skills/prompt-caching-strategy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prompt-caching-strategy", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install affaan-m/ECC prompt-caching-strategyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add affaan-m/ECC --skill prompt-caching-strategy -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/prompt-caching-strategy .github/skills/prompt-caching-strategy && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "prompt-caching-strategy" agent skill from https://github.com/affaan-m/ECC/tree/main/skills/prompt-caching-strategy into .github/skills/prompt-caching-strategy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prompt-caching-strategy", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add affaan-m/ECC --skill prompt-caching-strategy -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install affaan-m/ECC prompt-caching-strategy --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/prompt-caching-strategy .opencode/skills/prompt-caching-strategy && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "prompt-caching-strategy" agent skill from https://github.com/affaan-m/ECC/tree/main/skills/prompt-caching-strategy into .opencode/skills/prompt-caching-strategy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prompt-caching-strategy", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
prompt-caching-strategyOptimize LLM prompt caching hit rate to reduce API costs and improve latency.
Prompt Caching Strategy is an agent skill from affaan-m/ECC. Optimize LLM prompt caching hit rate to reduce API costs and improve latency. Analyzes prompt structure, identifies cacheable vs non-cacheable content, and recommends restructuring to maximize cache reuse. TRIGGER when: user says "prompt cache", "cache hit rate", "reduce LLM cost", "optimize token usage", "cheaper LLM calls", or asks about prompt caching strategies for OpenAI/Anthropic/other providers. DO NOT TRIGGER when: user just wants to optimize a single prompt's quality (use prompt-optimizer instead), or…
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM cost and token optimization. It works with OpenAI. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 4eb71d9. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
nodeFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
developers.openai.comai.google.devplatform.claude.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Prompt Caching Strategy loads about 2.5k tokens when it runs. Until then it costs about 146 tokens; SKILL.md has 989 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from affaan-m/ECC at commit 4eb71d9, republished under its MIT licence (© affaan-m). 989 words, ~2,528 tokens.
.claude/skills/prompt-caching-strategy/SKILL.md (or your agent's skills folder).Arrange reusable prompt content to improve cache reuse, then measure input cost and latency. Savings depend on model, platform, cache eligibility, writes, retention, reuse and output volume. Runtime cache hits and savings for this skill are unmeasured; the examples below validate arithmetic without provider calls.
Use for repeated API requests that share authorized instructions, tools or reference material. First identify the exact provider endpoint, model, service tier and caching mode. Consult its current documentation before selecting rates, minimum prefix length or retention. Caching does not reduce context-window use.
Keep reusable instructions, tool schemas and examples stable; place changing questions, timestamps and retrieved results after that content where the API allows. Respect the provider's serialization order: Claude uses tools, system, then messages. Matching prefixes need identical content, including tool schemas and relevant request settings. A changed prefix can prevent reuse from that point. A dynamic suffix may itself be cached later if reused unchanged.
Reusable prefix: approved tools, instructions, stable examples/reference
Changing suffix: current question, fresh retrieval results, request metadataChoose only relevant context. Padding prompts or broadening access just to meet an eligibility threshold needs its own cost and quality justification.
Authorize content access before cache lookup or reuse. Scope application-managed cache handles and accounting keys to the approved user, workspace, provider account, model and content revision. Share only material explicitly approved for that scope. A cache key is a routing/accounting hint, not an authorization gate or a sandbox. Provider cache isolation does not replace application access checks. Invalidate application references when permissions or content change; follow provider retention/deletion controls. Avoid logging private prompt contents. Context selection, capabilities, sandbox enforcement and evidence remain independent of cache optimization.
The linked documentation is authoritative; endpoint and SDK versions can expose different usage field names. Record that version alongside counters.
| Provider | Requirements to verify | Counters to record |
|---|---|---|
| OpenAI | Model-specific automatic/explicit modes, minimum length, retention and write/read prices; earlier and newer models differ | Responses usage.input_tokens_details.cached_tokens and, when present, cache_write_tokens; input and output totals. Chat Completions uses prompt_tokens_details.cached_tokens |
| Claude | Supported platform, model minimum, cache_control support and TTL; cache writes may cost more than ordinary input | input_tokens, cache_creation_input_tokens, cache_read_input_tokens, output_tokens; split writes by TTL when rates differ |
| Gemini | Model minimum and API: Interactions supports implicit caching; explicit cache objects require GenerateContent | Interactions usage.total_cached_tokens; GenerateContent usage_metadata cache/input/output counters plus explicit-cache token count and retained duration |
Use OpenAI's prompt caching guide and pricing for the selected model. For Chat Completions, check the usage schema. Derive ordinary input from the total minus cached reads and any separately reported writes; do not count the same tokens twice.
Use Claude's prompt caching guide
for current model/platform limits and TTL rates. Its input_tokens counts
ordinary input separately from cache reads and writes. A short prefix can produce
zero cache counters even when marked for caching. Do not assume every Claude
model or hosted platform uses the same minimum or read price.
Use Gemini's caching guide, GenerateContent caching guide and pricing. Explicit-cache storage is charged by token count and retained duration, separately from reads, ordinary input and output. Implicit hits are not guaranteed. Verify billing categories for the chosen endpoint rather than copying another provider's rates.
Normalize actual usage into nonoverlapping categories and use prices in the same currency per million tokens. Split write categories when TTL, modality, tier or model prices differ. Storage uses million-token-hours; other applicable charges such as tools and grounding need their own lines. Include all warmup, miss, expiration, refresh and retry requests in the measured period.
baseline = (ordinary + written + read) * ordinary_input_rate / 1e6
+ output * output_rate / 1e6
actual = ordinary * ordinary_input_rate / 1e6
+ written * cache_write_rate / 1e6
+ read * cache_read_rate / 1e6
+ output * output_rate / 1e6
+ storage_token_hours * storage_rate / 1e6
savings = baseline - actualThis counterfactual holds prompt and output volumes fixed. For explicit-cache creation, include any separately billed creation input in the write category; use the endpoint's documented rate. Set nonexistent categories to zero. Compare actual before/after invoices separately if warmup requests or output differ. Never apply a cache discount to the total bill or to output tokens.
The executable calculator takes normalized counters, not raw provider responses:
function compareCacheCost(usage, rates) {
const counts = ['uncached', 'written', 'read', 'output', 'storageTokenHours'];
const prices = ['input', 'write', 'read', 'output', 'storage'];
for (const [values, keys] of [[usage, counts], [rates, prices]]) {
for (const key of keys) {
if (!Number.isFinite(values[key]) || values[key] < 0) {
throw new Error(`${key} must be a finite nonnegative number`);
}
}
}
const outputCost = usage.output * rates.output / 1e6;
const storageCost = usage.storageTokenHours * rates.storage / 1e6;
const baseline = (usage.uncached + usage.written + usage.read)
* rates.input / 1e6 + outputCost;
const actual = (usage.uncached * rates.input + usage.written * rates.write
+ usage.read * rates.read) / 1e6 + outputCost + storageCost;
return { baseline, actual, savings: baseline - actual, outputCost, storageCost };
}These are hypothetical prices and counts, not provider quotes or measured hits. Ten calls each contain a 10,000-token stable prefix, 2,000-token changing suffix and 2,000 output tokens. Assume one prefix write and nine full prefix reads.
| Category | Tokens | Example rate per million | Cost |
|---|---|---|---|
| Ordinary input | 20,000 | $3.00 | $0.0600 |
| Cache write | 10,000 | $3.75 | $0.0375 |
| Cache read | 90,000 | $0.30 | $0.0270 |
| Output | 20,000 | $15.00 | $0.3000 |
| Storage | 0 token-hours | $0.00 | $0.0000 |
Baseline: 120,000 input tokens at $3 plus the same output = $0.6600. With caching: $0.4245, savings $0.2355. The write premium is included; the $0.3000 output bill remains unchanged. An output-heavy workload has the same absolute input savings under these assumptions, a smaller fraction of total cost.
If an explicit cache retains 10,000 tokens for two hours at a hypothetical $1 per million-token-hour, add $0.0200 storage. Actual cost becomes $0.4445. Use the selected Gemini endpoint's prices for a real estimate.
For a one-off call, that prefix write costs $0.0375 rather than $0.0300 ordinary input: $0.0075 more, with no read savings. For a 500-token prefix below a selected model's documented minimum, treat the request as uncached; claiming future hits without counters would overstate savings. Repeated expiry or changing prefixes can also turn expected savings into higher cost.
node tests/skills/prompt-caching-strategy.test.js in the ECC repository.© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/prompt-caching-strategy of affaan-m/ECC.
Open the folder on GitHubat commit 4eb71d9
Prompt Caching Strategy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Prompt Caching Strategy this skillaffaan-m/ECC | 276k | — | ~2.5k | Automated safety check: Pass | MIT | |
| Continue Enable DefaultsOnlyTerp/prompt-cache-skills | 114 | — | ~977 | Automated safety check: Pass | Custom licence | |
| Venice Chatveniceai/skills | 144 | — | ~5.9k | Automated safety check: Pass | MIT | |
| LLM Routercuriositech/some_claude_skills | 244 | — | ~1.7k | Automated safety check: Pass | MIT | |
| Agents Best PracticesDenisSergeevitch/agents-best-practices | 2.4k | — | ~7.4k | Automated safety check: Pass | MIT | |
| ClawRouter LLM GatewayBlockRunAI/ClawRouter | 6.6k | — | ~6.8k | Automated safety check: Pass | MIT |
OnlyTerp/prompt-cache-skills
Continue's prompt caching is opt-in via config and off by default.
veniceai/skills
Call POST /chat/completions on Venice. An agent skill from veniceai/skills.
curiositech/some_claude_skills
Selects the optimal LLM model and provider for each task based on complexity, cost budget, and capability requirements.
DenisSergeevitch/agents-best-practices
A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.
BlockRunAI/ClawRouter
Describes ClawRouter, a local proxy that forwards each LLM request to the blockrun.ai gateway, which routes to a cheaper capable model, paid by USDC wallet or API key credit.
OnlyTerp/prompt-cache-skills
Patch guide for adding a stable per-task prompt_cache_key to Cline's OpenAI native provider so cached token counts stop reading as zero.
affaan-m/ECC
Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.
affaan-m/ECC
Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.
affaan-m/ECC
Route broad documentation-governance requests to existing ECC skills and run an opt-in, read-only audit of mapped documentation roles, links, ADR indexes, and evidence references.
affaan-m/ECC
Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.
affaan-m/ECC
Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.
affaan-m/ECC
Set an ECC-specific frontend design direction for production UI work.
Works with
Categories
Optimize LLM prompt caching hit rate to reduce API costs and improve latency. Prompt Caching Strategy is an agent skill from affaan-m/ECC. Optimize LLM prompt caching hit rate to reduce API costs and improve latency.
Prompt Caching Strategy fits situations like: : user says prompt cache; reduce LLM cost; optimize token usage; cheaper LLM calls.
Run `npx skills add affaan-m/ECC --skill prompt-caching-strategy -a claude-code`. Or copy the skill folder (skills/prompt-caching-strategy in affaan-m/ECC) into .claude/skills/prompt-caching-strategy in your project. Claude Code loads it when a task matches its description.
Run `npx skills add affaan-m/ECC --skill prompt-caching-strategy -a codex`. Or copy the skill folder (skills/prompt-caching-strategy in affaan-m/ECC) into .agents/skills/prompt-caching-strategy in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill prompt-caching-strategy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prompt-caching-strategy, .gemini/skills/prompt-caching-strategy, .github/skills/prompt-caching-strategy and .opencode/skills/prompt-caching-strategy in your project.
Going by SKILL.md and its folder, Prompt Caching Strategy needs the command-line tools its instructions call (node).
SKILL.md names 3 domains. As links in the text: developers.openai.com, ai.google.dev and platform.claude.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Prompt Caching Strategy is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Prompt Caching Strategy: Continue Enable Defaults (OnlyTerp/prompt-cache-skills, 114 stars), Venice Chat (veniceai/skills, 144 stars), LLM Router (curiositech/some_claude_skills, 244 stars) and Agents Best Practices (DenisSergeevitch/agents-best-practices, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 276,111 GitHub stars. The repository holds 683 skills in this directory. The repository was last updated on October 10, 2026.
Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.