Agent skill

Prompt Caching

by sickn33 in sickn33/agentic-awesome-skills

Caching strategies for LLM prompts including Anthropic prompt caching, response caching, and CAG (Cache Augmented Generation)

MITAuto-check passedAI & LLM Engineering

Install Prompt Caching

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill prompt-caching -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills prompt-caching --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/prompt-caching .claude/skills/prompt-caching && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
prompt-caching
GitHub stars
47k
Used in
2 other repos
Token cost
~3.4k tokens
SKILL.md length
1,186 words
Files
1
Skills in repo
1,493
Repo updated
First seen
Licence
MIT

At a glance

Caching strategies for LLM prompts including Anthropic prompt caching, response caching, and CAG (Cache Augmented Generation)

  • Tasks that involve LLM cost and token optimization
  • SKILL.md covers Capabilities, Prerequisites, Scope and Ecosystem, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Caching

What it does

Prompt Caching is an agent skill from sickn33/agentic-awesome-skills. Caching strategies for LLM prompts including Anthropic prompt caching, response caching, and CAG (Cache Augmented Generation)

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM cost and token optimization and Caching. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve LLM cost and token optimization
  • Tasks that involve Caching

Example prompts

  • “/prompt-caching”

What it can do on your machine

Read from SKILL.md and the folder at commit 680176d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prompt Caching loads about 3.4k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 1,186 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit 680176d, republished under its MIT licence (© sickn33). 1,186 words, ~3,424 tokens.

Download SKILL.mdSave it as .claude/skills/prompt-caching/SKILL.md (or your agent's skills folder).
name
prompt-caching
description
Caching strategies for LLM prompts including Anthropic prompt caching, response caching, and CAG (Cache Augmented Generation)
risk
none
source
vibeship-spawner-skills (Apache 2.0)
date_added
2026-02-27

Prompt Caching

Caching strategies for LLM prompts including Anthropic prompt caching, response caching, and CAG (Cache Augmented Generation)

Capabilities

  • prompt-cache
  • response-cache
  • kv-cache
  • cag-patterns
  • cache-invalidation

Prerequisites

  • Knowledge: Caching fundamentals, LLM API usage, Hash functions
  • Skills_recommended: context-window-management

Scope

  • Does_not_cover: CDN caching, Database query caching, Static asset caching
  • Boundaries: Focus is LLM-specific caching, Covers prompt and response caching

Ecosystem

Primary_tools
  • Anthropic Prompt Caching - Native prompt caching in Claude API
  • Redis - In-memory cache for responses (ioredis on servers; @upstash/redis over HTTP on serverless, see upstash-redis)
  • OpenAI Caching - Automatic caching in OpenAI API

Patterns

Anthropic Prompt Caching

Use Claude's native prompt caching for repeated prefixes

When to use: Using Claude API with stable system prompts or context

import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

// Cache the stable parts of your prompt async function queryWithCaching(userQuery: string) { const response = await client.messages.create({ model: "claude-sonnet-4-20250514", max_tokens: 1024, system: [ { type: "text", text: LONG_SYSTEM_PROMPT, // Your detailed instructions cache_control: { type: "ephemeral" } // Cache this! }, { type: "text", text: KNOWLEDGE_BASE, // Large static context cache_control: { type: "ephemeral" } } ], messages: [ { role: "user", content: userQuery } // Dynamic part ] });

// Check cache usage
console.log(`Cache read: ${response.usage.cache_read_input_tokens}`);
console.log(`Cache write: ${response.usage.cache_creation_input_tokens}`);

return response;

}

// Cost savings: 90% reduction on cached tokens // Latency savings: Up to 2x faster

Response Caching

Cache full LLM responses for identical or similar queries

When to use: Same queries asked repeatedly

import { createHash } from 'crypto'; import Redis from 'ioredis';

const redis = new Redis(process.env.REDIS_URL); // Serverless/edge alternative without a persistent connection: // import { Redis } from '@upstash/redis'; const redis = Redis.fromEnv(); // then use redis.set(key, value, { ex: ttl })

class ResponseCache { private ttl = 3600; // 1 hour default

// Exact match caching
async getCached(prompt: string): Promise<string | null> {
    const key = this.hashPrompt(prompt);
    return await redis.get(`response:${key}`);
}

async setCached(prompt: string, response: string): Promise<void> {
    const key = this.hashPrompt(prompt);
    await redis.set(`response:${key}`, response, 'EX', this.ttl);
}

private hashPrompt(prompt: string): string {
    return createHash('sha256').update(prompt).digest('hex');
}

// Semantic similarity caching
async getSemanticallySimilar(
    prompt: string,
    threshold: number = 0.95
): Promise<string | null> {
    const embedding = await embed(prompt);
    const similar = await this.vectorCache.search(embedding, 1);

    if (similar.length && similar[0].similarity > threshold) {
        return await redis.get(`response:${similar[0].id}`);
    }
    return null;
}

// Temperature-aware caching
async getCachedWithParams(
    prompt: string,
    params: { temperature: number; model: string }
): Promise<string | null> {
    // Only cache low-temperature responses
    if (params.temperature > 0.5) return null;

    const key = this.hashPrompt(
        `${prompt}|${params.model}|${params.temperature}`
    );
    return await redis.get(`response:${key}`);
}

}

Cache Augmented Generation (CAG)

Pre-cache documents in prompt instead of RAG retrieval

When to use: Document corpus is stable and fits in context

// CAG: Pre-compute document context, cache in prompt // Better than RAG when: // - Documents are stable // - Total fits in context window // - Latency is critical

class CAGSystem { private cachedContext: string | null = null; private lastUpdate: number = 0;

async buildCachedContext(documents: Document[]): Promise<void> {
    // Pre-process and format documents
    const formatted = documents.map(d =>
        `## ${d.title}\n${d.content}`
    ).join('\n\n');

    // Store with timestamp
    this.cachedContext = formatted;
    this.lastUpdate = Date.now();
}

async query(userQuery: string): Promise<string> {
    // Use cached context directly in prompt
    const response = await client.messages.create({
        model: "claude-sonnet-4-20250514",
        max_tokens: 1024,
        system: [
            {
                type: "text",
                text: "You are a helpful assistant with access to the following documentation.",
                cache_control: { type: "ephemeral" }
            },
            {
                type: "text",
                text: this.cachedContext!,  // Pre-cached docs
                cache_control: { type: "ephemeral" }
            }
        ],
        messages: [{ role: "user", content: userQuery }]
    });

    return response.content[0].text;
}

// Periodic refresh
async refreshIfNeeded(documents: Document[]): Promise<void> {
    const stale = Date.now() - this.lastUpdate > 3600000;  // 1 hour
    if (stale) {
        await this.buildCachedContext(documents);
    }
}

}

// CAG vs RAG decision matrix: // | Factor | CAG Better | RAG Better | // |------------------|------------|------------| // | Corpus size | < 100K tokens | > 100K tokens | // | Update frequency | Low | High | // | Latency needs | Critical | Flexible | // | Query specificity| General | Specific |

Sharp Edges

Cache miss causes latency spike with additional overhead

Severity: HIGH

Situation: Slow response when cache miss, slower than no caching

Symptoms:

  • Slow responses on cache miss
  • Cache hit rate below 50%
  • Higher latency than uncached

Why this breaks: Cache check adds latency. Cache write adds more latency. Miss + overhead > no caching.

Recommended fix:

// Optimize for cache misses, not just hits

class OptimizedCache { async queryWithCache(prompt: string): Promise<string> { const cacheKey = this.hash(prompt);

    // Non-blocking cache check
    const cachedPromise = this.cache.get(cacheKey);
    const llmPromise = this.queryLLM(prompt);

    // Race: use cache if available before LLM returns
    const cached = await Promise.race([
        cachedPromise,
        sleep(50).then(() => null)  // 50ms cache timeout
    ]);

    if (cached) {
        // Cancel LLM request if possible
        return cached;
    }

    // Cache miss: continue with LLM
    const response = await llmPromise;

    // Async cache write (don't block response)
    this.cache.set(cacheKey, response).catch(console.error);

    return response;
}

}

// Alternative: Probabilistic caching // Only cache if query matches known high-frequency patterns class SelectiveCache { private patterns: Map<string, number> = new Map();

shouldCache(prompt: string): boolean {
    const pattern = this.extractPattern(prompt);
    const frequency = this.patterns.get(pattern) || 0;

    // Only cache high-frequency patterns
    return frequency > 10;
}

recordQuery(prompt: string): void {
    const pattern = this.extractPattern(prompt);
    this.patterns.set(pattern, (this.patterns.get(pattern) || 0) + 1);
}

}

Cached responses become incorrect over time

Severity: HIGH

Situation: Users get outdated or wrong information from cache

Symptoms:

  • Users report wrong information
  • Answers don't match current data
  • Complaints about outdated responses

Why this breaks: Source data changed. No cache invalidation. Long TTLs for dynamic data.

Recommended fix:

// Implement proper cache invalidation

class InvalidatingCache { // Version-based invalidation private cacheVersion = 1;

getCacheKey(prompt: string): string {
    return `v${this.cacheVersion}:${this.hash(prompt)}`;
}

invalidateAll(): void {
    this.cacheVersion++;
    // Old keys automatically become orphaned
}

// Content-hash invalidation
async setWithContentHash(
    key: string,
    response: string,
    sourceContent: string
): Promise<void> {
    const contentHash = this.hash(sourceContent);
    await this.cache.set(key, {
        response,
        contentHash,
        timestamp: Date.now()
    });
}

async getIfValid(
    key: string,
    currentSourceContent: string
): Promise<string | null> {
    const cached = await this.cache.get(key);
    if (!cached) return null;

    // Check if source content changed
    const currentHash = this.hash(currentSourceContent);
    if (cached.contentHash !== currentHash) {
        await this.cache.delete(key);
        return null;
    }

    return cached.response;
}

// Event-based invalidation
onSourceUpdate(sourceId: string): void {
    // Invalidate all caches that used this source
    this.invalidateByTag(`source:${sourceId}`);
}

}

Show full SKILL.md (293 more words)Show less
Prompt caching doesn't work due to prefix changes

Severity: MEDIUM

Situation: Cache misses despite similar prompts

Symptoms:

  • Cache hit rate lower than expected
  • Cache creation tokens high, read low
  • Similar prompts not hitting cache

Why this breaks: Anthropic caching requires exact prefix match. Timestamps or dynamic content in prefix. Different message order.

Recommended fix:

// Structure prompts for optimal caching

class CacheOptimizedPrompts { // WRONG: Dynamic content in cached prefix buildPromptBad(query: string): SystemMessage[] { return [ { type: "text", text: You are helpful. Current time: ${new Date()}, // BREAKS CACHE! cache_control: { type: "ephemeral" } } ]; }

// RIGHT: Static prefix, dynamic at end
buildPromptGood(query: string): SystemMessage[] {
    return [
        {
            type: "text",
            text: STATIC_SYSTEM_PROMPT,  // Never changes
            cache_control: { type: "ephemeral" }
        },
        {
            type: "text",
            text: STATIC_KNOWLEDGE_BASE,  // Rarely changes
            cache_control: { type: "ephemeral" }
        }
        // Dynamic content goes in messages, NOT system
    ];
}

// Prefix ordering matters
buildWithConsistentOrder(components: string[]): SystemMessage[] {
    // Sort components for consistent ordering
    const sorted = [...components].sort();
    return sorted.map((c, i) => ({
        type: "text",
        text: c,
        cache_control: i === sorted.length - 1
            ? { type: "ephemeral" }
            : undefined  // Only cache the full prefix
    }));
}

}

Validation Checks

Caching High Temperature Responses

Severity: WARNING

Message: Caching with high temperature. Responses are non-deterministic.

Fix action: Only cache responses with temperature <= 0.5

Cache Without TTL

Severity: WARNING

Message: Cache without TTL. May serve stale data indefinitely.

Fix action: Set appropriate TTL based on data freshness requirements

Dynamic Content in Cached Prefix

Severity: WARNING

Message: Dynamic content in cached prefix. Will cause cache misses.

Fix action: Move dynamic content outside of cache_control blocks

No Cache Metrics

Severity: INFO

Message: Cache without hit/miss tracking. Can't measure effectiveness.

Fix action: Add cache hit/miss metrics and logging

Collaboration

Delegation Triggers
  • context window|token -> context-window-management (Need context optimization)
  • rag|retrieval -> rag-implementation (Need retrieval system)
  • memory -> conversation-memory (Need memory persistence)
High-Performance LLM System

Skills: prompt-caching, context-window-management, rag-implementation

Workflow:

1. Analyze query patterns
2. Implement prompt caching for stable prefixes
3. Add response caching for frequent queries
4. Consider CAG for stable document sets
5. Monitor and optimize hit rates

Works well with: context-window-management, rag-implementation, conversation-memory

When to Use

  • User mentions or implies: prompt caching
  • User mentions or implies: cache prompt
  • User mentions or implies: response cache
  • User mentions or implies: cag
  • User mentions or implies: cache augmented

Limitations

  • Use this skill only when the task clearly matches the scope described above.
  • Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
  • Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/prompt-caching of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit 680176d

Used in 2 other repositories

We found 11 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 9, 2026.

Compare with similar skills

Prompt Caching next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prompt Caching compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prompt Caching this skillsickn33/agentic-awesome-skills47k2 repos~3.4kAutomated safety check: PassMIT
Continue Enable DefaultsOnlyTerp/prompt-cache-skills114—~977Automated safety check: PassCustom licence
Dt Obs GenaiDynatrace/dynatrace-for-ai162—~4.5kAutomated safety check: PassApache-2.0
Anth Performance Tuningjeremylongshore/tons-of-skills-marketplace2.8k—~1.9kAutomated safety check: PassMIT
LLM Cost OptimizationBagelHole/DevOps-Security-Agent-Skills1.1k—~2.2kAutomated safety check: PassMIT
Context Engineering Reviewmohitagw15856/pm-claude-skills1.4k—~1.4kAutomated safety check: PassMIT

Similar skills

  • Continue Enable Defaults

    OnlyTerp/prompt-cache-skills

    Continue's prompt caching is opt-in via config and off by default.

    114 GitHub stars~977 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Dt Obs Genai

    Dynatrace/dynatrace-for-ai

    Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup.

    162 GitHub stars~4.5k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Anth Performance Tuning

    jeremylongshore/tons-of-skills-marketplace

    Optimize Claude API performance with prompt caching, model selection, streaming, and latency reduction techniques.

    2.8k GitHub stars~1.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LLM Cost Optimization

    BagelHole/DevOps-Security-Agent-Skills

    Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

    1.1k GitHub stars~2.2k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Context Engineering Review

    mohitagw15856/pm-claude-skills

    Review what an LLM feature or agent actually puts in its context window — and find what's bloating, missing, or fighting itself.

    1.4k GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LLM Cost Latency Budget

    mohitagw15856/pm-claude-skills

    Model the cost and latency of an LLM feature before it ships and surprises the bill.

    1.4k GitHub stars~985 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,493 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Prompt Caching

What does Prompt Caching do?

Caching strategies for LLM prompts including Anthropic prompt caching, response caching, and CAG (Cache Augmented Generation). Prompt Caching is an agent skill from sickn33/agentic-awesome-skills.

When should I use Prompt Caching?

Prompt Caching fits situations like: tasks that involve LLM cost and token optimization; tasks that involve Caching.

How do I install Prompt Caching in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill prompt-caching -a claude-code`. Or copy the skill folder (skills/prompt-caching in sickn33/agentic-awesome-skills) into .claude/skills/prompt-caching in your project. Claude Code loads it when a task matches its description.

How do I install Prompt Caching in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill prompt-caching -a codex`. Or copy the skill folder (skills/prompt-caching in sickn33/agentic-awesome-skills) into .agents/skills/prompt-caching in your project. Codex loads it when a task matches its description.

Can I use Prompt Caching in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill prompt-caching -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prompt-caching, .gemini/skills/prompt-caching, .github/skills/prompt-caching and .opencode/skills/prompt-caching in your project.

What does Prompt Caching need to run?

SKILL.md names no scripts, command-line tools or credentials: Prompt Caching is instructions for the agent only.

Does Prompt Caching access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Prompt Caching safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prompt Caching use?

Prompt Caching is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prompt Caching use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Prompt Caching?

Skills that share tags, products or a category with Prompt Caching: Continue Enable Defaults (OnlyTerp/prompt-cache-skills, 114 stars), Dt Obs Genai (Dynatrace/dynatrace-for-ai, 162 stars), Anth Performance Tuning (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and LLM Cost Optimization (BagelHole/DevOps-Security-Agent-Skills, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prompt Caching?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,379 GitHub stars. The repository holds 1,493 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.