Agent skill

LLM Caching

by majiayu000 in majiayu000/claude-skill-registry

Implement multi-layer LLM caching with exact match, semantic similarity, and provider-side prompt caching.

MITAuto-check passedBackend & APIs

Install LLM Caching

skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill llm-caching -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/claude-skill-registry llm-caching --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-llm/llm-caching .claude/skills/llm-caching && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-caching
GitHub stars
666
Used in
3 other repos
Token cost
~2.6k tokens
SKILL.md length
244 words
Files
2
Skills in repo
1,273
Repo updated
First seen
Licence
MIT

At a glance

Implement multi-layer LLM caching with exact match, semantic similarity, and provider-side prompt caching.

  • Tasks that involve Caching
  • SKILL.md covers When to Use This Skill, Caching Layers, Layer 1: Exact Match Cache… and Layer 2: Semantic Cache…, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve LLM cost and token optimization

What it does

LLM Caching is an agent skill from majiayu000/claude-skill-registry. Implement multi-layer LLM caching with exact match, semantic similarity, and provider-side prompt caching. Reduce API costs by 30–70%, cut latency, and improve throughput using Redis, GPTCache, and provider caching APIs.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).

It sits in Backend & APIs, covering Caching and LLM cost and token optimization. It works with Redis. The repository describes itself as: Searchable Claude Code skills catalog with source-linked guides and generated registry artifacts. The licence is MIT.

When your agent uses it

  • Tasks that involve Caching
  • Tasks that involve LLM cost and token optimization

Example prompts

  • “/llm-caching”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 2d14a69. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Caching loads about 2.6k tokens when it runs. Until then it costs about 58 tokens; SKILL.md has 244 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/claude-skill-registry at commit 2d14a69, republished under its MIT licence (© majiayu000). 244 words, ~2,610 tokens.

Download SKILL.mdSave it as .claude/skills/llm-caching/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
llm-caching
description
Implement multi-layer LLM caching with exact match, semantic similarity, and provider-side prompt caching. Reduce API costs by 30–70%, cut latency, and improve throughput using Redis, GPTCache, and provider caching APIs.
license
MIT
metadata.author
devops-skills
metadata.version
1.0

LLM Caching

Cut LLM costs and latency with exact match, semantic, and provider-side caching layers.

When to Use This Skill

Use this skill when:

  • The same or similar queries are asked repeatedly (FAQ bots, support tools)
  • LLM API costs are growing and you need immediate savings
  • Serving high request volumes where repeated queries cause bottlenecks
  • Implementing prompt caching for long system prompts (Anthropic/OpenAI)
  • Building offline-capable AI features that need response persistence

Caching Layers

Request → Exact Cache → Semantic Cache → Provider Cache → LLM API
             ↓ hit            ↓ hit             ↓ hit
           instant          ~5ms           50-80% cheaper

Layer 1: Exact Match Cache (Redis)

python
import hashlib
import json
import redis
from openai import OpenAI

r = redis.Redis(host="localhost", port=6379, decode_responses=True)
client = OpenAI()

def build_cache_key(model: str, messages: list, temperature: float) -> str:
    """Deterministic key from request parameters."""
    payload = json.dumps({
        "model": model,
        "messages": messages,
        "temperature": temperature,
    }, sort_keys=True)
    return f"llm:exact:{hashlib.sha256(payload.encode()).hexdigest()}"

def cached_completion(model: str, messages: list, temperature: float = 0.0,
                      ttl: int = 3600) -> dict:
    key = build_cache_key(model, messages, temperature)

    # Check cache
    if cached := r.get(key):
        return json.loads(cached)

    # Call API
    response = client.chat.completions.create(
        model=model, messages=messages, temperature=temperature
    )
    result = response.model_dump()

    # Cache result (only cache deterministic responses)
    if temperature == 0.0:
        r.setex(key, ttl, json.dumps(result))

    return result

Layer 2: Semantic Cache (GPTCache)

python
from gptcache import cache, Config
from gptcache.adapter import openai
from gptcache.embedding import Onnx
from gptcache.manager import CacheBase, VectorBase, get_data_manager
from gptcache.similarity_evaluation.distance import SearchDistanceEvaluation

# Configure GPTCache with Qdrant backend
def init_gptcache(cache_obj, llm: str):
    onnx = Onnx()                              # local embedding model
    data_manager = get_data_manager(
        CacheBase("redis"),                    # metadata store
        VectorBase("qdrant",
                   host="localhost",
                   port=6333,
                   collection_name=f"llm-cache-{llm}",
                   dimension=onnx.dimension),
    )
    cache_obj.init(
        embedding_func=onnx.to_embeddings,
        data_manager=data_manager,
        similarity_evaluation=SearchDistanceEvaluation(),
        config=Config(similarity_threshold=0.80),  # 80% similarity = cache hit
    )

cache.set_openai_key()
init_gptcache(cache, "gpt-4o-mini")

# Now openai calls are automatically cached
response = openai.ChatCompletion.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "What is machine learning?"}],
)
# Second call with similar question ("Explain machine learning") → cache hit

Custom Semantic Cache (Production-Grade)

python
from sentence_transformers import SentenceTransformer
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams, PointStruct, Filter, FieldCondition, Range
import numpy as np
import uuid
import time

embed_model = SentenceTransformer("BAAI/bge-small-en-v1.5")  # fast, 33M params
qdrant = QdrantClient("http://localhost:6333")

CACHE_COLLECTION = "semantic-cache"
SIMILARITY_THRESHOLD = 0.88
CACHE_TTL_SECONDS = 86400  # 24h

# Create collection once
qdrant.create_collection(
    collection_name=CACHE_COLLECTION,
    vectors_config=VectorParams(size=384, distance=Distance.COSINE),
    on_disk_payload=True,
)

def semantic_cache_lookup(query: str, model: str) -> str | None:
    embedding = embed_model.encode(query).tolist()
    results = qdrant.query_points(
        collection_name=CACHE_COLLECTION,
        query=embedding,
        query_filter=Filter(must=[
            FieldCondition(key="model", match={"value": model}),
            FieldCondition(key="expires_at", range=Range(gte=time.time())),
        ]),
        limit=1,
        score_threshold=SIMILARITY_THRESHOLD,
    )
    if results.points:
        return results.points[0].payload["response"]
    return None

def semantic_cache_store(query: str, response: str, model: str):
    embedding = embed_model.encode(query).tolist()
    qdrant.upsert(
        collection_name=CACHE_COLLECTION,
        points=[PointStruct(
            id=str(uuid.uuid4()),
            vector=embedding,
            payload={
                "query": query,
                "response": response,
                "model": model,
                "created_at": time.time(),
                "expires_at": time.time() + CACHE_TTL_SECONDS,
            },
        )],
    )

def smart_llm_call(query: str, model: str = "gpt-4o-mini") -> dict:
    # 1. Semantic lookup
    if cached_response := semantic_cache_lookup(query, model):
        return {"response": cached_response, "source": "semantic_cache", "cost": 0}

    # 2. LLM call
    response = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": query}],
    )
    text = response.choices[0].message.content
    cost = litellm.completion_cost(response)

    # 3. Store in cache
    semantic_cache_store(query, text, model)

    return {"response": text, "source": "llm_api", "cost": cost}

Layer 3: Provider-Side Prompt Caching

python
# Anthropic — cache long system prompts (saves 90% on cached input tokens)
import anthropic

client = anthropic.Anthropic()

# Long system prompt — mark for caching
SYSTEM_PROMPT = open("knowledge-base.txt").read()  # e.g., 50k tokens

def call_with_prompt_cache(user_question: str) -> str:
    response = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=1024,
        system=[
            {"type": "text", "text": "You are a helpful assistant."},
            {
                "type": "text",
                "text": SYSTEM_PROMPT,
                "cache_control": {"type": "ephemeral"},  # cache this block
            }
        ],
        messages=[{"role": "user", "content": user_question}],
    )
    # Log cache efficiency
    usage = response.usage
    cache_savings = usage.cache_read_input_tokens * 0.9  # 90% discount on cached
    print(f"Cache hits: {usage.cache_read_input_tokens} tokens "
          f"(saved ~${cache_savings * 3.0 / 1_000_000:.4f})")
    return response.content[0].text

# OpenAI — automatic for repeated prefixes (≥1,024 tokens)
# No code change needed; cached tokens appear in usage.prompt_tokens_details
response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[
        {"role": "system", "content": LONG_SYSTEM_PROMPT},  # auto-cached
        {"role": "user", "content": user_question},
    ]
)
cached = response.usage.prompt_tokens_details.cached_tokens
print(f"OpenAI cached {cached} tokens")

Cache Warming

python
async def warm_cache(common_queries: list[str], model: str):
    """Pre-populate cache with known frequent queries."""
    import asyncio
    from openai import AsyncOpenAI

    aclient = AsyncOpenAI()

    async def warm_single(query: str):
        if not semantic_cache_lookup(query, model):
            response = await aclient.chat.completions.create(
                model=model,
                messages=[{"role": "user", "content": query}],
            )
            text = response.choices[0].message.content
            semantic_cache_store(query, text, model)
            print(f"Warmed: {query[:50]}...")

    await asyncio.gather(*[warm_single(q) for q in common_queries])

# Warm on startup
import asyncio
asyncio.run(warm_cache(FREQUENT_QUERIES, "gpt-4o-mini"))

Cache Metrics

python
from prometheus_client import Counter, Histogram

cache_hits = Counter("llm_cache_hits_total", "Cache hits", ["cache_layer", "model"])
cache_misses = Counter("llm_cache_misses_total", "Cache misses", ["model"])
cache_savings_usd = Counter("llm_cache_savings_usd_total", "USD saved by cache", ["model"])

# Use in your smart_llm_call function
if source == "semantic_cache":
    cache_hits.labels(cache_layer="semantic", model=model).inc()
    cache_savings_usd.labels(model=model).inc(estimated_cost)
else:
    cache_misses.labels(model=model).inc()

Redis Configuration for LLM Caching

bash
# redis.conf tuning for LLM cache workload
maxmemory 8gb
maxmemory-policy allkeys-lru    # evict least-recently-used when full
save ""                          # disable persistence (cache is ephemeral)
appendonly no
tcp-keepalive 60

Common Issues

IssueCauseFix
Low cache hit rateThreshold too strictLower SIMILARITY_THRESHOLD to 0.82–0.85
Stale cached responsesLong TTLUse topic-specific TTLs; invalidate on data updates
Cache serving wrong answersThreshold too looseRaise threshold or add model-name filtering
Redis OOMNo eviction policySet maxmemory + allkeys-lru
Slow semantic lookupLarge cache collectionAdd payload index on model + expires_at

Best Practices

  • Start with exact cache — zero cost, instant wins for identical queries.
  • Semantic threshold of 0.88–0.92 balances hit rate vs. accuracy; tune with your data.
  • Set per-model TTLs: longer for stable knowledge (1 week), shorter for news/events (1 hour).
  • Always filter by model name in semantic cache — different models give different answers.
  • Log cache hit rate as a KPI; target 30%+ for FAQ-style applications.

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/ai-llm/llm-caching of majiayu000/claude-skill-registry.

  • SKILL.md
  • metadata.json

Open the folder on GitHubat commit 2d14a69

Used in 3 other repositories

We found 7 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in majiayu000/claude-skill-registry, which our catalogue first saw on October 7, 2026.

Compare with similar skills

LLM Caching next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Caching compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Caching this skillmajiayu000/claude-skill-registry6663 repos~2.6kAutomated safety check: PassMIT
FastAPI-Redis SDK Developmentredis/fastapi-redis-sdk404—~2.5kAutomated safety check: NotesMIT
Cache Credits AnalyzerTsinHzl/kiro2cc-proxy164—~1.1kAutomated safety check: PassMIT
Prompt Cachingdavila7/claude-code-templates32k6 repos~452Automated safety check: PassMIT
Redis Patternsaffaan-m/ECC275k1 repos~3kAutomated safety check: PassMIT
Amazon Elasticacheaws/agent-toolkit-for-aws2.8k—~4.5kAutomated safety check: PassApache-2.0

Similar skills

  • FastAPI-Redis SDK Development

    redis/fastapi-redis-sdk

    Official

    Guides development on the fastapi-redis-sdk library itself - its connection lifecycle, dependency-injected caching, and async/sync bridging.

    404 GitHub stars~2.5k tokensUpdated 7 days ago
    Backend & APIsAuto-check: notes
  • Cache Credits Analyzer

    TsinHzl/kiro2cc-proxy

    分析 kiro2cc-proxy 访问日志,计算 Prompt Caching 节省的 credits。只要用户粘贴了含有"输入token 输出token 费用$ credits✓"格式的日志行,并询问节省了多少credits、缓存效率、cost分析等,立即使用此 skill。触发关键词:节省了多少credits、cache节省、分析日志、caching…

    164 GitHub stars~1.1k tokensUpdated today
    Backend & APIsAuto-check passed
  • Prompt Caching

    davila7/claude-code-templates

    Caching strategies for LLM prompts including Anthropic prompt caching, response caching, and CAG (Cache Augmented Generation) Use when: prompt caching, cache prompt, response cache, cag, cache…

    32k GitHub starsUsed in 6 repos~452 tokens
    Backend & APIsAuto-check passed
  • Redis Patterns

    affaan-m/ECC

    Redis data structure patterns, caching strategies, distributed locks, rate limiting, pub/sub, and connection management for production applications.

    275k GitHub starsUsed in 1 repo~3k tokens
    Backend & APIsAuto-check passed
  • Amazon Elasticache

    aws/agent-toolkit-for-aws

    Official

    Activate when developers have latent caching needs: slow API responses, database read bottlenecks, DynamoDB throttling or cost, RDS/Aurora scaling pressure, Bedrock latency or cost, or adding a…

    2.8k GitHub stars~4.5k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • LLM Gateway

    sickn33/agentic-awesome-skills

    Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Backend & APIsAuto-check passed

More from majiayu000/claude-skill-registry

All 1,273 skills in this repo
  • Deep Research

    majiayu000/claude-skill-registry

    Multi-source deep research using firecrawl and exa MCPs. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 6 repos~1.1k tokens
    Auto-check passed
  • Exa Search

    majiayu000/claude-skill-registry

    Neural search via Exa MCP for web, code, and company research.

    666 GitHub starsUsed in 5 repos~856 tokens
    Auto-check passed
  • Fal AI Media

    majiayu000/claude-skill-registry

    Unified media generation via fal.ai MCP — image, video, and audio.

    666 GitHub starsUsed in 5 repos~1.7k tokens
    Auto-check passed
  • Pyzotero

    majiayu000/claude-skill-registry

    Interact with Zotero reference management libraries using the pyzotero Python client.

    666 GitHub starsUsed in 5 repos~1.6k tokens
    Auto-check: notes
  • Bgpt Paper Search

    majiayu000/claude-skill-registry

    Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server.

    666 GitHub starsUsed in 4 repos~619 tokens
    Auto-check: notes
  • Bio Alignment Pairwise

    majiayu000/claude-skill-registry

    Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.

    666 GitHub starsUsed in 4 repos~1.7k tokens
    Auto-check passed

Works with

Questions about LLM Caching

What does LLM Caching do?

Implement multi-layer LLM caching with exact match, semantic similarity, and provider-side prompt caching. LLM Caching is an agent skill from majiayu000/claude-skill-registry. Implement multi-layer LLM caching with exact match, semantic similarity, and provider-side prompt caching.

When should I use LLM Caching?

LLM Caching fits situations like: tasks that involve Caching; tasks that involve LLM cost and token optimization.

How do I install LLM Caching in Claude Code?

Run `npx skills add majiayu000/claude-skill-registry --skill llm-caching -a claude-code`. Or copy the skill folder (skills/ai-llm/llm-caching in majiayu000/claude-skill-registry) into .claude/skills/llm-caching in your project. Claude Code loads it when a task matches its description.

How do I install LLM Caching in Codex?

Run `npx skills add majiayu000/claude-skill-registry --skill llm-caching -a codex`. Or copy the skill folder (skills/ai-llm/llm-caching in majiayu000/claude-skill-registry) into .agents/skills/llm-caching in your project. Codex loads it when a task matches its description.

Can I use LLM Caching in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/claude-skill-registry --skill llm-caching -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-caching, .gemini/skills/llm-caching, .github/skills/llm-caching and .opencode/skills/llm-caching in your project.

What does LLM Caching need to run?

SKILL.md names no scripts, command-line tools or credentials: LLM Caching is instructions for the agent only. Our summary lists: Python 3.

Does LLM Caching access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is LLM Caching safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Caching use?

LLM Caching is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Caching use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Caching?

Skills that share tags, products or a category with LLM Caching: FastAPI-Redis SDK Development (redis/fastapi-redis-sdk, 404 stars), Cache Credits Analyzer (TsinHzl/kiro2cc-proxy, 164 stars), Prompt Caching (davila7/claude-code-templates, 32k stars) and Redis Patterns (affaan-m/ECC, 275k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Caching?

majiayu000 (a GitHub user) maintains it in majiayu000/claude-skill-registry, which has 666 GitHub stars. The repository holds 1,273 skills in this directory. The repository was last updated on October 7, 2026.

Source: majiayu000/claude-skill-registry on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.