Agent skill

Caching Architecture

by majiayu000 in majiayu000/litellm-rs

LiteLLM-RS response caching architecture. An agent skill from majiayu000/litellm-rs.

MITAuto-check passedAI & LLM Engineering

Install Caching Architecture

skills CLI
$ npx skills add majiayu000/litellm-rs --skill caching-architecture -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/litellm-rs caching-architecture --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/litellm-rs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/caching-architecture .claude/skills/caching-architecture && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
caching-architecture
GitHub stars
118
Token cost
~2k tokens
SKILL.md length
543 words
Files
7
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

LiteLLM-RS response caching architecture. An agent skill from majiayu000/litellm-rs.

  • Works in 3 steps: Cache Key Determinism → TTL Strategy → Skip Caching — the Real Rules
  • Tuning gateway response caching — cache keys
  • SKILL.md covers Overview, Configuration, Request Flow and Best Practices, plus 1 more section
  • Needs BYPASS_CHAT_RESPONSE_CACHE_KEY

What it does

Caching Architecture is an agent skill from majiayu000/litellm-rs. LiteLLM-RS response caching architecture. Covers the two-tier deterministic cache (L1 in-memory + optional L2 Redis) behind LLMCache and DualCache, SHA-256 cache key generation with schema versioning, TTL and eviction policy, request-path wiring for chat completions and embeddings, cache statistics, and admin endpoints. Use when adding or tuning gateway response caching — cache keys, tiers, TTLs or eviction, cache metrics, or invalidation.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `reference/cache-key-generation.md`, `reference/cache-metrics.md` and `reference/in-memory-cache.md`).

It sits in AI & LLM Engineering, covering Caching, Embeddings and Model routing and gateways. It works with Redis, OpenAI and Rust. The repository describes itself as: Self-hosted Rust LLM gateway with OpenAI-compatible APIs, load balancing, failover, and a reusable Rust kernel. The licence is MIT.

When your agent uses it

  • Tuning gateway response caching — cache keys
  • Tasks that involve Caching
  • Tasks that involve Embeddings

Example prompts

  • “/caching-architecture”

Requirements

  • A credential in BYPASS_CHAT_RESPONSE_CACHE_KEY

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Cache Key Determinism
  2. TTL Strategy
  3. Skip Caching — the Real Rules

What it can do on your machine

Read from SKILL.md and the folder at commit ed3f4d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are rust and yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BYPASS_CHAT_RESPONSE_CACHE_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Caching Architecture loads about 2k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 543 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/litellm-rs at commit ed3f4d9, republished under its MIT licence (© majiayu000). 543 words, ~1,965 tokens.

Download SKILL.mdSave it as .claude/skills/caching-architecture/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
caching-architecture
description
LiteLLM-RS response caching architecture. Covers the two-tier deterministic cache (L1 in-memory + optional L2 Redis) behind LLMCache and DualCache, SHA-256 cache key generation with schema versioning, TTL and eviction policy, request-path wiring for chat completions and embeddings, cache statistics, and admin endpoints. Use when adding or tuning gateway response caching — cache keys, tiers, TTLs or eviction, cache metrics, or invalidation.

Caching Architecture Guide

Overview

LiteLLM-RS ships exactly one wired caching subsystem: an exact-match response cache for non-streaming chat completions and embeddings. It is a two-tier read-through cache, not a three-tier stack — the former unwired semantic (vector) module has been removed from unreleased source.

Request (non-streaming /v1/chat/completions, /v1/embeddings)
     │ lookup_chat / lookup_embedding (src/server/routes/ai/response_cache.rs)
     ▼
┌────────────────────────────────────────────────────────────┐
│ LLMCache  (src/core/cache/llm_cache.rs)                    │
│   chat_cache:      DualCache<CachedChatResponse>           │
│   embedding_cache: DualCache<CachedEmbeddingResponse>      │
└────────────────────────────────────────────────────────────┘
     │ per-key get / set
     ▼
┌────────────────────────────────────────────────────────────┐
│ DualCache<T>  (src/core/cache/dual.rs)                     │
│   L1  InMemoryCache<T> — DashMap, TTL, sampled eviction    │
│   L2  RedisCache<T>    — optional, backed by RedisPool     │
│   Read: L1 miss → L2 hit → repopulate L1                   │
│   Write: both tiers; L2 failure logs a warning, not fatal  │
└────────────────────────────────────────────────────────────┘
     │ miss
     ▼
LLM Provider → response stored back into both tiers
What Is Wired vs Not
CapabilityStatus
Exact-match response cache (chat + embeddings)Wired: AppState.response_cache, built by build_response_cache (src/server/state.rs:143)
Semantic similarity cacheRemoved: core::semantic_cache, cache.semantic_cache, and cache.similarity_threshold no longer exist in unreleased source; remove those config keys, including false/default values
Vector DB backendsStorage-only: QdrantStore implemented; weaviate/pinecone declared but return "not implemented yet" (src/storage/vector/backend.rs:29). Nothing connects them to caching at runtime
Cloud object-storage cachescore::cache::cloud (CloudCache trait; S3/GCS/Azure under feature s3) — not part of the request path

Configuration

yaml
cache:
  enabled: true               # default false; requires ttl > 0
  ttl: 3600                   # seconds; applied to chat AND embedding entries
  max_size: 1000              # max entries per in-memory layer

These are the only three fields (src/config/models/cache.rs:9, deny_unknown_fields). There is no l1/l2/l3 block, redis_url, prefix, exclude_models, or skip_streaming key.

  • Redis is not configured here: the cache reuses the gateway's Redis pool. Without a pool it runs memory-only (CacheMode::MemoryOnly).
  • enabled: true with ttl: 0 logs an error and leaves the cache off (src/server/state.rs:148).
  • Validation rejects enabled: true, ttl: 0 outright (src/config/validation/cache_validators.rs:12).

Request Flow

  1. POST /v1/chat/completions calls lookup_chat before routing (src/server/routes/ai/chat.rs:112). A hit passes ensure_chat_cache_pricing_gate and returns immediately.
  2. On a miss, the provider executes and store_chat writes the response (chat.rs:261). Embeddings do the same via lookup_embedding / store_embedding (src/server/routes/ai/embeddings.rs:98,283).
  3. Chat lookups and stores are skipped when the request carries a per-key budget, sets store: true, or was marked bypassed by an upstream handler (should_bypass_chat_cache, src/server/routes/ai/response_cache.rs:25). Embeddings have no such bypass conditions.
  4. Chat entries are scoped per caller: identity is api_key:{id} or user:{id}, optionally suffixed :max_tokens_per_request:{limit} (cache_identity, response_cache.rs:46). The key does not hash the separate client-supplied ChatCompletionRequest.user, so two requests from the same caller that differ only in that provider-facing field collide. Embedding entries are currently shared across callers: the route copies the identity into EmbeddingRequest.user, but LLMCache calls generate_embedding_key with no user_id, and that key does not hash request.user. Identical model/input embeddings therefore reuse one entry. Streaming chat requests are never cached (LLMCache::get_chat_response_with_user, src/core/cache/llm_cache.rs:280).
  5. Cache lookup errors are logged and treated as misses. Gateway startup constructs only Dual or MemoryOnly caches, and Dual suppresses Redis L2 write failures. A programmatically installed RedisOnly LLMCache differs: Redis store errors propagate through store_chat / store_embedding, so the route returns an error after the provider call succeeded.

Show full SKILL.md (159 more words)Show less

Best Practices

1. Cache Key Determinism

Keys come from free functions, not a generator struct. They hash a canonical-JSON payload (sorted keys, transport fields stripped) with SHA-256 under schema version v4:

rust
use litellm_rs::core::cache::{generate_chat_key, generate_chat_key_with_user};

let key = generate_chat_key(&request);                        // chat:gpt-4:v4:<64-hex>
let key = generate_chat_key_with_user(&request, Some(user));  // user-scoped variant

Do not add non-deterministic fields (request_id, stream, timestamps) — canonical_json_string already strips the known ones at the top level and inside extra_body. Details: reference/cache-key-generation.md.

2. TTL Strategy
rust
// LLMCacheConfig::default() (src/core/cache/llm_cache.rs:55)
chat_ttl: Duration::from_secs(3600),       // 1 hour
embedding_ttl: Duration::from_secs(86400), // 24 hours — embeddings are deterministic

At startup build_response_cache overrides both from cache.ttl, so per-tier TTL tuning requires code changes, not YAML.

3. Skip Caching — the Real Rules

The deterministic path already enforces its own skips; do not re-implement them:

rust
// src/core/cache/llm_cache.rs:280 — streaming requests are never cached
if request.stream.unwrap_or(false) {
    return Ok(None);
}

// src/server/routes/ai/response_cache.rs:25 — chat-only bypasses
fn should_bypass_chat_cache(request: &ChatCompletionRequest, context: &RequestContext) -> bool {
    context.metadata.get(BYPASS_CHAT_RESPONSE_CACHE_KEY).and_then(|v| v.as_bool()).unwrap_or(false)
        || context.api_key_budget_id().is_some()
        || request.store == Some(true)
}

The removed semantic-cache filters are not part of the current request path.


References

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files in .claude/skills/caching-architecture of majiayu000/litellm-rs.

  • SKILL.md
  • reference/cache-key-generation.md
  • reference/cache-metrics.md
  • reference/in-memory-cache.md
  • reference/redis-cache.md
  • reference/response-cache.md
  • reference/semantic-cache.md

Open the folder on GitHubat commit ed3f4d9

Compare with similar skills

Caching Architecture next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Caching Architecture compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Caching Architecture this skillmajiayu000/litellm-rs118—~2kAutomated safety check: PassMIT
Cognee Integrations Setuptopoteretes/cognee32k—~1kAutomated safety check: NotesApache-2.0
Embeddings via 9Routerdecolua/9router31k—~604Automated safety check: PassMIT
Using Ccproxy APIstarbaser/ccproxy350—~4kAutomated safety check: PassCustom licence
OmniRoute Inference Endpointsdiegosouzapw/OmniRoute75k—~5.5kAutomated safety check: PassMIT
Ax AIdosco/aithy107—~8.2kAutomated safety check: PassApache-2.0

Similar skills

  • Cognee Integrations Setup

    topoteretes/cognee

    Switches cognee's LLM, embedding, relational, vector and graph backends through environment variables, with the extras to install and the traps to avoid.

    32k GitHub stars~1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    31k GitHub stars~604 tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Using Ccproxy API

    starbaser/ccproxy

    Guides users through ccproxy as an OpenAI-compatible and Anthropic-compatible LLM API server with SDK integration, OAuth authentication, sentinel key substitution, model routing, and troubleshooting.

    350 GitHub stars~4k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • OmniRoute Inference Endpoints

    diegosouzapw/OmniRoute

    Documents OmniRoute's OpenAI-compatible endpoints for chat completions, embeddings, images, speech, transcription, moderation, rerank and the Responses API.

    75k GitHub stars~5.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ax AI

    dosco/aithy

    This skill helps an LLM generate correct AI provider setup and configuration code using @ax-llm/ax.

    107 GitHub stars~8.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Openrouter Caching Strategy

    jeremylongshore/tons-of-skills-marketplace

    Implement caching for OpenRouter API responses to reduce cost and latency.

    2.8k GitHub stars~2.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from majiayu000/litellm-rs

All 9 skills in this repo
  • Auth Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Authentication Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Config Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Configuration Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Error Handling

    majiayu000/litellm-rs

    LiteLLM-RS Error Handling Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Observability Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Provider Architecture

    majiayu000/litellm-rs

    LiteLLM-RS provider system in two tiers - data-driven OpenAI-compatible catalog entries auto-routed through OpenAILikeProvider, plus code-based provider modules implementing the LLMProvider trait…

    118 GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Routing Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Routing Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Works with

Questions about Caching Architecture

What does Caching Architecture do?

LiteLLM-RS response caching architecture. An agent skill from majiayu000/litellm-rs. Caching Architecture is an agent skill from majiayu000/litellm-rs. LiteLLM-RS response caching architecture.

When should I use Caching Architecture?

Caching Architecture fits situations like: tuning gateway response caching — cache keys; tasks that involve Caching; tasks that involve Embeddings.

How do I install Caching Architecture in Claude Code?

Run `npx skills add majiayu000/litellm-rs --skill caching-architecture -a claude-code`. Or copy the skill folder (.claude/skills/caching-architecture in majiayu000/litellm-rs) into .claude/skills/caching-architecture in your project. Claude Code loads it when a task matches its description.

How do I install Caching Architecture in Codex?

Run `npx skills add majiayu000/litellm-rs --skill caching-architecture -a codex`. Or copy the skill folder (.claude/skills/caching-architecture in majiayu000/litellm-rs) into .agents/skills/caching-architecture in your project. Codex loads it when a task matches its description.

Can I use Caching Architecture in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/litellm-rs --skill caching-architecture -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/caching-architecture, .gemini/skills/caching-architecture, .github/skills/caching-architecture and .opencode/skills/caching-architecture in your project.

What does Caching Architecture need to run?

Going by SKILL.md and its folder, Caching Architecture needs credentials named BYPASS_CHAT_RESPONSE_CACHE_KEY. Our summary lists: A credential in BYPASS_CHAT_RESPONSE_CACHE_KEY.

Does Caching Architecture access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Caching Architecture safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Caching Architecture use?

Caching Architecture is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Caching Architecture use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Caching Architecture?

Skills that share tags, products or a category with Caching Architecture: Cognee Integrations Setup (topoteretes/cognee, 32k stars), Embeddings via 9Router (decolua/9router, 31k stars), Using Ccproxy API (starbaser/ccproxy, 350 stars) and OmniRoute Inference Endpoints (diegosouzapw/OmniRoute, 75k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Caching Architecture?

majiayu000 (a GitHub user) maintains it in majiayu000/litellm-rs, which has 118 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 11, 2026.

Source: majiayu000/litellm-rs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.