Agent skill

Cache Policy Comparison

by benchflow-ai in benchflow-ai/skillsbench

Compare and implement eviction policies (LRU, LFU, FIFO, S3FIFO, ARC) for bounded-capacity caches.

Apache-2.0Auto-check passed

Install Cache Policy Comparison

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill cache-policy-comparison -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench cache-policy-comparison --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks/llm-prefix-cache-replay/environment/skills/cache-policy-comparison .claude/skills/cache-policy-comparison && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cache-policy-comparison
GitHub stars
1.8k
Token cost
~1.5k tokens
SKILL.md length
642 words
Files
1
Skills in repo
178
Repo updated
First seen
Licence
Apache-2.0

At a glance

Compare and implement eviction policies (LRU, LFU, FIFO, S3FIFO, ARC) for bounded-capacity caches.

  • Implementing an eviction policy for a buffer pool
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Writing a replay simulator that supports multiple policies

What it does

Cache Policy Comparison is an agent skill from benchflow-ai/skillsbench. Compare and implement eviction policies (LRU, LFU, FIFO, S3FIFO, ARC) for bounded-capacity caches. Use when choosing or implementing an eviction policy for a buffer pool, page cache, CDN edge, or LLM KV cache, or when writing a replay simulator that supports multiple policies. Clarifies recency vs frequency semantics, queue topology, saturating counters, ghost buffers, and the second-chance rule that distinguishes modern FIFO-family policies from classic LRU.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

When your agent uses it

  • Implementing an eviction policy for a buffer pool
  • Writing a replay simulator that supports multiple policies

Example prompts

  • “/cache-policy-comparison”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cache Policy Comparison loads about 1.5k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 642 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 642 words, ~1,485 tokens.

Download SKILL.mdSave it as .claude/skills/cache-policy-comparison/SKILL.md (or your agent's skills folder).
name
cache-policy-comparison
description
Compare and implement eviction policies (LRU, LFU, FIFO, S3FIFO, ARC) for bounded-capacity caches. Use when choosing or implementing an eviction policy for a buffer pool, page cache, CDN edge, or LLM KV cache, or when writing a replay simulator that supports multiple policies. Clarifies recency vs frequency semantics, queue topology, saturating counters, ghost buffers, and the second-chance rule that distinguishes modern FIFO-family policies from classic LRU.

Overview

An eviction policy decides which resident entry a cache removes when a new entry is admitted beyond capacity. Four policies cover almost every replay-and-measure task:

PolicyData structureOn hitOn admitEviction choice
LRUOrderedDictMove to tailAppend at tailPop head
LFU{key: freq} + insertion orderfreq[k] += 1freq[k] = 1Min freq, tiebreak by insertion order
FIFOOrderedDictNothingAppend at tailPop head
S3FIFOThree FIFO queues + freq[k]freq[k] = min(freq+1, cap)Admit to small; ghost-hit admits to mainSecond-chance on main; small drains to main/ghost

Each has subtleties that trip naive implementations.

LRU

Use an OrderedDict where the tail is the most-recently-accessed key. On hit, move_to_end. On miss + insert, append; pop from head if over capacity.

Most common bug: forgetting to update recency on a hit. Without the refresh, LRU degenerates to FIFO — hit rate drops substantially on any workload with recency structure.

python
from collections import OrderedDict

class LRU:
    def __init__(self, capacity):
        self.capacity = capacity
        self._d = OrderedDict()

    def contains(self, k): return k in self._d

    def access(self, k):
        if k in self._d:
            self._d.move_to_end(k)
        else:
            self._d[k] = None
            if len(self._d) > self.capacity:
                self._d.popitem(last=False)

LFU

Keep freq: dict[key, int] and a tie-breaker — an insertion counter is simplest and deterministic. On hit, increment freq[k]. On miss at capacity, evict min(freq) with ties broken by insertion order (oldest first).

Typical bugs:

  • No tie-breaker. min(freq.items(), key=lambda x: x[1])[0] has implementation-defined behaviour across interpreters and distributions. Always include a secondary key.
  • Frequency pollution. A block that was hot once and then went cold can linger forever because its freq is permanently above newcomers. Production systems add aging (periodic decay of freq) or combine with a recency signal (W-TinyLFU). Pure LFU is correct for the task as specified but fragile in practice.

FIFO

One queue, insertion order, no hit-time update. Useful as a lower-bound baseline.

Do NOT call it "LRU without hit update" — conceptually different even when implementations overlap. Hit on a FIFO cache is still a hit for accounting; the block just does not change rank.

S3FIFO

A modern FIFO-family policy (Yang et al., SOSP 2023) that matches or beats LRU on typical web and LLM workloads with a fraction of the bookkeeping cost — which is why recent production systems (Twitter, Google) have been switching to it. The full algorithm — three queues, saturating frequency counter, second-chance eviction on the main queue — is implemented in the prefix-cache-replay skill. Consult that skill if your task uses S3FIFO.

Show full SKILL.md (277 more words)Show less

Workload implications

  • Strong recency → LRU wins slightly.
  • Stable hot set with long tail (Zipf) → LFU or S3FIFO.
  • Nearly uniform random → all converge toward capacity / working_set hit rate.
  • Prefix-shared LLM workloads are mixed — shared prefixes are both recent and frequent, so LRU/LFU/S3FIFO typically sit within a few percent of each other at the same capacity, but they differ in which blocks remain resident at end-of-trace, and their miss-handling costs diverge. Measure, don't assume.

Comparing hit rates on a trace

Replay the same trace through each policy at identical capacity, record total_hit_tokens / total_prompt_tokens and the final resident set. Do not compare hit rate alone — also compare:

  • Final residency — how many unique blocks are resident at the end. Under S3FIFO this is often strictly less than capacity because ghost entries absorb the admission pressure.
  • Per-request hit-token distribution — two policies can have similar overall hit rate but very different per-request variance.
  • Admission effort — under policies with ghost structures, the bookkeeping cost per access is non-trivial.

Common mistakes

  • Reusing an LRU implementation when the task specifies S3FIFO (or vice versa). The final hit rate and residency will both differ; no partial credit for "close enough".
  • Making ghost count as resident, or treating a ghost hit as a hit for token accounting.
  • Forgetting to saturate freq — unbounded counters turn the main-queue second-chance loop into a spin.
  • Under LFU, using Python min(d.items(), key=d.get) without an explicit insertion-order tiebreaker.
  • Misordering admission and residency check. Always check h ∈ cache BEFORE applying the admission side effects of the current request, otherwise every request self-hits.
  • Final cache size off by small constants because you forgot to exclude ghost or you forgot to subtract the S-cap vs M-cap split.

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in tasks/llm-prefix-cache-replay/environment/skills/cache-policy-comparison of benchflow-ai/skillsbench.

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Cache Policy Comparison next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cache Policy Comparison compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cache Policy Comparison this skillbenchflow-ai/skillsbench1.8k—~1.5kAutomated safety check: PassApache-2.0
Cache Policy Hit Rate Comparisonben-manes/caffeine18k—~939Automated safety check: NotesApache-2.0
Prompt Cachingdavila7/claude-code-templates32k6 repos~452Automated safety check: PassMIT
Turborepo Cachingwshobson/agents40k9 repos~2kAutomated safety check: NotesMIT
OmniRoute LLM Cachediegosouzapw/OmniRoute74k1 repos~529Automated safety check: PassMIT
Cachingzebbern/claude-code-guide4.7k—~1.5kAutomated safety check: PassMIT

Similar skills

  • Compares cache eviction policies by hit rate across several cache sizes on a trace file using the Caffeine simulator, with CSV tables and a PNG chart.

    18k GitHub stars~939 tokensUpdated 2 days ago
    DevelopmentAuto-check: notes
  • Prompt Caching

    davila7/claude-code-templates

    Caching strategies for LLM prompts including Anthropic prompt caching, response caching, and CAG (Cache Augmented Generation) Use when: prompt caching, cache prompt, response cache, cag, cache…

    32k GitHub starsUsed in 6 repos~452 tokens
    Backend & APIsAuto-check passed
  • Turborepo Caching

    wshobson/agents

    Configures Turborepo pipelines and local or remote caching for monorepo builds, including Vercel remote cache, a self-hosted cache and cache-miss debugging.

    40k GitHub starsUsed in 9 repos~2k tokens
    DevelopmentAuto-check: notes
  • OmniRoute LLM Cache

    diegosouzapw/OmniRoute

    Documents OmniRoute's cache endpoints for reading cache statistics and clearing entries, statistics or the reasoning cache, with notes on TTL and similarity settings.

    74k GitHub starsUsed in 1 repo~529 tokens
    Backend & APIsAuto-check passed
  • Caching

    zebbern/claude-code-guide

    Caching strategies — invalidation, TTL guidelines, cache keys, cache layers, and when not to cache.

    4.7k GitHub stars~1.5k tokensUpdated today
    Backend & APIsAuto-check passed
  • Implementing Policy As Code With Open Policy Agent

    mukul975/Anthropic-Cybersecurity-Skills

    Implements policy-as-code enforcement with Open Policy Agent (OPA) and Gatekeeper for Kubernetes and CI/CD pipelines, covering writing Rego policies, deploying OPA Gatekeeper as a Kubernetes…

    34k GitHub stars~2.6k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check: notes

More from benchflow-ai/skillsbench

All 178 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Cache Policy Comparison

What does Cache Policy Comparison do?

Compare and implement eviction policies (LRU, LFU, FIFO, S3FIFO, ARC) for bounded-capacity caches. Cache Policy Comparison is an agent skill from benchflow-ai/skillsbench. Compare and implement eviction policies (LRU, LFU, FIFO, S3FIFO, ARC) for bounded-capacity caches.

When should I use Cache Policy Comparison?

Cache Policy Comparison fits situations like: implementing an eviction policy for a buffer pool; writing a replay simulator that supports multiple policies.

How do I install Cache Policy Comparison in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill cache-policy-comparison -a claude-code`. Or copy the skill folder (tasks/llm-prefix-cache-replay/environment/skills/cache-policy-comparison in benchflow-ai/skillsbench) into .claude/skills/cache-policy-comparison in your project. Claude Code loads it when a task matches its description.

How do I install Cache Policy Comparison in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill cache-policy-comparison -a codex`. Or copy the skill folder (tasks/llm-prefix-cache-replay/environment/skills/cache-policy-comparison in benchflow-ai/skillsbench) into .agents/skills/cache-policy-comparison in your project. Codex loads it when a task matches its description.

Can I use Cache Policy Comparison in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill cache-policy-comparison -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cache-policy-comparison, .gemini/skills/cache-policy-comparison, .github/skills/cache-policy-comparison and .opencode/skills/cache-policy-comparison in your project.

What does Cache Policy Comparison need to run?

SKILL.md names no scripts, command-line tools or credentials: Cache Policy Comparison is instructions for the agent only. Our summary lists: Python 3.

Does Cache Policy Comparison access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cache Policy Comparison safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cache Policy Comparison use?

Cache Policy Comparison is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cache Policy Comparison use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cache Policy Comparison?

Skills that share tags, products or a category with Cache Policy Comparison: Cache Policy Hit Rate Comparison (ben-manes/caffeine, 18k stars), Prompt Caching (davila7/claude-code-templates, 32k stars), Turborepo Caching (wshobson/agents, 40k stars) and OmniRoute LLM Cache (diegosouzapw/OmniRoute, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cache Policy Comparison?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,832 GitHub stars. The repository holds 178 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.