Agent skill

Content Hash Cache Pattern

by affaan-m in affaan-m/ECC

使用SHA-256内容哈希缓存昂贵的文件处理结果——路径无关、自动失效、服务层分离. An agent skill from affaan-m/ECC.

MITAuto-check passed

Install Content Hash Cache Pattern

skills CLI
$ npx skills add affaan-m/ECC --skill content-hash-cache-pattern -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC content-hash-cache-pattern --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/zh-CN/skills/content-hash-cache-pattern .claude/skills/content-hash-cache-pattern && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
content-hash-cache-pattern
GitHub stars
276k
Used in
3 other repos
Token cost
~1k tokens
SKILL.md length
79 words
Files
1
Skills in repo
673
Repo updated
First seen
Licence
MIT

At a glance

使用SHA-256内容哈希缓存昂贵的文件处理结果——路径无关、自动失效、服务层分离. An agent skill from affaan-m/ECC.

  • Works in 4 steps: 基于内容哈希的缓存键 → 用于缓存条目的冻结数据类 → 基于文件的缓存存储 → …
  • SKILL.md covers 何时激活, 核心模式, 关键设计决策 and 最佳实践, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Content Hash Cache Pattern is an agent skill from affaan-m/ECC. 使用SHA-256内容哈希缓存昂贵的文件处理结果——路径无关、自动失效、服务层分离。

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

Example prompts

  • “/content-hash-cache-pattern”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. 基于内容哈希的缓存键
  2. 用于缓存条目的冻结数据类
  3. 基于文件的缓存存储
  4. 服务层包装器(单一职责原则)

What it can do on your machine

Read from SKILL.md and the folder at commit ef648e0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Content Hash Cache Pattern loads about 1k tokens when it runs. Until then it costs about 17 tokens; SKILL.md has 79 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~17
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit ef648e0, republished under its MIT licence (© affaan-m). 79 words, ~1,029 tokens.

Download SKILL.mdSave it as .claude/skills/content-hash-cache-pattern/SKILL.md (or your agent's skills folder).
name
content-hash-cache-pattern
description
使用SHA-256内容哈希缓存昂贵的文件处理结果——路径无关、自动失效、服务层分离。
origin
ECC

内容哈希文件缓存模式

使用 SHA-256 内容哈希作为缓存键,缓存昂贵的文件处理结果(PDF 解析、文本提取、图像分析)。与基于路径的缓存不同,此方法在文件移动/重命名后仍然有效,并在内容更改时自动失效。

何时激活

  • 构建文件处理管道时(PDF、图像、文本提取)
  • 处理成本高且同一文件被重复处理时
  • 需要一个 --cache/--no-cache CLI 选项时
  • 希望在不修改现有纯函数的情况下为其添加缓存时

核心模式

1. 基于内容哈希的缓存键

使用文件内容(而非路径)作为缓存键:

python
import hashlib
from pathlib import Path

_HASH_CHUNK_SIZE = 65536  # 64KB chunks for large files

def compute_file_hash(path: Path) -> str:
    """SHA-256 of file contents (chunked for large files)."""
    if not path.is_file():
        raise FileNotFoundError(f"File not found: {path}")
    sha256 = hashlib.sha256()
    with open(path, "rb") as f:
        while True:
            chunk = f.read(_HASH_CHUNK_SIZE)
            if not chunk:
                break
            sha256.update(chunk)
    return sha256.hexdigest()

为什么使用内容哈希? 文件重命名/移动 = 缓存命中。内容更改 = 自动失效。无需索引文件。

2. 用于缓存条目的冻结数据类
python
from dataclasses import dataclass

@dataclass(frozen=True, slots=True)
class CacheEntry:
    file_hash: str
    source_path: str
    document: ExtractedDocument  # The cached result
3. 基于文件的缓存存储

每个缓存条目都存储为 {hash}.json —— 通过哈希实现 O(1) 查找,无需索引文件。

python
import json
from typing import Any

def write_cache(cache_dir: Path, entry: CacheEntry) -> None:
    cache_dir.mkdir(parents=True, exist_ok=True)
    cache_file = cache_dir / f"{entry.file_hash}.json"
    data = serialize_entry(entry)
    cache_file.write_text(json.dumps(data, ensure_ascii=False), encoding="utf-8")

def read_cache(cache_dir: Path, file_hash: str) -> CacheEntry | None:
    cache_file = cache_dir / f"{file_hash}.json"
    if not cache_file.is_file():
        return None
    try:
        raw = cache_file.read_text(encoding="utf-8")
        data = json.loads(raw)
        return deserialize_entry(data)
    except (json.JSONDecodeError, ValueError, KeyError):
        return None  # Treat corruption as cache miss
4. 服务层包装器(单一职责原则)

保持处理函数的纯净性。将缓存作为一个单独的服务层添加。

python
def extract_with_cache(
    file_path: Path,
    *,
    cache_enabled: bool = True,
    cache_dir: Path = Path(".cache"),
) -> ExtractedDocument:
    """Service layer: cache check -> extraction -> cache write."""
    if not cache_enabled:
        return extract_text(file_path)  # Pure function, no cache knowledge

    file_hash = compute_file_hash(file_path)

    # Check cache
    cached = read_cache(cache_dir, file_hash)
    if cached is not None:
        logger.info("Cache hit: %s (hash=%s)", file_path.name, file_hash[:12])
        return cached.document

    # Cache miss -> extract -> store
    logger.info("Cache miss: %s (hash=%s)", file_path.name, file_hash[:12])
    doc = extract_text(file_path)
    entry = CacheEntry(file_hash=file_hash, source_path=str(file_path), document=doc)
    write_cache(cache_dir, entry)
    return doc

关键设计决策

决策理由
SHA-256 内容哈希与路径无关,内容更改时自动失效
{hash}.json 文件命名O(1) 查找,无需索引文件
服务层包装器单一职责原则:提取功能保持纯净,缓存是独立的关注点
手动 JSON 序列化完全控制冻结数据类的序列化
损坏时返回 None优雅降级,在下次运行时重新处理
cache_dir.mkdir(parents=True)在首次写入时惰性创建目录

最佳实践

  • 哈希内容,而非路径 —— 路径会变,内容标识不变
  • 对大文件进行哈希时分块处理 —— 避免将整个文件加载到内存中
  • 保持处理函数的纯净性 —— 它们不应了解任何关于缓存的信息
  • 记录缓存命中/未命中,并使用截断的哈希值以便调试
  • 优雅地处理损坏 —— 将无效的缓存条目视为未命中,永不崩溃

应避免的反模式

python
# BAD: Path-based caching (breaks on file move/rename)
cache = {"/path/to/file.pdf": result}

# BAD: Adding cache logic inside the processing function (SRP violation)
def extract_text(path, *, cache_enabled=False, cache_dir=None):
    if cache_enabled:  # Now this function has two responsibilities
        ...

# BAD: Using dataclasses.asdict() with nested frozen dataclasses
# (can cause issues with complex nested types)
data = dataclasses.asdict(entry)  # Use manual serialization instead

适用场景

  • 文件处理管道(PDF 解析、OCR、文本提取、图像分析)
  • 受益于 --cache/--no-cache 选项的 CLI 工具
  • 跨多次运行出现相同文件的批处理
  • 在不修改现有纯函数的情况下为其添加缓存

不适用场景

  • 必须始终保持最新的数据(实时数据流)
  • 缓存条目可能极其庞大的情况(应考虑使用流式处理)
  • 结果依赖于文件内容之外参数的情况(例如,不同的提取配置)

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/zh-CN/skills/content-hash-cache-pattern of affaan-m/ECC.

Open the folder on GitHubat commit ef648e0

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in affaan-m/ECC, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Content Hash Cache Pattern next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Content Hash Cache Pattern compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Content Hash Cache Pattern this skillaffaan-m/ECC276k3 repos~1kAutomated safety check: PassMIT
Content Hash Cache Patternxu-xiang/everything-claude-code-zh2k—~994Automated safety check: PassMIT
Prompt Cachingsickn33/agentic-awesome-skills47k2 repos~3.4kAutomated safety check: PassMIT
Prompt Cachingdavila7/claude-code-templates32k5 repos~452Automated safety check: PassMIT
Turborepo Cachingwshobson/agents40k9 repos~2kAutomated safety check: NotesMIT
OmniRoute LLM Cachediegosouzapw/OmniRoute74k—~529Automated safety check: PassMIT

Similar skills

  • Content Hash Cache Pattern

    xu-xiang/everything-claude-code-zh

    使用 SHA-256 内容哈希缓存高昂的文件处理结果 —— 与路径无关、自动失效且服务层分离. An agent skill from xu-xiang/everything-claude-code-zh.

    2k GitHub stars~994 tokensUpdated 7 mo ago
    Auto-check passed
  • Prompt Caching

    sickn33/agentic-awesome-skills

    Caching strategies for LLM prompts including Anthropic prompt caching, response caching, and CAG (Cache Augmented Generation)

    47k GitHub starsUsed in 2 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Prompt Caching

    davila7/claude-code-templates

    Caching strategies for LLM prompts including Anthropic prompt caching, response caching, and CAG (Cache Augmented Generation) Use when: prompt caching, cache prompt, response cache, cag, cache…

    32k GitHub starsUsed in 5 repos~452 tokens
    Backend & APIsAuto-check passed
  • Turborepo Caching

    wshobson/agents

    Configures Turborepo pipelines and local or remote caching for monorepo builds, including Vercel remote cache, a self-hosted cache and cache-miss debugging.

    40k GitHub starsUsed in 9 repos~2k tokens
    DevelopmentAuto-check: notes
  • OmniRoute LLM Cache

    diegosouzapw/OmniRoute

    Documents OmniRoute's cache endpoints for reading cache statistics and clearing entries, statistics or the reasoning cache, with notes on TTL and similarity settings.

    74k GitHub stars~529 tokensUpdated today
    Backend & APIsAuto-check passed
  • Caching

    zebbern/claude-code-guide

    Caching strategies — invalidation, TTL guidelines, cache keys, cache layers, and when not to cache.

    4.7k GitHub stars~1.5k tokensUpdated today
    Backend & APIsAuto-check passed

More from affaan-m/ECC

All 673 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    276k GitHub starsUsed in 5 repos~1.9k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    276k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    276k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    276k GitHub stars~2.9k tokensUpdated 4 days ago
    Auto-check passed
  • Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

    276k GitHub starsUsed in 1 repo~623 tokens
    Auto-check passed
  • Instinct-based learning system that observes sessions via hooks, creates atomic instincts with confidence scoring, and evolves them into skills/commands/agents.

    276k GitHub stars~3.5k tokensUpdated 4 days ago
    Auto-check passed

Questions about Content Hash Cache Pattern

What does Content Hash Cache Pattern do?

使用SHA-256内容哈希缓存昂贵的文件处理结果——路径无关、自动失效、服务层分离. An agent skill from affaan-m/ECC. Content Hash Cache Pattern is an agent skill from affaan-m/ECC.

How do I install Content Hash Cache Pattern in Claude Code?

Run `npx skills add affaan-m/ECC --skill content-hash-cache-pattern -a claude-code`. Or copy the skill folder (docs/zh-CN/skills/content-hash-cache-pattern in affaan-m/ECC) into .claude/skills/content-hash-cache-pattern in your project. Claude Code loads it when a task matches its description.

How do I install Content Hash Cache Pattern in Codex?

Run `npx skills add affaan-m/ECC --skill content-hash-cache-pattern -a codex`. Or copy the skill folder (docs/zh-CN/skills/content-hash-cache-pattern in affaan-m/ECC) into .agents/skills/content-hash-cache-pattern in your project. Codex loads it when a task matches its description.

Can I use Content Hash Cache Pattern in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill content-hash-cache-pattern -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/content-hash-cache-pattern, .gemini/skills/content-hash-cache-pattern, .github/skills/content-hash-cache-pattern and .opencode/skills/content-hash-cache-pattern in your project.

What does Content Hash Cache Pattern need to run?

SKILL.md names no scripts, command-line tools or credentials: Content Hash Cache Pattern is instructions for the agent only. Our summary lists: Python 3.

Does Content Hash Cache Pattern access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Content Hash Cache Pattern safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Content Hash Cache Pattern use?

Content Hash Cache Pattern is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Content Hash Cache Pattern use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Content Hash Cache Pattern?

Skills that share tags, products or a category with Content Hash Cache Pattern: Content Hash Cache Pattern (xu-xiang/everything-claude-code-zh, 2k stars), Prompt Caching (sickn33/agentic-awesome-skills, 47k stars), Prompt Caching (davila7/claude-code-templates, 32k stars) and Turborepo Caching (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Content Hash Cache Pattern?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 275,546 GitHub stars. The repository holds 673 skills in this directory. The repository was last updated on October 5, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.