Agent skill

Content-Hash File Cache Pattern

by affaan-m in affaan-m/ECC

Caches slow file processing results in Python keyed by a SHA-256 hash of the file content, so renames still hit the cache and edits invalidate it automatically.

MITAuto-check passedDevelopment

Install Content-Hash File Cache Pattern

skills CLI
$ npx skills add affaan-m/ECC --skill content-hash-cache-pattern -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC content-hash-cache-pattern --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/content-hash-cache-pattern .claude/skills/content-hash-cache-pattern && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
content-hash-cache-pattern
GitHub stars
276k
Used in
5 other repos
Token cost
~1.4k tokens
SKILL.md length
330 words
Files
1
Skills in repo
673
Repo updated
First seen
Licence
MIT

At a glance

Caches slow file processing results in Python keyed by a SHA-256 hash of the file content, so renames still hit the cache and edits invalidate it automatically.

  • Works in 4 steps: Content-Hash Based Cache Key → Frozen Dataclass for Cache Entry → File-Based Cache Storage → …
  • Adding caching to a PDF, image or text extraction pipeline
  • SKILL.md covers When to Activate, Core Pattern, Key Design Decisions and Best Practices, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

This skill describes a caching pattern for expensive file work such as PDF parsing, text extraction and image analysis. Instead of keying the cache on a path, it hashes the file's content with SHA-256, so a moved or renamed file still hits the cache and a changed file misses automatically, with no index file to maintain.

The Python examples use hashlib to compute the key, a frozen slotted dataclass for each cache entry, and one `{hash}.json` file per entry for constant-time lookup. Caching sits in a separate service-layer wrapper such as `extract_with_cache` around a pure processing function, so the extraction code knows nothing about it. A corrupt entry returns `None` and counts as a miss, large files are hashed in chunks, and hits and misses are logged with truncated hashes. Path-based caches are listed as an anti-pattern.

When your agent uses it

  • Adding caching to a PDF, image or text extraction pipeline
  • Wrapping an existing pure function with a cache without editing it
  • Adding a --cache/--no-cache option to a command-line tool
  • Replacing path-based cache keys that break when files move

Example prompts

  • “Add a content-hash cache around our extract_text function so repeated PDF runs are skipped.”
  • “Give the CLI a --cache/--no-cache flag backed by SHA-256 keyed JSON files.”
  • “Our cache misses whenever invoices are renamed; switch it to key on file content.”

Requirements

  • Python

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Content-Hash Based Cache Key
  2. Frozen Dataclass for Cache Entry
  3. File-Based Cache Storage
  4. Service Layer Wrapper (SRP)

What it can do on your machine

Read from SKILL.md and the folder at commit ef648e0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Content-Hash File Cache Pattern loads about 1.4k tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 330 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit ef648e0, republished under its MIT licence (© affaan-m). 330 words, ~1,430 tokens.

Download SKILL.mdSave it as .claude/skills/content-hash-cache-pattern/SKILL.md (or your agent's skills folder).
name
content-hash-cache-pattern
description
Cache expensive file processing results using SHA-256 content hashes — path-independent, auto-invalidating, with service layer separation. Use when repeated file processing is slow and results should be cached and invalidated by content rather than path.
metadata.origin
ECC

Content-Hash File Cache Pattern

Cache expensive file processing results (PDF parsing, text extraction, image analysis) using SHA-256 content hashes as cache keys. Unlike path-based caching, this approach survives file moves/renames and auto-invalidates when content changes.

When to Activate

  • Building file processing pipelines (PDF, images, text extraction)
  • Processing cost is high and same files are processed repeatedly
  • Need a --cache/--no-cache CLI option
  • Want to add caching to existing pure functions without modifying them

Core Pattern

1. Content-Hash Based Cache Key

Use file content (not path) as the cache key:

python
import hashlib
from pathlib import Path

_HASH_CHUNK_SIZE = 65536  # 64KB chunks for large files

def compute_file_hash(path: Path) -> str:
    """SHA-256 of file contents (chunked for large files)."""
    if not path.is_file():
        raise FileNotFoundError(f"File not found: {path}")
    sha256 = hashlib.sha256()
    with open(path, "rb") as f:
        while True:
            chunk = f.read(_HASH_CHUNK_SIZE)
            if not chunk:
                break
            sha256.update(chunk)
    return sha256.hexdigest()

Why content hash? File rename/move = cache hit. Content change = automatic invalidation. No index file needed.

2. Frozen Dataclass for Cache Entry
python
from dataclasses import dataclass

@dataclass(frozen=True, slots=True)
class CacheEntry:
    file_hash: str
    source_path: str
    document: ExtractedDocument  # The cached result
3. File-Based Cache Storage

Each cache entry is stored as {hash}.json — O(1) lookup by hash, no index file required.

python
import json
from typing import Any

def write_cache(cache_dir: Path, entry: CacheEntry) -> None:
    cache_dir.mkdir(parents=True, exist_ok=True)
    cache_file = cache_dir / f"{entry.file_hash}.json"
    data = serialize_entry(entry)
    cache_file.write_text(json.dumps(data, ensure_ascii=False), encoding="utf-8")

def read_cache(cache_dir: Path, file_hash: str) -> CacheEntry | None:
    cache_file = cache_dir / f"{file_hash}.json"
    if not cache_file.is_file():
        return None
    try:
        raw = cache_file.read_text(encoding="utf-8")
        data = json.loads(raw)
        return deserialize_entry(data)
    except (json.JSONDecodeError, ValueError, KeyError):
        return None  # Treat corruption as cache miss
4. Service Layer Wrapper (SRP)

Keep the processing function pure. Add caching as a separate service layer.

python
def extract_with_cache(
    file_path: Path,
    *,
    cache_enabled: bool = True,
    cache_dir: Path = Path(".cache"),
) -> ExtractedDocument:
    """Service layer: cache check -> extraction -> cache write."""
    if not cache_enabled:
        return extract_text(file_path)  # Pure function, no cache knowledge

    file_hash = compute_file_hash(file_path)

    # Check cache
    cached = read_cache(cache_dir, file_hash)
    if cached is not None:
        logger.info("Cache hit: %s (hash=%s)", file_path.name, file_hash[:12])
        return cached.document

    # Cache miss -> extract -> store
    logger.info("Cache miss: %s (hash=%s)", file_path.name, file_hash[:12])
    doc = extract_text(file_path)
    entry = CacheEntry(file_hash=file_hash, source_path=str(file_path), document=doc)
    write_cache(cache_dir, entry)
    return doc

Key Design Decisions

DecisionRationale
SHA-256 content hashPath-independent, auto-invalidates on content change
{hash}.json file namingO(1) lookup, no index file needed
Service layer wrapperSRP: extraction stays pure, cache is a separate concern
Manual JSON serializationFull control over frozen dataclass serialization
Corruption returns NoneGraceful degradation, re-processes on next run
cache_dir.mkdir(parents=True)Lazy directory creation on first write

Best Practices

  • Hash content, not paths — paths change, content identity doesn't
  • Chunk large files when hashing — avoid loading entire files into memory
  • Keep processing functions pure — they should know nothing about caching
  • Log cache hit/miss with truncated hashes for debugging
  • Handle corruption gracefully — treat invalid cache entries as misses, never crash

Anti-Patterns to Avoid

python
# BAD: Path-based caching (breaks on file move/rename)
cache = {"/path/to/file.pdf": result}

# BAD: Adding cache logic inside the processing function (SRP violation)
def extract_text(path, *, cache_enabled=False, cache_dir=None):
    if cache_enabled:  # Now this function has two responsibilities
        ...

# BAD: Using dataclasses.asdict() with nested frozen dataclasses
# (can cause issues with complex nested types)
data = dataclasses.asdict(entry)  # Use manual serialization instead

When to Use

  • File processing pipelines (PDF parsing, OCR, text extraction, image analysis)
  • CLI tools that benefit from --cache/--no-cache options
  • Batch processing where the same files appear across runs
  • Adding caching to existing pure functions without modifying them

When NOT to Use

  • Data that must always be fresh (real-time feeds)
  • Cache entries that would be extremely large (consider streaming instead)
  • Results that depend on parameters beyond file content (e.g., different extraction configs)

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/content-hash-cache-pattern of affaan-m/ECC.

Open the folder on GitHubat commit ef648e0

Used in 5 other repositories

We found 12 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in affaan-m/ECC, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Content-Hash File Cache Pattern next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Content-Hash File Cache Pattern compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Content-Hash File Cache Pattern this skillaffaan-m/ECC276k5 repos~1.4kAutomated safety check: PassMIT
Keybase RPC Log Analysiskeybase/client9.3k—~3kAutomated safety check: PassBSD-3-Clause
Climber Step Minimizationben-manes/caffeine18k—~3kAutomated safety check: NotesApache-2.0
Performance CheckZeroDeng01/sublinkPro1.7k—~1.8kAutomated safety check: PassMIT
Mastering Python SkillSpillwaveSolutions/agent-brain119—~1.4kAutomated safety check: NotesMIT
Eviction Policy Regret Auditben-manes/caffeine18k—~16kAutomated safety check: NotesApache-2.0

Similar skills

  • Captures a clean Keybase service log and analyzes it for redundant, duplicated or looping RPCs, then checks whether a caching fix reduced the calls.

    9.3k GitHub stars~3k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Climber Step Minimization

    ben-manes/caffeine

    Prices each step of the window climber algorithm by disabling it in turn, to find steps that no longer earn their keep and branches that no longer fire.

    18k GitHub stars~3k tokensUpdated today
    DevelopmentAuto-check: notes
  • Performance Check

    ZeroDeng01/sublinkPro

    Checklist for reviewing code changes that touch queries, APIs, rendering, caching or algorithms for performance, scalability and resource-usage problems.

    1.7k GitHub stars~1.8k tokensUpdated today
    DevelopmentAuto-check passed
  • Mastering Python Skill

    SpillwaveSolutions/agent-brain

    Modern Python coaching covering language foundations through advanced production patterns.

    119 GitHub stars~1.4k tokensUpdated 20 days ago
    DevelopmentAuto-check: notes
  • Searches synthetic workloads for cases where Caffeine's adaptive eviction policy falls short of its achievable hit rate, then classifies each gap by cause.

    18k GitHub stars~16k tokensUpdated today
    DevelopmentAuto-check: notes
  • Performance Optimization

    ThibautBaissac/rails_ai_agents

    Identifies and fixes Rails performance issues including N+1 queries, slow queries, and memory problems.

    665 GitHub stars~1.2k tokensUpdated 4 mo ago
    DevelopmentAuto-check: notes

More from affaan-m/ECC

All 673 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    276k GitHub starsUsed in 5 repos~1.9k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    276k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    276k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    276k GitHub stars~2.9k tokensUpdated 4 days ago
    Auto-check passed
  • Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

    276k GitHub starsUsed in 1 repo~623 tokens
    Auto-check passed
  • Instinct-based learning system that observes sessions via hooks, creates atomic instincts with confidence scoring, and evolves them into skills/commands/agents.

    276k GitHub stars~3.5k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about Content-Hash File Cache Pattern

What does Content-Hash File Cache Pattern do?

Caches slow file processing results in Python keyed by a SHA-256 hash of the file content, so renames still hit the cache and edits invalidate it automatically. This skill describes a caching pattern for expensive file work such as PDF parsing, text extraction and image analysis. Instead of keying the cache on a path, it hashes the file's content with SHA-256, so a moved or renamed file still hits the cache and a changed file misses automatically, with no index file to maintain.

When should I use Content-Hash File Cache Pattern?

Content-Hash File Cache Pattern fits situations like: adding caching to a PDF, image or text extraction pipeline; wrapping an existing pure function with a cache without editing it; adding a --cache/--no-cache option to a command-line tool; replacing path-based cache keys that break when files move.

How do I install Content-Hash File Cache Pattern in Claude Code?

Run `npx skills add affaan-m/ECC --skill content-hash-cache-pattern -a claude-code`. Or copy the skill folder (skills/content-hash-cache-pattern in affaan-m/ECC) into .claude/skills/content-hash-cache-pattern in your project. Claude Code loads it when a task matches its description.

How do I install Content-Hash File Cache Pattern in Codex?

Run `npx skills add affaan-m/ECC --skill content-hash-cache-pattern -a codex`. Or copy the skill folder (skills/content-hash-cache-pattern in affaan-m/ECC) into .agents/skills/content-hash-cache-pattern in your project. Codex loads it when a task matches its description.

Can I use Content-Hash File Cache Pattern in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill content-hash-cache-pattern -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/content-hash-cache-pattern, .gemini/skills/content-hash-cache-pattern, .github/skills/content-hash-cache-pattern and .opencode/skills/content-hash-cache-pattern in your project.

What does Content-Hash File Cache Pattern need to run?

SKILL.md names no scripts, command-line tools or credentials: Content-Hash File Cache Pattern is instructions for the agent only. Our summary lists: Python.

Does Content-Hash File Cache Pattern access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Content-Hash File Cache Pattern safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Content-Hash File Cache Pattern use?

Content-Hash File Cache Pattern is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Content-Hash File Cache Pattern use?

About 1.4k tokens (SKILL.md is roughly 5.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Content-Hash File Cache Pattern?

Skills that share tags, products or a category with Content-Hash File Cache Pattern: Keybase RPC Log Analysis (keybase/client, 9.3k stars), Climber Step Minimization (ben-manes/caffeine, 18k stars), Performance Check (ZeroDeng01/sublinkPro, 1.7k stars) and Mastering Python Skill (SpillwaveSolutions/agent-brain, 119 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Content-Hash File Cache Pattern?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 275,546 GitHub stars. The repository holds 673 skills in this directory. The repository was last updated on October 5, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.