Agent skill

Prompt Caching

by Archive228 in Archive228/loopkit

Cache the parts of the prompt that don't change so a long-running loop stops paying full price on every turn.

MITAuto-check passedAI & LLM Engineering

Install Prompt Caching

skills CLI
$ npx skills add Archive228/loopkit --skill prompt-caching -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Archive228/loopkit prompt-caching --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/prompt-caching .claude/skills/prompt-caching && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
prompt-caching
GitHub stars
755
Token cost
~735 tokens
SKILL.md length
395 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Cache the parts of the prompt that don't change so a long-running loop stops paying full price on every turn.

  • Works in 4 steps: System prompt — mark the end of it as a… → Tool definitions — cache immediately… → Large stable docs — repo map, style… → …
  • The system prompt
  • SKILL.md covers Where to put breakpoints, TTL choice, Staleness rules — cache… and When NOT to cache, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Prompt Caching is an agent skill from Archive228/loopkit. Cache the parts of the prompt that don't change so a long-running loop stops paying full price on every turn. Use when the system prompt, tool defs, or reference docs are stable across many turns.

Its SKILL.md is about 740 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM cost and token optimization, Prompt engineering and Responsive design. The repository describes itself as: 33 battle-tested skills + minimal .claude harness for any coding agent (Claude Code, Cursor, Codex, Gemini CLI). The licence is MIT.

When your agent uses it

  • The system prompt
  • Reference docs are stable across many turns

Example prompts

  • “/prompt-caching”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. System prompt — mark the end of it as a breakpoint if it's >1024 tokens (Sonnet/Opus) or >2048 (Haiku).
  2. Tool definitions — cache immediately after, if the tool set is stable across the loop.
  3. Large stable docs — repo map, style guide, spec — before any turn-specific user text.
  4. User message stem — only if the same preamble repeats every turn.

What it can do on your machine

Read from SKILL.md and the folder at commit 5ae033e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prompt Caching loads about 735 tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 395 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~735

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Archive228/loopkit at commit 5ae033e, republished under its MIT licence (© Archive228). 395 words, ~735 tokens.

Download SKILL.mdSave it as .claude/skills/prompt-caching/SKILL.md (or your agent's skills folder).
name
prompt-caching
description
Cache the parts of the prompt that don't change so a long-running loop stops paying full price on every turn. Use when the system prompt, tool defs, or reference docs are stable across many turns.
when_to_use
long session, repeated large context, cost climbing, system prompt >1024 tokens, unchanged tool definitions across turns

Prompt Caching

Every turn of a Plan→Act→Verify loop resends the same system prompt, the same tool definitions, and (usually) the same reference docs. Without cache breakpoints you pay full input price on all of it, every turn. With them, cached reads cost ~10% of the write.

Where to put breakpoints

Cache from the top of the prompt down. The cache is prefix-matched — a break in the middle invalidates everything after it.

  1. System prompt — mark the end of it as a breakpoint if it's >1024 tokens (Sonnet/Opus) or >2048 (Haiku).
  2. Tool definitions — cache immediately after, if the tool set is stable across the loop.
  3. Large stable docs — repo map, style guide, spec — before any turn-specific user text.
  4. User message stem — only if the same preamble repeats every turn.

Everything past the last breakpoint is billed fresh every turn. That's fine — that's where the changing content goes.

TTL choice

  • 5-minute cache (default) — for tight loops where turns are seconds apart. Free to write.
  • 1-hour cache — for slow loops (human in the loop, background jobs). Write cost is higher; break-even is ~2 hits.

Pick 5m unless you know turns are minutes apart.

Staleness rules — cache invalidation is silent

  • Any byte change above the breakpoint invalidates the cache from that point.
  • Reordering tools or messages counts as change.
  • Trailing whitespace counts as change.
  • A different model version counts as change.

If cost isn't dropping, log the cache-hit metric. Do not assume.

Show full SKILL.md (154 more words)Show less

When NOT to cache

  • Prompt <1024 tokens — below the minimum block size, no savings.
  • One-shot calls — no reuse, cache write is wasted.
  • Highly dynamic system prompt (per-user templating) — cache misses will exceed hits.

Red flags

  • Cost graph flat after adding breakpoints — you're invalidating on every turn. Diff two consecutive requests byte-for-byte above the breakpoint.
  • Breakpoint after user message — pointless; the user message changes every turn.
  • Four+ breakpoints — max is four; extras are ignored silently.
  • Caching a prompt that gets edited mid-session — one edit above the cut wipes all downstream savings.

The math

Cache write ≈ 1.25× normal input. Cache read ≈ 0.1× normal input. So a stable 20K-token prefix hit N times: N=1 costs more than no-cache; N=2 breaks even; N=10 costs ~15% of no-cache. Long loops win big; short chats lose.

Cache is the single biggest cost lever on a long-running agent. Set it once at the start of the loop, verify hits, forget it.

© Archive228, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/prompt-caching of Archive228/loopkit.

Open the folder on GitHubat commit 5ae033e

Compare with similar skills

Prompt Caching next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prompt Caching compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prompt Caching this skillArchive228/loopkit755—~735Automated safety check: PassMIT
Context Auditundefined-ui/second-brain-os1k—~802Automated safety check: PassMIT
Senior Prompt Engineeralirezarezvani/claude-skills28k1 repos~2.5kAutomated safety check: PassMIT
Prompt Governancealirezarezvani/claude-skills28k—~2.8kAutomated safety check: PassMIT
LLM Cost Optimizeralirezarezvani/claude-skills28k—~2.9kAutomated safety check: PassMIT
Prompt Engineering InterviewerPrepLabsAI/InterviewMentor112—~5kAutomated safety check: PassMIT

Similar skills

  • Context Audit

    undefined-ui/second-brain-os

    Audit an agent's context layout against the four places: system prompt, tools, history, tail.

    1k GitHub stars~802 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Senior Prompt Engineer

    alirezarezvani/claude-skills

    A skill your agent uses when the user asks to optimize prompts, design prompt templates, evaluate LLM outputs with an eval set, measure RAG retrieval quality, validate agent/tool configurations…

    28k GitHub starsUsed in 1 repo~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Prompt Governance

    alirezarezvani/claude-skills

    A skill your agent uses when managing prompts in production at scale: versioning prompts, running A/B tests on prompts, building prompt registries, preventing prompt regressions, or creating eval…

    28k GitHub stars~2.8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLM Cost Optimizer

    alirezarezvani/claude-skills

    Use proactively whenever LLM API costs come up -- or should.

    28k GitHub stars~2.9k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Prompt Engineering Interviewer

    PrepLabsAI/InterviewMentor

    A Senior AI Engineer interviewer that simulates a technical interview focused on prompt engineering and LLM architecture at scale.

    112 GitHub stars~5k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Prompt Caching Patterns

    softspark/ai-toolkit

    Anthropic API prompt caching: TTL, breakpoints, stacking, invalidation, hit rate.

    179 GitHub stars~1.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from Archive228/loopkit

All 43 skills in this repo
  • Hitl Escalate

    Archive228/loopkit

    Escalate blocked runs to a human via configured channel or fallback to BLOCKED.md and exit the loop.

    755 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Structured Output

    Archive228/loopkit

    Get JSON out of the model reliably. An agent skill from Archive228/loopkit.

    755 GitHub stars~830 tokensUpdated 2 mo ago
    Auto-check passed
  • Using Loopkit

    Archive228/loopkit

    A skill your agent uses when starting any conversation in a loopkit-enabled project - establishes how to find and use loopkit's 49 skills, requiring skill invocation before ANY response including…

    755 GitHub stars~1.4k tokensUpdated 2 mo ago
    Auto-check passed
  • Active Memory Reminder

    Archive228/loopkit

    Before compaction Loopkit extracts decisions into claude-decisions.json (machine-readable).

    755 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Eval Harness

    Archive228/loopkit

    Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.

    755 GitHub stars~876 tokensUpdated 2 mo ago
    Auto-check passed
  • Feature List JSON

    Archive228/loopkit

    Enumerate every end-to-end feature as strict JSON entries with passes:false, editable-passes-only discipline, and priority order.

    755 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Prompt Caching

What does Prompt Caching do?

Cache the parts of the prompt that don't change so a long-running loop stops paying full price on every turn. Prompt Caching is an agent skill from Archive228/loopkit. Cache the parts of the prompt that don't change so a long-running loop stops paying full price on every turn.

When should I use Prompt Caching?

Prompt Caching fits situations like: the system prompt; reference docs are stable across many turns.

How do I install Prompt Caching in Claude Code?

Run `npx skills add Archive228/loopkit --skill prompt-caching -a claude-code`. Or copy the skill folder (skills/prompt-caching in Archive228/loopkit) into .claude/skills/prompt-caching in your project. Claude Code loads it when a task matches its description.

How do I install Prompt Caching in Codex?

Run `npx skills add Archive228/loopkit --skill prompt-caching -a codex`. Or copy the skill folder (skills/prompt-caching in Archive228/loopkit) into .agents/skills/prompt-caching in your project. Codex loads it when a task matches its description.

Can I use Prompt Caching in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Archive228/loopkit --skill prompt-caching -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prompt-caching, .gemini/skills/prompt-caching, .github/skills/prompt-caching and .opencode/skills/prompt-caching in your project.

What does Prompt Caching need to run?

SKILL.md names no scripts, command-line tools or credentials: Prompt Caching is instructions for the agent only.

Does Prompt Caching access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Prompt Caching safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prompt Caching use?

Prompt Caching is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prompt Caching use?

About 735 tokens (SKILL.md is roughly 2.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Prompt Caching?

Skills that share tags, products or a category with Prompt Caching: Context Audit (undefined-ui/second-brain-os, 1k stars), Senior Prompt Engineer (alirezarezvani/claude-skills, 28k stars), Prompt Governance (alirezarezvani/claude-skills, 28k stars) and LLM Cost Optimizer (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prompt Caching?

Archive228 (a GitHub user) maintains it in Archive228/loopkit, which has 755 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on July 14, 2026.

Source: Archive228/loopkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.