Agent skill

Cache Efficiency

by hoangsonww in hoangsonww/Claude-Code-Agent-Monitor

Analyze prompt-cache effectiveness for Claude Code usage from the Agent Monitor dashboard — cache hit rate (totalcacheread / (totalcacheread + totalinput)), cachewrite vs cacheread reuse, cache-read…

MITAuto-check passedAI & LLM Engineering

Install Cache Efficiency

skills CLI
$ npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill cache-efficiency -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hoangsonww/Claude-Code-Agent-Monitor cache-efficiency --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hoangsonww/Claude-Code-Agent-Monitor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ccam-analytics/skills/cache-efficiency .claude/skills/cache-efficiency && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cache-efficiency
GitHub stars
1k
Token cost
~966 tokens
SKILL.md length
393 words
Files
2
Skills in repo
78
Repo updated
First seen
Licence
MIT

At a glance

Analyze prompt-cache effectiveness for Claude Code usage from the Agent Monitor dashboard — cache hit rate (totalcacheread / (totalcacheread + totalinput)), cachewrite vs cacheread reuse, cache-read…

  • Works in 5 steps: Fleet Cache Hit Rate → Write vs Read Reuse → Cache Spend Split → …
  • Diagnosing cache spend
  • SKILL.md covers Input, Data Sources, Report Sections and Output
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cache Efficiency is an agent skill from hoangsonww/Claude-Code-Agent-Monitor. Analyze prompt-cache effectiveness for Claude Code usage from the Agent Monitor dashboard — cache hit rate (totalcacheread / (totalcacheread + totalinput)), cachewrite vs cacheread reuse, cache-read vs cache-write spend, and the sessions with the poorest reuse. Pulls token totals from /api/analytics, per-session detail from /api/sessions, and dollar splits from /api/pricing/cost. Use when diagnosing cache spend or deciding whether prompt caching is paying off.

Its SKILL.md is about 970 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in AI & LLM Engineering, covering LLM cost and token optimization. The repository describes itself as: 🚀 A real-time monitoring dashboard for Claude Code & Codex, built with SQLite3, Node.js, Express, React, Vite, TailwindCSS, & WebSockets. It tracks sessions, agent activity… The licence is MIT.

When your agent uses it

  • Diagnosing cache spend
  • Deciding whether prompt caching is paying off

Example prompts

  • “/cache-efficiency”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Fleet Cache Hit Rate
  2. Write vs Read Reuse
  3. Cache Spend Split
  4. Sessions With Poor Reuse
  5. Recommendations

What it can do on your machine

Read from SKILL.md and the folder at commit e0f4a1a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cache Efficiency loads about 966 tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 393 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~966

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hoangsonww/Claude-Code-Agent-Monitor at commit e0f4a1a, republished under its MIT licence (© hoangsonww). 393 words, ~966 tokens.

Download SKILL.mdSave it as .claude/skills/cache-efficiency/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
cache-efficiency
description
Analyze prompt-cache effectiveness for Claude Code usage from the Agent Monitor dashboard — cache hit rate (total_cache_read / (total_cache_read + total_input)), cache_write vs cache_read reuse, cache-read vs cache-write spend, and the sessions with the poorest reuse. Pulls token totals from /api/analytics, per-session detail from /api/sessions, and dollar splits from /api/pricing/cost. Use when diagnosing cache spend or deciding whether prompt caching is paying off.

Cache Efficiency

Diagnose whether prompt caching is actually saving money, and where it is not.

Input

The user provides: $ARGUMENTS

This may be: empty (analyze the whole fleet), "today" / "this week" / a date range, a session ID to scope the analysis, or a target like "hit rate > 80%". When empty, analyze all data from /api/analytics.

Data Sources

EndpointReturns
GET /api/analyticstokens.total_input, tokens.total_output, tokens.total_cache_read, tokens.total_cache_write (baselines pre-summed), plus daily_sessions
GET /api/sessions?limit=200Session list — each has model, cwd, started_at, ended_at, inline cost, metadata (JSON: usage_extras with cache token detail)
GET /api/sessions/{id}Full session detail with nested agents and events, for drill-down on a flagged session
GET /api/pricing/cost{ total_cost, breakdown: [{ model, input_tokens, output_tokens, cache_read_tokens, cache_write_tokens, cost, matched_rule }] } — used to price cache read vs write spend
How cache economics work
cache_hit_rate   = total_cache_read / (total_cache_read + total_input)
cache_reuse      = total_cache_read / total_cache_write
cache_read_cost  = (cache_read_tokens  / 1M) × cache_read_per_mtok
cache_write_cost = (cache_write_tokens / 1M) × cache_write_per_mtok

Cache writes cost more per token than cache reads (e.g. Sonnet $3.75 write vs $0.30 read per Mtok), and writes are billed even if the cached block is never reused. The payoff only arrives on subsequent reads — so a healthy fleet shows cache_read_tokens far exceeding cache_write_tokens. When cache_reuse < 1, you are paying to cache context you barely re-read.

Token counts are effective totals = current + baseline (baselines preserve pre-compaction tokens).

Report Sections

1. Fleet Cache Hit Rate

From /api/analytics: compute cache_hit_rate × 100. State raw total_cache_read and total_input. Benchmark: >70% strong, 40–70% moderate, <40% weak prompt-cache utilization.

2. Write vs Read Reuse

Compute cache_reuse = total_cache_read / total_cache_write. Show both token counts. Flag if reuse < 1 (writing more cache than is ever read back).

Show full SKILL.md (146 more words)Show less
3. Cache Spend Split

From /api/pricing/cost breakdown, sum cache_read_cost and cache_write_cost across all models. Show the dollar split and what fraction of total cost is cache-write overhead vs cache-read savings.

4. Sessions With Poor Reuse

From /api/sessions?limit=200, parse metadata.usage_extras for per-session cache read/write where available; rank sessions by lowest read/write reuse (and by cache_write-heavy cost). List the worst 10 with model, cost, and reuse ratio. Use /api/sessions/{id} to drill into any single flagged session.

5. Recommendations
  • Sessions where cache_write >> cache_read: short or one-shot sessions rarely recoup cache writes — note them.
  • Stable, repeated context (system prompts, large files) should be cached once and reused; high churn defeats caching.
  • Estimate the dollar impact of raising the hit rate to the next benchmark tier.

Output

Structured Markdown with tables. Currency as USD to 4 decimal places; rates as $/Mtok; percentages with ▲/▼ for any trend. Token counts with thousands separators.

© hoangsonww, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in plugins/ccam-analytics/skills/cache-efficiency of hoangsonww/Claude-Code-Agent-Monitor.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit e0f4a1a

Compare with similar skills

Cache Efficiency next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cache Efficiency compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cache Efficiency this skillhoangsonww/Claude-Code-Agent-Monitor1k—~966Automated safety check: PassMIT
Context Compressionguanyang/open-agent-hub9732 repos~4.6kAutomated safety check: PassMIT
Bounty Hunter1sadjlk/bounty-hunter-skill2821 repos~761Automated safety check: PassMIT
Skill Shortenerluongnv89/asm953—~3.8kAutomated safety check: NotesMIT
Fleet Auditoralexgreensh/token-optimizer2.5k—~1.7kAutomated safety check: PassCustom licence
Context Auditundefined-ui/second-brain-os999—~810Automated safety check: PassMIT

Similar skills

  • Context Compression

    guanyang/open-agent-hub

    This skill should be used when long-running agent sessions need context compression, structured summarization, compaction, token-per-task optimization, or durable handoff summaries that preserve…

    973 GitHub starsUsed in 2 repos~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Bounty Hunter

    1sadjlk/bounty-hunter-skill

    A professional AI bounty hunter persona named Atlas. An agent skill from 1sadjlk/bounty-hunter-skill.

    282 GitHub starsUsed in 1 repo~761 tokens
    AI & LLM EngineeringAuto-check passed
  • Skill Shortener

    luongnv89/asm

    Refactor a too-long SKILL.md by progressive disclosure: measure token cost, classify every section KEEP/CUT/MOVE, shorten the body into references/ and scripts/, verify nothing was lost.

    953 GitHub stars~3.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Fleet Auditor

    alexgreensh/token-optimizer

    Cross-system agent token/cost audit (Claude Code, Codex, OpenClaw, Hermes, OpenCode): idle burns, model misrouting, config bloat, with dollar savings.

    2.5k GitHub stars~1.7k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Context Audit

    undefined-ui/second-brain-os

    Audit an agent's context layout against the four places: system prompt, tools, history, tail.

    999 GitHub stars~810 tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Headroom

    momori777/Artemis

    SmartCrusher + CCR context compression — crunch large JSON arrays, tool outputs, and search results to save tokens.

    377 GitHub stars~562 tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed

More from hoangsonww/Claude-Code-Agent-Monitor

All 78 skills in this repo
  • Budget Set

    hoangsonww/Claude-Code-Agent-Monitor

    Define a spend budget for Claude Code and, optionally, create a cost alert rule that fires when usage crosses the limit, via POST /api/alerts/rules on the Agent Monitor dashboard.

    1k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Version Release

    hoangsonww/Claude-Code-Agent-Monitor

    Choose and apply the correct semantic version bump for this repository.

    1k GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Cost Breakdown

    hoangsonww/Claude-Code-Agent-Monitor

    Break down Claude Code costs using the Agent Monitor pricing engine.

    1k GitHub starsUsed in 1 repo~845 tokens
    Auto-check passed
  • Hook Diagnostics

    hoangsonww/Claude-Code-Agent-Monitor

    Diagnose Claude Code hook installation, delivery, and ingestion issues.

    1k GitHub starsUsed in 1 repo~676 tokens
    Auto-check passed
  • Memory Review

    hoangsonww/Claude-Code-Agent-Monitor

    Review the file-based memory store via the Agent Monitor Config Explorer API: the user and project CLAUDE.md plus per-project auto-memory files under ~/.claude/projects/<slug/memory/.md.

    1k GitHub starsUsed in 1 repo~970 tokens
    Auto-check passed
  • Transcript Grep

    hoangsonww/Claude-Code-Agent-Monitor

    Search a Claude Code session transcript for a string or regex pattern and show every matching message with surrounding context.

    1k GitHub starsUsed in 1 repo~706 tokens
    Auto-check passed

Questions about Cache Efficiency

What does Cache Efficiency do?

Analyze prompt-cache effectiveness for Claude Code usage from the Agent Monitor dashboard — cache hit rate (totalcacheread / (totalcacheread + totalinput)), cachewrite vs cacheread reuse, cache-read…. Cache Efficiency is an agent skill from hoangsonww/Claude-Code-Agent-Monitor. Analyze prompt-cache effectiveness for Claude Code usage from the Agent Monitor dashboard — cache hit rate (totalcacheread / (totalcacheread + totalinput)), cachewrite vs cacheread reuse, cache-read vs cache-write spend, and the sessions with the poorest reuse.

When should I use Cache Efficiency?

Cache Efficiency fits situations like: diagnosing cache spend; deciding whether prompt caching is paying off.

How do I install Cache Efficiency in Claude Code?

Run `npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill cache-efficiency -a claude-code`. Or copy the skill folder (plugins/ccam-analytics/skills/cache-efficiency in hoangsonww/Claude-Code-Agent-Monitor) into .claude/skills/cache-efficiency in your project. Claude Code loads it when a task matches its description.

How do I install Cache Efficiency in Codex?

Run `npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill cache-efficiency -a codex`. Or copy the skill folder (plugins/ccam-analytics/skills/cache-efficiency in hoangsonww/Claude-Code-Agent-Monitor) into .agents/skills/cache-efficiency in your project. Codex loads it when a task matches its description.

Can I use Cache Efficiency in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill cache-efficiency -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cache-efficiency, .gemini/skills/cache-efficiency, .github/skills/cache-efficiency and .opencode/skills/cache-efficiency in your project.

What does Cache Efficiency need to run?

SKILL.md names no scripts, command-line tools or credentials: Cache Efficiency is instructions for the agent only.

Does Cache Efficiency access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cache Efficiency safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cache Efficiency use?

Cache Efficiency is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cache Efficiency use?

About 966 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cache Efficiency?

Skills that share tags, products or a category with Cache Efficiency: Context Compression (guanyang/open-agent-hub, 973 stars), Bounty Hunter (1sadjlk/bounty-hunter-skill, 282 stars), Skill Shortener (luongnv89/asm, 953 stars) and Fleet Auditor (alexgreensh/token-optimizer, 2.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cache Efficiency?

hoangsonww (a GitHub user) maintains it in hoangsonww/Claude-Code-Agent-Monitor, which has 1,049 GitHub stars. The repository holds 78 skills in this directory. The repository was last updated on October 6, 2026.

Source: hoangsonww/Claude-Code-Agent-Monitor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.