Agent skill

Caveman Evidence Review

by JuliusBrussee in JuliusBrussee/caveman

Read-only review of Caveman Cloud data to explain where LLM spend goes: cost, score, workflows, traces, latency, errors, routing and verified savings.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Caveman Evidence Review

skills CLI
$ npx skills add JuliusBrussee/caveman --skill caveman-evidence-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install JuliusBrussee/caveman caveman-evidence-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/JuliusBrussee/caveman.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/caveman-evidence-review .claude/skills/caveman-evidence-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
caveman-evidence-review
GitHub stars
111k
Used in
1 other repo
Token cost
~927 tokens
SKILL.md length
362 words
Files
1
Skills in repo
18
Repo updated
First seen
Licence
Apache-2.0

At a glance

Read-only review of Caveman Cloud data to explain where LLM spend goes: cost, score, workflows, traces, latency, errors, routing and verified savings.

  • Works in 5 steps: Load context → Establish baseline → Test the leading explanation with traces → …
  • Finding where LLM spend is going in a project
  • SKILL.md covers Hard rules, Step 1 — Load context, Step 2 — Establish baseline and Step 3 — Test the leading…, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

A read-only review of Caveman Cloud evidence that answers what Caveman found and where LLM spend goes: cost, Cave Score, workflows, traces, latency, errors, routing and savings. The agent builds conclusions from current Caveman data rather than repository guesses and never starts, approves, cancels or rolls back an experiment. Four numbers must stay in separate buckets and never be added or relabeled: measured list-price cost, inferred daily headroom, verified ledger savings and evidence cost.

Further rules: do not fetch prompt, completion, tool or artifact payloads unless you ask for payload review, scope every read to the selected project, treat empty results as no current signal rather than zero cost, and cite trace ids and time windows. The flow loads context through `caveman_context` or `caveman cloud whoami` and stops if login or a project is missing, then reads reports for overview, costs, score, workflows and verified savings plus a ranked plan, tests the leading explanation with bounded trace searches grouped by workflow, model or session against a control, and inspects a few representative traces. The excerpt ends in that last step.

When your agent uses it

  • Finding where LLM spend is going in a project
  • Explaining a change in cost, latency or error rate using traces
  • Checking verified savings against inferred headroom

Example prompts

  • “What did Caveman find about our LLM spend this week?”
  • “Which workflows drive the highest cost, and what do the traces say about retries?”
  • “Compare the verified savings with the inferred daily headroom for this project.”

Requirements

  • A Caveman Cloud login with a project selected
  • The Caveman MCP tools or the caveman CLI

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Load context
  2. Establish baseline
  3. Test the leading explanation with traces
  4. Inspect representative traces
  5. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 2e08b91. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Caveman Evidence Review loads about 927 tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 362 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~927

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from JuliusBrussee/caveman at commit 2e08b91, republished under its Apache-2.0 licence (© JuliusBrussee). 362 words, ~927 tokens.

Download SKILL.mdSave it as .claude/skills/caveman-evidence-review/SKILL.md (or your agent's skills folder).
name
caveman-evidence-review
description
Read-only review of Caveman Cloud evidence: cost, Cave Score, workflows, traces, latency, errors, routing, savings. Use when asked what Caveman found or where LLM spend goes.

Review Caveman evidence

Act as a read-only operator. Build conclusions from current Caveman data, not from repository guesses. Never start, approve, cancel, or roll back an experiment from this skill.

Hard rules

  1. Keep these buckets separate:
    • measured provider-complete list-price cost;
    • inferred daily headroom;
    • verified ledger savings;
    • evidence cost. Never add or relabel them.
  2. Do not fetch prompt, completion, tool, or artifact payloads unless the user explicitly asks for payload review. Metadata, spans, timing, models, token counts, status, and optimizer attribution are enough for the default review.
  3. Scope every read to the project selected by Caveman context. Never supply an organization id.
  4. Empty results are evidence of no current signal, not zero cost or zero risk.
  5. Cite trace ids and exact time windows used. Do not claim a cause from an aggregate alone.

Step 1 — Load context

Prefer MCP:

text
caveman_context {}

CLI fallback:

bash
caveman cloud whoami
caveman cloud projects list

Stop if login or project selection is missing. Ask the user to run caveman login or select a project; never guess.

Step 2 — Establish baseline

Use caveman_report for:

  • overview
  • costs
  • score
  • workflows
  • verified_savings

Then use caveman_plan for ranked daily headroom. If question is narrow, skip unrelated reports. Read shortest set that can answer it.

CLI fallback:

bash
caveman cloud costs
caveman cloud score
caveman cloud plan --json

State report window and basis before interpreting direction.

Show full SKILL.md (153 more words)Show less

Step 3 — Test the leading explanation with traces

Use caveman_trace_search. Choose a bounded window and closed filters: workflow, agent, model, provider, error code, runtime mode, cache status, optimization id, status class, token/cost/latency bounds, compression, or monitor verdict.

Useful groupings:

  • workflow — find jobs driving cost or failures;
  • model — compare model mix;
  • session — isolate retry or loop behavior;
  • ungrouped — identify exact traces.

Compare a suspect cohort with a control cohort or earlier bounded window. Do not infer causality from one expensive trace.

CLI fallback:

bash
caveman cloud traces search \
  --workflow <slug> \
  --from <RFC3339> \
  --to <RFC3339> \
  --sort total_cost_usd \
  --dir desc \
  --limit 25

Step 4 — Inspect representative traces

Call caveman_trace_get for a small number of high-signal trace ids. Inspect request and span metadata, latency, status, token counts, cache state, applied optimizers, and model route. Keep payload retrieval off.

CLI fallback:

bash
caveman cloud traces show <trace-id> --spans

Step 5 — Report

Use this shape:

text
## Caveman evidence review

Scope: <project> · <from> to <to>
Measured cost: <value and basis>
Verified savings: <ledger value, kept separate>
Inferred headroom: <per-day band, kept separate>

Findings:
1. <finding> — <aggregate evidence> — traces <ids>
2. <finding> — <aggregate evidence> — traces <ids>

Unproven:
- <plausible explanation lacking a control, trace, or eval>

Next read-only check:
- <one bounded query>

Possible action:
- <proposal only; use caveman-manage for read-only lifecycle review and safety gate>

If data is missing, name missing signal and stop at strongest supported statement. Never turn a catalog subtotal into an invoice or an experiment result into verified savings.

© JuliusBrussee, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/caveman-evidence-review of JuliusBrussee/caveman.

Open the folder on GitHubat commit 2e08b91

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in JuliusBrussee/caveman, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Caveman Evidence Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Caveman Evidence Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Caveman Evidence Review this skillJuliusBrussee/caveman111k1 repos~927Automated safety check: PassApache-2.0
CodexBar Usage Readersteipete/CodexBar22k—~320Automated safety check: PassMIT
Analyzing Claude Code Sessionsamd/gaia1.6k—~2.3kAutomated safety check: PassMIT
OmniRoute Cost and Usage CLIdiegosouzapw/OmniRoute75k—~693Automated safety check: PassMIT
Langfuse and LLM Gateway LogsKonghaYao/peri229—~4.3kAutomated safety check: NotesApache-2.0
Claude Code Daily Cost Reporttombelieber/claude-view111—~3kAutomated safety check: PassMIT

Similar skills

  • CodexBar Usage Reader

    steipete/CodexBar

    CodexBar read. Provider usage, limits, credits, config health. JSON. No writes.

    22k GitHub stars~320 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

    1.6k GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • OmniRoute Cost and Usage CLI

    diegosouzapw/OmniRoute

    View cost breakdowns, token usage, and call logs from the CLI. Filter by provider, model, or date range. Export usage reports and inspect per-connection…

    75k GitHub stars~693 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Queries Langfuse traces, prompts, datasets and sessions, and analyzes local LLM gateway logs for requests, context growth, token use and cache hits.

    229 GitHub stars~4.3k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Claude Code Daily Cost Report

    tombelieber/claude-view

    Shows today's Claude Code spending through the claude-view MCP server, with total cost, running sessions and a per-session breakdown, and other date ranges on request.

    111 GitHub stars~3k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Sentry Instrument

    getsentry/sentry-for-ai

    Official

    Instrument an application with Sentry — detect the platform, install and initialize the SDK if needed, and wire up any signal — error monitoring, tracing/performance, logging, metrics, profiling…

    268 GitHub stars~3.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from JuliusBrussee/caveman

All 18 skills in this repo
  • Caveman Workflow Labeler

    JuliusBrussee/caveman

    Finds every LLM workflow in a repository, proposes a labeling table and, once you agree, wires labels so Caveman Cloud groups spend per workflow.

    111k GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • Caveman Experiment Manager

    JuliusBrussee/caveman

    Reads the state and results of Caveman Cloud experiments and reports one recommendation or a block, without changing an experiment's lifecycle itself.

    111k GitHub starsUsed in 1 repo~975 tokens
    Auto-check passed
  • Caveman Optimization Evaluator

    JuliusBrussee/caveman

    Turns a Caveman report-only optimization observation into one minimal code change and a paired baseline evaluation, after the operator picks which to pursue.

    111k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Caveman Gateway Setup

    JuliusBrussee/caveman

    Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.

    111k GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check: warnings
  • Caveman Help Card

    JuliusBrussee/caveman

    Quick-reference card for the three caveman skills and their commands. Trigger: /caveman-help or "caveman help".

    111k GitHub stars~690 tokensUpdated 2 days ago
    Auto-check passed
  • Caveman Mode

    JuliusBrussee/caveman

    Switches the agent to a terse reply style that gives the answer first, drops filler, and keeps every technical fact, command and number exact.

    111k GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed

Questions about Caveman Evidence Review

What does Caveman Evidence Review do?

Read-only review of Caveman Cloud data to explain where LLM spend goes: cost, score, workflows, traces, latency, errors, routing and verified savings. A read-only review of Caveman Cloud evidence that answers what Caveman found and where LLM spend goes: cost, Cave Score, workflows, traces, latency, errors, routing and savings. The agent builds conclusions from current Caveman data rather than repository guesses and never starts, approves, cancels or rolls back an experiment.

When should I use Caveman Evidence Review?

Caveman Evidence Review fits situations like: finding where LLM spend is going in a project; explaining a change in cost, latency or error rate using traces; checking verified savings against inferred headroom.

How do I install Caveman Evidence Review in Claude Code?

Run `npx skills add JuliusBrussee/caveman --skill caveman-evidence-review -a claude-code`. Or copy the skill folder (skills/caveman-evidence-review in JuliusBrussee/caveman) into .claude/skills/caveman-evidence-review in your project. Claude Code loads it when a task matches its description.

How do I install Caveman Evidence Review in Codex?

Run `npx skills add JuliusBrussee/caveman --skill caveman-evidence-review -a codex`. Or copy the skill folder (skills/caveman-evidence-review in JuliusBrussee/caveman) into .agents/skills/caveman-evidence-review in your project. Codex loads it when a task matches its description.

Can I use Caveman Evidence Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JuliusBrussee/caveman --skill caveman-evidence-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/caveman-evidence-review, .gemini/skills/caveman-evidence-review, .github/skills/caveman-evidence-review and .opencode/skills/caveman-evidence-review in your project.

What does Caveman Evidence Review need to run?

SKILL.md names no scripts, command-line tools or credentials: Caveman Evidence Review is instructions for the agent only. Our summary lists: A Caveman Cloud login with a project selected; The Caveman MCP tools or the caveman CLI.

Does Caveman Evidence Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Caveman Evidence Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Caveman Evidence Review use?

Caveman Evidence Review is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Caveman Evidence Review use?

About 927 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Caveman Evidence Review?

Skills that share tags, products or a category with Caveman Evidence Review: CodexBar Usage Reader (steipete/CodexBar, 22k stars), Analyzing Claude Code Sessions (amd/gaia, 1.6k stars), OmniRoute Cost and Usage CLI (diegosouzapw/OmniRoute, 75k stars) and Langfuse and LLM Gateway Logs (KonghaYao/peri, 229 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Caveman Evidence Review?

JuliusBrussee (a GitHub user) maintains it in JuliusBrussee/caveman, which has 110,815 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 9, 2026.

Source: JuliusBrussee/caveman on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.