Agent skill

Cost Counterfactual

by ruvnet in ruvnet/ruflo

Multi-baseline counterfactual cost analysis. An agent skill from ruvnet/ruflo.

MITAuto-check: notes

Install Cost Counterfactual

skills CLI
$ npx skills add ruvnet/ruflo --skill cost-counterfactual -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ruvnet/ruflo cost-counterfactual --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ruvnet/ruflo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ruflo-cost-tracker/skills/cost-counterfactual .claude/skills/cost-counterfactual && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cost-counterfactual
GitHub stars
74k
Token cost
~786 tokens
SKILL.md length
293 words
Files
1
Skills in repo
265
Repo updated
First seen
Licence
MIT

At a glance

Multi-baseline counterfactual cost analysis. An agent skill from ruvnet/ruflo.

  • Works in 6 steps: Read all session-* records from the… → Apply --since window filter (default… → Sum tokens across byModel[*] entries for… → …
  • SKILL.md covers Algorithm, Smoke transcript (2 sessions:…, How to read negative savings and When to use, plus 1 more section
  • Calls jq

What it does

Cost Counterfactual is an agent skill from ruvnet/ruflo. Multi-baseline counterfactual cost analysis. Compares actual session spend to hypothetical always-haiku / always-sonnet / always-opus routing baselines. Answers "is the routing earning its keep?" Negative savings flag over-escalation; positive savings quantify the router's win.

Its SKILL.md is about 790 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: 🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory…. The licence is MIT.

Example prompts

  • “is the routing earning its keep?”
  • “/cost-counterfactual”

Requirements

  • Pre-approved tools (allowed-tools): Bash

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Read all session-* records from the cost-tracking namespace.
  2. Apply --since window filter (default all-time).
  3. Sum tokens across byModel[*] entries for each session.
  4. For each requested baseline (default: all three)
  5. Compute savings = counterfactualUsd − actualUsd.
  6. Emit per-baseline totals + savings % across the comparison set.

What it can do on your machine

Read from SKILL.md and the folder at commit 6c04654. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cost Counterfactual loads about 786 tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 293 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~786

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ruvnet/ruflo at commit 6c04654, republished under its MIT licence (© ruvnet). 293 words, ~786 tokens.

Download SKILL.mdSave it as .claude/skills/cost-counterfactual/SKILL.md (or your agent's skills folder).
name
cost-counterfactual
description
Multi-baseline counterfactual cost analysis. Compares actual session spend to hypothetical always-haiku / always-sonnet / always-opus routing baselines. Answers "is the routing earning its keep?" Negative savings flag over-escalation; positive savings quantify the router's win.
allowed-tools
Bash
argument-hint
[--since 7d] [--baseline always-haiku|always-sonnet|always-opus|all] [--format table|json]

Multi-baseline counterfactual cost analysis. Pairs with the existing observability surface:

  • cost-budget-check — "have we crossed a threshold?" (reactive)
  • cost-projection — "when will we cross a threshold?" (predictive)
  • cost-counterfactual — "is the routing earning its keep?" (comparative) ← this one

Algorithm

  1. Read all session-* records from the cost-tracking namespace.
  2. Apply --since window filter (default all-time).
  3. Sum tokens across byModel[*] entries for each session.
  4. For each requested baseline (default: all three):
    • counterfactualUsd = (input × tier.input + output × tier.output + cache_write × tier.cache_write + cache_read × tier.cache_read) / 1M
  5. Compute savings = counterfactualUsd − actualUsd.
  6. Emit per-baseline totals + savings % across the comparison set.

Smoke transcript (2 sessions: 50K haiku tokens + 50K sonnet tokens)

| Sessions considered | 2 |
| Total input tokens  | 100,000 |
| Actual spend        | $0.162500 |

| Baseline           | Hypothetical | Actual    | Savings    | %       |
| `always-haiku`     | $0.025000    | $0.162500 | -$0.137500 | -550.00% |
| `always-sonnet`    | $0.300000    | $0.162500 | +$0.137500 |   45.83% |
| `always-opus`      | $1.500000    | $0.162500 | +$1.337500 |   89.17% |

How to read negative savings

A negative always-haiku result means the router chose more-expensive models than haiku on tasks haiku could have handled. That's an over-escalation signal:

  • Maybe qualityBar is set too high
  • Maybe the sonnet/opus session was warranted by complexity but the baseline doesn't know that
  • Run cost optimize (or inspect specific sessions via cost conversation) to investigate

Positive savings quantify the router's win against that baseline. The most informative number is usually always-sonnet — it's the standard "safe default" baseline most teams would pick if they didn't have routing.

When to use

  • Quarterly cost review: "We saved $X vs always-Sonnet — here's the proof."
  • CI gate: cost counterfactual --format json | jq '.baselines[1].savingsPct > 30' — fail builds if routing isn't saving ≥30% vs sonnet baseline (workload-shift detector).
  • Routing-config validation: When introducing a new qualityBar or cost-ceiling, re-run counterfactual to confirm savings didn't regress.

Stationarity caveat

Like all counterfactual analyses, this assumes the same tokens at the same complexity would have produced the same outcome from the baseline model. That's an upper bound — the baseline might have failed and required retries, which the math doesn't capture. Treat the numbers as a quality-blind ceiling.

© ruvnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/ruflo-cost-tracker/skills/cost-counterfactual of ruvnet/ruflo.

Open the folder on GitHubat commit 6c04654

Compare with similar skills

Cost Counterfactual next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cost Counterfactual compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cost Counterfactual this skillruvnet/ruflo74k—~786Automated safety check: NotesMIT
Model Cost Comparemergisi/awesome-openclaw-agents4k—~1kAutomated safety check: PassMIT
Model Cost Comparenyldn/claude-octopus4.2k1 repos~718Automated safety check: PassMIT
Cost Trackingaffaan-m/ECC277k1 repos~1.3kAutomated safety check: PassMIT
Cost Trackingaffaan-m/ECC276k—~829Automated safety check: PassMIT
Exploring LLM CostsPostHog/posthog40k—~2.9kAutomated safety check: PassCustom licence

Similar skills

  • Model Cost Compare

    mergisi/awesome-openclaw-agents

    Trigger when the user asks which model to use, wants to compare model costs, says "what's cheapest for this task", "should I use Opus or Sonnet", "can a smaller model handle this", or…

    4k GitHub stars~1k tokensUpdated 14 days ago
    AI & LLM EngineeringAuto-check passed
  • Model Cost Compare

    nyldn/claude-octopus

    Starter: compare model costs for a described task — maps task shape to the cheapest adequate seat and shows the price spread

    4.2k GitHub starsUsed in 1 repo~718 tokens
    Auto-check passed
  • Cost Tracking

    affaan-m/ECC

    Track and report Claude Code token usage, spending, and budgets from the local ECC cost-tracker metrics log.

    277k GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Cost Tracking

    affaan-m/ECC

    ローカルのコスト追跡データベースからClaude Codeのトークン使用量、支出、予算を追跡・レポートします。コスト、支出、使用量、トークン、予算、またはプロジェクト、ツール、セッション、日付によるコスト内訳について質問する場合に使用します。

    276k GitHub stars~829 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Exploring LLM Costs

    PostHog/posthog

    Official

    Investigate LLM spend in PostHog — total cost over time, cost by model, provider, user, trace, or custom dimension, token and cache-hit economics, and cost regressions.

    40k GitHub stars~2.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Analyze Cloud Costs

    langfuse/langfuse

    Analyze Langfuse Cloud infrastructure cost structure using Metabase cost marts.

    36k GitHub stars~672 tokensUpdated today
    DevOps & CloudAuto-check passed

More from ruvnet/ruflo

All 265 skills in this repo
  • Stores, searches, and retrieves successful patterns with HNSW-indexed semantic search so agents can reuse past solutions instead of relearning them.

    74k GitHub starsUsed in 1 repo~830 tokens
    Auto-check passed
  • Sets up and drives Ruflo, an npm-installed orchestration layer for multi-agent swarms, persistent memory, routing, hooks and its MCP tool catalog.

    74k GitHub starsUsed in 1 repo~975 tokens
    Auto-check passed
  • Runs claude-flow CLI security scans for input validation, path traversal, SQL injection, XSS, hardcoded secrets and known CVEs, and writes an audit report.

    74k GitHub starsUsed in 1 repo~823 tokens
    Auto-check passed
  • Applies the SPARC method (specification, pseudocode, architecture, refinement, completion) with 17 specialized modes and multi-agent orchestration, from research to deployment.

    74k GitHub starsUsed in 1 repo~829 tokens
    Auto-check passed
  • Coordinates a hierarchical swarm of specialized agents through the claude-flow CLI for work that spans several files or modules at once.

    74k GitHub starsUsed in 1 repo~779 tokens
    Auto-check passed
  • Finds models on the Hugging Face router that lack descriptions in chat-ui's prod.yaml and dev.yaml, researches each one and adds short descriptions.

    74k GitHub starsUsed in 1 repo~600 tokens
    Auto-check passed

Questions about Cost Counterfactual

What does Cost Counterfactual do?

Multi-baseline counterfactual cost analysis. An agent skill from ruvnet/ruflo. Cost Counterfactual is an agent skill from ruvnet/ruflo. Multi-baseline counterfactual cost analysis.

How do I install Cost Counterfactual in Claude Code?

Run `npx skills add ruvnet/ruflo --skill cost-counterfactual -a claude-code`. Or copy the skill folder (plugins/ruflo-cost-tracker/skills/cost-counterfactual in ruvnet/ruflo) into .claude/skills/cost-counterfactual in your project. Claude Code loads it when a task matches its description.

How do I install Cost Counterfactual in Codex?

Run `npx skills add ruvnet/ruflo --skill cost-counterfactual -a codex`. Or copy the skill folder (plugins/ruflo-cost-tracker/skills/cost-counterfactual in ruvnet/ruflo) into .agents/skills/cost-counterfactual in your project. Codex loads it when a task matches its description.

Can I use Cost Counterfactual in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ruvnet/ruflo --skill cost-counterfactual -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cost-counterfactual, .gemini/skills/cost-counterfactual, .github/skills/cost-counterfactual and .opencode/skills/cost-counterfactual in your project.

What does Cost Counterfactual need to run?

Going by SKILL.md and its folder, Cost Counterfactual needs the command-line tools its instructions call (jq). Its frontmatter pre-approves these tools: Bash.

Does Cost Counterfactual access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cost Counterfactual safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Cost Counterfactual use?

Cost Counterfactual is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cost Counterfactual use?

About 786 tokens (SKILL.md is roughly 3.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cost Counterfactual?

Skills that share tags, products or a category with Cost Counterfactual: Model Cost Compare (mergisi/awesome-openclaw-agents, 4k stars), Model Cost Compare (nyldn/claude-octopus, 4.2k stars), Cost Tracking (affaan-m/ECC, 277k stars) and Cost Tracking (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cost Counterfactual?

ruvnet (a GitHub user) maintains it in ruvnet/ruflo, which has 74,222 GitHub stars. The repository holds 265 skills in this directory. The repository was last updated on October 10, 2026.

Source: ruvnet/ruflo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.