Agent skill

Cost Benchmark

by ruvnet in ruvnet/ruflo

Run the corpus benchmark — booster locally, optional Gemini/Sonnet/Opus baselines — and persist a verifiable measured-vs-claimed table

MITAuto-check: notesAI & LLM Engineering

Install Cost Benchmark

skills CLI
$ npx skills add ruvnet/ruflo --skill cost-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ruvnet/ruflo cost-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ruvnet/ruflo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ruflo-cost-tracker/skills/cost-benchmark .claude/skills/cost-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cost-benchmark
GitHub stars
74k
Token cost
~745 tokens
SKILL.md length
253 words
Files
1
Skills in repo
264
Repo updated
First seen
Licence
MIT

At a glance

Run the corpus benchmark — booster locally, optional Gemini/Sonnet/Opus baselines — and persist a verifiable measured-vs-claimed table

  • Works in 4 steps: Run the bench from v3/ (where… → Inspect the markdown summary printed to… → Persisted output lands at → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers When to use, Steps, Smoke gates and Env overrides, plus 1 more section
  • Calls node and gcloud; needs GOOGLE_AI_API_KEY and ANTHROPIC_API_KEY

What it does

Cost Benchmark is an agent skill from ruvnet/ruflo. Run the corpus benchmark — booster locally, optional Gemini/Sonnet/Opus baselines — and persist a verifiable measured-vs-claimed table

Its SKILL.md is about 750 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. It works with OpenAI. The repository describes itself as: 🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory…. The licence is MIT.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/cost-benchmark”

Requirements

  • A credential in GOOGLE_AI_API_KEY
  • A credential in ANTHROPIC_API_KEY
  • Pre-approved tools (allowed-tools): Bash

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Run the bench from v3/ (where agent-booster resolves)
  2. Inspect the markdown summary printed to stdout. The gate metric is winRate (Tier 1 cases). Adversarial cases are tracked separately as…
  3. Persisted output lands at
  4. Read it back in subsequent skills (e.g. cost-report step 2 reads latest.json for live tier-spend numbers).

What it can do on your machine

Read from SKILL.md and the folder at commit 6051f67. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node
    • gcloud

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gcloud, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GOOGLE_AI_API_KEY
    • ANTHROPIC_API_KEY
    • BENCH_LLM_API_KEY
    • BENCH_ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cost Benchmark loads about 745 tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 253 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~37
When it runs · the whole SKILL.md, loaded when a task matches
~745

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ruvnet/ruflo at commit 6051f67, republished under its MIT licence (© ruvnet). 253 words, ~745 tokens.

Download SKILL.mdSave it as .claude/skills/cost-benchmark/SKILL.md (or your agent's skills folder).
name
cost-benchmark
description
Run the corpus benchmark — booster locally, optional Gemini/Sonnet/Opus baselines — and persist a verifiable measured-vs-claimed table
allowed-tools
Bash
argument-hint
[--llm] [--anthropic]

Cost Benchmark

Runs scripts/bench.mjs against the structural+adversarial corpus and writes per-case + summary results to docs/benchmarks/runs/. This is the verification gate that backs every measurable claim in cost-booster-edit / cost-booster-route.

When to use

  • Before publishing a release — verify booster win rate didn't regress.
  • After expanding bench/booster-corpus.json — confirm new cases route correctly.
  • When auditing a "claimed upstream" tag — flip it to "verified" once the bench supports it.
  • On a cost question ("is Sonnet 4.6 cheaper than Opus 4.7 for these tasks?") — re-run with BENCH_ANTHROPIC=1.

Steps

  1. Run the bench from v3/ (where agent-booster resolves):

    bash
    ( cd v3 && node ../plugins/ruflo-cost-tracker/scripts/bench.mjs )                  # booster only — free, ~85 ms
    ( cd v3 && BENCH_LLM_BASELINE=1 node ../plugins/ruflo-cost-tracker/scripts/bench.mjs ) # + Gemini 2.0 Flash (cheap)
    ( cd v3 && BENCH_LLM_BASELINE=1 BENCH_ANTHROPIC=1 \
         node ../plugins/ruflo-cost-tracker/scripts/bench.mjs )                          # + Sonnet 4.6 + Opus 4.7
  2. Inspect the markdown summary printed to stdout. The gate metric is winRate (Tier 1 cases). Adversarial cases are tracked separately as escalationRate.

  3. Persisted output lands at:

    • docs/benchmarks/runs/latest.json — pointer to the most recent run
    • docs/benchmarks/runs/<ISO-timestamp>.json — historical record
  4. Read it back in subsequent skills (e.g. cost-report step 2 reads latest.json for live tier-spend numbers).

Smoke gates

  • winRate ≥ 0.80 on Tier 1 cases (smoke step 23). Lower the threshold by editing scripts/smoke.sh.
  • escalationRate is reported but ungated — adversarial cases are diagnostic.

Env overrides

Env varDefaultPurpose
BENCH_LLM_BASELINEunset=1 runs the OpenAI-compat baseline
BENCH_LLM_MODELmodels/gemini-2.0-flashOverride the OpenAI-compat model
BENCH_LLM_BASE_URLGemini OpenAI shimOverride endpoint
BENCH_ANTHROPICunset=1 runs Anthropic baseline (Sonnet 4.6 + Opus 4.7)
BENCH_ANTHROPIC_MODELSclaude-sonnet-4-6,claude-opus-4-7Comma-separated Claude IDs
BENCH_OUTtimestamped fileOverride output path
BENCH_QUIET=1unsetSuppress markdown summary

API keys auto-pulled from gcloud secrets (GOOGLE_AI_API_KEY, ANTHROPIC_API_KEY); override with BENCH_LLM_API_KEY / BENCH_ANTHROPIC_API_KEY.

Cross-references

ADR-0002 §"Decision 1" / §"Riskiest assumption" · cost-booster-edit/SKILL.md (verification table consumes this skill's output) · cost-report/SKILL.md step 2 (reads runs/latest.json).

© ruvnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/ruflo-cost-tracker/skills/cost-benchmark of ruvnet/ruflo.

Open the folder on GitHubat commit 6051f67

Compare with similar skills

Cost Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cost Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cost Benchmark this skillruvnet/ruflo74k—~745Automated safety check: NotesMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k8 repos~1.7kAutomated safety check: PassMIT
Codebase Managementgiancarloerra/SocratiCode3.3k1 repos~1.8kAutomated safety check: PassAGPL-3.0
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Codebase Management

    giancarloerra/SocratiCode

    Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.

    3.3k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Docs Planner

    strands-agents/harness-sdk

    Identify documentation gaps and prioritize the docs backlog.

    8.7k GitHub stars~821 tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from ruvnet/ruflo

All 264 skills in this repo
  • Stores, searches, and retrieves successful patterns with HNSW-indexed semantic search so agents can reuse past solutions instead of relearning them.

    74k GitHub starsUsed in 3 repos~830 tokens
    Auto-check passed
  • Runs claude-flow CLI security scans for input validation, path traversal, SQL injection, XSS, hardcoded secrets and known CVEs, and writes an audit report.

    74k GitHub starsUsed in 2 repos~823 tokens
    Auto-check passed
  • Applies the SPARC method (specification, pseudocode, architecture, refinement, completion) with 17 specialized modes and multi-agent orchestration, from research to deployment.

    74k GitHub starsUsed in 2 repos~829 tokens
    Auto-check passed
  • Coordinates a hierarchical swarm of specialized agents through the claude-flow CLI for work that spans several files or modules at once.

    74k GitHub starsUsed in 2 repos~779 tokens
    Auto-check passed
  • Sets up and drives Ruflo, an npm-installed orchestration layer for multi-agent swarms, persistent memory, routing, hooks and its MCP tool catalog.

    74k GitHub starsUsed in 1 repo~975 tokens
    Auto-check passed
  • Agent Coordination

    ruvnet/ruflo

    Reference for spawning, listing, monitoring and stopping agents with claude-flow commands, with agent type families, routing codes and coordination tips.

    74k GitHub starsUsed in 2 repos~519 tokens
    Auto-check passed

Works with

Questions about Cost Benchmark

What does Cost Benchmark do?

Run the corpus benchmark — booster locally, optional Gemini/Sonnet/Opus baselines — and persist a verifiable measured-vs-claimed table. Cost Benchmark is an agent skill from ruvnet/ruflo.

When should I use Cost Benchmark?

Cost Benchmark fits situations like: AI & LLM Engineering work in your project.

How do I install Cost Benchmark in Claude Code?

Run `npx skills add ruvnet/ruflo --skill cost-benchmark -a claude-code`. Or copy the skill folder (plugins/ruflo-cost-tracker/skills/cost-benchmark in ruvnet/ruflo) into .claude/skills/cost-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Cost Benchmark in Codex?

Run `npx skills add ruvnet/ruflo --skill cost-benchmark -a codex`. Or copy the skill folder (plugins/ruflo-cost-tracker/skills/cost-benchmark in ruvnet/ruflo) into .agents/skills/cost-benchmark in your project. Codex loads it when a task matches its description.

Can I use Cost Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ruvnet/ruflo --skill cost-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cost-benchmark, .gemini/skills/cost-benchmark, .github/skills/cost-benchmark and .opencode/skills/cost-benchmark in your project.

What does Cost Benchmark need to run?

Going by SKILL.md and its folder, Cost Benchmark needs the command-line tools its instructions call (node and gcloud) and credentials named GOOGLE_AI_API_KEY, ANTHROPIC_API_KEY, BENCH_LLM_API_KEY and BENCH_ANTHROPIC_API_KEY. Our summary lists: A credential in GOOGLE_AI_API_KEY; A credential in ANTHROPIC_API_KEY. Its frontmatter pre-approves these tools: Bash.

Does Cost Benchmark access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cost Benchmark safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Cost Benchmark use?

Cost Benchmark is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cost Benchmark use?

About 745 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cost Benchmark?

Skills that share tags, products or a category with Cost Benchmark: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Codebase Management (giancarloerra/SocratiCode, 3.3k stars) and Azure AI Projects Python SDK (microsoft/skills, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cost Benchmark?

ruvnet (a GitHub user) maintains it in ruvnet/ruflo, which has 74,089 GitHub stars. The repository holds 264 skills in this directory. The repository was last updated on October 8, 2026.

Source: ruvnet/ruflo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.