Agent skill

Add a remindb Bench Scenario

by radimsem in radimsem/remindb

Adds a new row to remindb's token-savings benchmark table by writing a scenario function that measures the naive shell-tool token cost against the remindb tool-call cost for the same task.

MITAuto-check passedDevelopment

Install Add a remindb Bench Scenario

skills CLI
$ npx skills add radimsem/remindb --skill add-bench-scenario -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install radimsem/remindb add-bench-scenario --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/radimsem/remindb.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/add-bench-scenario .claude/skills/add-bench-scenario && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-bench-scenario
GitHub stars
129
Token cost
~1.7k tokens
SKILL.md length
713 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

Adds a new row to remindb's token-savings benchmark table by writing a scenario function that measures the naive shell-tool token cost against the remindb tool-call cost for the same task.

  • Works in 3 steps: Use tokens.Estimate(string) for both… → Compare to a believable naive baseline.… → Use callTool(ctx, s, name, args) (the…
  • Adding a new comparison row to remindb's token-savings benchmark
  • SKILL.md covers Where it lands, The scenario function shape, Wiring into Run and Naming rows, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The internal/bench package is remindb's external-facing benchmark, comparing how many tokens an agent would spend doing a task with remindb's tools against doing the same task with grep, cat or find, rendered as a token-savings table through text/tabwriter and invoked from the remindb bench CLI subcommand. This is explicitly not the Go testing.B benchmark surface, which lives separately as Benchmark functions in bench_test.go files with its own rules.

A typical scenario touches two files, or three if it needs a new CLI flag: a new benXxx function added to internal/bench/scenarios.go, wiring that function into Run inside bench.go, and optionally a new flag surfaced on the bench subcommand. Every scenario function returns a scenarioResult with a name, a naive token count and a remindb token count, measured with the shared tokens.Estimate helper on both sides so the comparison uses one tokenizer; a scenario that runs over several inputs instead returns a slice of results, one per input. The naive baseline must reflect what an agent would actually do without remindb, such as find plus cat rather than wc -c.

When your agent uses it

  • Adding a new comparison row to remindb's token-savings benchmark
  • Benchmarking a new remindb tool against its naive shell-command equivalent
  • Wiring a new benchmark scenario into the bench CLI subcommand

Example prompts

  • “Add a benchmark scenario comparing remindb's search tool against grep plus cat.”
  • “Wire a new token-savings benchmark for the fetch tool into bench.Run.”
  • “Add a CLI flag so I can pass a custom query list into the search benchmark.”

Requirements

  • Go with the remindb repository checked out

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Use tokens.Estimate(string) for both sides. Apples-to-apples comparison only works if the same tokenizer measures both.
  2. Compare to a believable naive baseline. "What would the agent actually do without remindb?" — find + cat * for tree (not wc -c), grep +…
  3. Use callTool(ctx, s, name, args) (the package helper in scenarios.go), not raw s.CallTool. It strips the *mcp.CallToolResult boilerplate…

What it can do on your machine

Read from SKILL.md and the folder at commit 977b31c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are go).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add a remindb Bench Scenario loads about 1.7k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 713 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from radimsem/remindb at commit 977b31c, republished under its MIT licence (© radimsem). 713 words, ~1,722 tokens.

Download SKILL.mdSave it as .claude/skills/add-bench-scenario/SKILL.md (or your agent's skills folder).
name
add-bench-scenario
description
Use when adding a new scenario to remindb's benchmark suite — symptoms include "compare X tool against grep/cat", "add a token-savings benchmark for Y", "extend `internal/bench/scenarios.go`", "wire a new scenario into `bench.Run`", or any task that adds a row to the `scenario / naive (tok) / remindb (tok) / saved` output table. Distinct from Go `testing.B` benchmarks in `*_bench_test.go`.

Add a benchmark scenario

internal/bench/ is remindb's external-facing benchmark — it compares "how many tokens does an agent consume to do X via remindb tools" against "how many would they consume doing X with grep / cat / find". The output is a token-savings table rendered via text/tabwriter. It's invoked from the remindb bench CLI subcommand and exercised end-to-end by scripts/bench-agents.sh.

This is not the Go testing.B benchmark surface — those live as Benchmark* functions inside pkg/*/bench_test.go and have their own discipline.

Where it lands

Two files for a typical scenario, three if it needs CLI flag plumbing.

FileWhat changes
internal/bench/scenarios.goNew benchXxx(ctx, session, srcDir, ...) (scenarioResult, error) function
internal/bench/bench.goWire the new scenario into Run, append to results
cmd/remindb/... (only if new flag needed)Surface a new flag on the bench subcommand and pass it into bench.Config

The scenario function shape

Every scenario implements the same contract: produce one (or more) scenarioResult{name, naiveTok, remindbTok} by measuring two paths to the same answer — the naive path (token count of what an agent would have to read using shell tools) and the remindb path (token count of the tool's response).

Mirror benchTree, benchSearch, or benchFetch in scenarios.go:

go
func benchExample(ctx context.Context, s *gomcp.ClientSession, srcDir string, budget int) (scenarioResult, error) {
    // Naive path: what the agent would read using shell tools.
    naive := tokens.Estimate(naiveContent(srcDir))

    // remindb path: call the tool and count its response.
    text, err := callTool(ctx, s, "MemoryExample", map[string]any{
        "budget": budget,
    })
    if err != nil {
        return scenarioResult{}, err
    }

    return scenarioResult{
        name:       "example",
        naiveTok:   naive,
        remindbTok: tokens.Estimate(text),
    }, nil
}

A scenario that runs over multiple inputs (like benchSearch over a query list) returns []scenarioResult instead — one row per input, named with the input as suffix ("search:rate-limit").

Three contracts the function must honor:

  1. Use tokens.Estimate(string) for both sides. Apples-to-apples comparison only works if the same tokenizer measures both.
  2. Compare to a believable naive baseline. "What would the agent actually do without remindb?" — find + cat * for tree (not wc -c), grep + cat <matches> for search (not just grep). The naive must reflect the real alternative.
  3. Use callTool(ctx, s, name, args) (the package helper in scenarios.go), not raw s.CallTool. It strips the *mcp.CallToolResult boilerplate and returns the text content directly.

Wiring into Run

bench.go:Run calls each scenario in sequence and appends to results. Add your call in the natural sequence — Tree first (orientation), then Search, then Fetch, then Delta, then your scenario:

go
r, err = benchExample(ctx, session, stage.srcDir, cfg.Budget)
if err != nil {
    return err
}
results = append(results, r)

Or for a multi-result scenario:

go
rs, err := benchExample(ctx, session, stage.srcDir, cfg.Budget)
if err != nil {
    return err
}
results = append(results, rs...)

The renderer handles the rest — the table column count and total-row computation are static.

Naming rows

The name field becomes the leftmost column. Use lowercase, hyphenated, prefixed by the tool category if the scenario family has multiple rows:

  • tree (single)
  • search:rate-limit (one of many search queries)
  • fetch (single)
  • delta (single)

Stay under ~30 chars; shorten(s, 30) is the existing helper used by benchSearch.

When to add a new flag

Most scenarios reuse cfg.Budget and cfg.Queries. Add a new field to bench.Config only when the scenario needs an input that isn't already there — and even then, default it sensibly so existing bench-agents.sh runs don't break.

Show full SKILL.md (269 more words)Show less

Quick reference

1. internal/bench/scenarios.go        (benchXxx function with naive + remindb paths)
2. internal/bench/bench.go            (call site in Run, append to results)
3. (optional) cmd/remindb/...         (new flag on the bench subcommand)
4. go build ./... && ./remindb bench --db <path> --dir <path>
5. scripts/bench-agents.sh            (full suite — confirms the new row appears for every agent)

Common mistakes

  • Comparing to a trivial naive baseline. "remindb tree returns 200 tokens, naive ls returns 50 tokens — we're 4x worse" misses the point. The naive is what the agent actually does: find . -type f + reading every file. Use the existing helpers (listDirFiles, countDirTokens, grepDir, sumFileTokens).
  • Using a different tokenizer for one side. Both sides must go through tokens.Estimate. Counting bytes, lines, or words on one side breaks the comparison.
  • Forgetting tokens.Estimate is over a string, not []byte. Convert byte slices first; mixing produces a confusing compile error since tokens.Estimate is overloaded by neither Go nor remindb.
  • Failing the entire bench on one scenario error. The current contract is fail-fast — Run returns the first error. If your scenario has expected partial-failure modes (skipped queries, missing fixtures), handle them inside the scenario and emit a scenarioResult{name: "example:skipped", ...} instead of returning err.
  • Running benches against the live DB. stageBench copies the DB and source tree to /tmp before running so the user's real DB isn't touched. If you reach for cfg.DBPath directly inside a scenario, you're operating on the original — pass stage.dbPath instead.
  • Not updating bench-agents.sh query lists. New scenarios that take per-agent inputs won't be exercised by the CI-style script unless the per-agent map gets entries. For non-input scenarios (like a new fetch variant), no script change is needed.

Cross-references

  • .claude/rules/go-concise.md — error wrapping, named locals (naive, remindbTok)
  • .claude/skills/add-mcp-tool/SKILL.md — when adding a tool that should be benchmarked, do that skill first; the bench scenario consumes the tool you wired
  • internal/bench/render.go — the table renderer; rarely needs changes but read it once if you're adding a column rather than a row

© radimsem, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/add-bench-scenario of radimsem/remindb.

Open the folder on GitHubat commit 977b31c

Compare with similar skills

Add a remindb Bench Scenario next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add a remindb Bench Scenario compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add a remindb Bench Scenario this skillradimsem/remindb129—~1.7kAutomated safety check: PassMIT
Gograph Go Repository Intelligenceozgurcd/gograph229—~4.8kAutomated safety check: NotesMIT
Grepaiyoanbernabeu/grepai1.9k—~1.1kAutomated safety check: PassMIT
Context Handling in fp-goIBM/fp-go2k—~4.2kAutomated safety check: PassApache-2.0
Releaseyoanbernabeu/grepai1.9k—~918Automated safety check: PassMIT
Typed Dependencies with fp-go EffectIBM/fp-go2k—~4.3kAutomated safety check: PassApache-2.0

Similar skills

  • Gives an agent working in a Go codebase a structural view through a local MCP server: call graphs, blast-radius and impact analysis, and bounded first-call exploration.

    229 GitHub stars~4.8k tokensUpdated today
    DevelopmentAuto-check: notes
  • Grepai

    yoanbernabeu/grepai

    Semantic code search and call-graph tracing. An agent skill from yoanbernabeu/grepai.

    1.9k GitHub stars~1.1k tokensUpdated 19 days ago
    DevelopmentAuto-check passed
  • Official

    Covers handling Go's context.Context idiomatically in fp-go code: reading and scoping context through operators, timeouts, cancellation and converting ctx-first functions.

    2k GitHub stars~4.2k tokensUpdated today
    DevelopmentAuto-check passed
  • Release

    yoanbernabeu/grepai

    Create a new release for grepai. An agent skill from yoanbernabeu/grepai.

    1.9k GitHub stars~918 tokensUpdated 19 days ago
    DevelopmentAuto-check passed
  • Teaches an agent to write fp-go v2 services with the Effect type, carrying dependencies in its type parameter instead of in context.Context or parameters.

    2k GitHub stars~4.3k tokensUpdated today
    DevelopmentAuto-check passed
  • Local Dev

    SmilyOrg/photofield

    Run, test, and debug the photofield server locally. An agent skill from SmilyOrg/photofield.

    608 GitHub stars~2.4k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed

More from radimsem/remindb

All 11 skills in this repo
  • remindb Memory Writer

    radimsem/remindb

    Guides an agent writing to a remindb memory server: structured notes go in as files parsed into a node tree, single text facts go in through MemoryWrite.

    129 GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed
  • remindb Read Path

    radimsem/remindb

    Covers the read-side tools of the remindb MCP server, a SQLite-backed agent memory, for orienting, searching, resyncing and tracing relations.

    129 GitHub stars~2.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Adds a new Go fuzz target to remindb following its seed-corpus discipline: one target per logical surface, discovered automatically by its FuzzXxx function name.

    129 GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Explains how to add an end-to-end test scenario to remindb, choosing between a direct API test and an MCP test and using the shared helpers and fixtures.

    129 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Checklist for adding a new Memory-prefixed tool to remindb's Go MCP server: tool file, registration, test, skill docs, plus the locking and logging rules.

    129 GitHub stars~2.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Add a remindb Parser

    radimsem/remindb

    Walks through adding a new file format to remindb's Go parser package: the parser file, the ParseBytes case, table tests and fuzz seeds.

    129 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Add a remindb Bench Scenario

What does Add a remindb Bench Scenario do?

Adds a new row to remindb's token-savings benchmark table by writing a scenario function that measures the naive shell-tool token cost against the remindb tool-call cost for the same task. The internal/bench package is remindb's external-facing benchmark, comparing how many tokens an agent would spend doing a task with remindb's tools against doing the same task with grep, cat or find, rendered as a token-savings table through text/tabwriter and invoked from the remindb bench CLI subcommand.go files with its own rules.

When should I use Add a remindb Bench Scenario?

Add a remindb Bench Scenario fits situations like: adding a new comparison row to remindb's token-savings benchmark; benchmarking a new remindb tool against its naive shell-command equivalent; wiring a new benchmark scenario into the bench CLI subcommand.

How do I install Add a remindb Bench Scenario in Claude Code?

Run `npx skills add radimsem/remindb --skill add-bench-scenario -a claude-code`. Or copy the skill folder (.claude/skills/add-bench-scenario in radimsem/remindb) into .claude/skills/add-bench-scenario in your project. Claude Code loads it when a task matches its description.

How do I install Add a remindb Bench Scenario in Codex?

Run `npx skills add radimsem/remindb --skill add-bench-scenario -a codex`. Or copy the skill folder (.claude/skills/add-bench-scenario in radimsem/remindb) into .agents/skills/add-bench-scenario in your project. Codex loads it when a task matches its description.

Can I use Add a remindb Bench Scenario in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add radimsem/remindb --skill add-bench-scenario -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-bench-scenario, .gemini/skills/add-bench-scenario, .github/skills/add-bench-scenario and .opencode/skills/add-bench-scenario in your project.

What does Add a remindb Bench Scenario need to run?

SKILL.md names no scripts, command-line tools or credentials: Add a remindb Bench Scenario is instructions for the agent only. Our summary lists: Go with the remindb repository checked out.

Does Add a remindb Bench Scenario access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Add a remindb Bench Scenario safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add a remindb Bench Scenario use?

Add a remindb Bench Scenario is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add a remindb Bench Scenario use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add a remindb Bench Scenario?

Skills that share tags, products or a category with Add a remindb Bench Scenario: Gograph Go Repository Intelligence (ozgurcd/gograph, 229 stars), Grepai (yoanbernabeu/grepai, 1.9k stars), Context Handling in fp-go (IBM/fp-go, 2k stars) and Release (yoanbernabeu/grepai, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add a remindb Bench Scenario?

radimsem (a GitHub user) maintains it in radimsem/remindb, which has 129 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on August 3, 2026.

Source: radimsem/remindb on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.