Agent skill

Benchmark Memory

by Abilityai in Abilityai/cornelius

Systematic benchmarking framework for Local Brain Search memory system with LLM-as-judge scoring

MITAuto-check: notesAI & LLM Engineering

Install Benchmark Memory

skills CLI
$ npx skills add Abilityai/cornelius --skill benchmark-memory -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Abilityai/cornelius benchmark-memory --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Abilityai/cornelius.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/benchmark-memory .claude/skills/benchmark-memory && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-memory
GitHub stars
109
Token cost
~2.4k tokens
SKILL.md length
731 words
Files
1
Skills in repo
56
Repo updated
First seen
Licence
MIT

At a glance

Systematic benchmarking framework for Local Brain Search memory system with LLM-as-judge scoring

  • Works in 4 steps: Measure retrieval quality objectively… → Compare different configuration settings… → Identify optimal parameters for… → …
  • Tasks that involve LLM evaluation
  • SKILL.md covers Purpose, Design Principles, State Dependencies and Prerequisites, plus 8 more sections
  • Calls python, claude and pip; needs ANTHROPIC_API_KEY

What it does

Benchmark Memory is an agent skill from Abilityai/cornelius. Systematic benchmarking framework for Local Brain Search memory system with LLM-as-judge scoring

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: AI-powered second brain template for Claude Code + Obsidian. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM evaluation

Example prompts

  • “/benchmark-memory”

Requirements

  • Python 3
  • A credential in ANTHROPIC_API_KEY
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Glob, Grep, Task

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Measure retrieval quality objectively using LLM-as-judge scoring
  2. Compare different configuration settings (spreading vs static, parameter sweeps)
  3. Identify optimal parameters for different query types (factual, conceptual, synthesis)
  4. Generate reproducible results against frozen test datasets

What it can do on your machine

Read from SKILL.md and the folder at commit fd5e9a4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Glob
    • Grep
    • Task

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • claude
    • pip
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Memory loads about 2.4k tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 731 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~28
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Glob, Grep, Task

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Abilityai/cornelius at commit fd5e9a4, republished under its MIT licence (© Abilityai). 731 words, ~2,377 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-memory/SKILL.md (or your agent's skills folder).
name
benchmark-memory
description
Systematic benchmarking framework for Local Brain Search memory system with LLM-as-judge scoring
allowed-tools
Bash, Read, Write, Glob, Grep, Task
automation
gated
user-invocable
true

Benchmark Memory System

Systematic benchmarking framework to measure retrieval quality, compare configurations, and identify optimal parameters for the Local Brain Search memory system.

Purpose

  1. Measure retrieval quality objectively using LLM-as-judge scoring
  2. Compare different configuration settings (spreading vs static, parameter sweeps)
  3. Identify optimal parameters for different query types (factual, conceptual, synthesis)
  4. Generate reproducible results against frozen test datasets

Design Principles

  • Contained: Skill + sub-agent + bundled scripts
  • Reproducible: Test against frozen Brain snapshot
  • Automated: LLM-as-judge for relevance scoring
  • Analyzable: CSV output for analysis

State Dependencies

SourceLocationReadWrite
Brain snapshot.claude/skills/benchmark-memory/snapshots/YesYes
Query sets.claude/skills/benchmark-memory/query-sets/YesYes
Benchmark results.claude/skills/benchmark-memory/results/YesYes
Analysis reports.claude/skills/benchmark-memory/analysis/NoYes
Memory systemresources/local-brain-search/YesNo

Prerequisites

  • Local Brain Search system indexed (resources/local-brain-search/data/brain.faiss)
  • Python venv at resources/local-brain-search/venv/ with search dependencies
  • Claude Code CLI installed and authenticated (for LLM-as-judge scoring via headless mode)
LLM-as-Judge Scoring

This skill uses Claude Code headless mode (claude -p) for LLM relevance scoring, not a separate API key. This means:

  • No ANTHROPIC_API_KEY environment variable needed
  • Uses your existing Claude Code authentication
  • Default model: sonnet (good quality) - can also use haiku (faster/cheaper) or opus
  • JSON output via prompt engineering for reliable scoring

To verify Claude Code is available:

bash
claude --version
Installing Dependencies

Dependencies are installed in the local-brain-search venv:

bash
cd resources/local-brain-search
source venv/bin/activate
pip install pandas tqdm  # anthropic not required - uses Claude Code headless

Sub-Commands

/benchmark-memory setup

Create a frozen Brain snapshot and build its index.

bash
cd .claude/skills/benchmark-memory/scripts
./run_benchmark.sh --list-snapshots  # Check existing
python3 create_snapshot.py           # Create new snapshot

What it does:

  1. Creates snapshot directory with date stamp
  2. Copies Brain folder (excluding .obsidian, .trash)
  3. Builds FAISS index for the snapshot
  4. Creates SNAPSHOT-INFO.md with metadata
/benchmark-memory create-queries [--count N]

Generate or manage test query sets.

bash
cd .claude/skills/benchmark-memory/scripts
python build_query_set.py --count 50 --output ../query-sets/core-50.json

Query Categories (50 total):

CategoryCountExample
Factual10"What is dopamine?"
Conceptual10"How does motivation work?"
Synthesis15"Connect Buddhism and neuroscience"
Temporal5"Recent notes about AI agents"
Needle5"Note about intermittent reinforcement"
Broad5"Identity"
/benchmark-memory run [--config CONFIG] [--snapshot SNAPSHOT]

Execute benchmark with specified configuration.

bash
cd .claude/skills/benchmark-memory/scripts
./run_benchmark.sh --config focused --snapshot brain-snapshot-2026-02-18
./run_benchmark.sh --dry-run --config focused  # Preview without execution
./run_benchmark.sh --list-configs              # List available configs

Configurations:

  • focused: 15 key configurations (recommended for initial benchmarking)
  • single:CONFIG_NAME: Run single configuration
  • all: Full parameter sweep (expensive)

Estimated cost per run:

  • 50 queries x 10 results x 15 configs = 7,500 LLM scores
  • Using Sonnet (default): ~$75
  • Using Haiku: ~$7.50
/benchmark-memory analyze [--results FILE]

Generate analysis summary from benchmark results.

bash
cd .claude/skills/benchmark-memory/scripts
python analyze_results.py --results ../results/benchmark-*.csv

Outputs:

  • Summary by configuration
  • Summary by query category
  • Best config per intent
  • Recommendations

Workflow

Step 1: Setup (one-time)
/benchmark-memory setup
    |
    v
Step 2: Create Queries (one-time)
/benchmark-memory create-queries --count 50
    |
    v
Step 3: Run Benchmark (per experiment)
/benchmark-memory run --config focused
    |
    v
Step 4: Analyze Results
/benchmark-memory analyze

Metrics Collected

Performance Metrics
MetricDescription
latency_msQuery execution time
iterationsSpreading iterations used
convergedWhether spreading converged
Quality Metrics
MetricRangeDescription
Precision@K0-1Fraction of results that are relevant
Recall@K0-1Fraction of relevant notes found
MRR0-1Mean Reciprocal Rank
NDCG@K0-1Ranking quality with position discount
Avg Score0-3Average LLM relevance score
Show full SKILL.md (299 more words)Show less
LLM-as-Judge Scoring Scale
ScoreLabelDefinition
0IrrelevantNo connection to query
1TangentialLoosely related
2RelevantAddresses the query
3Highly RelevantDirectly answers the query

Configurations to Test

Baseline
  • static_baseline: Traditional vector search
  • spreading_default: Spreading activation with defaults
Parameter Sweeps
  • Iteration count: 2, 5, 7
  • Inhibition strength: 0.1, 0.3, 0.5
  • Temporal decay: 0.8, 0.9, 0.95
  • Q-weight: 0.0, 0.3, 0.5
Optimized Combinations
  • synthesis_optimized: max_iterations=7, inhibition=0.1
  • factual_optimized: max_iterations=2, inhibition=0.5
  • balanced_optimized: max_iterations=5, inhibition=0.2, decay=0.85

Output Files

Results CSV
results/benchmark-YYYY-MM-DD-HHMMSS.csv

Schema:

timestamp,config_name,query_id,query_category,mode,max_iterations,
inhibition_strength,latency_ms,result_1_note,result_1_score,...,
precision_at_5,precision_at_10,mrr,ndcg_at_10,avg_score
Analysis Report
analysis/report-YYYY-MM-DD.md

Error Recovery

Partial benchmark run

Results are appended incrementally. Resume by running with --resume:

bash
python run_benchmark.py --config focused --resume
API rate limits

Built-in retry with exponential backoff. Adjust --delay if needed:

bash
python run_benchmark.py --config focused --delay 1.0
Invalid snapshot

Re-create snapshot:

bash
python create_snapshot.py --force

Success Criteria

  • Snapshot creation produces valid index
  • Query set covers all 6 categories (50+ queries)
  • LLM judge produces consistent scores (>80% agreement on re-run)
  • All 15 focused configs can be benchmarked
  • Results CSV is valid and analyzable
  • Analysis identifies best config per query type

Expected Insights

After running benchmarks, answer these questions:

  1. Does spreading beat static? For which query types?
  2. What's the optimal iteration count for synthesis queries?
  3. Does high inhibition help factual queries?
  4. Does q_weight > 0 improve results over time?
  5. What settings work best for each query type?

Cost Management

StrategySavingsTrade-off
Score top 5 only50%Less data on long-tail
Use Haiku judge90%Slightly less accurate
Cache scoresVariableOnly for unchanged retrieval

Recommendation: Start with Haiku judge, validate sample against Sonnet.

Directory Structure

.claude/skills/benchmark-memory/
├── SKILL.md                    # This file
├── requirements.txt            # Python dependencies
├── scripts/
│   ├── run_benchmark.sh        # Wrapper script (uses venv Python)
│   ├── create_snapshot.py      # Create frozen Brain snapshot
│   ├── build_query_set.py      # Generate/manage query test set
│   ├── run_benchmark.py        # Execute benchmark with config
│   ├── score_results.py        # LLM-as-judge scoring
│   ├── compute_metrics.py      # Calculate evaluation metrics
│   └── analyze_results.py      # Generate analysis summary
├── configs/
│   ├── focused_configs.json    # Test configurations (15 configs)
│   └── judge_prompt.txt        # LLM judge prompt template
├── snapshots/                  # Frozen Brain copies
├── query-sets/                 # Test queries (core-50.json included)
├── results/                    # Benchmark CSVs
└── analysis/                   # Analysis reports

Tested Components

ComponentStatusNotes
Snapshot creation✅ WorksCreates snapshot with FAISS index + graph
Query set✅ Works50 queries across 6 categories
Static search✅ WorksTraditional vector similarity
Spreading search✅ WorksMulti-iteration activation
15 configs✅ WorksFocused parameter sweep
LLM-as-judge✅ WorksUses Claude Code headless mode (claude -p)
Results CSVReadyIncremental writes, resume support

© Abilityai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/benchmark-memory of Abilityai/cornelius.

Open the folder on GitHubat commit fd5e9a4

Compare with similar skills

Benchmark Memory next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Memory compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Memory this skillAbilityai/cornelius109—~2.4kAutomated safety check: NotesMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from Abilityai/cornelius

All 56 skills in this repo
  • Nano Banana Image Generator

    Abilityai/cornelius

    Generate images using Google's Nano Banana (Gemini 2.5 Flash Image).

    109 GitHub stars~1.2k tokensUpdated 15 days ago
    Auto-check: notes
  • Changelog Protocol

    Abilityai/cornelius

    Protocol for creating dated changelog files after significant agent sessions.

    109 GitHub stars~555 tokensUpdated 15 days ago
    Auto-check passed
  • Create Article

    Abilityai/cornelius

    Create long-form articles from knowledge base insights. An agent skill from Abilityai/cornelius.

    109 GitHub stars~2.1k tokensUpdated 15 days ago
    Auto-check: notes
  • Epistemic Classification

    Abilityai/cornelius

    Framework for distinguishing research findings from hypotheses and speculative synthesis.

    109 GitHub stars~1.5k tokensUpdated 15 days ago
    Auto-check passed
  • Get Youtube Transcript

    Abilityai/cornelius

    Extract the transcript from a YouTube video by URL or video ID.

    109 GitHub stars~525 tokensUpdated 15 days ago
    Auto-check: notes
  • Insight Capture Format

    Abilityai/cornelius

    Standard format for capturing and documenting insights in the knowledge base.

    109 GitHub stars~616 tokensUpdated 15 days ago
    Auto-check passed

Questions about Benchmark Memory

What does Benchmark Memory do?

Systematic benchmarking framework for Local Brain Search memory system with LLM-as-judge scoring. Benchmark Memory is an agent skill from Abilityai/cornelius.

When should I use Benchmark Memory?

Benchmark Memory fits situations like: tasks that involve LLM evaluation.

How do I install Benchmark Memory in Claude Code?

Run `npx skills add Abilityai/cornelius --skill benchmark-memory -a claude-code`. Or copy the skill folder (.claude/skills/benchmark-memory in Abilityai/cornelius) into .claude/skills/benchmark-memory in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Memory in Codex?

Run `npx skills add Abilityai/cornelius --skill benchmark-memory -a codex`. Or copy the skill folder (.claude/skills/benchmark-memory in Abilityai/cornelius) into .agents/skills/benchmark-memory in your project. Codex loads it when a task matches its description.

Can I use Benchmark Memory in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Abilityai/cornelius --skill benchmark-memory -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-memory, .gemini/skills/benchmark-memory, .github/skills/benchmark-memory and .opencode/skills/benchmark-memory in your project.

What does Benchmark Memory need to run?

Going by SKILL.md and its folder, Benchmark Memory needs the command-line tools its instructions call (python, claude, pip and python3) and credentials named ANTHROPIC_API_KEY. Our summary lists: Python 3; A credential in ANTHROPIC_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write, Glob, Grep, Task.

Does Benchmark Memory access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Benchmark Memory safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Memory use?

Benchmark Memory is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark Memory use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark Memory?

Skills that share tags, products or a category with Benchmark Memory: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Memory?

Abilityai (a GitHub organization) maintains it in Abilityai/cornelius, which has 109 GitHub stars. The repository holds 56 skills in this directory. The repository was last updated on September 22, 2026.

Source: Abilityai/cornelius on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.