LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Systematic benchmarking framework for Local Brain Search memory system with LLM-as-judge scoring
$ npx skills add Abilityai/cornelius --skill benchmark-memory -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Abilityai/cornelius benchmark-memory --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Abilityai/cornelius.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/benchmark-memory .claude/skills/benchmark-memory && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchmark-memory" agent skill from https://github.com/Abilityai/cornelius/tree/main/.claude/skills/benchmark-memory into .claude/skills/benchmark-memory/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-memory", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Abilityai/cornelius/tree/main/.claude/skills/benchmark-memoryType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Abilityai/cornelius --skill benchmark-memory -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Abilityai/cornelius benchmark-memory --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Abilityai/cornelius.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/benchmark-memory .agents/skills/benchmark-memory && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchmark-memory" agent skill from https://github.com/Abilityai/cornelius/tree/main/.claude/skills/benchmark-memory into .agents/skills/benchmark-memory/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-memory", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Abilityai/cornelius --skill benchmark-memory -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Abilityai/cornelius benchmark-memory --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Abilityai/cornelius.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/benchmark-memory .cursor/skills/benchmark-memory && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchmark-memory" agent skill from https://github.com/Abilityai/cornelius/tree/main/.claude/skills/benchmark-memory into .cursor/skills/benchmark-memory/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-memory", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Abilityai/cornelius.git --path .claude/skills/benchmark-memory--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Abilityai/cornelius --skill benchmark-memory -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Abilityai/cornelius benchmark-memory --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Abilityai/cornelius.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/benchmark-memory .gemini/skills/benchmark-memory && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchmark-memory" agent skill from https://github.com/Abilityai/cornelius/tree/main/.claude/skills/benchmark-memory into .gemini/skills/benchmark-memory/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-memory", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Abilityai/cornelius benchmark-memoryInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Abilityai/cornelius --skill benchmark-memory -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Abilityai/cornelius.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/benchmark-memory .github/skills/benchmark-memory && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-memory" agent skill from https://github.com/Abilityai/cornelius/tree/main/.claude/skills/benchmark-memory into .github/skills/benchmark-memory/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-memory", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Abilityai/cornelius --skill benchmark-memory -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Abilityai/cornelius benchmark-memory --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Abilityai/cornelius.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/benchmark-memory .opencode/skills/benchmark-memory && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-memory" agent skill from https://github.com/Abilityai/cornelius/tree/main/.claude/skills/benchmark-memory into .opencode/skills/benchmark-memory/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-memory", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchmark-memorySystematic benchmarking framework for Local Brain Search memory system with LLM-as-judge scoring
Benchmark Memory is an agent skill from Abilityai/cornelius. Systematic benchmarking framework for Local Brain Search memory system with LLM-as-judge scoring
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: AI-powered second brain template for Claude Code + Obsidian. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit fd5e9a4. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadWriteGlobGrepTaskFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonclaudepippython3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
ANTHROPIC_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Benchmark Memory loads about 2.4k tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 731 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, Read, Write, Glob, Grep, TaskAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Abilityai/cornelius at commit fd5e9a4, republished under its MIT licence (© Abilityai). 731 words, ~2,377 tokens.
.claude/skills/benchmark-memory/SKILL.md (or your agent's skills folder).Systematic benchmarking framework to measure retrieval quality, compare configurations, and identify optimal parameters for the Local Brain Search memory system.
| Source | Location | Read | Write |
|---|---|---|---|
| Brain snapshot | .claude/skills/benchmark-memory/snapshots/ | Yes | Yes |
| Query sets | .claude/skills/benchmark-memory/query-sets/ | Yes | Yes |
| Benchmark results | .claude/skills/benchmark-memory/results/ | Yes | Yes |
| Analysis reports | .claude/skills/benchmark-memory/analysis/ | No | Yes |
| Memory system | resources/local-brain-search/ | Yes | No |
resources/local-brain-search/data/brain.faiss)resources/local-brain-search/venv/ with search dependenciesThis skill uses Claude Code headless mode (claude -p) for LLM relevance scoring, not a separate API key. This means:
ANTHROPIC_API_KEY environment variable neededsonnet (good quality) - can also use haiku (faster/cheaper) or opusTo verify Claude Code is available:
claude --versionDependencies are installed in the local-brain-search venv:
cd resources/local-brain-search
source venv/bin/activate
pip install pandas tqdm # anthropic not required - uses Claude Code headless/benchmark-memory setupCreate a frozen Brain snapshot and build its index.
cd .claude/skills/benchmark-memory/scripts
./run_benchmark.sh --list-snapshots # Check existing
python3 create_snapshot.py # Create new snapshotWhat it does:
/benchmark-memory create-queries [--count N]Generate or manage test query sets.
cd .claude/skills/benchmark-memory/scripts
python build_query_set.py --count 50 --output ../query-sets/core-50.jsonQuery Categories (50 total):
| Category | Count | Example |
|---|---|---|
| Factual | 10 | "What is dopamine?" |
| Conceptual | 10 | "How does motivation work?" |
| Synthesis | 15 | "Connect Buddhism and neuroscience" |
| Temporal | 5 | "Recent notes about AI agents" |
| Needle | 5 | "Note about intermittent reinforcement" |
| Broad | 5 | "Identity" |
/benchmark-memory run [--config CONFIG] [--snapshot SNAPSHOT]Execute benchmark with specified configuration.
cd .claude/skills/benchmark-memory/scripts
./run_benchmark.sh --config focused --snapshot brain-snapshot-2026-02-18
./run_benchmark.sh --dry-run --config focused # Preview without execution
./run_benchmark.sh --list-configs # List available configsConfigurations:
focused: 15 key configurations (recommended for initial benchmarking)single:CONFIG_NAME: Run single configurationall: Full parameter sweep (expensive)Estimated cost per run:
/benchmark-memory analyze [--results FILE]Generate analysis summary from benchmark results.
cd .claude/skills/benchmark-memory/scripts
python analyze_results.py --results ../results/benchmark-*.csvOutputs:
Step 1: Setup (one-time)
/benchmark-memory setup
|
v
Step 2: Create Queries (one-time)
/benchmark-memory create-queries --count 50
|
v
Step 3: Run Benchmark (per experiment)
/benchmark-memory run --config focused
|
v
Step 4: Analyze Results
/benchmark-memory analyze| Metric | Description |
|---|---|
latency_ms | Query execution time |
iterations | Spreading iterations used |
converged | Whether spreading converged |
| Metric | Range | Description |
|---|---|---|
| Precision@K | 0-1 | Fraction of results that are relevant |
| Recall@K | 0-1 | Fraction of relevant notes found |
| MRR | 0-1 | Mean Reciprocal Rank |
| NDCG@K | 0-1 | Ranking quality with position discount |
| Avg Score | 0-3 | Average LLM relevance score |
| Score | Label | Definition |
|---|---|---|
| 0 | Irrelevant | No connection to query |
| 1 | Tangential | Loosely related |
| 2 | Relevant | Addresses the query |
| 3 | Highly Relevant | Directly answers the query |
static_baseline: Traditional vector searchspreading_default: Spreading activation with defaultssynthesis_optimized: max_iterations=7, inhibition=0.1factual_optimized: max_iterations=2, inhibition=0.5balanced_optimized: max_iterations=5, inhibition=0.2, decay=0.85results/benchmark-YYYY-MM-DD-HHMMSS.csvSchema:
timestamp,config_name,query_id,query_category,mode,max_iterations,
inhibition_strength,latency_ms,result_1_note,result_1_score,...,
precision_at_5,precision_at_10,mrr,ndcg_at_10,avg_scoreanalysis/report-YYYY-MM-DD.mdResults are appended incrementally. Resume by running with --resume:
python run_benchmark.py --config focused --resumeBuilt-in retry with exponential backoff. Adjust --delay if needed:
python run_benchmark.py --config focused --delay 1.0Re-create snapshot:
python create_snapshot.py --forceAfter running benchmarks, answer these questions:
| Strategy | Savings | Trade-off |
|---|---|---|
| Score top 5 only | 50% | Less data on long-tail |
| Use Haiku judge | 90% | Slightly less accurate |
| Cache scores | Variable | Only for unchanged retrieval |
Recommendation: Start with Haiku judge, validate sample against Sonnet.
.claude/skills/benchmark-memory/
├── SKILL.md # This file
├── requirements.txt # Python dependencies
├── scripts/
│ ├── run_benchmark.sh # Wrapper script (uses venv Python)
│ ├── create_snapshot.py # Create frozen Brain snapshot
│ ├── build_query_set.py # Generate/manage query test set
│ ├── run_benchmark.py # Execute benchmark with config
│ ├── score_results.py # LLM-as-judge scoring
│ ├── compute_metrics.py # Calculate evaluation metrics
│ └── analyze_results.py # Generate analysis summary
├── configs/
│ ├── focused_configs.json # Test configurations (15 configs)
│ └── judge_prompt.txt # LLM judge prompt template
├── snapshots/ # Frozen Brain copies
├── query-sets/ # Test queries (core-50.json included)
├── results/ # Benchmark CSVs
└── analysis/ # Analysis reports| Component | Status | Notes |
|---|---|---|
| Snapshot creation | ✅ Works | Creates snapshot with FAISS index + graph |
| Query set | ✅ Works | 50 queries across 6 categories |
| Static search | ✅ Works | Traditional vector similarity |
| Spreading search | ✅ Works | Multi-iteration activation |
| 15 configs | ✅ Works | Focused parameter sweep |
| LLM-as-judge | ✅ Works | Uses Claude Code headless mode (claude -p) |
| Results CSV | Ready | Incremental writes, resume support |
© Abilityai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/benchmark-memory of Abilityai/cornelius.
Open the folder on GitHubat commit fd5e9a4
Benchmark Memory next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Benchmark Memory this skillAbilityai/cornelius | 109 | — | ~2.4k | Automated safety check: Notes | MIT | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
Abilityai/cornelius
Generate images using Google's Nano Banana (Gemini 2.5 Flash Image).
Abilityai/cornelius
Protocol for creating dated changelog files after significant agent sessions.
Abilityai/cornelius
Create long-form articles from knowledge base insights. An agent skill from Abilityai/cornelius.
Abilityai/cornelius
Framework for distinguishing research findings from hypotheses and speculative synthesis.
Abilityai/cornelius
Extract the transcript from a YouTube video by URL or video ID.
Abilityai/cornelius
Standard format for capturing and documenting insights in the knowledge base.
Categories
Systematic benchmarking framework for Local Brain Search memory system with LLM-as-judge scoring. Benchmark Memory is an agent skill from Abilityai/cornelius.
Benchmark Memory fits situations like: tasks that involve LLM evaluation.
Run `npx skills add Abilityai/cornelius --skill benchmark-memory -a claude-code`. Or copy the skill folder (.claude/skills/benchmark-memory in Abilityai/cornelius) into .claude/skills/benchmark-memory in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Abilityai/cornelius --skill benchmark-memory -a codex`. Or copy the skill folder (.claude/skills/benchmark-memory in Abilityai/cornelius) into .agents/skills/benchmark-memory in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Abilityai/cornelius --skill benchmark-memory -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-memory, .gemini/skills/benchmark-memory, .github/skills/benchmark-memory and .opencode/skills/benchmark-memory in your project.
Going by SKILL.md and its folder, Benchmark Memory needs the command-line tools its instructions call (python, claude, pip and python3) and credentials named ANTHROPIC_API_KEY. Our summary lists: Python 3; A credential in ANTHROPIC_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write, Glob, Grep, Task.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Benchmark Memory is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Benchmark Memory: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Abilityai (a GitHub organization) maintains it in Abilityai/cornelius, which has 109 GitHub stars. The repository holds 56 skills in this directory. The repository was last updated on September 22, 2026.
Source: Abilityai/cornelius on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.