Agent skill

Test Memory System

by Abilityai in Abilityai/cornelius

Comprehensive testing playbook for Local Brain Search memory improvements (Phases 1, 3, 4)

MITAuto-check: notes

Install Test Memory System

skills CLI
$ npx skills add Abilityai/cornelius --skill test-memory-system -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Abilityai/cornelius test-memory-system --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Abilityai/cornelius.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/test-memory-system .claude/skills/test-memory-system && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-memory-system
GitHub stars
109
Token cost
~3.4k tokens
SKILL.md length
781 words
Files
1
Skills in repo
51
Repo updated
First seen
Licence
MIT

At a glance

Comprehensive testing playbook for Local Brain Search memory improvements (Phases 1, 3, 4)

  • Works in 12 steps: Environment Verification → Test Intent Classification → Test Static vs Spreading Search → …
  • SKILL.md covers Purpose, State Dependencies, Prerequisites and Inputs, plus 5 more sections
  • Calls python and pip

What it does

Test Memory System is an agent skill from Abilityai/cornelius. Comprehensive testing playbook for Local Brain Search memory improvements (Phases 1, 3, 4)

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: AI-powered second brain template for Claude Code + Obsidian. The licence is MIT.

Example prompts

  • “/test-memory-system”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Glob, Grep

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Environment Verification
  2. Test Intent Classification
  3. Test Static vs Spreading Search
  4. Test Intent-Adaptive Spreading
  5. Test Lateral Inhibition
  6. Test Usage-Based Learning Status
  7. Test Usage Event Tracking
  8. Test Q-Value Updates
  9. Test Q-Value Ranking Adjustment
  10. Test No-Track Mode
  11. Test Configuration Loading
  12. Performance Benchmarks

What it can do on your machine

Read from SKILL.md and the folder at commit fd5e9a4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Glob
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Memory System loads about 3.4k tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 781 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Glob, Grep

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Abilityai/cornelius at commit fd5e9a4, republished under its MIT licence (© Abilityai). 781 words, ~3,388 tokens.

Download SKILL.mdSave it as .claude/skills/test-memory-system/SKILL.md (or your agent's skills folder).
name
test-memory-system
description
Comprehensive testing playbook for Local Brain Search memory improvements (Phases 1, 3, 4)
allowed-tools
Bash, Read, Write, Glob, Grep
automation
manual
user-invocable
true
metadata.version
1.0
metadata.created
2026-02-18
metadata.author
Cornelius
metadata.related-systems
local-brain-search, SYNAPSE-memory-architecture

Test Memory System

Purpose

Comprehensively test the new memory improvements implemented in Local Brain Search:

  • Phase 1: Intent Classification
  • Phase 3: Spreading Activation
  • Phase 4: Usage-Based Learning (Q-values)

State Dependencies

SourceLocationReadWriteDescription
Memory Configresources/local-brain-search/memory_config.py✓Central configuration
Q-Valuesresources/local-brain-search/data/q_values.json✓Learned preferences
Usage Historyresources/local-brain-search/data/usage_history.jsonl✓Event tracking
FAISS Indexresources/local-brain-search/data/brain.faiss✓Vector index
Graphresources/local-brain-search/data/brain_graph.pkl✓Note graph
Test Reportresources/memory-test-reports/✓Output location

Prerequisites

  • Virtual environment activated: source resources/local-brain-search/venv/bin/activate
  • FAISS index built: data/brain.faiss exists
  • Graph built: data/brain_graph.pkl exists
  • Dependencies installed: pip install -r requirements.txt

Inputs

  • None (runs full test suite)

Process

Step 1: Environment Verification

Verify all components are installed and accessible.

bash
cd resources/local-brain-search
source venv/bin/activate

# Check files exist
ls -la data/brain.faiss data/brain_metadata.pkl data/brain_graph.pkl

# Check Python dependencies
python -c "import faiss; import networkx; import sentence_transformers; print('All dependencies OK')"

# Check new modules exist
python -c "from intent import classify_intent; from spreading import spreading_activation; from learning import get_learning_stats; print('All new modules OK')"

Expected: All files exist, all imports succeed.


Step 2: Test Intent Classification

Test that queries are correctly classified into four intent types.

bash
cd resources/local-brain-search
source venv/bin/activate
python intent.py

Test Cases:

QueryExpected IntentConfidence
"What is dopamine?"factual>70%
"How does motivation work?"conceptual>70%
"Connect dopamine and Buddhism"synthesis>70%
"Recent notes about AI agents"temporal>70%
"Dopamine"factual>50%
"relationship between identity and belief"synthesis>70%
"patterns across neuroscience and AI"synthesis>80%

Record Results:

  • Total queries tested: X
  • Correctly classified: Y
  • Accuracy: Y/X %

Compare results between static (traditional) and spreading activation modes.

bash
cd resources/local-brain-search

# Test 1: Static search
./run_search.sh "dopamine and motivation" --mode static --limit 5 --json | python -m json.tool

# Test 2: Spreading search (same query)
./run_search.sh "dopamine and motivation" --mode spreading --limit 5 --json | python -m json.tool

# Test 3: Synthesis query (where spreading should excel)
./run_search.sh "connect Buddhism neuroscience consciousness" --mode static --limit 5 --json
./run_search.sh "connect Buddhism neuroscience consciousness" --mode spreading --limit 5 --json

Expected Results:

  • Static and spreading should produce DIFFERENT result sets
  • Spreading mode should find cross-domain connections
  • Spreading should reduce "hub dominance" (e.g., Dopamine note appearing in everything)

Record:

  • Result overlap: X out of 5 (should be 0-2 for synthesis queries)
  • Spreading iterations: usually 3-5
  • Did spreading find non-obvious connections? Y/N

Step 4: Test Intent-Adaptive Spreading

Verify that different intents produce different spreading parameters.

bash
cd resources/local-brain-search

# Factual query (should use minimal spreading)
./run_search.sh "What is dopamine?" --mode spreading --json 2>&1 | grep -E "(iterations|Intent)"

# Synthesis query (should use maximum spreading)
./run_search.sh "connect dopamine Buddhism identity" --mode spreading --json 2>&1 | grep -E "(iterations|Intent)"

Expected:

  • Factual: max_iterations=2, inhibition_strength=0.5
  • Conceptual: max_iterations=5, inhibition_strength=0.2
  • Synthesis: max_iterations=7, inhibition_strength=0.1
  • Temporal: max_iterations=3, temporal_decay=0.7

Step 5: Test Lateral Inhibition

Verify that hub notes don't dominate results.

Scope note: Dopamine.md was the canonical hub example on the whole-graph fingerprint (DI-inflated). Scope enforcement is now ON (since 2026-06-25; see SCOPE-IMPLEMENTATION-PLAN.md), so the core-scoped hubs are MOC - Eight-Circuit / Decision Making / Gilbert / Tetlock - use one of those as the "known hub" probe. Dopamine remains a valid permanent note and search target either way - only its centrality ranking changed (it is no longer a top hub).

bash
cd resources/local-brain-search

# Search for topic where Dopamine.md is a known hub
./run_search.sh "motivation reward behavior" --mode spreading --limit 10 --json

# Check if Dopamine.md dominates or if diverse results appear

Expected:

  • Top 10 results should include notes beyond the immediate Dopamine cluster
  • Lateral inhibition suppresses over-represented clusters
  • Results should show diversity across topics

Record:

  • Number of results from Dopamine cluster: X/10
  • Number of distinct topic clusters represented: Y

Step 6: Test Usage-Based Learning Status

Check current learning system state.

bash
cd resources/local-brain-search

# Check learning status
./run_learning.sh status

# View top notes by Q-value
./run_learning.sh top --limit 10

# Export full learning data
./run_learning.sh export --output /tmp/learning_export.json

Expected:

  • Learning enabled: True
  • Events tracked with proper structure
  • Q-values in reasonable range (-1.0 to 2.0)

Record:

  • Total events tracked: X
  • Unique notes with Q-values: Y
  • Average Q-value: Z

Step 7: Test Usage Event Tracking

Verify that search operations are being tracked.

bash
cd resources/local-brain-search

# Count events before
BEFORE=$(wc -l < data/usage_history.jsonl)

# Run a search
./run_search.sh "test query for tracking" --mode spreading --limit 5

# Count events after
AFTER=$(wc -l < data/usage_history.jsonl)

echo "Events before: $BEFORE, after: $AFTER, new: $((AFTER - BEFORE))"

# Verify event structure
tail -5 data/usage_history.jsonl | python -m json.tool

Expected:

  • New events should be logged (5 for limit=5)
  • Events should have: timestamp, note_id, query, query_intent, event_type, position, session_id, mode

Show full SKILL.md (319 more words)Show less
Step 8: Test Q-Value Updates

Verify that different event types produce appropriate Q-value changes.

bash
cd resources/local-brain-search

# Get current Q-value for a test note
cat data/q_values.json | python -c "import json,sys; d=json.load(sys.stdin); print(d.get('02-Permanent/Dopamine.md', 'not found'))"

# Log a read event
./run_learning.sh log read "02-Permanent/Dopamine.md" --query "test"

# Check Q-value increased
cat data/q_values.json | python -c "import json,sys; d=json.load(sys.stdin); print(d.get('02-Permanent/Dopamine.md', 'not found'))"

Expected Q-value changes:

  • retrieved: +0.0 (no change)
  • read: +0.5 base reward
  • referenced: +1.0 base reward
  • linked: +1.5 base reward

Note: Actual changes are modulated by learning_rate (0.1) and position factor.


Step 9: Test Q-Value Ranking Adjustment

Verify that Q-values influence search result ranking.

bash
cd resources/local-brain-search

# First, artificially boost a note's Q-value
python -c "
import json
with open('data/q_values.json', 'r') as f:
    q = json.load(f)
q['02-Permanent/Identity.md'] = 1.5  # High Q-value
with open('data/q_values.json', 'w') as f:
    json.dump(q, f, indent=2)
print('Q-value set for Identity.md')
"

# Search for something where Identity.md is somewhat relevant
./run_search.sh "belief systems self" --mode spreading --limit 10 --json

# Check if Identity.md ranks higher due to Q-value boost

Expected:

  • Notes with high Q-values should rank higher (30% weight by default)
  • Q-value boost = 1.0 + (q_value * 0.3)

Step 10: Test No-Track Mode

Verify that --no-track flag prevents usage logging.

bash
cd resources/local-brain-search

# Count events before
BEFORE=$(wc -l < data/usage_history.jsonl)

# Run search with --no-track
./run_search.sh "no track test" --mode static --no-track

# Count events after
AFTER=$(wc -l < data/usage_history.jsonl)

echo "Events should be same: before=$BEFORE, after=$AFTER"

Expected: No new events logged when --no-track is used.


Step 11: Test Configuration Loading

Verify memory_config.py is the single source of truth.

bash
cd resources/local-brain-search

# Print current configuration
python memory_config.py

# Verify spreading uses config
python -c "
from memory_config import MEMORY_CONFIG
print('Spreading max_iterations:', MEMORY_CONFIG['spreading']['max_iterations'])
print('Learning enabled:', MEMORY_CONFIG['learning']['enabled'])
print('Q-weight:', MEMORY_CONFIG['learning']['q_weight'])
"

Expected:

  • All configuration centralized in memory_config.py
  • No hardcoded values in search.py, spreading.py, learning.py

Step 12: Performance Benchmarks

Measure search latency for both modes.

bash
cd resources/local-brain-search

# Benchmark static search
time (for i in {1..5}; do ./run_search.sh "dopamine" --mode static --limit 10 --no-track > /dev/null; done)

# Benchmark spreading search
time (for i in {1..5}; do ./run_search.sh "dopamine" --mode spreading --limit 10 --no-track > /dev/null; done)

Expected:

  • Static search: ~100-200ms per query
  • Spreading search: ~300-500ms per query
  • Spreading should be <2x slower than static

Step 13: Edge Case Testing

Test boundary conditions.

bash
cd resources/local-brain-search

# Empty query
./run_search.sh "" --mode spreading --limit 5 2>&1

# Very long query
./run_search.sh "$(python -c 'print("dopamine " * 100)')" --mode spreading --limit 5 2>&1

# Non-existent topic
./run_search.sh "xyznonexistenttopicxyz" --mode spreading --limit 5 --json

# Unicode query
./run_search.sh "意識 consciousness" --mode spreading --limit 5 --json

Expected:

  • Empty query: Graceful error or default behavior
  • Long query: Should truncate or handle gracefully
  • Non-existent: Return empty or low-confidence results
  • Unicode: Should not crash

Generate Test Report

After running all tests, generate a comprehensive report.

bash
REPORT_DIR="resources/memory-test-reports"
REPORT_FILE="$REPORT_DIR/test-report-$(date +%Y-%m-%d-%H%M).md"
mkdir -p "$REPORT_DIR"
Report Template
markdown
# Memory System Test Report

**Date:** YYYY-MM-DD HH:MM
**Tester:** [name]
**System Version:** Cornelius v01.25

## Summary

| Component | Status | Notes |
|-----------|--------|-------|
| Intent Classification | ✓/✗ | |
| Spreading Activation | ✓/✗ | |
| Lateral Inhibition | ✓/✗ | |
| Usage Tracking | ✓/✗ | |
| Q-Value Learning | ✓/✗ | |
| Configuration Loading | ✓/✗ | |
| Performance | ✓/✗ | |

## Detailed Results

### Intent Classification
- Accuracy: X%
- Failures: [list]

### Spreading vs Static
- Result overlap: X/5
- Spreading found cross-domain connections: Y/N
- Example good result: [describe]

### Lateral Inhibition
- Hub dominance reduced: Y/N
- Diversity improved: Y/N

### Learning System
- Events tracked: X
- Q-values updated correctly: Y/N
- Ranking adjustment working: Y/N

### Performance
- Static search avg: Xms
- Spreading search avg: Xms
- Acceptable: Y/N

## Issues Found

1. [Issue description]
   - Severity: High/Medium/Low
   - Steps to reproduce: [steps]
   - Suggested fix: [suggestion]

## Recommendations

1. [Recommendation]
2. [Recommendation]

## Next Steps

- [ ] Address issues found
- [ ] Move to Phase 2 (Extended Graph) if all passes
- [ ] Enable spreading as default mode if stable

Completion Checklist

  • Step 1: Environment verified
  • Step 2: Intent classification tested (>85% accuracy)
  • Step 3: Static vs spreading compared
  • Step 4: Intent-adaptive spreading verified
  • Step 5: Lateral inhibition working
  • Step 6: Learning status checked
  • Step 7: Usage tracking verified
  • Step 8: Q-value updates working
  • Step 9: Ranking adjustment confirmed
  • Step 10: No-track mode verified
  • Step 11: Configuration centralized
  • Step 12: Performance acceptable
  • Step 13: Edge cases handled
  • Test report generated and saved

Error Recovery

If tests fail:

  1. Import errors: Check requirements.txt installed in venv
  2. File not found: Run python index_brain.py to rebuild index
  3. Learning errors: Run ./run_learning.sh reset --confirm to reset
  4. Config errors: Check memory_config.py syntax

© Abilityai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/test-memory-system of Abilityai/cornelius.

Open the folder on GitHubat commit fd5e9a4

Compare with similar skills

Test Memory System next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Memory System compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Memory System this skillAbilityai/cornelius109—~3.4kAutomated safety check: NotesMIT
Brainjeremylongshore/tons-of-skills-marketplace2.8k—~1.9kAutomated safety check: PassApache-2.0
Skill Improversickn33/agentic-awesome-skills47k2 repos~1.5kAutomated safety check: PassMIT
Brain Taxonomistgarrytan/gbrain31k—~2.3kAutomated safety check: PassMIT
Improve Oaselastic/kibana21k—~4.8kAutomated safety check: PassCustom licence
Brainbrycewang-stanford/Awesome-Journal-Skills1.2k—~2kAutomated safety check: PassMIT

Similar skills

  • Brain

    jeremylongshore/tons-of-skills-marketplace

    Answers questions about your own systems, notes, decisions, runbooks, and conventions from your governed knowledge brain, returning a qmd:// citation for every claim — receipts, not recall.

    2.8k GitHub stars~1.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Skill Improver

    sickn33/agentic-awesome-skills

    Iteratively improve a Claude Code skill using the skill-reviewer agent until it meets quality standards.

    47k GitHub starsUsed in 2 repos~1.5k tokens
    Agent WorkflowsAuto-check passed
  • Brain Taxonomist

    garrytan/gbrain

    Filing gate for ALL brain writes. An agent skill from garrytan/gbrain.

    31k GitHub stars~2.3k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Improve Oas

    elastic/kibana

    Official

    Add or improve OpenAPI descriptions, examples, and code samples for a Kibana API area.

    21k GitHub stars~4.8k tokensUpdated today
    Backend & APIsAuto-check passed
  • Brain

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when targeting Brain or deciding whether a clinical-neurology or translational-neuroscience study fits this venue.

    1.2k GitHub stars~2k tokensUpdated 11 days ago
    Research & ScienceAuto-check passed
  • Brain

    majiayu000/claude-skill-registry

    A skill your agent uses when the user says brain, project knowledge, what do we know about, open tasks, or recent decisions, or runs /brain or /brain-update.

    666 GitHub starsUsed in 1 repo~1.5k tokens
    DatabasesAuto-check passed

More from Abilityai/cornelius

All 51 skills in this repo
  • Nano Banana Image Generator

    Abilityai/cornelius

    Generate images using Google's Nano Banana (Gemini 2.5 Flash Image).

    109 GitHub stars~1.2k tokensUpdated 16 days ago
    Auto-check: notes
  • Changelog Protocol

    Abilityai/cornelius

    Protocol for creating dated changelog files after significant agent sessions.

    109 GitHub stars~555 tokensUpdated 16 days ago
    Auto-check passed
  • Create Article

    Abilityai/cornelius

    Create long-form articles from knowledge base insights. An agent skill from Abilityai/cornelius.

    109 GitHub stars~2.1k tokensUpdated 16 days ago
    Auto-check: notes
  • Epistemic Classification

    Abilityai/cornelius

    Framework for distinguishing research findings from hypotheses and speculative synthesis.

    109 GitHub stars~1.5k tokensUpdated 16 days ago
    Auto-check passed
  • Get Youtube Transcript

    Abilityai/cornelius

    Extract the transcript from a YouTube video by URL or video ID.

    109 GitHub stars~525 tokensUpdated 16 days ago
    Auto-check: notes
  • Insight Capture Format

    Abilityai/cornelius

    Standard format for capturing and documenting insights in the knowledge base.

    109 GitHub stars~616 tokensUpdated 16 days ago
    Auto-check passed

Questions about Test Memory System

What does Test Memory System do?

Comprehensive testing playbook for Local Brain Search memory improvements (Phases 1, 3, 4). Test Memory System is an agent skill from Abilityai/cornelius.

How do I install Test Memory System in Claude Code?

Run `npx skills add Abilityai/cornelius --skill test-memory-system -a claude-code`. Or copy the skill folder (.claude/skills/test-memory-system in Abilityai/cornelius) into .claude/skills/test-memory-system in your project. Claude Code loads it when a task matches its description.

How do I install Test Memory System in Codex?

Run `npx skills add Abilityai/cornelius --skill test-memory-system -a codex`. Or copy the skill folder (.claude/skills/test-memory-system in Abilityai/cornelius) into .agents/skills/test-memory-system in your project. Codex loads it when a task matches its description.

Can I use Test Memory System in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Abilityai/cornelius --skill test-memory-system -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-memory-system, .gemini/skills/test-memory-system, .github/skills/test-memory-system and .opencode/skills/test-memory-system in your project.

What does Test Memory System need to run?

Going by SKILL.md and its folder, Test Memory System needs the command-line tools its instructions call (python and pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read, Write, Glob, Grep.

Does Test Memory System access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Memory System safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Test Memory System use?

Test Memory System is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Memory System use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Memory System?

Skills that share tags, products or a category with Test Memory System: Brain (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Skill Improver (sickn33/agentic-awesome-skills, 47k stars), Brain Taxonomist (garrytan/gbrain, 31k stars) and Improve Oas (elastic/kibana, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Memory System?

Abilityai (a GitHub organization) maintains it in Abilityai/cornelius, which has 109 GitHub stars. The repository holds 51 skills in this directory. The repository was last updated on September 22, 2026.

Source: Abilityai/cornelius on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.