Agent skill

Benchmark Context

by Signet-AI in Signet-AI/signetai

Automatically benchmark your custom memory implementation against established systems like Supermemory.

Custom licenceAuto-check: notesAgent Workflows

Install Benchmark Context

skills CLI
$ npx skills add Signet-AI/signetai --skill benchmark-context -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Signet-AI/signetai benchmark-context --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Signet-AI/signetai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/memorybench/skills/memorybench .claude/skills/benchmark-context && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-context
GitHub stars
304
Token cost
~3k tokens
SKILL.md length
1,292 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
Custom licence

At a glance

Automatically benchmark your custom memory implementation against established systems like Supermemory.

  • Works in 12 steps: Provider Name → Memory Code Location → Benchmark Dataset → …
  • Agent Workflows work in your project
  • SKILL.md covers What This Skill Does, When to Use This Skill, How It Works and Initial Questions, plus 8 more sections
  • Calls bun and git; reaches github.com

What it does

Benchmark Context is an agent skill from Signet-AI/signetai. Automatically benchmark your custom memory implementation against established systems like Supermemory. Set up a public benchmark, or create your own. Compare solutions against quality, latency, features and cost, easily, with a simple UI and CLI.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows. The repository describes itself as: Sync and store memories, shared identity files (AGENTS.md, CLAUDE.md), session transcripts, institutional knowledge, and secrets between all of your favorite harnesses and models.

When your agent uses it

  • Agent Workflows work in your project

Example prompts

  • “/benchmark-context”

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Provider Name
  2. Memory Code Location
  3. Benchmark Dataset
  4. Comparison Targets (Multi-select)
  5. Test Size
  6. Verify Environment
  7. Use a pinned MemoryBench checkout
  8. Gather User Input
  9. Analyze User's Memory Code
  10. Generate Provider Code
  11. Register Provider
  12. Configure Environment

What it can do on your machine

Read from SKILL.md and the folder at commit aa4c499. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bun
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Context loads about 3k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 1,292 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:58
    - Creates `.env.local` with required API keys
  • NoteMentions a .env fileSKILL.md:142
    ├── .env.local                              # Your API keys
  • NoteMentions a .env fileSKILL.md:179
    not initialized"** - Check API keys in `.env.local`
  • NoteMentions a .env fileSKILL.md:259
    Create or update `memorybench/.env.local` with provided values.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 1,292 words (~3,047 tokens).

“Automatically benchmark your custom memory implementation against established systems like Supermemory, Mem0, and Zep.”

— opening of SKILL.md by Signet-AI, Custom licence
name
benchmark-context

Read the full SKILL.md on GitHub

Files

Just SKILL.md in memorybench/skills/memorybench of Signet-AI/signetai.

Open the folder on GitHubat commit aa4c499

Compare with similar skills

Benchmark Context next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Context compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Context this skillSignet-AI/signetai304—~3kAutomated safety check: NotesCustom licence
Claude Code Agent Developmentanthropics/claude-plugins-official38k7 repos~2.8kAutomated safety check: PassApache-2.0
Copilot Session Failure Analysisdotnet/maui23k—~3.4kAutomated safety check: PassMIT
Mem0 CLI Memory Commandsmem0ai/mem067k—~2kAutomated safety check: NotesApache-2.0
Create Agentvectorize-io/hindsight47k—~1.1kAutomated safety check: PassMIT
Google Antigravity SDKgoogle-antigravity/antigravity-sdk-python3.7k—~2.1kAutomated safety check: NotesApache-2.0

Similar skills

  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 7 repos~2.8k tokens
    Agent WorkflowsAuto-check passed
  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Adds, searches, lists, updates and deletes memories on the Mem0 platform from the terminal with the mem0 command, including a JSON mode built for agents.

    67k GitHub stars~2k tokensUpdated today
    Agent WorkflowsAuto-check: notes
  • Create Agent

    vectorize-io/hindsight

    Create a new Hindsight-powered subagent with long-term memory.

    47k GitHub stars~1.1k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Google Antigravity SDK

    google-antigravity/antigravity-sdk-python

    Design, implement, and debug autonomous AI agents and multi-agent systems using the Google Antigravity (AGY) SDK.

    3.7k GitHub stars~2.1k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check: notes
  • Yao Meta Skill

    yaojingang/yao-meta-skill

    Create, improve, or evaluate an existing skill from workflows, prompts, SOPs, scripts.

    2.7k GitHub stars~768 tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed

More from Signet-AI/signetai

All 9 skills in this repo
  • Dreaming Development

    Signet-AI/signetai

    A skill your agent uses for Signet Dreaming development: inspect existing source, semantic-memory, retrieval, and inference architecture before changing it; prevent duplicate modules and ad-hoc…

    304 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Benchmarking

    Signet-AI/signetai

    Benchmark Signet memory with MemoryBench: regression-check Dreaming and recall changes, compare models, diagnose score drops, and record results; NOT for the prompt-submit latency benchmark.

    304 GitHub stars~1.5k tokensUpdated today
    Auto-check: notes
  • Onboarding

    Signet-AI/signetai

    Interactive interview to set up your Signet workspace (~5-10 minutes).

    304 GitHub stars~6.3k tokensUpdated today
    Auto-check passed
  • Agents Md Sync

    Signet-AI/signetai

    Prune and route AGENTS.md trees for durable, high-signal context.

    304 GitHub stars~918 tokensUpdated today
    Auto-check passed
  • Dreaming

    Signet-AI/signetai

    Maintain Signet's living ontology and memory substrate from transcripts, memory artifacts, source artifacts, notes, summaries, and imported records.

    304 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Provider SDK Upgrades

    Signet-AI/signetai

    Upgrade provider SDKs and refresh bundled model presets; NOT for unrelated provider architecture.

    304 GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Benchmark Context

What does Benchmark Context do?

Automatically benchmark your custom memory implementation against established systems like Supermemory. Benchmark Context is an agent skill from Signet-AI/signetai. Automatically benchmark your custom memory implementation against established systems like Supermemory.

When should I use Benchmark Context?

Benchmark Context fits situations like: agent Workflows work in your project.

How do I install Benchmark Context in Claude Code?

Run `npx skills add Signet-AI/signetai --skill benchmark-context -a claude-code`. Or copy the skill folder (memorybench/skills/memorybench in Signet-AI/signetai) into .claude/skills/benchmark-context in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Context in Codex?

Run `npx skills add Signet-AI/signetai --skill benchmark-context -a codex`. Or copy the skill folder (memorybench/skills/memorybench in Signet-AI/signetai) into .agents/skills/benchmark-context in your project. Codex loads it when a task matches its description.

Can I use Benchmark Context in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Signet-AI/signetai --skill benchmark-context -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-context, .gemini/skills/benchmark-context, .github/skills/benchmark-context and .opencode/skills/benchmark-context in your project.

What does Benchmark Context need to run?

Going by SKILL.md and its folder, Benchmark Context needs the command-line tools its instructions call (bun and git).

Does Benchmark Context access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Benchmark Context safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Context use?

Benchmark Context has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Benchmark Context use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark Context?

Skills that share tags, products or a category with Benchmark Context: Claude Code Agent Development (anthropics/claude-plugins-official, 38k stars), Copilot Session Failure Analysis (dotnet/maui, 23k stars), Mem0 CLI Memory Commands (mem0ai/mem0, 67k stars) and Create Agent (vectorize-io/hindsight, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Context?

Signet-AI (a GitHub organization) maintains it in Signet-AI/signetai, which has 304 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 8, 2026.

Source: Signet-AI/signetai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.