Agent skill

Benchmark Methodology

by jgravelle in jgravelle/jcodemunch-mcp

Rules for measuring or quoting any benchmark number in this repo.

Custom licenceAuto-check passedAgent Workflows

Install Benchmark Methodology

skills CLI
$ npx skills add jgravelle/jcodemunch-mcp --skill benchmark-methodology -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jgravelle/jcodemunch-mcp benchmark-methodology --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jgravelle/jcodemunch-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/benchmark-methodology .claude/skills/benchmark-methodology && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-methodology
GitHub stars
2.7k
Token cost
~438 tokens
SKILL.md length
191 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
Custom licence

At a glance

Rules for measuring or quoting any benchmark number in this repo.

  • Works in 8 steps: Never hand-type a number. Every figure… → Per row, never per total (F-13): the… → Five mirrors move together (Practice 4):… → …
  • Tasks that involve Changelog and release notes
  • Calls uv
  • Tasks that involve MCP servers

What it does

Benchmark Methodology is an agent skill from jgravelle/jcodemunch-mcp. Rules for measuring or quoting any benchmark number in this repo. Load before running the bench tier, writing a delta into a PR or CHANGELOG, or touching anything under benchmarks/.

Its SKILL.md is about 440 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Changelog and release notes and MCP servers. It works with Model Context Protocol. The repository describes itself as: Cut AI token costs 95%+ on code exploration. The leading MCP server for precise, symbol-level GitHub code retrieval via tree-sitter AST. Works with Claude Code, Cursor & any MCP…

When your agent uses it

  • Tasks that involve Changelog and release notes
  • Tasks that involve MCP servers

Example prompts

  • “Use the benchmark-methodology skill to rule for measuring or quoting any benchmark number in this repo”
  • “/benchmark-methodology”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Never hand-type a number. Every figure comes from a run in THIS
  2. Per row, never per total (F-13): the total hid one cause behind
  3. Five mirrors move together (Practice 4): results.md,
  4. The reference is captured where the gate runs (F-13)
  5. Deterministic configuration is the bench tier itself: --offline,
  6. A refusal is not a zero: an absent value prints n/a.
  7. Floors live only in harness/thresholds.json; read one with
  8. benchmarks/schema_baseline.json is the only source for schema-token

What it can do on your machine

Read from SKILL.md and the folder at commit 526da05. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Methodology loads about 438 tokens when it runs. Until then it costs about 51 tokens; SKILL.md has 191 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~51
When it runs · the whole SKILL.md, loaded when a task matches
~438

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 191 words (~438 tokens).

“Authority: docs/harness/ARCHAEOLOGY.md (the benchmark rows), benchmarks/METHODOLOGY.md, benchmarks/REPRODUCING.md, CLAUDE.md Practice 4, docs/harness/FINDINGS.md F-10, F-13, F-17.”

— opening of SKILL.md by jgravelle, Custom licence
name
benchmark-methodology

Read the full SKILL.md on GitHub

Files

Just SKILL.md in .claude/skills/benchmark-methodology of jgravelle/jcodemunch-mcp.

Open the folder on GitHubat commit 526da05

Compare with similar skills

Benchmark Methodology next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Methodology compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Methodology this skilljgravelle/jcodemunch-mcp2.7k—~438Automated safety check: PassCustom licence
Project Releaseswimmwatch/cloakbrowser-mcp164—~1.9kAutomated safety check: PassMIT
ReleasePrefectHQ/fastmcp28k—~2.9kAutomated safety check: PassApache-2.0
Reading Livekit Docslivekit-examples/agent-starter-python2641 repos~1.1kAutomated safety check: PassMIT
MCP Server Trello Releasedelorenj/mcp-server-trello445—~3.5kAutomated safety check: PassMIT
Githits Releasegithits-com/githits-cli115—~2.8kAutomated safety check: PassApache-2.0

Similar skills

  • Project Release

    swimmwatch/cloakbrowser-mcp

    Prepare, publish, verify, or recover a cloakbrowser-mcp release only when the user explicitly requests release work.

    164 GitHub stars~1.9k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Release

    PrefectHQ/fastmcp

    Cut a FastMCP release end to end. An agent skill from PrefectHQ/fastmcp.

    28k GitHub stars~2.9k tokensUpdated today
    DevelopmentAuto-check passed
  • Reading Livekit Docs

    livekit-examples/agent-starter-python

    Looks up current LiveKit facts (API signatures, CLI flags, config options, model and provider support, SDK changelogs, pricing) from the docs instead of answering from memory.

    264 GitHub starsUsed in 1 repo~1.1k tokens
    DevelopmentAuto-check passed
  • MCP Server Trello Release

    delorenj/mcp-server-trello

    Canonical build → release → tag → publish procedure for the @delorenj/mcp-server-trello repo.

    445 GitHub stars~3.5k tokensUpdated 17 days ago
    DevelopmentAuto-check passed
  • Githits Release

    githits-com/githits-cli

    A skill your agent uses when maintaining the GitHits changelog or preparing, reviewing, or executing a GitHits release.

    115 GitHub stars~2.8k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Suede Launch Packaging

    JasonColapietro/suede-creator-skills

    Suede-owned launch packaging and install verification. An agent skill from JasonColapietro/suede-creator-skills.

    127 GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check passed

More from jgravelle/jcodemunch-mcp

All 10 skills in this repo
  • Release

    jgravelle/jcodemunch-mcp

    Publishing a jMunch release (jcodemunch-mcp, jdocmunch-mcp, jdatamunch-mcp, jragmunch-cli), reviewing/merging/closing PRs, and responding to the community.

    2.7k GitHub stars~6.5k tokensUpdated today
    Auto-check passed
  • Observatory

    jgravelle/jcodemunch-mcp

    Context for the jcodemunch-observatory weekly scorecard — what it tracks, the Monday 06:00 UTC cron, why scores must be pulled live rather than transcribed, the NestJS grade story, and the…

    2.7k GitHub stars~782 tokensUpdated today
    Auto-check passed
  • PR Description

    jgravelle/jcodemunch-mcp

    The PR title and body template, and the rules for any comment drafted for a reporter or contributor.

    2.7k GitHub stars~409 tokensUpdated today
    Auto-check passed
  • Changelog Format

    jgravelle/jcodemunch-mcp

    The shape of a CHANGELOG.md entry and heading in this repo. An agent skill from jgravelle/jcodemunch-mcp.

    2.7k GitHub stars~320 tokensUpdated today
    Auto-check passed
  • Claude Md Budget

    jgravelle/jcodemunch-mcp

    How to keep CLAUDE.md under its character Floor without deleting a rule: measure sections first, rotate by derivability, two edits per rotation.

    2.7k GitHub stars~426 tokensUpdated today
    Auto-check passed
  • Mechanism Not Instance

    jgravelle/jcodemunch-mcp

    The two questions that separate a fix for the reported spelling from a fix for the property.

    2.7k GitHub stars~332 tokensUpdated today
    Auto-check passed

Questions about Benchmark Methodology

What does Benchmark Methodology do?

Rules for measuring or quoting any benchmark number in this repo. Benchmark Methodology is an agent skill from jgravelle/jcodemunch-mcp. Rules for measuring or quoting any benchmark number in this repo.

When should I use Benchmark Methodology?

Benchmark Methodology fits situations like: tasks that involve Changelog and release notes; tasks that involve MCP servers.

How do I install Benchmark Methodology in Claude Code?

Run `npx skills add jgravelle/jcodemunch-mcp --skill benchmark-methodology -a claude-code`. Or copy the skill folder (.claude/skills/benchmark-methodology in jgravelle/jcodemunch-mcp) into .claude/skills/benchmark-methodology in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Methodology in Codex?

Run `npx skills add jgravelle/jcodemunch-mcp --skill benchmark-methodology -a codex`. Or copy the skill folder (.claude/skills/benchmark-methodology in jgravelle/jcodemunch-mcp) into .agents/skills/benchmark-methodology in your project. Codex loads it when a task matches its description.

Can I use Benchmark Methodology in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jgravelle/jcodemunch-mcp --skill benchmark-methodology -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-methodology, .gemini/skills/benchmark-methodology, .github/skills/benchmark-methodology and .opencode/skills/benchmark-methodology in your project.

What does Benchmark Methodology need to run?

Going by SKILL.md and its folder, Benchmark Methodology needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Benchmark Methodology access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Benchmark Methodology safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Methodology use?

Benchmark Methodology has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Benchmark Methodology use?

About 438 tokens (SKILL.md is roughly 1.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark Methodology?

Skills that share tags, products or a category with Benchmark Methodology: Project Release (swimmwatch/cloakbrowser-mcp, 164 stars), Release (PrefectHQ/fastmcp, 28k stars), Reading Livekit Docs (livekit-examples/agent-starter-python, 264 stars) and MCP Server Trello Release (delorenj/mcp-server-trello, 445 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Methodology?

jgravelle (a GitHub user) maintains it in jgravelle/jcodemunch-mcp, which has 2,749 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 10, 2026.

Source: jgravelle/jcodemunch-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.