Run OhMyCode benchmarks — score any provider/model with token tracking.

MITAuto-check passedAI & LLM Engineering

Install Bench

skills CLI
$ npx skills add AlphaLab-USTC/OhMyCode --skill bench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AlphaLab-USTC/OhMyCode bench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AlphaLab-USTC/OhMyCode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/bench .claude/skills/bench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bench
GitHub stars
131
Token cost
~941 tokens
SKILL.md length
326 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Run OhMyCode benchmarks — score any provider/model with token tracking.

  • Works in 3 steps: Run the Benchmark → Read Results → Analyze Failures
  • User wants to benchmark
  • SKILL.md covers When to Use, Input, Step 1 — Run the Benchmark and Step 2 — Read Results, plus 5 more sections
  • Calls python3

What it does

Bench is an agent skill from AlphaLab-USTC/OhMyCode. Run OhMyCode benchmarks — score any provider/model with token tracking. Use when user wants to benchmark, evaluate, test performance, or compare models.

Its SKILL.md is about 940 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM observability. The repository describes itself as: Minimal and Customizable CC-Style Coding Agent. The licence is MIT.

When your agent uses it

  • User wants to benchmark
  • Test performance

Example prompts

  • “/bench”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Run the Benchmark
  2. Read Results
  3. Analyze Failures

What it can do on your machine

Read from SKILL.md and the folder at commit 4d1bb28. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bench loads about 941 tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 326 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~941

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AlphaLab-USTC/OhMyCode at commit 4d1bb28, republished under its MIT licence (© AlphaLab-USTC). 326 words, ~941 tokens.

Download SKILL.mdSave it as .claude/skills/bench/SKILL.md (or your agent's skills folder).
name
bench
description
Run OhMyCode benchmarks — score any provider/model with token tracking. Use when user wants to benchmark, evaluate, test performance, or compare models.

OhMyCode Benchmark

One-command benchmarking: run 8 SWE-bench-style coding tasks through OhMyCode, track token usage (in/out), and produce a scorecard.

Works with any provider and model — uses whatever is configured in ~/.ohmycode/config.json or overridden via CLI args.

When to Use

  • User says "run benchmark", "bench", "score", "evaluate", "test performance"
  • User wants to compare models or providers
  • After major code changes, to verify agent capabilities still work

Input

$ARGUMENTS — optional filters and overrides.

ArgumentExampleEffect
(empty)/benchFull suite, current config
task filter/bench fib,bugOnly matching tasks
--dry-run/bench --dry-runValidate task definitions without LLM

Step 1 — Run the Benchmark

bash
python3 benchmarks/run_bench.py $ARGUMENTS 2>&1 | tee bench_run.log
Override provider/model for comparison
bash
# Test with a different model
python3 benchmarks/run_bench.py --provider openai --model gpt-4o-mini 2>&1 | tee bench_run.log

# Test with Anthropic
python3 benchmarks/run_bench.py --provider anthropic --model claude-sonnet-4-20250514 2>&1 | tee bench_run.log

# Test with custom endpoint
python3 benchmarks/run_bench.py --base-url http://localhost:8080/v1 --api-key test 2>&1 | tee bench_run.log

Step 2 — Read Results

The harness outputs:

  1. Terminal report — table with per-task pass/fail, tokens, time
  2. bench_results.json — machine-readable results for comparison

Key metrics to report:

  • Score: X/8 tasks passed
  • Tokens in: total prompt tokens across all tasks
  • Tokens out: total completion tokens across all tasks
  • Total tokens: in + out
  • Time: wall-clock seconds

Step 3 — Analyze Failures

If any tasks failed:

  1. Read the reason column in the report
  2. Check bench_results.json for the error field
  3. Classify: is it a model capability issue, or an OhMyCode bug?
  4. For OhMyCode bugs → fix and re-run /bench (closed-loop)

Benchmark Tasks

#TaskCategoryWhat It Tests
1fibonaccicode-genCreate a function from spec
2bug-fix-roundbug-fixFind and fix an off-by-one error
3test-generationtest-genWrite tests for existing code
4refactor-preserverefactorImprove code without breaking tests
5grep-replacetool-useMulti-file search and replace
6stack-modulecode-genCreate module + tests from scratch
7type-error-fixbug-fixFix a TypeError in existing code
8code-comprehensioncomprehensionRead code and explain the algorithm

Adding New Tasks

Edit benchmarks/suite.py:

python
BenchTask(
    name="your-task",
    category="bug-fix",
    prompt="The task description for the agent...",
    setup=lambda d: (d / "code.py").write_text("..."),  # prepare files
    validate=lambda d: (True, "reason"),                 # check result
    max_turns=10,
)

Then append to BENCH_SUITE list.

Model Comparison Workflow

To compare two models:

bash
python3 benchmarks/run_bench.py --model gpt-4o 2>&1 | tee bench_gpt4o.log
cp bench_results.json bench_gpt4o.json

python3 benchmarks/run_bench.py --model gpt-4o-mini 2>&1 | tee bench_mini.log
cp bench_results.json bench_mini.json

Then compare the JSON files for score and token efficiency.

  • /run-tests — run unit tests (the benchmark runs these too as Phase 1)
  • /gen-tests — generate tests (task #3 in the benchmark tests this capability)

© AlphaLab-USTC, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/bench of AlphaLab-USTC/OhMyCode.

Open the folder on GitHubat commit 4d1bb28

Compare with similar skills

Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bench this skillAlphaLab-USTC/OhMyCode131—~941Automated safety check: PassMIT
Phx Examplesoliver-kriska/claude-elixir-phoenix565—~1kAutomated safety check: PassMIT
Phx Quickoliver-kriska/claude-elixir-phoenix565—~805Automated safety check: PassMIT
Phx Quickoliver-kriska/claude-elixir-phoenix565—~811Automated safety check: PassMIT
Langfuse Integration Pagelangfuse/langfuse-docs246—~3.7kAutomated safety check: PassMIT
Langfuselangfuse/skills299—~2.1kAutomated safety check: NotesMIT

Similar skills

  • Phx Examples

    oliver-kriska/claude-elixir-phoenix

    Provide Phoenix, LiveView, Ecto, OTP, or Oban examples. An agent skill from oliver-kriska/claude-elixir-phoenix.

    565 GitHub stars~1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Phx Quick

    oliver-kriska/claude-elixir-phoenix

    Implement small Phoenix changes without planning — add validations; Use for single-file edits under 50 lines.

    565 GitHub stars~805 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Phx Quick

    oliver-kriska/claude-elixir-phoenix

    Implement small Phoenix changes without planning — add validations, update routes, fix components, create migrations.

    565 GitHub stars~811 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Langfuse Integration Page

    langfuse/langfuse-docs

    Create a new Langfuse integration page in the langfuse-docs repo.

    246 GitHub stars~3.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Langfuse

    langfuse/skills

    Interact with Langfuse and access its documentation: tracing, monitoring, creating datasets, running experiments, and evaluating AI applications.

    299 GitHub stars~2.1k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check: notes
  • Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

    1.6k GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from AlphaLab-USTC/OhMyCode

All 10 skills in this repo
  • Add Feature

    AlphaLab-USTC/OhMyCode

    Guide for adding a new feature to OhMyCode. An agent skill from AlphaLab-USTC/OhMyCode.

    131 GitHub stars~995 tokensUpdated 6 mo ago
    Auto-check passed
  • Add Provider

    AlphaLab-USTC/OhMyCode

    Guide for adding a new LLM provider to OhMyCode. An agent skill from AlphaLab-USTC/OhMyCode.

    131 GitHub stars~888 tokensUpdated 6 mo ago
    Auto-check passed
  • Add Tool

    AlphaLab-USTC/OhMyCode

    Guide for adding a new tool to OhMyCode. An agent skill from AlphaLab-USTC/OhMyCode.

    131 GitHub stars~758 tokensUpdated 6 mo ago
    Auto-check passed
  • Customize Response Style

    AlphaLab-USTC/OhMyCode

    Guide for customizing OhMyCode's terminal output style and rendering.

    131 GitHub stars~1.1k tokensUpdated 6 mo ago
    Auto-check passed
  • Customize System Prompt

    AlphaLab-USTC/OhMyCode

    Guide for customizing OhMyCode's AI personality and behavior via system prompt.

    131 GitHub stars~819 tokensUpdated 6 mo ago
    Auto-check passed
  • Debug Ohmycode

    AlphaLab-USTC/OhMyCode

    Guide for debugging OhMyCode issues. An agent skill from AlphaLab-USTC/OhMyCode.

    131 GitHub stars~1.3k tokensUpdated 6 mo ago
    Auto-check passed

Questions about Bench

What does Bench do?

Run OhMyCode benchmarks — score any provider/model with token tracking. Bench is an agent skill from AlphaLab-USTC/OhMyCode. Run OhMyCode benchmarks — score any provider/model with token tracking.

When should I use Bench?

Bench fits situations like: user wants to benchmark; test performance.

How do I install Bench in Claude Code?

Run `npx skills add AlphaLab-USTC/OhMyCode --skill bench -a claude-code`. Or copy the skill folder (.claude/skills/bench in AlphaLab-USTC/OhMyCode) into .claude/skills/bench in your project. Claude Code loads it when a task matches its description.

How do I install Bench in Codex?

Run `npx skills add AlphaLab-USTC/OhMyCode --skill bench -a codex`. Or copy the skill folder (.claude/skills/bench in AlphaLab-USTC/OhMyCode) into .agents/skills/bench in your project. Codex loads it when a task matches its description.

Can I use Bench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AlphaLab-USTC/OhMyCode --skill bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bench, .gemini/skills/bench, .github/skills/bench and .opencode/skills/bench in your project.

What does Bench need to run?

Going by SKILL.md and its folder, Bench needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Bench access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bench safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bench use?

Bench is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bench use?

About 941 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bench?

Skills that share tags, products or a category with Bench: Phx Examples (oliver-kriska/claude-elixir-phoenix, 565 stars), Phx Quick (oliver-kriska/claude-elixir-phoenix, 565 stars), Phx Quick (oliver-kriska/claude-elixir-phoenix, 565 stars) and Langfuse Integration Page (langfuse/langfuse-docs, 246 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bench?

AlphaLab-USTC (a GitHub user) maintains it in AlphaLab-USTC/OhMyCode, which has 131 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on April 2, 2026.

Source: AlphaLab-USTC/OhMyCode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.