Agent skill

Benchmaker

by DanMcInerney in DanMcInerney/orchflows

Build benchmarks of realistic, hard agent tasks, researched from real work, admitting only verified, shortcut-resistant, calibrated tasks and reporting how well the suite separates systems.

MITAuto-check passedGame Development

Install Benchmaker

skills CLI
$ npx skills add DanMcInerney/orchflows --skill benchmaker -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install DanMcInerney/orchflows benchmaker --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/DanMcInerney/orchflows.git skills-src && mkdir -p .claude/skills && cp -r skills-src/example-workflows/benchmaker/skills/benchmaker .claude/skills/benchmaker && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmaker
GitHub stars
117
Token cost
~2k tokens
SKILL.md length
1,116 words
Files
80 (incl. scripts)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Build benchmarks of realistic, hard agent tasks, researched from real work, admitting only verified, shortcut-resistant, calibrated tasks and reporting how well the suite separates systems.

  • Works in 7 steps: Claim. Inspect the target,… → Research the work. Fan out web research… → Source. Apply common and Make… → …
  • Game Development work in your project
  • Runs Python scripts from its folder

What it does

Benchmaker is an agent skill from DanMcInerney/orchflows. Build benchmarks of realistic, hard agent tasks, researched from real work, admitting only verified, shortcut-resistant, calibrated tasks and reporting how well the suite separates systems.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 82 other files, including scripts (for example `agents/openai.yaml`, `scripts/benchkit/INTERFACE.md` and `scripts/benchkit/__init__.py`).

It sits in Game Development. The repository describes itself as: 2 skills, composable into workflows, which can then be built into more complex workflows. The licence is MIT.

When your agent uses it

  • Game Development work in your project

Example prompts

  • “/benchmaker”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Claim. Inspect the target, representative user work, research lessons and relevant primary precedents. State the solver's goal: the work…
  2. Research the work. Fan out web research guided by the solver's goal: where people do this work, what real instances look like, how it goes…
  3. Source. Apply common and Make benchmarking guidance. Draw more candidate work than the suite needs from the catalog and the target's own…
  4. Build and admit. Authors build each candidate under the benchmark contract: public instruction and interface, environment, reference…
  5. Measure. Freeze the admitted suite. Run the comparison set and the known-order check with predeclared repeats, measuring the actual target…
  6. Review and repair. Apply shared:review-revise-once with benchmarking and task guidance, using a fresh reviewer who neither authored…
  7. Deliver. Return the suite, commands, card, research catalog, rejection log, per-family results with denominators, measured time and cost…

What it can do on your machine

Read from SKILL.md and the folder at commit 8d16eb3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 16 files in scripts/ (Python, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmaker loads about 2k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 1,116 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from DanMcInerney/orchflows at commit 8d16eb3, republished under its MIT licence (© DanMcInerney). 1,116 words, ~2,015 tokens.

Download SKILL.mdSave it as .claude/skills/benchmaker/SKILL.md (or your agent's skills folder). This skill also uses 79 other files; get the full folder from GitHub.
name
benchmaker
description
Build benchmarks of realistic, hard agent tasks, researched from real work, admitting only verified, shortcut-resistant, calibrated tasks and reporting how well the suite separates systems.
disable-model-invocation
true

Apply library context. Accept a target: path, callable, command, endpoint, native workflow, skill or capability description. Also accept an optional claim, comparisons, budget and output location. Infer ordinary choices; clarify only material claim, permission or budget decisions. Without an executable target, a named representative stands in for it as the quality card defines; a description with no executable representative supports only a draft.

A benchmark samples realistic, hard, valuable work and separates systems that differ in the claimed ability. Most candidate tasks should fail admission. A smaller budget admits fewer tasks and reaches a lower stage; it never weakens tasks or skips admission checks. Spend follows earned evidence: a task version that fails a criterion gets no further checks, so costly judgments go to tasks that cheaper evidence has not disqualified.

The coordinator owns every assignment and every target, comparison, known-order check, calibration and adversary launch. Other agents return results and execution requests without delegating. Composing targets run in separate top-level sessions under core docs/hosts.md#workflow-trials. Target, comparison, known-order check and calibration attempts run through the package's runner, which enforces their caps; deadlines given to native helpers are requests, so record overruns.

Before launching, declare and show planned research workers, candidates, per-task revision limit, repeats, calibration and known-order check launches for admission and measurement, concurrency, authoring and attempt time, deadlines and estimated spend. When caller limits cannot fit building or solving work at the difficulty target, say so before launching. A wall-clock limit covers the whole build: plan when measurement, review and delivery must start for each to finish inside it. When earlier work runs long or stops converging, reduce scope to fewer tasks or families, or a lower stage, so later steps still run. Caller limits on launches or sessions count every launch they name. When a limit names only target runs, report authoring, auditing, adversary, calibration and known-order check launches separately. Record measurement conditions separately from caller limits; failed attempts consume execution limits. Record each launch and each task's admission state durably before dependent work, so an interrupted build resumes without relaunching completed work and can still deliver a draft.

  1. Claim. Inspect the target, representative user work, research lessons and relevant primary precedents. State the solver's goal: the work it does, for whom, and what success looks like. Record the claim entries of the quality card, including the systems the benchmark must measure; ask the user when neither the request nor the target settles them. Choose the comparison set, known-order check, calibration systems and, without an executable target, its representative under benchmarking guidance and the card.

  2. Research the work. Fan out web research guided by the solver's goal: where people do this work, what real instances look like, how it goes wrong, what makes real cases hard, and the most complex real work people attempt, including work beyond today's strongest systems. Give each fresh research worker a distinct scope with research guidance; workers return sources and findings, not tasks. Gather them into a catalog of realistic scenarios and hard cases, recording sources, capture dates and reuse constraints. Include the real material that could become tasks, such as repositories, issues with their fixes, datasets, incident reports and practitioner accounts. Report uncovered ground as a gap. Without web access, research what the caller supplies and record the gap.

  3. Source. Apply common and Make benchmarking guidance. Draw more candidate work than the suite needs from the catalog and the target's own failures: real work first, then reconstructions from real material, and synthetic work only as a labeled fallback. Each candidate states its value, why it is hard for the claimed ability and an estimated expert time. Spread candidates across families, independent source groups and the catalog's range of difficulty, including its hardest real work, before expanding any one.

  4. Build and admit. Authors build each candidate under the benchmark contract: public instruction and interface, environment, reference solution and verifier. Freeze each built task, then gather its admission evidence:

    • run the reference and trivial attempts through the verifier;
    • a fresh reviewer who did not build the task, holding the research catalog, judges whether each reconstructed or synthetic task could occur in practice;
    • a fresh auditor with no authoring history or evaluator material solves from public inputs. The coordinator keeps a copy of each saved outcome; an auditor covering several tasks saves all of them before any disclosure. It then receives the evaluator material and reports unstated requirements, rejected valid outcomes and accepted wrong ones; where the host cannot continue it, a fresh reviewer given the saved outcome does this, recorded as a condition;
    • run labeled outcomes through the verifier: an independent valid alternative, which may be the auditor's when valid, and plausible wrong outcomes, including broken copies of the reference;
    • a fresh adversary with the solver's material and access plus the exploit catalogue in research, and never evaluator material, tries to earn credit without doing the work, and tries again after each repair its findings prompt;
    • the calibration systems and the known-order check attempt the task until more attempts could not change its placement. Measurement waits until every admitted task has this evidence.

    Auditors and adversaries may cover several tasks they did not author. When the suite holds many short items, audit, adversary and realism judgments may cover each family through a recorded sample. Revise or reject by the admission criteria; a revision reruns the checks it affects and counts toward the per-task limit.

  5. Measure. Freeze the admitted suite. Run the comparison set and the known-order check with predeclared repeats, measuring the actual target or its representative. Compute the card. Classify failures from transcripts as ability, task defect, grading, infrastructure, refusal, cut-off or unknown, and audit passing transcripts, or a recorded sample of them, for unearned credit.

  6. Review and repair. Apply shared:review-revise-once with benchmarking and task guidance, using a fresh reviewer who neither authored, audited nor reviewed the realism of tasks. Its inputs are the frozen suite, card, rejection log, admission evidence and a transcript sample that includes successes. Scope repairs and checks to the package. Review repairs count toward each task's revision limit. Changed tasks receive new identities and rerun affected admission and measurement within remaining allowance. A repaired task that fails re-admission leaves the suite, which then receives a new identity, or stays draft.

  7. Deliver. Return the suite, commands, card, research catalog, rejection log, per-family results with denominators, measured time and cost, requested and achieved stage, and gaps. State whether measured headroom and separation support the claim. Missing required execution, judgment or review, or an unresolved validity defect, leaves affected claims draft. An interrupted or limit-stopped build delivers these items as a draft from its records. Name the next unmet stage.

© DanMcInerney, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 79 other files (scripts) in example-workflows/benchmaker/skills/benchmaker of DanMcInerney/orchflows.

  • SKILL.md
  • agents/openai.yaml
  • scripts/benchkit/INTERFACE.md
  • scripts/benchkit/__init__.py
  • scripts/benchkit/adapters.py
  • scripts/benchkit/aggregate.py
  • scripts/benchkit/cli.py
  • scripts/benchkit/identity.py
  • scripts/benchkit/launch.py
  • scripts/benchkit/records.py
  • scripts/benchkit/runner.py
  • scripts/benchkit/schedule.py
  • scripts/benchkit/selfcheck.py
  • scripts/benchkit/shapes.py
  • scripts/benchkit/stage.py
  • scripts/benchkit/stats.py
  • scripts/benchkit/stopping.py
  • scripts/benchkit/suite.py
  • … and 62 more

Open the folder on GitHubat commit 8d16eb3

Compare with similar skills

Benchmaker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmaker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmaker this skillDanMcInerney/orchflows117—~2kAutomated safety check: PassMIT
Image to Three.js Modelimg2threejs/img2threejs18k1 repos~8.2kAutomated safety check: PassApache-2.0
Web CloneJane-xiaoer/claude-skill-web-clone1k1 repos~2.7kAutomated safety check: PassMIT
Threejs Game Directormajidmanzarpour/threejs-game-skills2.4k—~2.2kAutomated safety check: PassMIT
Game Asset Generatorhtdt/godogen7.1k—~2.8kAutomated safety check: PassMIT
Threejs Gameplay Systemsvalkor-ai/loom1.2k1 repos~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Image to Three.js Model

    img2threejs/img2threejs

    Rebuilds the object in a reference image as a procedural, animation-ready Three.js model written entirely in code, using staged sculpting with quality checks.

    18k GitHub starsUsed in 1 repo~8.2k tokens
    Game DevelopmentAuto-check passed
  • Web Clone

    Jane-xiaoer/claude-skill-web-clone

    网站复刻 / 克隆方法论。USE WHEN 用户说 复刻网站、克隆网站、clone website、抄个站、仿站、 照着这个站做一个、reproduce site、还原某个网页效果、把这个站搬下来改成我的、 复刻某个交互/WebGL/Canvas/Three.js 效果。提供「先拿真源码 → 判路径 → 逆向拆解 → 搭工程 → 替换内容」的可移植决策树,覆盖静态站 /…

    1k GitHub starsUsed in 1 repo~2.7k tokens
    Game DevelopmentAuto-check passed
  • Threejs Game Director

    majidmanzarpour/threejs-game-skills

    Entrypoint for building, upgrading, and finishing Three.js browser games.

    2.4k GitHub stars~2.2k tokensUpdated 9 days ago
    Game DevelopmentAuto-check passed
  • Generates game art from text prompts: PNG images, GLB 3D models, rigged characters, animations and sprites, with background removal.

    7.1k GitHub stars~2.8k tokensUpdated 6 days ago
    Game DevelopmentAuto-check passed
  • Build and iterate playable Three.js game systems: starter scaffold, architecture, design briefs, core loops, level and encounter design, entities, input, camera, collision and physics, scoring…

    1.2k GitHub starsUsed in 1 repo~1.4k tokens
    Game DevelopmentAuto-check passed
  • Uloop Execute Dynamic Code

    CyberAgentGameEntertainment/NovaShader

    Execute C with Unity APIs when existing uloop tools cannot inspect or edit enough.

    1.6k GitHub starsUsed in 1 repo~1.9k tokens
    Game DevelopmentAuto-check passed

More from DanMcInerney/orchflows

All 12 skills in this repo
  • Evolve

    DanMcInerney/orchflows

    Iteratively improve any artifact and its improvement harness, inventing evaluation when needed; supports bounded runs, tournaments and continuous resumable search.

    117 GitHub stars~805 tokensUpdated 3 days ago
    Auto-check passed
  • 3D Browser Game

    DanMcInerney/orchflows

    Build a complete Three.js game or bounded production phase through mechanics experiments, Blender assets, QA and independent playtests.

    117 GitHub stars~1.5k tokensUpdated 3 days ago
    Auto-check passed
  • Design Loop

    DanMcInerney/orchflows

    Develop a project endgoal through N bounded cycles of brainstorm and research, design, implementation, comparison testing and analysis.

    117 GitHub stars~870 tokensUpdated 3 days ago
    Auto-check passed
  • Export Workflow

    DanMcInerney/orchflows

    Export an Orchflows workflow as a standalone native skill with a portability report and bounded trial.

    117 GitHub stars~674 tokensUpdated 3 days ago
    Auto-check passed
  • Gauntlet Loop

    DanMcInerney/orchflows

    Drive an ambitious goal past a real-world quality bar by splitting it into pieces, each looped through a builder and a fresh, harsh, blind critic until ours reaches the bar or the caller stops.

    117 GitHub stars~1.2k tokensUpdated 3 days ago
    Auto-check passed
  • Orchflows Review

    DanMcInerney/orchflows

    Review Orchflows against its purpose and design principles (architecture, workflow design, wording and bugs) and apply reviewed improvements on unmerged branches.

    117 GitHub stars~1k tokensUpdated 3 days ago
    Auto-check passed

Questions about Benchmaker

What does Benchmaker do?

Build benchmarks of realistic, hard agent tasks, researched from real work, admitting only verified, shortcut-resistant, calibrated tasks and reporting how well the suite separates systems. Benchmaker is an agent skill from DanMcInerney/orchflows. Build benchmarks of realistic, hard agent tasks, researched from real work, admitting only verified, shortcut-resistant, calibrated tasks and reporting how well the suite separates systems.

When should I use Benchmaker?

Benchmaker fits situations like: game Development work in your project.

How do I install Benchmaker in Claude Code?

Run `npx skills add DanMcInerney/orchflows --skill benchmaker -a claude-code`. Or copy the skill folder (example-workflows/benchmaker/skills/benchmaker in DanMcInerney/orchflows) into .claude/skills/benchmaker in your project. Claude Code loads it when a task matches its description.

How do I install Benchmaker in Codex?

Run `npx skills add DanMcInerney/orchflows --skill benchmaker -a codex`. Or copy the skill folder (example-workflows/benchmaker/skills/benchmaker in DanMcInerney/orchflows) into .agents/skills/benchmaker in your project. Codex loads it when a task matches its description.

Can I use Benchmaker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add DanMcInerney/orchflows --skill benchmaker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmaker, .gemini/skills/benchmaker, .github/skills/benchmaker and .opencode/skills/benchmaker in your project.

What does Benchmaker need to run?

Going by SKILL.md and its folder, Benchmaker needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Benchmaker access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Benchmaker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Benchmaker use?

Benchmaker is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmaker use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmaker?

Skills that share tags, products or a category with Benchmaker: Image to Three.js Model (img2threejs/img2threejs, 18k stars), Web Clone (Jane-xiaoer/claude-skill-web-clone, 1k stars), Threejs Game Director (majidmanzarpour/threejs-game-skills, 2.4k stars) and Game Asset Generator (htdt/godogen, 7.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmaker?

DanMcInerney (a GitHub user) maintains it in DanMcInerney/orchflows, which has 117 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 4, 2026.

Source: DanMcInerney/orchflows on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.