Agent skill

Benchmark Runner

by RConsortium in RConsortium/pharma-skills

Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs.

MITAuto-check passedAI & LLM Engineering

Install Benchmark Runner

skills CLI
$ npx skills add RConsortium/pharma-skills --skill benchmark-runner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RConsortium/pharma-skills benchmark-runner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RConsortium/pharma-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/_automation/benchmark-runner .claude/skills/benchmark-runner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-runner
GitHub stars
118
Token cost
~5.3k tokens
SKILL.md length
912 words
Files
13 (incl. scripts)
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs.

  • Works in 2 steps: Agent A Run (With Skill) → Agent B Run + Scoring
  • Someone says run benchmarks
  • SKILL.md covers Model Selection — Series Label…, Routine Setup (one-time), GitHub Access — Use Whichever… and Phase Detection — Run Before…, plus 7 more sections
  • Runs Python and Shell scripts from its folder; calls gh, claude and python3; reaches docs.github.com; needs GH_TOKEN and GITHUB_TOKEN

What it does

Benchmark Runner is an agent skill from RConsortium/pharma-skills. Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs. without skill using matched isolated sessions, and post scored results to the linked GitHub issue. Use whenever someone says "run benchmarks", "compare skill performance", "eval the skills", or wants to measure whether a skill improves output quality.

Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts (for example `CLAUDE_CODE_ROUTINE.md`, `README.md` and `runs/README.md`).

It sits in AI & LLM Engineering, covering LLM evaluation. It works with GitHub. The repository describes itself as: A collection of agent skills for BioPharma use cases GSDBench Intake https://rconsortium.github.io/pharma-skills/gsdbench-intake/. The licence is MIT.

When your agent uses it

  • Someone says run benchmarks
  • Compare skill performance
  • Eval the skills
  • Wants to measure whether a skill improves output quality

Example prompts

  • “run benchmarks”
  • “compare skill performance”
  • “eval the skills”
  • “/benchmark-runner”

Requirements

  • Python 3
  • A Bash shell
  • A credential in GITHUB_TOKEN

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Agent A Run (With Skill)
  2. Agent B Run + Scoring

What it can do on your machine

Read from SKILL.md and the folder at commit ae5d83b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 5 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • gh
    • claude
    • python3
    • bash
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • docs.github.com

    Also links to:

    • claude.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GH_TOKEN
    • GITHUB_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Runner loads about 5.3k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 912 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from RConsortium/pharma-skills at commit ae5d83b, republished under its MIT licence (© RConsortium). 912 words, ~5,298 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-runner/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
benchmark-runner
description
Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs. without skill using matched isolated sessions, and post scored results to the linked GitHub issue. Use whenever someone says "run benchmarks", "compare skill performance", "eval the skills", or wants to measure whether a skill improves output quality.

Skill Benchmark Runner

Benchmark every evaluation case in the _automation/evals/ directory of the RConsortium/pharma-skills repository. Each routine invocation is one of two short phases (~20 min each). The routine inspects GitHub issue comments on startup to decide which phase to execute — no configuration needed, no commits required, no repo write access required from the human user.

Repository: RConsortium/pharma-skills (https://github.com/RConsortium/pharma-skills)


Model Selection — Series Label and Sub-Agent Model

{CURRENT_MODEL_NAME} throughout this skill is the benchmark series label, fixed to the literal string Claude Routine. Use it verbatim everywhere it appears: get_next_eval.py --model, record_run_result.py --model, the marker JSON ("model"), and the report metadata. The scripts normalise this label (lowercase, strip spaces/punctuation) for deduplication, so all runs group into one consistent series.

Sub-agent model:

  • Agent A and Agent B are launched without an explicit --model flag (claude -p ...), so each sub-agent inherits this session's default model — the same model orchestrating the routine. This keeps the measurement clean (bare model ± skill, identical model on both sides) without ever writing a concrete model ID into a repository artifact. Never hardcode a concrete model ID anywhere in this skill or its outputs.
  • The prompt files constructed in Step 2 / Step 6 and the CLAUDE_CODE_MAX_OUTPUT_TOKENS setting are the only information passed from the parent session to a sub-agent. Do not forward any other parent context: no conversation history, no additional environment variables, no eval assertions, no scoring prompt, no blinded map. This prevents the orchestrator's context from leaking into either candidate.

Routine Setup (one-time)

Create a single routine at claude.ai/code/routines:

FieldValue
PromptRead _automation/benchmark-runner/SKILL.md and execute.
RepositoryRConsortium/pharma-skills
Schedule0 1,6 * * * (1 AM and 6 AM UTC — 5 h gap, matches rolling usage window)

That is all. The skill determines its own phase on every invocation.


GitHub Access — Use Whichever Method Works

Throughout this skill you will read issue comments, post issue comments, and create release assets. Use whatever method is available in your environment — pick the one that works without prompting:

MethodBest whenNotes
mcp__github__* MCP toolsRunning inside Claude Code with the GitHub MCP serverNo token required; preferred when available
gh CLI (gh issue view, gh release upload, etc.)Running locally with gh authenticatedConcise, supports all operations
REST API via curlAnywhere with GH_TOKEN / GITHUB_TOKEN setUniversal fallback; use for release-asset upload (no MCP equivalent)
Provider-specific GitHub tools (Codex, Gemini, etc.)Running under another agent CLIUse whatever the host provides

Reason about which method to use; do not enforce a rigid order. If one fails, try another. Always confirm the operation succeeded (e.g., the comment URL came back, the asset was uploaded) before continuing.

For release-asset upload there is currently no MCP tool — use gh release upload or curl POST to the upload URL.


Phase Detection — Run Before Any Other Step

Scan all benchmark eval issues to find any that are waiting for Phase 2 (Agent B + scoring) for the model you are running:

  1. List all eval files: ls _automation/evals/*.json — extract each id field (e.g. github-issue-27 → issue #27).

  2. For each issue number, fetch comments using whichever GitHub access method is available (see above). Scan each comment body for a <!-- BENCHMARK_PARTIAL: marker.

  3. Filter and evaluate each BENCHMARK_PARTIAL marker found:

    • Skip if state.model does not match {CURRENT_MODEL_NAME}. This is critical: a partial run by another user on a different model belongs to that user. Only pick up partials matching your current model.
    • Skip if a later comment on the same issue contains a matching <!-- BENCHMARK_COMPLETE: {"eval_id":"{same}","model":"{same}" — Phase 2 already finished for that combination.
    • Otherwise → Phase 2 candidate. Extract the JSON from the marker (see format below) and note the partial comment id.
  4. Decision:

    • One or more Phase 2 candidates found → pick the oldest (earliest created_at) → enter Phase 2 with that state.
    • No candidates for your model → enter Phase 1.

BENCHMARK_PARTIAL marker format (hidden HTML comment embedded in the issue comment body):

<!-- BENCHMARK_PARTIAL: {"eval_id":"github-issue-27","model":"Claude Routine","skill_sha":"b5ede6a...","issue_number":27,"blinded_map":{"candidate_1":"output_B","candidate_2":"output_A"},"agent_a_asset_url":"https://github.com/RConsortium/pharma-skills/releases/download/benchmark-results/benchmark_agent_a_github-issue-27.zip","run_date":"2026-05-03T06:00Z","tokens_a":199382,"partial_comment_id":4367060533} -->

Phase 1 — Agent A Run (With Skill)

Runs when no Phase 2 candidate is found. Executes Agent A, archives its output, and posts a partial comment that holds state for Phase 2.

Step 0 — R Environment Pre-flight

Always run first. Idempotent — safe to re-run.

bash
bash _automation/benchmark-runner/scripts/setup_r_env.sh

Exits non-zero on failure — stop and report the error. Do not proceed.

R packages installed: jsonlite, digest, gsDesign, gsDesign2, lrstat, graphicalMCP, eventPred, ggplot2

Step 1 — Discover Next Eval
bash
python3 _automation/benchmark-runner/scripts/get_next_eval.py --model {CURRENT_MODEL_NAME}
  • STATUS: UP_TO_DATE → all evals complete for this model+SHA. Exit.
  • JSON output → parse to a temp file; extract _skill_name, _skill_sha, _skill_content, _bundled_resources, _prompt_a, _blinded_scoring_map, and the issue number from id.

Optional flags:

bash
--runner-id {YOUR_NAME}           # stable per-person ordering
--priority-issue github-issue-{N} # force a specific eval
Step 2 — Run Agent A (With Skill)

Create the working directory:

bash
mkdir -p /tmp/benchmark_{id}/agent_A/output_A

Stage bundled resource files to disk (progressive disclosure — files read on demand, not embedded in the prompt):

python
import os, json
agent_a_dir = "/tmp/benchmark_{id}/agent_A"
for rel_path, content in eval_case["_bundled_resources"].items():
    if rel_path == "SKILL.md":
        continue
    dest = os.path.join(agent_a_dir, rel_path)
    os.makedirs(os.path.dirname(dest), exist_ok=True)
    with open(dest, "w", encoding="utf-8") as f:
        f.write(content)

Write prompt_A.txt — _skill_content (SKILL.md) followed by _prompt_a only. No bundled resource content in the prompt:

python
prompt_a = eval_case["_skill_content"] + "\n\n" + eval_case["_prompt_a"]
with open(os.path.join(agent_a_dir, "prompt_A.txt"), "w", encoding="utf-8") as f:
    f.write(prompt_a)

Launch Agent A:

bash
export CLAUDE_CODE_MAX_OUTPUT_TOKENS=64000
cd /tmp/benchmark_{id}/agent_A && \
  cat prompt_A.txt | claude -p \
  --allowedTools "Bash,Read,Write,Edit,Glob" \
  --output-format json > agent_A_run.json 2>&1

Note: export is required — a prefix (VAR=val cat ... | claude) only sets the variable for cat, not for the claude process receiving the pipe.

--output-format json emits a single JSON object when the agent finishes — resilient to long-running agents and session timeouts.

When Agent A returns, extract token count:

python
import json
d = json.load(open("/tmp/benchmark_{id}/agent_A/agent_A_run.json"))
u = d.get("usage", {})
tokens_a = u.get("input_tokens", 0) + u.get("cache_creation_input_tokens", 0) + u.get("output_tokens", 0)
is_error_a = d.get("is_error", False)

Record in runs.json:

bash
python3 _automation/benchmark-runner/scripts/record_run_result.py \
  --eval-id {id} --model {CURRENT_MODEL_NAME} \
  --status partial_a --tokens-a {tokens_a}
Step 3 — Archive Agent A Output

Create the zip:

bash
cd /tmp/benchmark_{id} && zip -r benchmark_agent_a_{eval_id}.zip \
  agent_A/output_A/ agent_A/agent_A_run.json

Upload to the benchmark-results GitHub release as a named asset. The release must already exist (create it once if needed). Use whichever method works in your environment — examples below; pick what works:

  • gh CLI (simplest if available):
    bash
    gh release view benchmark-results --repo RConsortium/pharma-skills \
      || gh release create benchmark-results --repo RConsortium/pharma-skills \
         --prerelease --title "Automated Benchmark Results" --notes "Rolling release."
    gh release upload benchmark-results /tmp/benchmark_{id}/benchmark_agent_a_{eval_id}.zip \
      --repo RConsortium/pharma-skills --clobber
  • REST API via curl (when only GH_TOKEN is available):
    bash
    # Get-or-create release, then POST to its upload_url with the zip as data-binary.
    # See https://docs.github.com/en/rest/releases for the exact endpoints.
  • MCP does not currently expose a release-asset upload tool — use gh or curl for the upload step. Comment posting and reading can still use MCP.

Construct the asset download URL (used in the partial comment state):

https://github.com/RConsortium/pharma-skills/releases/download/benchmark-results/benchmark_agent_a_{eval_id}.zip

If no upload method works (no gh, no token), skip the upload and set agent_a_asset_url: null in the partial state. Phase 2 will detect the null URL and re-run Agent A for that eval — wasteful but correct.

Show full SKILL.md (643 more words)Show less
Step 4 — Post Partial Comment

Write the partial comment body to /tmp/partial_comment_{eval_id}.md:

markdown
## Automated Benchmark Results — `{_skill_name}` 🟡 In Progress

### Run Metadata

| Field | Value |
|---|---|
| **Eval ID** | `{id}` |
| **Run date** | {YYYY-MM-DD HH:MM UTC} |
| **Model** | `{CURRENT_MODEL_NAME}` |
| **Skill version** | `{_skill_sha[:7]}` |
| **Phase** | 1 of 2 complete — Agent A (with skill) finished |

Agent A has completed. Agent B (without skill) will run in the next scheduled window (~5 h).
Results will be updated here automatically.

<!-- BENCHMARK_PARTIAL: {"eval_id":"{id}","model":"{CURRENT_MODEL_NAME}","skill_sha":"{_skill_sha}","issue_number":{N},"blinded_map":{_blinded_scoring_map},"agent_a_asset_url":"{asset_url}","run_date":"{ISO8601}","tokens_a":{tokens_a}} -->

Post it using whichever GitHub access method is available (see "GitHub Access" above). The partial comment id returned by the API is not needed for Phase 2 (Phase 2 discovers it by scanning), but log it for debugging.

Phase 1 is complete. Print this summary to the user before exiting:

✓ Phase 1 complete — Agent A finished for {eval_id} ({model})
  • Output archived: {asset_url}
  • Partial comment: {comment_url}
  • Tokens used: {tokens_a:,}

NEXT STEP — Phase 2 (Agent B + scoring):
  • If running as a scheduled routine: nothing to do. The next scheduled
    invocation (≥5 h from now, after the rolling usage window resets) will
    detect this partial state automatically and run Phase 2.
  • If running manually: re-invoke this skill any time. It will
    detect the BENCHMARK_PARTIAL marker on issue #{N} and run Phase 2 to
    completion.

Then exit cleanly.


Phase 2 — Agent B Run + Scoring

Runs when a BENCHMARK_PARTIAL state is found in a GitHub issue comment. Loads Agent A's output, runs Agent B, scores both, posts the full result.

Step 5 — Load Partial State

Parse the BENCHMARK_PARTIAL JSON from the comment body found during Phase Detection:

python
import re, json
marker_re = re.compile(r'<!-- BENCHMARK_PARTIAL: ({.*?}) -->', re.DOTALL)
m = marker_re.search(comment_body)
state = json.loads(m.group(1))
# state keys: eval_id, model, skill_sha, issue_number, blinded_map,
#             agent_a_asset_url, run_date, tokens_a

Also reload the full eval case (for assertions, scoring prompt, prompt_b):

bash
# Send stdout (the eval-case JSON) to the file and stderr (warnings such as
# the >100 KB bundle notice) to a separate log. Do NOT use `2>&1` here — it
# would merge warning lines into the JSON file and break the json.load below.
python3 _automation/benchmark-runner/scripts/get_next_eval.py \
  --model {state["model"]} \
  --priority-issue {state["eval_id"]} \
  > /tmp/eval_case_{id}.json 2>/tmp/eval_case_{id}.log

Restore Agent A's output. If agent_a_asset_url is set, download and unzip it. Use whichever method works:

bash
mkdir -p /tmp/benchmark_{id}/agent_A/output_A

# Option A — gh CLI:
gh release download benchmark-results --repo RConsortium/pharma-skills \
  --pattern "benchmark_agent_a_{eval_id}.zip" --dir /tmp/benchmark_{id}/

# Option B — curl (release assets are public for public repos; token only needed for private):
curl -L "{agent_a_asset_url}" -o /tmp/benchmark_{id}/benchmark_agent_a_{eval_id}.zip

# Then unzip:
cd /tmp/benchmark_{id} && unzip -q benchmark_agent_a_{eval_id}.zip

If agent_a_asset_url is null (Phase 1 could not upload), re-run Agent A from scratch using the same procedure as Phase 1 Step 2 before continuing.

Step 6 — Run Agent B (Without Skill)
bash
mkdir -p /tmp/benchmark_{id}/agent_B/output_B

Write prompt_B.txt — contains only _prompt_b. No skill content, no resource files:

python
with open("/tmp/benchmark_{id}/agent_B/prompt_B.txt", "w") as f:
    f.write(eval_case["_prompt_b"])

Launch Agent B:

bash
export CLAUDE_CODE_MAX_OUTPUT_TOKENS=64000
cd /tmp/benchmark_{id}/agent_B && \
  cat prompt_B.txt | claude -p \
  --allowedTools "Bash,Read,Write,Edit,Glob" \
  --output-format json > agent_B_run.json 2>&1

Extract token count and record:

python
d = json.load(open("/tmp/benchmark_{id}/agent_B/agent_B_run.json"))
u = d.get("usage", {})
tokens_b = u.get("input_tokens", 0) + u.get("cache_creation_input_tokens", 0) + u.get("output_tokens", 0)
is_error_b = d.get("is_error", False)
bash
python3 _automation/benchmark-runner/scripts/record_run_result.py \
  --eval-id {state["eval_id"]} --model {state["model"]} \
  --status completed --tokens-b {tokens_b}
Step 7 — Score Blinded Outputs

Copy outputs per state["blinded_map"] to /tmp/benchmark_{id}/scoring/:

bash
mkdir -p /tmp/benchmark_{id}/scoring/candidate_1 /tmp/benchmark_{id}/scoring/candidate_2
# blinded_map: {"candidate_1": "output_B", "candidate_2": "output_A"} (or reversed)
cp -r /tmp/benchmark_{id}/agent_{X}/output_{X}/. /tmp/benchmark_{id}/scoring/candidate_1/
cp -r /tmp/benchmark_{id}/agent_{Y}/output_{Y}/. /tmp/benchmark_{id}/scoring/candidate_2/

For each candidate, evaluate every assertion in the eval case:

  • Pass — clearly met
  • Partial — partially met
  • Fail — not met

Score = (passes + 0.5 × partials) / total_assertions

Then unblind using state["blinded_map"] to map candidate scores back to "With Skill" and "Without Skill".

Step 8 — Format Full Report

Write /tmp/benchmark_comment_{skill}_{eval_id}.md:

markdown
## Automated Benchmark Results — `{_skill_name}`

### Run Metadata

| Field | Value |
|---|---|
| **Eval ID** | `{id}` |
| **Run date** | {YYYY-MM-DD HH:MM UTC} |
| **Model** | `{model}` |
| **Skill version** | `{skill_sha[:7]}` |
| **Triggered by** | Scheduled |

### Scorecard

| Metric | With Skill | Without Skill |
|---|---|---|
| **Score** | {score_A} ({pct_A}%) | {score_B} ({pct_B}%) |
| **Assertions** | {pass_A} Pass · {partial_A} Partial · {fail_A} Fail | {pass_B} Pass · {partial_B} Partial · {fail_B} Fail |
| **Skills loaded** | 1 | 0 |
| **Execution time** | {time_A} min | {time_B} min |
| **Token usage** | {tokens_a} | {tokens_b} |
| **{Key Metric 1}** | {value_A1} | {value_B1} |
| **{Key Metric 2}** | {value_A2} | {value_B2} |

### Key Observations

- {2-4 bullet points comparing both agents}

### Verdict

{1-2 sentence overall verdict}

---

## Technical Details & Artifacts

<details>
<summary>View Assertion Breakdown, Code Artifacts, and Logs</summary>

### Assertion Breakdown

| Assertion | With Skill | Without Skill |
|---|---|---|
| {assertion_text_1} | {Pass/Partial/Fail} | {Pass/Partial/Fail} |

### Debugging Information

#### Agent A (With Skill)
- **Total Turns:** {num_turns from agent_A_run.json}
- **Errors/Retries:** {is_error value, or "None"}

#### Agent B (Without Skill)
- **Total Turns:** {num_turns from agent_B_run.json}
- **Errors/Retries:** {is_error value, or "None"}

### Detailed Artifacts

**Agent A Output:** [Download Agent A Archive]({agent_a_asset_url})

#### Agent A (With Skill)
{Key output files — .R, .json, text summaries}

#### Agent B (Without Skill)
{Key output files}

</details>

---
<!-- BENCHMARK_COMPLETE: {"eval_id":"{id}","model":"{model}","skill_sha":"{skill_sha}"} -->
*Posted automatically by `benchmark-runner` · Repo: https://github.com/RConsortium/pharma-skills*

Note the <!-- BENCHMARK_COMPLETE: --> marker at the bottom — this tells future Phase Detection scans that Phase 2 is done for this eval+model+sha.

Step 9 — Post Full Results Comment

Post as a new comment using whichever GitHub access method is available (see "GitHub Access" above). The new comment carries the BENCHMARK_COMPLETE marker; the partial comment can stay in place — future Phase Detection scans will see the COMPLETE marker on a later comment and skip the partial.

If you prefer to also edit the partial comment to mark it as superseded (cleaner timeline), use whatever update method works in your environment (gh api PATCH, REST PATCH with GH_TOKEN, etc.). Optional — not required for correctness.

Phase 2 is complete. Print this summary to the user before exiting:

✓ Phase 2 complete — full benchmark posted for {eval_id} ({model})
  • Score: With Skill {pct_A}% · Without Skill {pct_B}%
  • Comment: {comment_url}
  • Tokens — A: {tokens_a:,} · B: {tokens_b:,}

Then exit cleanly.


Execution Flow

EVERY ROUTINE INVOCATION:

  Phase Detection
    │
    ├─ BENCHMARK_PARTIAL found (no BENCHMARK_COMPLETE for same eval+model) ──► Phase 2
    │     Step 5: load state from comment + restore Agent A output
    │     Step 6: run Agent B (without skill)
    │     Step 7: score blinded
    │     Step 8: format full report
    │     Step 9: post full results comment (with BENCHMARK_COMPLETE marker)
    │     EXIT
    │
    └─ No partial found ──► Phase 1
          Step 0: R pre-flight
          Step 1: get_next_eval.py → if UP_TO_DATE, EXIT
          Step 2: run Agent A (with skill)
          Step 3: archive + upload Agent A output
          Step 4: post partial comment (with BENCHMARK_PARTIAL marker + state JSON)
          EXIT

Notes on Model Name

{CURRENT_MODEL_NAME} is the fixed series label Claude Routine — see "Model Selection — Series Label and Sub-Agent Model" at the top of this skill. Sub-agents inherit the host's default model rather than being pinned with --model, so no concrete model ID is recorded in issue comments or release assets. The dedup logic normalises the label, so every run posted under Claude Routine groups into one series.

Notes on Distributed Selection

When several people run the same model, set distinct --runner-id values. The dispatcher hashes runner-id + model + UTC minute + eval-id + skill-SHA to spread different runners across different pending evals. Runners starting in the same minute may collide; the GitHub issue-comment deduplication (checking for BENCHMARK_COMPLETE markers) prevents redundant Phase 1 runs.

Notes on Rate Limits

If Agent A or Agent B hits a usage rate limit mid-run (is_error: true, result contains "You've hit your limit"):

  • Agent A rate-limited in Phase 1: record status: error_a_rate_limited in runs.json, do NOT post a partial comment, exit. The next Phase 1 invocation will retry.
  • Agent B rate-limited in Phase 2: record status: error_b_rate_limited, do NOT post a full results comment. Leave the BENCHMARK_PARTIAL comment in place so the next Phase 2 invocation retries Agent B. Include a note in the partial comment body edit if possible.

Success Criteria

  • One phase executed per invocation (~20 min each)
  • State persists entirely in GitHub issue comments — no commits, no repo write access needed from the human user
  • Blinded scoring: _blinded_scoring_map is never visible to the scorer
  • Deduplication: BENCHMARK_COMPLETE marker prevents re-running finished evals
  • Results posted on the correct GitHub issue with full assertion breakdown

© RConsortium, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts) in _automation/benchmark-runner of RConsortium/pharma-skills.

  • SKILL.md
  • CLAUDE_CODE_ROUTINE.md
  • LICENSE
  • README.md
  • image/devcontainer.png
  • runs/.gitkeep
  • runs/README.md
  • scripts/generate_dashboard.py
  • scripts/get_next_eval.py
  • scripts/post_issue_comment.py
  • scripts/record_run_result.py
  • scripts/setup_r_env.sh
  • vscode-codespaces-setup.md

Open the folder on GitHubat commit ae5d83b

Compare with similar skills

Benchmark Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Runner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Runner this skillRConsortium/pharma-skills118—~5.3kAutomated safety check: PassMIT
Eval Creator CIpskoett/pskoett-ai-skills311—~2.2kAutomated safety check: PassNone
Octocode Benchmark Runnerbgauryy/octocode949—~2.1kAutomated safety check: PassMIT
AI Project Copilotsun461941-hub/ai-project-copilot100—~3kAutomated safety check: PassMIT
Frontierharness Evalfrontier-harness-eval/eval297—~8kAutomated safety check: PassNone
SWE Benchmark Task Adderory/lumen305—~497Automated safety check: PassCustom licence

Similar skills

  • Eval Creator CI

    pskoett/pskoett-ai-skills

    [Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows).

    311 GitHub stars~2.2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

    949 GitHub stars~2.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • AI Project Copilot

    sun461941-hub/ai-project-copilot

    A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…

    100 GitHub stars~3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Frontierharness Eval

    frontier-harness-eval/eval

    Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes.

    297 GitHub stars~8k tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check passed
  • Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.

    305 GitHub stars~497 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Managed Deep Agents

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.

    1.3k GitHub stars~8.7k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check: notes

More from RConsortium/pharma-skills

All 13 skills in this repo
  • Rounding

    RConsortium/pharma-skills

    Audit R code that prepares CSR/TLF statistics for SAS-compatible rounding compliance (ties away from zero, round-once-at-display, fixed trailing-zero precision).

    118 GitHub stars~3.8k tokensUpdated 4 days ago
    Auto-check passed
  • Issue To Eval

    RConsortium/pharma-skills

    Converts one or more GitHub Issues into standardized benchmark data using automated scripts.

    118 GitHub stars~463 tokensUpdated 4 days ago
    Auto-check passed
  • Weekly Summary

    RConsortium/pharma-skills

    Generate a concise weekly progress summary for the pharmaskills repository.

    118 GitHub stars~546 tokensUpdated 4 days ago
    Auto-check passed
  • Admiral Adae

    RConsortium/pharma-skills

    Derives an ADaM Adverse Events Analysis Dataset (ADAE) using the {admiral} R package and pharmaverse ecosystem.

    118 GitHub stars~3.9k tokensUpdated 4 days ago
    Auto-check passed
  • Admiral Adsl

    RConsortium/pharma-skills

    Derives an ADaM Subject-Level Analysis Dataset (ADSL) using the {admiral} R package and pharmaverse ecosystem.

    118 GitHub stars~4.4k tokensUpdated 4 days ago
    Auto-check passed
  • Admiral Bds

    RConsortium/pharma-skills

    Derives ADaM Basic Data Structure (BDS) datasets using the {admiral} R package.

    118 GitHub stars~3.6k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about Benchmark Runner

What does Benchmark Runner do?

Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs. Benchmark Runner is an agent skill from RConsortium/pharma-skills. Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs.

When should I use Benchmark Runner?

Benchmark Runner fits situations like: someone says run benchmarks; compare skill performance; eval the skills; wants to measure whether a skill improves output quality.

How do I install Benchmark Runner in Claude Code?

Run `npx skills add RConsortium/pharma-skills --skill benchmark-runner -a claude-code`. Or copy the skill folder (_automation/benchmark-runner in RConsortium/pharma-skills) into .claude/skills/benchmark-runner in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Runner in Codex?

Run `npx skills add RConsortium/pharma-skills --skill benchmark-runner -a codex`. Or copy the skill folder (_automation/benchmark-runner in RConsortium/pharma-skills) into .agents/skills/benchmark-runner in your project. Codex loads it when a task matches its description.

Can I use Benchmark Runner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RConsortium/pharma-skills --skill benchmark-runner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-runner, .gemini/skills/benchmark-runner, .github/skills/benchmark-runner and .opencode/skills/benchmark-runner in your project.

What does Benchmark Runner need to run?

Going by SKILL.md and its folder, Benchmark Runner needs Python and a shell for the scripts in its folder, the command-line tools its instructions call (gh, claude, python3, bash and curl) and credentials named GH_TOKEN and GITHUB_TOKEN. Our summary lists: Python 3; A Bash shell; A credential in GITHUB_TOKEN.

Does Benchmark Runner access the network?

SKILL.md names 2 domains. In commands or code: docs.github.com; the agent is likely to contact it when it follows the instructions. As links in the text: claude.ai. This is read from the text; nothing was executed.

Is Benchmark Runner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Benchmark Runner use?

Benchmark Runner is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark Runner use?

About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark Runner?

Skills that share tags, products or a category with Benchmark Runner: Eval Creator CI (pskoett/pskoett-ai-skills, 311 stars), Octocode Benchmark Runner (bgauryy/octocode, 949 stars), AI Project Copilot (sun461941-hub/ai-project-copilot, 100 stars) and Frontierharness Eval (frontier-harness-eval/eval, 297 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Runner?

RConsortium (a GitHub organization) maintains it in RConsortium/pharma-skills, which has 118 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 4, 2026.

Source: RConsortium/pharma-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.