Eval Creator CI
pskoett/pskoett-ai-skills
[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows).
Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs.
$ npx skills add RConsortium/pharma-skills --skill benchmark-runner -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install RConsortium/pharma-skills benchmark-runner --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/RConsortium/pharma-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/_automation/benchmark-runner .claude/skills/benchmark-runner && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchmark-runner" agent skill from https://github.com/RConsortium/pharma-skills/tree/main/_automation/benchmark-runner into .claude/skills/benchmark-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-runner", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/RConsortium/pharma-skills/tree/main/_automation/benchmark-runnerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add RConsortium/pharma-skills --skill benchmark-runner -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install RConsortium/pharma-skills benchmark-runner --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RConsortium/pharma-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/_automation/benchmark-runner .agents/skills/benchmark-runner && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchmark-runner" agent skill from https://github.com/RConsortium/pharma-skills/tree/main/_automation/benchmark-runner into .agents/skills/benchmark-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-runner", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add RConsortium/pharma-skills --skill benchmark-runner -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install RConsortium/pharma-skills benchmark-runner --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RConsortium/pharma-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/_automation/benchmark-runner .cursor/skills/benchmark-runner && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchmark-runner" agent skill from https://github.com/RConsortium/pharma-skills/tree/main/_automation/benchmark-runner into .cursor/skills/benchmark-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-runner", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/RConsortium/pharma-skills.git --path _automation/benchmark-runner--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add RConsortium/pharma-skills --skill benchmark-runner -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install RConsortium/pharma-skills benchmark-runner --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RConsortium/pharma-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/_automation/benchmark-runner .gemini/skills/benchmark-runner && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchmark-runner" agent skill from https://github.com/RConsortium/pharma-skills/tree/main/_automation/benchmark-runner into .gemini/skills/benchmark-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-runner", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install RConsortium/pharma-skills benchmark-runnerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add RConsortium/pharma-skills --skill benchmark-runner -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/RConsortium/pharma-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/_automation/benchmark-runner .github/skills/benchmark-runner && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-runner" agent skill from https://github.com/RConsortium/pharma-skills/tree/main/_automation/benchmark-runner into .github/skills/benchmark-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-runner", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add RConsortium/pharma-skills --skill benchmark-runner -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install RConsortium/pharma-skills benchmark-runner --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RConsortium/pharma-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/_automation/benchmark-runner .opencode/skills/benchmark-runner && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-runner" agent skill from https://github.com/RConsortium/pharma-skills/tree/main/_automation/benchmark-runner into .opencode/skills/benchmark-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-runner", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchmark-runnerAuto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs.
Benchmark Runner is an agent skill from RConsortium/pharma-skills. Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs. without skill using matched isolated sessions, and post scored results to the linked GitHub issue. Use whenever someone says "run benchmarks", "compare skill performance", "eval the skills", or wants to measure whether a skill improves output quality.
Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts (for example `CLAUDE_CODE_ROUTINE.md`, `README.md` and `runs/README.md`).
It sits in AI & LLM Engineering, covering LLM evaluation. It works with GitHub. The repository describes itself as: A collection of agent skills for BioPharma use cases GSDBench Intake https://rconsortium.github.io/pharma-skills/gsdbench-intake/. The licence is MIT.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ae5d83b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 5 files in scripts/ (Python and Shell), which the agent can run.
Shell commands in SKILL.md call:
ghclaudepython3bashcurlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
docs.github.comAlso links to:
claude.aiFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GH_TOKENGITHUB_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Benchmark Runner loads about 5.3k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 912 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from RConsortium/pharma-skills at commit ae5d83b, republished under its MIT licence (© RConsortium). 912 words, ~5,298 tokens.
.claude/skills/benchmark-runner/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.Benchmark every evaluation case in the _automation/evals/ directory of the RConsortium/pharma-skills repository. Each routine invocation is one of two short phases (~20 min each). The routine inspects GitHub issue comments on startup to decide which phase to execute — no configuration needed, no commits required, no repo write access required from the human user.
Repository: RConsortium/pharma-skills (https://github.com/RConsortium/pharma-skills)
{CURRENT_MODEL_NAME} throughout this skill is the benchmark series label, fixed to the literal string Claude Routine. Use it verbatim everywhere it appears: get_next_eval.py --model, record_run_result.py --model, the marker JSON ("model"), and the report metadata. The scripts normalise this label (lowercase, strip spaces/punctuation) for deduplication, so all runs group into one consistent series.
Sub-agent model:
--model flag (claude -p ...), so each sub-agent inherits this session's default model — the same model orchestrating the routine. This keeps the measurement clean (bare model ± skill, identical model on both sides) without ever writing a concrete model ID into a repository artifact. Never hardcode a concrete model ID anywhere in this skill or its outputs.CLAUDE_CODE_MAX_OUTPUT_TOKENS setting are the only information passed from the parent session to a sub-agent. Do not forward any other parent context: no conversation history, no additional environment variables, no eval assertions, no scoring prompt, no blinded map. This prevents the orchestrator's context from leaking into either candidate.Create a single routine at claude.ai/code/routines:
| Field | Value |
|---|---|
| Prompt | Read _automation/benchmark-runner/SKILL.md and execute. |
| Repository | RConsortium/pharma-skills |
| Schedule | 0 1,6 * * * (1 AM and 6 AM UTC — 5 h gap, matches rolling usage window) |
That is all. The skill determines its own phase on every invocation.
Throughout this skill you will read issue comments, post issue comments, and create release assets. Use whatever method is available in your environment — pick the one that works without prompting:
| Method | Best when | Notes |
|---|---|---|
mcp__github__* MCP tools | Running inside Claude Code with the GitHub MCP server | No token required; preferred when available |
gh CLI (gh issue view, gh release upload, etc.) | Running locally with gh authenticated | Concise, supports all operations |
REST API via curl | Anywhere with GH_TOKEN / GITHUB_TOKEN set | Universal fallback; use for release-asset upload (no MCP equivalent) |
| Provider-specific GitHub tools (Codex, Gemini, etc.) | Running under another agent CLI | Use whatever the host provides |
Reason about which method to use; do not enforce a rigid order. If one fails, try another. Always confirm the operation succeeded (e.g., the comment URL came back, the asset was uploaded) before continuing.
For release-asset upload there is currently no MCP tool — use gh release upload or curl POST to the upload URL.
Scan all benchmark eval issues to find any that are waiting for Phase 2 (Agent B + scoring) for the model you are running:
List all eval files: ls _automation/evals/*.json — extract each id field (e.g. github-issue-27 → issue #27).
For each issue number, fetch comments using whichever GitHub access method is available (see above). Scan each comment body for a <!-- BENCHMARK_PARTIAL: marker.
Filter and evaluate each BENCHMARK_PARTIAL marker found:
state.model does not match {CURRENT_MODEL_NAME}. This is critical: a partial run by another user on a different model belongs to that user. Only pick up partials matching your current model.<!-- BENCHMARK_COMPLETE: {"eval_id":"{same}","model":"{same}" — Phase 2 already finished for that combination.id.Decision:
created_at) → enter Phase 2 with that state.BENCHMARK_PARTIAL marker format (hidden HTML comment embedded in the issue comment body):
<!-- BENCHMARK_PARTIAL: {"eval_id":"github-issue-27","model":"Claude Routine","skill_sha":"b5ede6a...","issue_number":27,"blinded_map":{"candidate_1":"output_B","candidate_2":"output_A"},"agent_a_asset_url":"https://github.com/RConsortium/pharma-skills/releases/download/benchmark-results/benchmark_agent_a_github-issue-27.zip","run_date":"2026-05-03T06:00Z","tokens_a":199382,"partial_comment_id":4367060533} -->Runs when no Phase 2 candidate is found. Executes Agent A, archives its output, and posts a partial comment that holds state for Phase 2.
Always run first. Idempotent — safe to re-run.
bash _automation/benchmark-runner/scripts/setup_r_env.shExits non-zero on failure — stop and report the error. Do not proceed.
R packages installed:
jsonlite,digest,gsDesign,gsDesign2,lrstat,graphicalMCP,eventPred,ggplot2
python3 _automation/benchmark-runner/scripts/get_next_eval.py --model {CURRENT_MODEL_NAME}STATUS: UP_TO_DATE → all evals complete for this model+SHA. Exit._skill_name, _skill_sha, _skill_content, _bundled_resources, _prompt_a, _blinded_scoring_map, and the issue number from id.Optional flags:
--runner-id {YOUR_NAME} # stable per-person ordering
--priority-issue github-issue-{N} # force a specific evalCreate the working directory:
mkdir -p /tmp/benchmark_{id}/agent_A/output_AStage bundled resource files to disk (progressive disclosure — files read on demand, not embedded in the prompt):
import os, json
agent_a_dir = "/tmp/benchmark_{id}/agent_A"
for rel_path, content in eval_case["_bundled_resources"].items():
if rel_path == "SKILL.md":
continue
dest = os.path.join(agent_a_dir, rel_path)
os.makedirs(os.path.dirname(dest), exist_ok=True)
with open(dest, "w", encoding="utf-8") as f:
f.write(content)Write prompt_A.txt — _skill_content (SKILL.md) followed by _prompt_a only. No bundled resource content in the prompt:
prompt_a = eval_case["_skill_content"] + "\n\n" + eval_case["_prompt_a"]
with open(os.path.join(agent_a_dir, "prompt_A.txt"), "w", encoding="utf-8") as f:
f.write(prompt_a)Launch Agent A:
export CLAUDE_CODE_MAX_OUTPUT_TOKENS=64000
cd /tmp/benchmark_{id}/agent_A && \
cat prompt_A.txt | claude -p \
--allowedTools "Bash,Read,Write,Edit,Glob" \
--output-format json > agent_A_run.json 2>&1Note:
exportis required — a prefix (VAR=val cat ... | claude) only sets the variable forcat, not for theclaudeprocess receiving the pipe.
--output-format json emits a single JSON object when the agent finishes — resilient to long-running agents and session timeouts.
When Agent A returns, extract token count:
import json
d = json.load(open("/tmp/benchmark_{id}/agent_A/agent_A_run.json"))
u = d.get("usage", {})
tokens_a = u.get("input_tokens", 0) + u.get("cache_creation_input_tokens", 0) + u.get("output_tokens", 0)
is_error_a = d.get("is_error", False)Record in runs.json:
python3 _automation/benchmark-runner/scripts/record_run_result.py \
--eval-id {id} --model {CURRENT_MODEL_NAME} \
--status partial_a --tokens-a {tokens_a}Create the zip:
cd /tmp/benchmark_{id} && zip -r benchmark_agent_a_{eval_id}.zip \
agent_A/output_A/ agent_A/agent_A_run.jsonUpload to the benchmark-results GitHub release as a named asset. The release must already exist (create it once if needed). Use whichever method works in your environment — examples below; pick what works:
gh CLI (simplest if available):gh release view benchmark-results --repo RConsortium/pharma-skills \
|| gh release create benchmark-results --repo RConsortium/pharma-skills \
--prerelease --title "Automated Benchmark Results" --notes "Rolling release."
gh release upload benchmark-results /tmp/benchmark_{id}/benchmark_agent_a_{eval_id}.zip \
--repo RConsortium/pharma-skills --clobbercurl (when only GH_TOKEN is available):# Get-or-create release, then POST to its upload_url with the zip as data-binary.
# See https://docs.github.com/en/rest/releases for the exact endpoints.gh or curl for the upload step. Comment posting and reading can still use MCP.Construct the asset download URL (used in the partial comment state):
https://github.com/RConsortium/pharma-skills/releases/download/benchmark-results/benchmark_agent_a_{eval_id}.zipIf no upload method works (no gh, no token), skip the upload and set agent_a_asset_url: null in the partial state. Phase 2 will detect the null URL and re-run Agent A for that eval — wasteful but correct.
Write the partial comment body to /tmp/partial_comment_{eval_id}.md:
## Automated Benchmark Results — `{_skill_name}` 🟡 In Progress
### Run Metadata
| Field | Value |
|---|---|
| **Eval ID** | `{id}` |
| **Run date** | {YYYY-MM-DD HH:MM UTC} |
| **Model** | `{CURRENT_MODEL_NAME}` |
| **Skill version** | `{_skill_sha[:7]}` |
| **Phase** | 1 of 2 complete — Agent A (with skill) finished |
Agent A has completed. Agent B (without skill) will run in the next scheduled window (~5 h).
Results will be updated here automatically.
<!-- BENCHMARK_PARTIAL: {"eval_id":"{id}","model":"{CURRENT_MODEL_NAME}","skill_sha":"{_skill_sha}","issue_number":{N},"blinded_map":{_blinded_scoring_map},"agent_a_asset_url":"{asset_url}","run_date":"{ISO8601}","tokens_a":{tokens_a}} -->Post it using whichever GitHub access method is available (see "GitHub Access" above). The partial comment id returned by the API is not needed for Phase 2 (Phase 2 discovers it by scanning), but log it for debugging.
Phase 1 is complete. Print this summary to the user before exiting:
✓ Phase 1 complete — Agent A finished for {eval_id} ({model})
• Output archived: {asset_url}
• Partial comment: {comment_url}
• Tokens used: {tokens_a:,}
NEXT STEP — Phase 2 (Agent B + scoring):
• If running as a scheduled routine: nothing to do. The next scheduled
invocation (≥5 h from now, after the rolling usage window resets) will
detect this partial state automatically and run Phase 2.
• If running manually: re-invoke this skill any time. It will
detect the BENCHMARK_PARTIAL marker on issue #{N} and run Phase 2 to
completion.Then exit cleanly.
Runs when a BENCHMARK_PARTIAL state is found in a GitHub issue comment. Loads Agent A's output, runs Agent B, scores both, posts the full result.
Parse the BENCHMARK_PARTIAL JSON from the comment body found during Phase Detection:
import re, json
marker_re = re.compile(r'<!-- BENCHMARK_PARTIAL: ({.*?}) -->', re.DOTALL)
m = marker_re.search(comment_body)
state = json.loads(m.group(1))
# state keys: eval_id, model, skill_sha, issue_number, blinded_map,
# agent_a_asset_url, run_date, tokens_aAlso reload the full eval case (for assertions, scoring prompt, prompt_b):
# Send stdout (the eval-case JSON) to the file and stderr (warnings such as
# the >100 KB bundle notice) to a separate log. Do NOT use `2>&1` here — it
# would merge warning lines into the JSON file and break the json.load below.
python3 _automation/benchmark-runner/scripts/get_next_eval.py \
--model {state["model"]} \
--priority-issue {state["eval_id"]} \
> /tmp/eval_case_{id}.json 2>/tmp/eval_case_{id}.logRestore Agent A's output. If agent_a_asset_url is set, download and unzip it. Use whichever method works:
mkdir -p /tmp/benchmark_{id}/agent_A/output_A
# Option A — gh CLI:
gh release download benchmark-results --repo RConsortium/pharma-skills \
--pattern "benchmark_agent_a_{eval_id}.zip" --dir /tmp/benchmark_{id}/
# Option B — curl (release assets are public for public repos; token only needed for private):
curl -L "{agent_a_asset_url}" -o /tmp/benchmark_{id}/benchmark_agent_a_{eval_id}.zip
# Then unzip:
cd /tmp/benchmark_{id} && unzip -q benchmark_agent_a_{eval_id}.zipIf agent_a_asset_url is null (Phase 1 could not upload), re-run Agent A from scratch using the same procedure as Phase 1 Step 2 before continuing.
mkdir -p /tmp/benchmark_{id}/agent_B/output_BWrite prompt_B.txt — contains only _prompt_b. No skill content, no resource files:
with open("/tmp/benchmark_{id}/agent_B/prompt_B.txt", "w") as f:
f.write(eval_case["_prompt_b"])Launch Agent B:
export CLAUDE_CODE_MAX_OUTPUT_TOKENS=64000
cd /tmp/benchmark_{id}/agent_B && \
cat prompt_B.txt | claude -p \
--allowedTools "Bash,Read,Write,Edit,Glob" \
--output-format json > agent_B_run.json 2>&1Extract token count and record:
d = json.load(open("/tmp/benchmark_{id}/agent_B/agent_B_run.json"))
u = d.get("usage", {})
tokens_b = u.get("input_tokens", 0) + u.get("cache_creation_input_tokens", 0) + u.get("output_tokens", 0)
is_error_b = d.get("is_error", False)python3 _automation/benchmark-runner/scripts/record_run_result.py \
--eval-id {state["eval_id"]} --model {state["model"]} \
--status completed --tokens-b {tokens_b}Copy outputs per state["blinded_map"] to /tmp/benchmark_{id}/scoring/:
mkdir -p /tmp/benchmark_{id}/scoring/candidate_1 /tmp/benchmark_{id}/scoring/candidate_2
# blinded_map: {"candidate_1": "output_B", "candidate_2": "output_A"} (or reversed)
cp -r /tmp/benchmark_{id}/agent_{X}/output_{X}/. /tmp/benchmark_{id}/scoring/candidate_1/
cp -r /tmp/benchmark_{id}/agent_{Y}/output_{Y}/. /tmp/benchmark_{id}/scoring/candidate_2/For each candidate, evaluate every assertion in the eval case:
Score = (passes + 0.5 × partials) / total_assertions
Then unblind using state["blinded_map"] to map candidate scores back to "With Skill" and "Without Skill".
Write /tmp/benchmark_comment_{skill}_{eval_id}.md:
## Automated Benchmark Results — `{_skill_name}`
### Run Metadata
| Field | Value |
|---|---|
| **Eval ID** | `{id}` |
| **Run date** | {YYYY-MM-DD HH:MM UTC} |
| **Model** | `{model}` |
| **Skill version** | `{skill_sha[:7]}` |
| **Triggered by** | Scheduled |
### Scorecard
| Metric | With Skill | Without Skill |
|---|---|---|
| **Score** | {score_A} ({pct_A}%) | {score_B} ({pct_B}%) |
| **Assertions** | {pass_A} Pass · {partial_A} Partial · {fail_A} Fail | {pass_B} Pass · {partial_B} Partial · {fail_B} Fail |
| **Skills loaded** | 1 | 0 |
| **Execution time** | {time_A} min | {time_B} min |
| **Token usage** | {tokens_a} | {tokens_b} |
| **{Key Metric 1}** | {value_A1} | {value_B1} |
| **{Key Metric 2}** | {value_A2} | {value_B2} |
### Key Observations
- {2-4 bullet points comparing both agents}
### Verdict
{1-2 sentence overall verdict}
---
## Technical Details & Artifacts
<details>
<summary>View Assertion Breakdown, Code Artifacts, and Logs</summary>
### Assertion Breakdown
| Assertion | With Skill | Without Skill |
|---|---|---|
| {assertion_text_1} | {Pass/Partial/Fail} | {Pass/Partial/Fail} |
### Debugging Information
#### Agent A (With Skill)
- **Total Turns:** {num_turns from agent_A_run.json}
- **Errors/Retries:** {is_error value, or "None"}
#### Agent B (Without Skill)
- **Total Turns:** {num_turns from agent_B_run.json}
- **Errors/Retries:** {is_error value, or "None"}
### Detailed Artifacts
**Agent A Output:** [Download Agent A Archive]({agent_a_asset_url})
#### Agent A (With Skill)
{Key output files — .R, .json, text summaries}
#### Agent B (Without Skill)
{Key output files}
</details>
---
<!-- BENCHMARK_COMPLETE: {"eval_id":"{id}","model":"{model}","skill_sha":"{skill_sha}"} -->
*Posted automatically by `benchmark-runner` · Repo: https://github.com/RConsortium/pharma-skills*Note the <!-- BENCHMARK_COMPLETE: --> marker at the bottom — this tells future Phase Detection scans that Phase 2 is done for this eval+model+sha.
Post as a new comment using whichever GitHub access method is available (see "GitHub Access" above). The new comment carries the BENCHMARK_COMPLETE marker; the partial comment can stay in place — future Phase Detection scans will see the COMPLETE marker on a later comment and skip the partial.
If you prefer to also edit the partial comment to mark it as superseded (cleaner timeline), use whatever update method works in your environment (gh api PATCH, REST PATCH with GH_TOKEN, etc.). Optional — not required for correctness.
Phase 2 is complete. Print this summary to the user before exiting:
✓ Phase 2 complete — full benchmark posted for {eval_id} ({model})
• Score: With Skill {pct_A}% · Without Skill {pct_B}%
• Comment: {comment_url}
• Tokens — A: {tokens_a:,} · B: {tokens_b:,}Then exit cleanly.
EVERY ROUTINE INVOCATION:
Phase Detection
│
├─ BENCHMARK_PARTIAL found (no BENCHMARK_COMPLETE for same eval+model) ──► Phase 2
│ Step 5: load state from comment + restore Agent A output
│ Step 6: run Agent B (without skill)
│ Step 7: score blinded
│ Step 8: format full report
│ Step 9: post full results comment (with BENCHMARK_COMPLETE marker)
│ EXIT
│
└─ No partial found ──► Phase 1
Step 0: R pre-flight
Step 1: get_next_eval.py → if UP_TO_DATE, EXIT
Step 2: run Agent A (with skill)
Step 3: archive + upload Agent A output
Step 4: post partial comment (with BENCHMARK_PARTIAL marker + state JSON)
EXIT{CURRENT_MODEL_NAME} is the fixed series label Claude Routine — see "Model Selection — Series Label and Sub-Agent Model" at the top of this skill. Sub-agents inherit the host's default model rather than being pinned with --model, so no concrete model ID is recorded in issue comments or release assets. The dedup logic normalises the label, so every run posted under Claude Routine groups into one series.
When several people run the same model, set distinct --runner-id values. The dispatcher hashes runner-id + model + UTC minute + eval-id + skill-SHA to spread different runners across different pending evals. Runners starting in the same minute may collide; the GitHub issue-comment deduplication (checking for BENCHMARK_COMPLETE markers) prevents redundant Phase 1 runs.
If Agent A or Agent B hits a usage rate limit mid-run (is_error: true, result contains "You've hit your limit"):
status: error_a_rate_limited in runs.json, do NOT post a partial comment, exit. The next Phase 1 invocation will retry.status: error_b_rate_limited, do NOT post a full results comment. Leave the BENCHMARK_PARTIAL comment in place so the next Phase 2 invocation retries Agent B. Include a note in the partial comment body edit if possible._blinded_scoring_map is never visible to the scorerBENCHMARK_COMPLETE marker prevents re-running finished evals© RConsortium, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 12 other files (scripts) in _automation/benchmark-runner of RConsortium/pharma-skills.
Open the folder on GitHubat commit ae5d83b
Benchmark Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Benchmark Runner this skillRConsortium/pharma-skills | 118 | — | ~5.3k | Automated safety check: Pass | MIT | |
| Eval Creator CIpskoett/pskoett-ai-skills | 311 | — | ~2.2k | Automated safety check: Pass | None | |
| Octocode Benchmark Runnerbgauryy/octocode | 949 | — | ~2.1k | Automated safety check: Pass | MIT | |
| AI Project Copilotsun461941-hub/ai-project-copilot | 100 | — | ~3k | Automated safety check: Pass | MIT | |
| Frontierharness Evalfrontier-harness-eval/eval | 297 | — | ~8k | Automated safety check: Pass | None | |
| SWE Benchmark Task Adderory/lumen | 305 | — | ~497 | Automated safety check: Pass | Custom licence |
pskoett/pskoett-ai-skills
[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows).
bgauryy/octocode
Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.
sun461941-hub/ai-project-copilot
A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…
frontier-harness-eval/eval
Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes.
ory/lumen
Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.
langchain-ai/langchain-skills
INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.
RConsortium/pharma-skills
Audit R code that prepares CSR/TLF statistics for SAS-compatible rounding compliance (ties away from zero, round-once-at-display, fixed trailing-zero precision).
RConsortium/pharma-skills
Converts one or more GitHub Issues into standardized benchmark data using automated scripts.
RConsortium/pharma-skills
Generate a concise weekly progress summary for the pharmaskills repository.
RConsortium/pharma-skills
Derives an ADaM Adverse Events Analysis Dataset (ADAE) using the {admiral} R package and pharmaverse ecosystem.
RConsortium/pharma-skills
Derives an ADaM Subject-Level Analysis Dataset (ADSL) using the {admiral} R package and pharmaverse ecosystem.
RConsortium/pharma-skills
Derives ADaM Basic Data Structure (BDS) datasets using the {admiral} R package.
Works with
Categories
Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs. Benchmark Runner is an agent skill from RConsortium/pharma-skills. Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs.
Benchmark Runner fits situations like: someone says run benchmarks; compare skill performance; eval the skills; wants to measure whether a skill improves output quality.
Run `npx skills add RConsortium/pharma-skills --skill benchmark-runner -a claude-code`. Or copy the skill folder (_automation/benchmark-runner in RConsortium/pharma-skills) into .claude/skills/benchmark-runner in your project. Claude Code loads it when a task matches its description.
Run `npx skills add RConsortium/pharma-skills --skill benchmark-runner -a codex`. Or copy the skill folder (_automation/benchmark-runner in RConsortium/pharma-skills) into .agents/skills/benchmark-runner in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RConsortium/pharma-skills --skill benchmark-runner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-runner, .gemini/skills/benchmark-runner, .github/skills/benchmark-runner and .opencode/skills/benchmark-runner in your project.
Going by SKILL.md and its folder, Benchmark Runner needs Python and a shell for the scripts in its folder, the command-line tools its instructions call (gh, claude, python3, bash and curl) and credentials named GH_TOKEN and GITHUB_TOKEN. Our summary lists: Python 3; A Bash shell; A credential in GITHUB_TOKEN.
SKILL.md names 2 domains. In commands or code: docs.github.com; the agent is likely to contact it when it follows the instructions. As links in the text: claude.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Benchmark Runner is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Benchmark Runner: Eval Creator CI (pskoett/pskoett-ai-skills, 311 stars), Octocode Benchmark Runner (bgauryy/octocode, 949 stars), AI Project Copilot (sun461941-hub/ai-project-copilot, 100 stars) and Frontierharness Eval (frontier-harness-eval/eval, 297 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
RConsortium (a GitHub organization) maintains it in RConsortium/pharma-skills, which has 118 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 4, 2026.
Source: RConsortium/pharma-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.