Agent skill

Benchmark Summary

by RConsortium in RConsortium/pharma-skills

Generate a combined benchmark analysis for the group-sequential-design skill by reading all benchmark GitHub issues, selecting the latest completed run per issue, and producing a structured…

No licenceAuto-check passed

Install Benchmark Summary

skills CLI
$ npx skills add RConsortium/pharma-skills --skill benchmark-summary -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RConsortium/pharma-skills benchmark-summary --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RConsortium/pharma-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/_automation/benchmark-summary .claude/skills/benchmark-summary && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-summary
GitHub stars
118
Token cost
~1.9k tokens
SKILL.md length
672 words
Files
1
Skills in repo
13
Repo updated
First seen
Licence
None found

At a glance

Generate a combined benchmark analysis for the group-sequential-design skill by reading all benchmark GitHub issues, selecting the latest completed run per issue, and producing a structured…

  • Works in 5 steps: Discover all benchmark issues → Fetch results for each issue → Select the latest completed run per issue → …
  • The user asks to update the benchmark summary
  • SKILL.md covers Step 1 — Discover all…, Step 2 — Fetch results for…, Step 3 — Select the latest… and Step 4 — Build the…, plus 2 more sections
  • Calls gh and jq

What it does

Benchmark Summary is an agent skill from RConsortium/pharma-skills. Generate a combined benchmark analysis for the group-sequential-design skill by reading all benchmark GitHub issues, selecting the latest completed run per issue, and producing a structured three-section report (summary table + overall scorecard + failure pattern analysis). Use this skill whenever the user asks to update the benchmark summary, generate the benchmark analysis, summarize skill vs no-skill results, add failure patterns, or produce benchmarkanalysisYYYY-MM-DD.md. Always invoke for any request to…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with GitHub. The repository describes itself as: A collection of agent skills for BioPharma use cases GSDBench Intake https://rconsortium.github.io/pharma-skills/gsdbench-intake/.

When your agent uses it

  • The user asks to update the benchmark summary
  • Generate the benchmark analysis
  • Summarize skill vs no-skill results
  • Add failure patterns

Example prompts

  • “/benchmark-summary”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Discover all benchmark issues
  2. Fetch results for each issue
  3. Select the latest completed run per issue
  4. Build the three-section document
  5. Save locally and post to GitHub

What it can do on your machine

Read from SKILL.md and the folder at commit ae5d83b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Summary loads about 1.9k tokens when it runs. Until then it costs about 147 tokens; SKILL.md has 672 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~147
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 672 words (~1,898 tokens).

“Produce the combined benchmark analysis document for RConsortium/pharma-skills. This involves reading benchmark results that live as comments on GitHub issues, selecting the right run for each, and synthesizing them into a structured report.”

— opening of SKILL.md by RConsortium
name
benchmark-summary

Read the full SKILL.md on GitHub

Files

Just SKILL.md in _automation/benchmark-summary of RConsortium/pharma-skills.

Open the folder on GitHubat commit ae5d83b

Compare with similar skills

Benchmark Summary next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Summary compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Summary this skillRConsortium/pharma-skills118—~1.9kAutomated safety check: PassNone
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Diagnosing Superpowers Sessionsobra/superpowers296k3 repos~1.7kAutomated safety check: PassMIT
GitHub Deep Researchbytedance/deer-flow83k5 repos~1.3kAutomated safety check: PassMIT
Greplooponyx-dot-app/onyx32k4 repos~3.3kAutomated safety check: PassMIT
Update V8 Versionopeninterpreter/openinterpreter69k2 repos~845Automated safety check: PassApache-2.0

Similar skills

  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    296k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    83k GitHub starsUsed in 5 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Greploop

    onyx-dot-app/onyx

    Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.

    32k GitHub starsUsed in 4 repos~3.3k tokens
    DevelopmentAuto-check passed
  • Update V8 Version

    openinterpreter/openinterpreter

    Bumps the pinned v8 and rusty_v8 versions in Codex, validates the release-candidate path with the v8-canary check, and traces failures to upstream build changes.

    69k GitHub starsUsed in 2 repos~845 tokens
    DevOps & CloudAuto-check passed
  • Last30days

    mvanhorn/last30days-skill

    Research what people actually say about any topic in the last 30 days.

    64k GitHub stars~7.8k tokensUpdated today
    Research & ScienceAuto-check: notes

More from RConsortium/pharma-skills

All 13 skills in this repo
  • Rounding

    RConsortium/pharma-skills

    Audit R code that prepares CSR/TLF statistics for SAS-compatible rounding compliance (ties away from zero, round-once-at-display, fixed trailing-zero precision).

    118 GitHub stars~3.8k tokensUpdated 4 days ago
    Auto-check passed
  • Issue To Eval

    RConsortium/pharma-skills

    Converts one or more GitHub Issues into standardized benchmark data using automated scripts.

    118 GitHub stars~463 tokensUpdated 4 days ago
    Auto-check passed
  • Weekly Summary

    RConsortium/pharma-skills

    Generate a concise weekly progress summary for the pharmaskills repository.

    118 GitHub stars~546 tokensUpdated 4 days ago
    Auto-check passed
  • Admiral Adae

    RConsortium/pharma-skills

    Derives an ADaM Adverse Events Analysis Dataset (ADAE) using the {admiral} R package and pharmaverse ecosystem.

    118 GitHub stars~3.9k tokensUpdated 4 days ago
    Auto-check passed
  • Admiral Adsl

    RConsortium/pharma-skills

    Derives an ADaM Subject-Level Analysis Dataset (ADSL) using the {admiral} R package and pharmaverse ecosystem.

    118 GitHub stars~4.4k tokensUpdated 4 days ago
    Auto-check passed
  • Admiral Bds

    RConsortium/pharma-skills

    Derives ADaM Basic Data Structure (BDS) datasets using the {admiral} R package.

    118 GitHub stars~3.6k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about Benchmark Summary

What does Benchmark Summary do?

Generate a combined benchmark analysis for the group-sequential-design skill by reading all benchmark GitHub issues, selecting the latest completed run per issue, and producing a structured…. Benchmark Summary is an agent skill from RConsortium/pharma-skills. Generate a combined benchmark analysis for the group-sequential-design skill by reading all benchmark GitHub issues, selecting the latest completed run per issue, and producing a structured three-section report (summary table + overall scorecard + failure pattern analysis).

When should I use Benchmark Summary?

Benchmark Summary fits situations like: the user asks to update the benchmark summary; generate the benchmark analysis; summarize skill vs no-skill results; add failure patterns.

How do I install Benchmark Summary in Claude Code?

Run `npx skills add RConsortium/pharma-skills --skill benchmark-summary -a claude-code`. Or copy the skill folder (_automation/benchmark-summary in RConsortium/pharma-skills) into .claude/skills/benchmark-summary in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Summary in Codex?

Run `npx skills add RConsortium/pharma-skills --skill benchmark-summary -a codex`. Or copy the skill folder (_automation/benchmark-summary in RConsortium/pharma-skills) into .agents/skills/benchmark-summary in your project. Codex loads it when a task matches its description.

Can I use Benchmark Summary in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RConsortium/pharma-skills --skill benchmark-summary -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-summary, .gemini/skills/benchmark-summary, .github/skills/benchmark-summary and .opencode/skills/benchmark-summary in your project.

What does Benchmark Summary need to run?

Going by SKILL.md and its folder, Benchmark Summary needs the command-line tools its instructions call (gh and jq).

Does Benchmark Summary access the network?

SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Benchmark Summary safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Summary use?

No licence was found for Benchmark Summary or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Benchmark Summary use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark Summary?

Skills that share tags, products or a category with Benchmark Summary: PR Babysitter (openinterpreter/openinterpreter, 69k stars), Diagnosing Superpowers Sessions (obra/superpowers, 296k stars), GitHub Deep Research (bytedance/deer-flow, 83k stars) and Greploop (onyx-dot-app/onyx, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Summary?

RConsortium (a GitHub organization) maintains it in RConsortium/pharma-skills, which has 118 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 4, 2026.

Source: RConsortium/pharma-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.