Agent skill

Bioprobench

by PKU-YuanGroup in PKU-YuanGroup/OpenAI4S

Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.

MITAuto-check passed

Install Bioprobench

skills CLI
$ npx skills add PKU-YuanGroup/OpenAI4S --skill bioprobench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PKU-YuanGroup/OpenAI4S bioprobench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PKU-YuanGroup/OpenAI4S.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/bioprobench .claude/skills/bioprobench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bioprobench
GitHub stars
622
Token cost
~2.3k tokens
SKILL.md length
858 words
Files
13 (incl. scripts)
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.

  • SKILL.md covers The input contract is the…, Getting the data, Scoring an existing response… and Generating responses first, plus 2 more sections
  • Runs Python scripts from its folder; reaches github.com and api.deepseek.com

What it does

Bioprobench is an agent skill from PKU-YuanGroup/OpenAI4S. Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts (for example `README.md`, `README_zh.md` and `Scripts/LLM-as-a-judge_for_REA-ERR.py`).

The repository describes itself as: Open-source AI agent for scientific research. Analyze data in Python/R with Claude, GPT, Gemini, and more. The licence is MIT.

Example prompts

  • “/bioprobench”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 4a72e87. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • api.deepseek.com

    Also links to:

    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bioprobench loads about 2.3k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 858 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from PKU-YuanGroup/OpenAI4S at commit 4a72e87, republished under its MIT licence (© PKU-YuanGroup). 858 words, ~2,317 tokens.

Download SKILL.mdSave it as .claude/skills/bioprobench/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
bioprobench
description
Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.
origin
openai4s
category
model-evaluation
metadata.display-name
BioProBench
metadata.upstream
https://github.com/YuyangSunshine/bioprotocolbench

BioProBench — protocol understanding and reasoning

Biological protocols are where a plausible-sounding answer becomes a failed experiment: a wrong dosage, a swapped step, an unflagged hazard. BioProBench scores a model on five tasks over real wet-lab protocols, roughly 5,000 instances in the full release.

TaskWhat it measuresMetrics returned
PQAProtocol question answering — reagents, dosages, parametersAccuracy, Brier_Score, Failed_Rate
ORDStep ordering — reconstructing procedural sequenceExact_Match, Kendall_Tau, Failed_Rate
ERRError correction — is this modified step validaccuracy, precision, recall, f1, failed_rate
GENProtocol generation — synthesising stepsBLEU, METEOR, ROUGE-L, KW_F1, Step_Recall, Redundancy_Penalty, Failed_Rate
REA-ERRError reasoning, graded by an LLM judgeConsistency, Failure_Rate, Total, Failed, Total_Items, Judged, Unjudged, Coverage

Metric key casing differs per task — ERR returns lowercase keys, the rest are capitalised. Read them off the table above rather than guessing.

The input contract is the thing that bites

run_bioprobench_eval does not take a plain model-output file and compare it against a separate answer key. It takes one file that already has the ground truth merged into each record alongside the model's response. The upstream inference scripts produce exactly this, by adding a generated_response key to each benchmark record in place.

Hand it a file containing only model outputs and it does not raise: every record simply fails to parse and the metrics come back at zero. The envelope says so — status is "failed" when nothing scored and "partial" when some records dropped out — but still check Failed_Rate on every run. A rate of 1.0 means the input contract was violated, not that the model scored zero.

Required keys per record, per task:

TaskModel output keyGround-truth key(s)
PQAgenerated_responseanswer
ORDgenerated_responsewrong_steps, correct_steps
ERRgenerated_responseis_correct (true/false, 1/0, or "true"/"false")
GENgenerated_responseoutput (string, or list of reference steps)
REA-ERRLLM_judgenone — the judgment text is itself the signal

REA-ERR scores only records that carry an LLM_judge key. Consistency and Failure_Rate are over that judged subset, so read Unjudged and Coverage alongside them — a partly-judged file reports status: "partial". Populate LLM_judge with run_llm_judge_evaluation first.

Except for REA-ERR, the parser wants the answer wrapped in [ANSWER_START] … [ANSWER_END]. PQA additionally expects answer & confidence inside those tags; without the & a trailing token is read as the confidence only when it looks like one (0-100, optional %) and something is left over for the answer, otherwise the record counts as failed rather than being scored against a fabricated confidence. ORD expects a Python list literal of indices that is a genuine permutation of range(len(wrong_steps)) — anything else counts that one record as failed and leaves the rest of the run intact. Anything before a </think> or [/INST] marker is stripped first. See data/sample_pqa_output.json for two records in the exact expected shape.

Getting the data

Only the two-record sample above ships here. Download the full benchmark from Hugging Face (https://huggingface.co/BioProBench) into your working directory before running a real evaluation.

Show full SKILL.md (382 more words)Show less

Scoring an existing response file

python
from bioprobench.kernel import run_bioprobench_eval

result = run_bioprobench_eval(
    task_name="PQA",  # 'PQA' | 'ORD' | 'ERR' | 'GEN' | 'REA-ERR'
    response_file_path="/path/to/PQA_test_o3-mini.json",
)

On success the return value is:

python
{"task": "PQA", "status": "success",
 "metrics": {"Accuracy": 1.0, "Brier_Score": 0.021, "Failed_Rate": 0.0}}

status is "success" only when every record scored; "partial" when some were dropped (or, for REA-ERR, left unjudged); "failed" when none scored.

On a missing file, an unknown task name, a file that is not a non-empty JSON list of records, or any exception raised inside the evaluator, it returns a dict with no status field, so branch on "error" in result:

python
{"error": "File not found: /path/to/PQA_test_o3-mini.json"}

{"error": "out.json contains no response records, so there is nothing to"
          " score for task PQA.",
 "task": "PQA"}

{"error": "Evaluation failed during execution of task GEN on out.json:"
          " KeyError: 'output'",
 "task": "GEN", "traceback": "Traceback (most recent call last): …"}

Completion goes through host.submit_output, which takes the structured output and a required list of 1-4 completed-action bullets:

python
if "error" in result:
    raise RuntimeError(result["error"])

metrics = result["metrics"]
host.submit_output(
    result,
    [
        f"Scored the model's BioProBench {result['task']} responses.",
        f"Reported accuracy {metrics['Accuracy']:.3f} at a"
        f" {metrics['Failed_Rate']:.1%} parse-failure rate.",
    ],
)

Generating responses first

trigger_batch_inference runs the upstream inference scripts over a benchmark file and writes <TASK>_test_<model>_api.json (or ..._local.json) into the current working directory. task_name and model_name are reduced to a safe filename component first, so a slash-bearing model id lands beside the others rather than somewhere else on disk. It returns {"status": ..., "message": ...} and {"error": ...} when mode="API" is missing a key.

python
from bioprobench.kernel import trigger_batch_inference

trigger_batch_inference(
    task_name="PQA",
    test_file_path="/path/to/PQA_test.json",
    mode="API",            # "API" (OpenAI-compatible) or "Local" (HuggingFace)
    model_name="o3-mini",
    api_key="...",         # required when mode="API"
)

Both modes leave the machine: "API" sends every prompt to the configured endpoint, "Local" loads a HuggingFace checkpoint onto cuda:0. Neither is sandboxed by this skill, and the API key you pass is handed straight to the client.

Two smaller helpers:

python
from bioprobench.kernel import get_task_prompt, run_llm_judge_evaluation

# Standardised prompt for one record. Returns the prompt, or an error *string*
# beginning "Error loading prompt format:" — it does not raise.
prompt = get_task_prompt(sample, "PQA")

# Populate `LLM_judge` for REA-ERR. Returns the judgment text, or an error
# string. Requires the `openai` package and a live endpoint.
judgment = run_llm_judge_evaluation(
    sample,
    api_key="...",
    base_url="https://api.deepseek.com",
    model_name="deepseek-chat",
)

Dependencies and network access

Two things to know before the first call.

Importing kernel.py always works. Heavy dependencies are resolved per task, at call time — so ORD, ERR and REA-ERR run on a bare kernel, and a task whose dependencies are missing raises an ImportError naming exactly the packages it needs. None of them ship in the default OpenAI4S venv.

Only GEN touches the network for corpora. It fetches the NLTK tokenizer (punkt_tab, or punkt on nltk < 3.8.2) and wordnet on first use. A failure there is not fatal: the fetch is reported under a NLTK_Data_Warnings metric key. trigger_batch_inference and run_llm_judge_evaluation are the calls that send your data off-machine.

What each task needs:

TaskRuntime requirements
ORD, ERRStandard library only
PQAnumpy, scikit-learn
GENnumpy, scikit-learn, nltk, rouge_score, keybert, sentence_transformers, plus a first-call download of the all-mpnet-base-v2 and all-MiniLM-L6-v2 sentence-transformer checkpoints. Slowest task by a wide margin.
REA-ERRStandard library to score; openai and a live endpoint to produce the judgments
trigger_batch_inferenceopenai for "API"; transformers, torch, and a GPU for "Local"

Citation

Adapted from the upstream BioProBench project (https://github.com/YuyangSunshine/bioprotocolbench). If you use this benchmark, cite the original work:

bibtex
@article{liu2025bioprobench,
  title={BioProBench: Comprehensive Dataset and Benchmark in Biological Protocol Understanding and Reasoning},
  author={Liu, Yuyang and Lv, Liuzhenghao and Zhang, Xiancheng and Yuan, Li and Tian, Yonghong},
  journal={ICML},
  url={https://github.com/YuyangSunshine/bioprotocolbench/tree/main},
  year={2026}
}

© PKU-YuanGroup, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts) in skills/bioprobench of PKU-YuanGroup/OpenAI4S.

  • SKILL.md
  • README.md
  • README_zh.md
  • Scripts/LLM-as-a-judge_for_REA-ERR.py
  • Scripts/README.md
  • Scripts/README_zh.md
  • Scripts/generate_response.py
  • Scripts/generate_response_local.py
  • Scripts/prompt_format.py
  • data/README.md
  • data/README_zh.md
  • data/sample_pqa_output.json
  • kernel.py

Open the folder on GitHubat commit 4a72e87

Compare with similar skills

Bioprobench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bioprobench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bioprobench this skillPKU-YuanGroup/OpenAI4S622—~2.3kAutomated safety check: PassMIT
Benchmarkaffaan-m/ECC277k3 repos~654Automated safety check: PassMIT
Benchmarkaffaan-m/ECC276k—~412Automated safety check: PassMIT
Benchmarkaffaan-m/ECC276k—~330Automated safety check: PassMIT
Claw Scoreopenclaw/openclaw392k—~2.5kAutomated safety check: PassMIT
Benchmark Methodologyaffaan-m/ECC277k1 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • Benchmark

    affaan-m/ECC

    Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…

    277k GitHub starsUsed in 3 repos~654 tokens
    Frontend & DesignAuto-check passed
  • Benchmark

    affaan-m/ECC

    このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.

    276k GitHub stars~412 tokensUpdated yesterday
    Auto-check passed
  • Benchmark

    affaan-m/ECC

    使用此技能测量性能基线,检测PR前后的回归,并比较堆栈替代方案。

    276k GitHub stars~330 tokensUpdated yesterday
    Auto-check passed
  • Claw Score

    openclaw/openclaw

    Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.

    392k GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Score a scoped competitor set into comparable profile cards: nine weighted dimensions (positioning, voice, visual craft, offer packaging, evidence, enterprise-readiness, thought leadership, pricing…

    277k GitHub starsUsed in 1 repo~2.5k tokens
    EducationAuto-check passed
  • Benchmark

    androidx/androidx

    Benchmarking and improving the performance of Jetpack Compose.

    6.1k GitHub stars~1.1k tokensUpdated yesterday
    MobileAuto-check passed

More from PKU-YuanGroup/OpenAI4S

All 17 skills in this repo
  • Single Cell Rna Analysis

    PKU-YuanGroup/OpenAI4S

    Reproducible Scanpy workflow for human or mouse 10x scRNA-seq and snRNA-seq count matrices: single-sample descriptive QC, clustering and annotation, or comparative donor-aware pseudobulk DE and Milo…

    622 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Reaction Atom Mapping

    PKU-YuanGroup/OpenAI4S

    Map atoms and changed bonds for a complete reaction with RXNMapper.

    622 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Reaction Forward Prediction

    PKU-YuanGroup/OpenAI4S

    Predict ranked products from reactants and reagents with ReactionT5v2-forward; use for outcome prediction or round-trip recovery.

    622 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Reaction Yield Estimation

    PKU-YuanGroup/OpenAI4S

    Estimate yield for a fully specified reactant/reagent/product record with ReactionT5v2-yield.

    622 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Rfdiffusion

    PKU-YuanGroup/OpenAI4S

    Generate de novo protein backbones with RFdiffusion for protein-target binders, hotspot-conditioned interfaces, motif scaffolding, partial diffusion, or symmetric assemblies.

    622 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Single Step Retrosynthesis

    PKU-YuanGroup/OpenAI4S

    Generate ranked one-step precursor sets for a product with RetroChimera; use for disconnection ideas or expansion-policy calls.

    622 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed

Questions about Bioprobench

What does Bioprobench do?

Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses. Bioprobench is an agent skill from PKU-YuanGroup/OpenAI4S. Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.

How do I install Bioprobench in Claude Code?

Run `npx skills add PKU-YuanGroup/OpenAI4S --skill bioprobench -a claude-code`. Or copy the skill folder (skills/bioprobench in PKU-YuanGroup/OpenAI4S) into .claude/skills/bioprobench in your project. Claude Code loads it when a task matches its description.

How do I install Bioprobench in Codex?

Run `npx skills add PKU-YuanGroup/OpenAI4S --skill bioprobench -a codex`. Or copy the skill folder (skills/bioprobench in PKU-YuanGroup/OpenAI4S) into .agents/skills/bioprobench in your project. Codex loads it when a task matches its description.

Can I use Bioprobench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PKU-YuanGroup/OpenAI4S --skill bioprobench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bioprobench, .gemini/skills/bioprobench, .github/skills/bioprobench and .opencode/skills/bioprobench in your project.

What does Bioprobench need to run?

Going by SKILL.md and its folder, Bioprobench needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Bioprobench access the network?

SKILL.md names 3 domains. In commands or code: github.com and api.deepseek.com; the agent is likely to contact these when it follows the instructions. As links in the text: huggingface.co. This is read from the text; nothing was executed.

Is Bioprobench safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Bioprobench use?

Bioprobench is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bioprobench use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bioprobench?

Skills that share tags, products or a category with Bioprobench: Benchmark (affaan-m/ECC, 277k stars), Benchmark (affaan-m/ECC, 276k stars), Benchmark (affaan-m/ECC, 276k stars) and Claw Score (openclaw/openclaw, 392k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bioprobench?

PKU-YuanGroup (a GitHub organization) maintains it in PKU-YuanGroup/OpenAI4S, which has 622 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 9, 2026.

Source: PKU-YuanGroup/OpenAI4S on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.