Benchmark
affaan-m/ECC
Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…
Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.
$ npx skills add PKU-YuanGroup/OpenAI4S --skill bioprobench -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install PKU-YuanGroup/OpenAI4S bioprobench --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/PKU-YuanGroup/OpenAI4S.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/bioprobench .claude/skills/bioprobench && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bioprobench" agent skill from https://github.com/PKU-YuanGroup/OpenAI4S/tree/main/skills/bioprobench into .claude/skills/bioprobench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bioprobench", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/PKU-YuanGroup/OpenAI4S/tree/main/skills/bioprobenchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add PKU-YuanGroup/OpenAI4S --skill bioprobench -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install PKU-YuanGroup/OpenAI4S bioprobench --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PKU-YuanGroup/OpenAI4S.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/bioprobench .agents/skills/bioprobench && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bioprobench" agent skill from https://github.com/PKU-YuanGroup/OpenAI4S/tree/main/skills/bioprobench into .agents/skills/bioprobench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bioprobench", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PKU-YuanGroup/OpenAI4S --skill bioprobench -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install PKU-YuanGroup/OpenAI4S bioprobench --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PKU-YuanGroup/OpenAI4S.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/bioprobench .cursor/skills/bioprobench && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bioprobench" agent skill from https://github.com/PKU-YuanGroup/OpenAI4S/tree/main/skills/bioprobench into .cursor/skills/bioprobench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bioprobench", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/PKU-YuanGroup/OpenAI4S.git --path skills/bioprobench--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add PKU-YuanGroup/OpenAI4S --skill bioprobench -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install PKU-YuanGroup/OpenAI4S bioprobench --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PKU-YuanGroup/OpenAI4S.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/bioprobench .gemini/skills/bioprobench && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bioprobench" agent skill from https://github.com/PKU-YuanGroup/OpenAI4S/tree/main/skills/bioprobench into .gemini/skills/bioprobench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bioprobench", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install PKU-YuanGroup/OpenAI4S bioprobenchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add PKU-YuanGroup/OpenAI4S --skill bioprobench -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/PKU-YuanGroup/OpenAI4S.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/bioprobench .github/skills/bioprobench && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bioprobench" agent skill from https://github.com/PKU-YuanGroup/OpenAI4S/tree/main/skills/bioprobench into .github/skills/bioprobench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bioprobench", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PKU-YuanGroup/OpenAI4S --skill bioprobench -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install PKU-YuanGroup/OpenAI4S bioprobench --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PKU-YuanGroup/OpenAI4S.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/bioprobench .opencode/skills/bioprobench && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bioprobench" agent skill from https://github.com/PKU-YuanGroup/OpenAI4S/tree/main/skills/bioprobench into .opencode/skills/bioprobench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bioprobench", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bioprobenchScore an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.
Bioprobench is an agent skill from PKU-YuanGroup/OpenAI4S. Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts (for example `README.md`, `README_zh.md` and `Scripts/LLM-as-a-judge_for_REA-ERR.py`).
The repository describes itself as: Open-source AI agent for scientific research. Analyze data in Python/R with Claude, GPT, Gemini, and more. The licence is MIT.
Read from SKILL.md and the folder at commit 4a72e87. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comapi.deepseek.comAlso links to:
huggingface.coFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bioprobench loads about 2.3k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 858 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from PKU-YuanGroup/OpenAI4S at commit 4a72e87, republished under its MIT licence (© PKU-YuanGroup). 858 words, ~2,317 tokens.
.claude/skills/bioprobench/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.Biological protocols are where a plausible-sounding answer becomes a failed experiment: a wrong dosage, a swapped step, an unflagged hazard. BioProBench scores a model on five tasks over real wet-lab protocols, roughly 5,000 instances in the full release.
| Task | What it measures | Metrics returned |
|---|---|---|
PQA | Protocol question answering — reagents, dosages, parameters | Accuracy, Brier_Score, Failed_Rate |
ORD | Step ordering — reconstructing procedural sequence | Exact_Match, Kendall_Tau, Failed_Rate |
ERR | Error correction — is this modified step valid | accuracy, precision, recall, f1, failed_rate |
GEN | Protocol generation — synthesising steps | BLEU, METEOR, ROUGE-L, KW_F1, Step_Recall, Redundancy_Penalty, Failed_Rate |
REA-ERR | Error reasoning, graded by an LLM judge | Consistency, Failure_Rate, Total, Failed, Total_Items, Judged, Unjudged, Coverage |
Metric key casing differs per task — ERR returns lowercase keys, the rest are
capitalised. Read them off the table above rather than guessing.
run_bioprobench_eval does not take a plain model-output file and compare
it against a separate answer key. It takes one file that already has the
ground truth merged into each record alongside the model's response. The
upstream inference scripts produce exactly this, by adding a
generated_response key to each benchmark record in place.
Hand it a file containing only model outputs and it does not raise: every
record simply fails to parse and the metrics come back at zero. The envelope
says so — status is "failed" when nothing scored and "partial" when some
records dropped out — but still check Failed_Rate on every run. A rate of
1.0 means the input contract was violated, not that the model scored zero.
Required keys per record, per task:
| Task | Model output key | Ground-truth key(s) |
|---|---|---|
PQA | generated_response | answer |
ORD | generated_response | wrong_steps, correct_steps |
ERR | generated_response | is_correct (true/false, 1/0, or "true"/"false") |
GEN | generated_response | output (string, or list of reference steps) |
REA-ERR | LLM_judge | none — the judgment text is itself the signal |
REA-ERR scores only records that carry an LLM_judge key. Consistency and
Failure_Rate are over that judged subset, so read Unjudged and Coverage
alongside them — a partly-judged file reports status: "partial". Populate
LLM_judge with run_llm_judge_evaluation first.
Except for REA-ERR, the parser wants the answer wrapped in
[ANSWER_START] … [ANSWER_END]. PQA additionally expects
answer & confidence inside those tags; without the & a trailing token is
read as the confidence only when it looks like one (0-100, optional %) and
something is left over for the answer, otherwise the record counts as failed
rather than being scored against a fabricated confidence. ORD expects a
Python list literal of indices that is a genuine permutation of
range(len(wrong_steps)) — anything else counts that one record as failed and
leaves the rest of the run intact. Anything before a </think> or
[/INST] marker is stripped first. See
data/sample_pqa_output.json for two records in
the exact expected shape.
Only the two-record sample above ships here. Download the full benchmark from Hugging Face (https://huggingface.co/BioProBench) into your working directory before running a real evaluation.
from bioprobench.kernel import run_bioprobench_eval
result = run_bioprobench_eval(
task_name="PQA", # 'PQA' | 'ORD' | 'ERR' | 'GEN' | 'REA-ERR'
response_file_path="/path/to/PQA_test_o3-mini.json",
)On success the return value is:
{"task": "PQA", "status": "success",
"metrics": {"Accuracy": 1.0, "Brier_Score": 0.021, "Failed_Rate": 0.0}}status is "success" only when every record scored; "partial" when some
were dropped (or, for REA-ERR, left unjudged); "failed" when none scored.
On a missing file, an unknown task name, a file that is not a non-empty JSON
list of records, or any exception raised inside the evaluator, it returns a dict
with no status field, so branch on "error" in result:
{"error": "File not found: /path/to/PQA_test_o3-mini.json"}
{"error": "out.json contains no response records, so there is nothing to"
" score for task PQA.",
"task": "PQA"}
{"error": "Evaluation failed during execution of task GEN on out.json:"
" KeyError: 'output'",
"task": "GEN", "traceback": "Traceback (most recent call last): …"}Completion goes through host.submit_output, which takes the structured output
and a required list of 1-4 completed-action bullets:
if "error" in result:
raise RuntimeError(result["error"])
metrics = result["metrics"]
host.submit_output(
result,
[
f"Scored the model's BioProBench {result['task']} responses.",
f"Reported accuracy {metrics['Accuracy']:.3f} at a"
f" {metrics['Failed_Rate']:.1%} parse-failure rate.",
],
)trigger_batch_inference runs the upstream inference scripts over a benchmark
file and writes <TASK>_test_<model>_api.json (or ..._local.json) into the
current working directory. task_name and model_name are reduced to a safe
filename component first, so a slash-bearing model id lands beside the others
rather than somewhere else on disk. It returns {"status": ..., "message": ...} and
{"error": ...} when mode="API" is missing a key.
from bioprobench.kernel import trigger_batch_inference
trigger_batch_inference(
task_name="PQA",
test_file_path="/path/to/PQA_test.json",
mode="API", # "API" (OpenAI-compatible) or "Local" (HuggingFace)
model_name="o3-mini",
api_key="...", # required when mode="API"
)Both modes leave the machine: "API" sends every prompt to the configured
endpoint, "Local" loads a HuggingFace checkpoint onto cuda:0. Neither is
sandboxed by this skill, and the API key you pass is handed straight to the
client.
Two smaller helpers:
from bioprobench.kernel import get_task_prompt, run_llm_judge_evaluation
# Standardised prompt for one record. Returns the prompt, or an error *string*
# beginning "Error loading prompt format:" — it does not raise.
prompt = get_task_prompt(sample, "PQA")
# Populate `LLM_judge` for REA-ERR. Returns the judgment text, or an error
# string. Requires the `openai` package and a live endpoint.
judgment = run_llm_judge_evaluation(
sample,
api_key="...",
base_url="https://api.deepseek.com",
model_name="deepseek-chat",
)Two things to know before the first call.
Importing kernel.py always works. Heavy dependencies are resolved per
task, at call time — so ORD, ERR and REA-ERR run on a bare kernel, and a
task whose dependencies are missing raises an ImportError naming exactly the
packages it needs. None of them ship in the default OpenAI4S venv.
Only GEN touches the network for corpora. It fetches the NLTK tokenizer
(punkt_tab, or punkt on nltk < 3.8.2) and wordnet on first use. A failure
there is not fatal: the fetch is reported under a NLTK_Data_Warnings metric
key. trigger_batch_inference and run_llm_judge_evaluation are the calls that
send your data off-machine.
What each task needs:
| Task | Runtime requirements |
|---|---|
ORD, ERR | Standard library only |
PQA | numpy, scikit-learn |
GEN | numpy, scikit-learn, nltk, rouge_score, keybert, sentence_transformers, plus a first-call download of the all-mpnet-base-v2 and all-MiniLM-L6-v2 sentence-transformer checkpoints. Slowest task by a wide margin. |
REA-ERR | Standard library to score; openai and a live endpoint to produce the judgments |
trigger_batch_inference | openai for "API"; transformers, torch, and a GPU for "Local" |
Adapted from the upstream BioProBench project (https://github.com/YuyangSunshine/bioprotocolbench). If you use this benchmark, cite the original work:
@article{liu2025bioprobench,
title={BioProBench: Comprehensive Dataset and Benchmark in Biological Protocol Understanding and Reasoning},
author={Liu, Yuyang and Lv, Liuzhenghao and Zhang, Xiancheng and Yuan, Li and Tian, Yonghong},
journal={ICML},
url={https://github.com/YuyangSunshine/bioprotocolbench/tree/main},
year={2026}
}© PKU-YuanGroup, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 12 other files (scripts) in skills/bioprobench of PKU-YuanGroup/OpenAI4S.
Open the folder on GitHubat commit 4a72e87
Bioprobench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bioprobench this skillPKU-YuanGroup/OpenAI4S | 622 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Benchmarkaffaan-m/ECC | 277k | 3 repos | ~654 | Automated safety check: Pass | MIT | |
| Benchmarkaffaan-m/ECC | 276k | — | ~412 | Automated safety check: Pass | MIT | |
| Benchmarkaffaan-m/ECC | 276k | — | ~330 | Automated safety check: Pass | MIT | |
| Claw Scoreopenclaw/openclaw | 392k | — | ~2.5k | Automated safety check: Pass | MIT | |
| Benchmark Methodologyaffaan-m/ECC | 277k | 1 repos | ~2.5k | Automated safety check: Pass | MIT |
affaan-m/ECC
Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…
affaan-m/ECC
このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.
affaan-m/ECC
使用此技能测量性能基线,检测PR前后的回归,并比较堆栈替代方案。
openclaw/openclaw
Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.
affaan-m/ECC
Score a scoped competitor set into comparable profile cards: nine weighted dimensions (positioning, voice, visual craft, offer packaging, evidence, enterprise-readiness, thought leadership, pricing…
androidx/androidx
Benchmarking and improving the performance of Jetpack Compose.
PKU-YuanGroup/OpenAI4S
Reproducible Scanpy workflow for human or mouse 10x scRNA-seq and snRNA-seq count matrices: single-sample descriptive QC, clustering and annotation, or comparative donor-aware pseudobulk DE and Milo…
PKU-YuanGroup/OpenAI4S
Map atoms and changed bonds for a complete reaction with RXNMapper.
PKU-YuanGroup/OpenAI4S
Predict ranked products from reactants and reagents with ReactionT5v2-forward; use for outcome prediction or round-trip recovery.
PKU-YuanGroup/OpenAI4S
Estimate yield for a fully specified reactant/reagent/product record with ReactionT5v2-yield.
PKU-YuanGroup/OpenAI4S
Generate de novo protein backbones with RFdiffusion for protein-target binders, hotspot-conditioned interfaces, motif scaffolding, partial diffusion, or symmetric assemblies.
PKU-YuanGroup/OpenAI4S
Generate ranked one-step precursor sets for a product with RetroChimera; use for disconnection ideas or expansion-policy calls.
Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses. Bioprobench is an agent skill from PKU-YuanGroup/OpenAI4S. Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.
Run `npx skills add PKU-YuanGroup/OpenAI4S --skill bioprobench -a claude-code`. Or copy the skill folder (skills/bioprobench in PKU-YuanGroup/OpenAI4S) into .claude/skills/bioprobench in your project. Claude Code loads it when a task matches its description.
Run `npx skills add PKU-YuanGroup/OpenAI4S --skill bioprobench -a codex`. Or copy the skill folder (skills/bioprobench in PKU-YuanGroup/OpenAI4S) into .agents/skills/bioprobench in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PKU-YuanGroup/OpenAI4S --skill bioprobench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bioprobench, .gemini/skills/bioprobench, .github/skills/bioprobench and .opencode/skills/bioprobench in your project.
Going by SKILL.md and its folder, Bioprobench needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md names 3 domains. In commands or code: github.com and api.deepseek.com; the agent is likely to contact these when it follows the instructions. As links in the text: huggingface.co. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Bioprobench is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bioprobench: Benchmark (affaan-m/ECC, 277k stars), Benchmark (affaan-m/ECC, 276k stars), Benchmark (affaan-m/ECC, 276k stars) and Claw Score (openclaw/openclaw, 392k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
PKU-YuanGroup (a GitHub organization) maintains it in PKU-YuanGroup/OpenAI4S, which has 622 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 9, 2026.
Source: PKU-YuanGroup/OpenAI4S on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.