Advanced Evaluation
guanyang/open-agent-hub
This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…
A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection.
$ npx skills add Aperivue/medsci-skills --skill design-ai-benchmarking -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Aperivue/medsci-skills design-ai-benchmarking --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/design-ai-benchmarking .claude/skills/design-ai-benchmarking && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "design-ai-benchmarking" agent skill from https://github.com/Aperivue/medsci-skills/tree/main/skills/design-ai-benchmarking into .claude/skills/design-ai-benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "design-ai-benchmarking", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Aperivue/medsci-skills/tree/main/skills/design-ai-benchmarkingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Aperivue/medsci-skills --skill design-ai-benchmarking -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Aperivue/medsci-skills design-ai-benchmarking --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/design-ai-benchmarking .agents/skills/design-ai-benchmarking && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "design-ai-benchmarking" agent skill from https://github.com/Aperivue/medsci-skills/tree/main/skills/design-ai-benchmarking into .agents/skills/design-ai-benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "design-ai-benchmarking", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Aperivue/medsci-skills --skill design-ai-benchmarking -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Aperivue/medsci-skills design-ai-benchmarking --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/design-ai-benchmarking .cursor/skills/design-ai-benchmarking && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "design-ai-benchmarking" agent skill from https://github.com/Aperivue/medsci-skills/tree/main/skills/design-ai-benchmarking into .cursor/skills/design-ai-benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "design-ai-benchmarking", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Aperivue/medsci-skills.git --path skills/design-ai-benchmarking--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Aperivue/medsci-skills --skill design-ai-benchmarking -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Aperivue/medsci-skills design-ai-benchmarking --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/design-ai-benchmarking .gemini/skills/design-ai-benchmarking && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "design-ai-benchmarking" agent skill from https://github.com/Aperivue/medsci-skills/tree/main/skills/design-ai-benchmarking into .gemini/skills/design-ai-benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "design-ai-benchmarking", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Aperivue/medsci-skills design-ai-benchmarkingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Aperivue/medsci-skills --skill design-ai-benchmarking -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/design-ai-benchmarking .github/skills/design-ai-benchmarking && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "design-ai-benchmarking" agent skill from https://github.com/Aperivue/medsci-skills/tree/main/skills/design-ai-benchmarking into .github/skills/design-ai-benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "design-ai-benchmarking", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Aperivue/medsci-skills --skill design-ai-benchmarking -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Aperivue/medsci-skills design-ai-benchmarking --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/design-ai-benchmarking .opencode/skills/design-ai-benchmarking && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "design-ai-benchmarking" agent skill from https://github.com/Aperivue/medsci-skills/tree/main/skills/design-ai-benchmarking into .opencode/skills/design-ai-benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "design-ai-benchmarking", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
design-ai-benchmarkingA skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection.
Design AI Benchmarking is an agent skill from Aperivue/medsci-skills. Use when designing a study that benchmarks AI systems against a human-expert panel, before data collection. Plans the arms, decoupled rubrics with anchors, planted calibration probes, reviewer panel, inter-rater reliability targets and LLM-as-judge versus human adjudication.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `references/anchor_rotate_reader_allocation.md`, `references/benchmark_export_schema.json` and `references/elicitation_rubric_template.md`).
It sits in Education, covering LLM evaluation, Quizzes and assessments and Performance reviews. The repository describes itself as: Agent Skills for medical research — literature search, reporting-guideline & citation checks, statistics, publication figures, submission. Works with Claude Code, Codex, Cursor &… The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3b14ae2. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Design AI Benchmarking loads about 2.4k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 75 tokens; SKILL.md has 1,135 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Aperivue/medsci-skills at commit 3b14ae2, republished under its MIT licence (© Aperivue). 1,135 words, ~2,402 tokens.
.claude/skills/design-ai-benchmarking/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.## AI-Benchmark Design Review
Evaluation question: ...
Arms / systems compared: ...
Reference (human-expert panel): ...
Unit of rating: (item / case / output)
### Rubric (decoupled dimensions)
- dimension -> construct -> anchors (1..k)
### Calibration probes (blinded, randomized)
- positive-control / known-bad / instability / mechanism-contradiction
### Reviewer panel
- n reviewers, metadata captured, per-reviewer randomized order
### Reliability plan
- IRR target on the anchor set of real items (ICC form + 95% CI) + control-item hit rate (a competence check, reported separately)
### Judge strategy
- human-as-judge / LLM-as-judge / both + adjudication rule
### Validity risks
1. ...
### Minimal fixes
- ...
### Decision
- Ready to collect / Needs rubric revision / Needs arm or judge redesignReviewer ratings, reference labels, probe outcomes and agreement statistics come only from collected
rating records. Never invent them: a reported ICC, kappa or score with no underlying rating record is
the failure this skill exists to prevent. Cite a reference only with a /search-lit-confirmed DOI or
PMID; mark any other [UNVERIFIED - NEEDS MANUAL CHECK]. Flag an unconfirmed clinical definition,
diagnostic criterion or guideline recommendation [VERIFY] and ask the user.
Pin down, in writing:
Gate: Present the reconstructed evaluation question, arms, and reference to the user and confirm before designing the rubric. A wrong reconstruction misdirects the entire benchmark.
${CLAUDE_SKILL_DIR}/references/elicitation_rubric_template.md (dimensions, anchors,
probe flavors).Plant a few deliberate control items, blinded and randomized across raters (record who received which
via a probe_arm flag), to anchor the scale, measure rater drift/fatigue, and audit the rubric and
pipeline. Four flavors:
Probes are planted or adjudicated, never fabricated to fit a hypothesis.
Gate: Present the panel composition, stratification, and randomization plan for user review before recruitment is finalized.
/analyze-stats).When the item pool exceeds what one reader can rate in a session, do not make every reader rate
every item (that caps the pool at the per-reader limit). Use anchor-and-rotate (an anchor set
plus rotating incomplete blocks — not a balanced incomplete block design, since reader pairs are
not co-rated equally often): all readers rate a shared anchor set of real items (which carries
the inter-rater reliability; the planted controls carry only the competence check) plus a
rotating unique block each. The binding
constraint is usually the number of available expert readers, so solve the reverse problem (largest
pool for R readers) to size the must-rate set. Pre-specify anchor membership, raters-per-item, and the
rotation seed before rating. Read ${CLAUDE_SKILL_DIR}/references/anchor_rotate_reader_allocation.md
for the formulas, trade-offs, and a stdlib implementation.
Write the machine-readable rating record as a JSON schema, starting from
${CLAUDE_SKILL_DIR}/references/benchmark_export_schema.json: per-item ratings on every rubric
dimension, free-text justifications, follow-up flags, the probe_arm flag, reviewer id and metadata,
item order, and timing.
Gate: Present the final rubric, probe set, panel plan, judge strategy, and export schema together; collect explicit user approval before any rating begins, because changes after collection starts compromise the comparison.
/analyze-stats for ICC (stated form) / Krippendorff's α / kappa (two raters) / DeLong, agreement sample size, and effect-size real-world
translation of the benchmark results/check-reporting for STARD-AI, CLAIM, or TRIPOD+AI item-level reporting once the design is locked/design-study when the broader study around the benchmark (cohort logic, analysis unit,
comparator) also needs review/peer-review or /self-review only after ratings exist and a manuscript is being assessed© Aperivue, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/design-ai-benchmarking of Aperivue/medsci-skills.
Open the folder on GitHubat commit 3b14ae2
Design AI Benchmarking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Design AI Benchmarking this skillAperivue/medsci-skills | 333 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Advanced Evaluationguanyang/open-agent-hub | 977 | 2 repos | ~4.2k | Automated safety check: Pass | MIT | |
| Advanced Evaluationaiskillstore/marketplace | 433 | 3 repos | ~4.2k | Automated safety check: Pass | None | |
| Woo AI Smokewoocommerce/woocommerce-ios | 358 | — | ~7.4k | Automated safety check: Notes | GPL-2.0 | |
| Agentic Evaluation Frameworkborghei/Claude-Skills | 891 | — | ~1.9k | Automated safety check: Pass | MIT | |
| AI Eval Planmohitagw15856/pm-claude-skills | 1.4k | — | ~996 | Automated safety check: Pass | MIT |
guanyang/open-agent-hub
This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…
aiskillstore/marketplace
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise…
woocommerce/woocommerce-ios
Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.
borghei/Claude-Skills
This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality".
mohitagw15856/pm-claude-skills
Design an evaluation plan for an LLM or AI feature before shipping it.
benchflow-ai/benchflow
SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.
Aperivue/medsci-skills
A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.
Aperivue/medsci-skills
A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).
Aperivue/medsci-skills
A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.
Aperivue/medsci-skills
A skill your agent uses when each author needs an ICMJE Conflict of Interest disclosure form (coidisclosure.docx) for submission.
Aperivue/medsci-skills
A skill your agent uses when an institutional Word form (.doc/.docx IRB protocol, ethics application, grant template) must be filled without breaking its styles, tables, fonts or page layout.
Aperivue/medsci-skills
A skill your agent uses when looking for research topics a longitudinal cohort database can answer (NHIS, UK Biobank, an institutional EMR or registry).
Categories
A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection. Design AI Benchmarking is an agent skill from Aperivue/medsci-skills. Use when designing a study that benchmarks AI systems against a human-expert panel, before data collection.
Design AI Benchmarking fits situations like: designing a study that benchmarks AI systems against a human-expert panel; before data collection.
Run `npx skills add Aperivue/medsci-skills --skill design-ai-benchmarking -a claude-code`. Or copy the skill folder (skills/design-ai-benchmarking in Aperivue/medsci-skills) into .claude/skills/design-ai-benchmarking in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Aperivue/medsci-skills --skill design-ai-benchmarking -a codex`. Or copy the skill folder (skills/design-ai-benchmarking in Aperivue/medsci-skills) into .agents/skills/design-ai-benchmarking in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Aperivue/medsci-skills --skill design-ai-benchmarking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/design-ai-benchmarking, .gemini/skills/design-ai-benchmarking, .github/skills/design-ai-benchmarking and .opencode/skills/design-ai-benchmarking in your project.
Going by SKILL.md and its folder, Design AI Benchmarking needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Design AI Benchmarking is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Design AI Benchmarking: Advanced Evaluation (guanyang/open-agent-hub, 977 stars), Advanced Evaluation (aiskillstore/marketplace, 433 stars), Woo AI Smoke (woocommerce/woocommerce-ios, 358 stars) and Agentic Evaluation Framework (borghei/Claude-Skills, 891 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Aperivue (a GitHub organization) maintains it in Aperivue/medsci-skills, which has 333 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 5, 2026.
Source: Aperivue/medsci-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.