Agent skill

Evaluate Model

by PKU-YuanGroup in PKU-YuanGroup/OpenAI4S

Evaluate binary classification or regression models with confusion-matrix metrics, tie-aware ROC AUC, regression errors, and deterministic bootstrap confidence intervals; emphasizes held-out data…

MITAuto-check passedData & Analytics

Install Evaluate Model

skills CLI
$ npx skills add PKU-YuanGroup/OpenAI4S --skill evaluate-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PKU-YuanGroup/OpenAI4S evaluate-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PKU-YuanGroup/OpenAI4S.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluate-model .claude/skills/evaluate-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluate-model
GitHub stars
620
Token cost
~583 tokens
SKILL.md length
217 words
Files
4
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Evaluate binary classification or regression models with confusion-matrix metrics, tie-aware ROC AUC, regression errors, and deterministic bootstrap confidence intervals; emphasizes held-out data…

  • Works in 6 steps: Confirm that evaluation examples and… → Choose a primary metric from the… → Compare against a simple baseline and… → …
  • Tasks that involve Machine learning
  • SKILL.md covers Workflow, Import and run and Reporting contract
  • Runs Python scripts from its folder

What it does

Evaluate Model is an agent skill from PKU-YuanGroup/OpenAI4S. Evaluate binary classification or regression models with confusion-matrix metrics, tie-aware ROC AUC, regression errors, and deterministic bootstrap confidence intervals; emphasizes held-out data, uncertainty, baselines, and subgroup checks.

Its SKILL.md is about 580 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `README.md`, `README_zh.md` and `kernel.py`).

It sits in Data & Analytics, covering Machine learning and Statistics. The repository describes itself as: Open-source AI agent for scientific research. Analyze data in Python/R with Claude, GPT, Gemini, and more. The licence is MIT.

When your agent uses it

  • Tasks that involve Machine learning
  • Tasks that involve Statistics

Example prompts

  • “/evaluate-model”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Confirm that evaluation examples and related entities were never used for
  2. Choose a primary metric from the scientific or operational question; do not
  3. Compare against a simple baseline and report uncertainty, not only a point
  4. Inspect clinically or scientifically relevant subgroups and failure modes.
  5. Keep probability quality, ranking quality, and thresholded decisions
  6. Save predictions with stable sample IDs so every aggregate is auditable.

What it can do on your machine

Read from SKILL.md and the folder at commit 4a72e87. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluate Model loads about 583 tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 217 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~583

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PKU-YuanGroup/OpenAI4S at commit 4a72e87, republished under its MIT licence (© PKU-YuanGroup). 217 words, ~583 tokens.

Download SKILL.mdSave it as .claude/skills/evaluate-model/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
evaluate-model
description
Evaluate binary classification or regression models with confusion-matrix metrics, tie-aware ROC AUC, regression errors, and deterministic bootstrap confidence intervals; emphasizes held-out data, uncertainty, baselines, and subgroup checks.
origin
openai4s
category
model-evaluation

Evaluate a model

Use this skill when comparing predictive models, selecting a threshold, or reporting held-out performance. Compute metrics only after defining the target, unit of analysis, split boundary, baseline, and decision cost.

Workflow

  1. Confirm that evaluation examples and related entities were never used for fitting, preprocessing, feature selection, or threshold selection.
  2. Choose a primary metric from the scientific or operational question; do not select it after inspecting test results.
  3. Compare against a simple baseline and report uncertainty, not only a point estimate.
  4. Inspect clinically or scientifically relevant subgroups and failure modes.
  5. Keep probability quality, ranking quality, and thresholded decisions separate. Accuracy is rarely sufficient for an imbalanced target.
  6. Save predictions with stable sample IDs so every aggregate is auditable.

Import and run

python
from importlib import import_module

metrics = import_module("evaluate-model.kernel")
classification = metrics.binary_classification_metrics(
    y_true, scores=probabilities, threshold=0.35
)
regression = metrics.regression_metrics(y_true_continuous, predictions)
interval = metrics.bootstrap_ci(per_sample_losses, resamples=2000, seed=42)

The binary helper reports the confusion matrix, accuracy, precision, recall, specificity, F1, balanced accuracy, and tie-aware ROC AUC when scores are provided. Undefined ratios are returned as None, not silently replaced by zero.

Reporting contract

State the split strategy, sample and group counts, prevalence, threshold source, primary metric with interval, baseline, subgroup caveats, and all exclusions. Bootstrap intervals describe sampling variability under the resampling unit; they do not correct leakage, dataset shift, measurement error, or dependence between observations. Use grouped or clustered resampling when rows are not independent.

© PKU-YuanGroup, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in skills/evaluate-model of PKU-YuanGroup/OpenAI4S.

  • SKILL.md
  • README.md
  • README_zh.md
  • kernel.py

Open the folder on GitHubat commit 4a72e87

Compare with similar skills

Evaluate Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluate Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluate Model this skillPKU-YuanGroup/OpenAI4S620—~583Automated safety check: PassMIT
scikit-survival Time-to-Event Modelingdavila7/claude-code-templates32k11 repos~3.7kAutomated safety check: PassMIT
Data Scientistdavila7/claude-code-templates32k8 repos~2.6kAutomated safety check: PassMIT
Scientific Toolkit SkillzLanqing/codex-claude-academic-skills4.7k—~1.2kAutomated safety check: PassMIT
Data Sciencetravisjneuman/.claude101—~2.3kAutomated safety check: PassMIT
Regression Analysis Modelingliangdabiao/claude-data-analysis-ultra-main290—~1.7kAutomated safety check: NotesNone

Similar skills

  • scikit-survival Time-to-Event Modeling

    davila7/claude-code-templates

    Fits and evaluates survival models with scikit-survival: Cox models, Random Survival Forests, boosting, survival SVMs, concordance index, Brier score and competing risks.

    32k GitHub starsUsed in 11 repos~3.7k tokens
    Data & AnalyticsAuto-check passed
  • Data Scientist

    davila7/claude-code-templates

    Expert data scientist for advanced analytics, machine learning, and statistical modeling.

    32k GitHub starsUsed in 8 repos~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Scientific Toolkit Skill

    zLanqing/codex-claude-academic-skills

    Research computing toolkit for optoelectronic information science and engineering, MATLAB/Octave, Python scientific analysis, signal processing, image processing, statistics, simulation…

    4.7k GitHub stars~1.2k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Data Science

    travisjneuman/.claude

    Data science and analytics expertise for statistical analysis, machine learning pipelines, data governance, business intelligence, predictive modeling, and analytics strategy.

    101 GitHub stars~2.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Regression Analysis Modeling

    liangdabiao/claude-data-analysis-ultra-main

    Perform comprehensive regression analysis and predictive modeling using linear regression, decision trees, and random forests.

    290 GitHub stars~1.7k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes
  • Performing Regression Analysis

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill empowers AI assistant to perform regression analysis and modeling using the regression-analysis-tool plugin.

    2.8k GitHub stars~944 tokensUpdated today
    Data & AnalyticsAuto-check passed

More from PKU-YuanGroup/OpenAI4S

All 17 skills in this repo
  • Single Cell Rna Analysis

    PKU-YuanGroup/OpenAI4S

    Reproducible Scanpy workflow for human or mouse 10x scRNA-seq and snRNA-seq count matrices: single-sample descriptive QC, clustering and annotation, or comparative donor-aware pseudobulk DE and Milo…

    620 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Bioprobench

    PKU-YuanGroup/OpenAI4S

    Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.

    620 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Reaction Atom Mapping

    PKU-YuanGroup/OpenAI4S

    Map atoms and changed bonds for a complete reaction with RXNMapper.

    620 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Reaction Forward Prediction

    PKU-YuanGroup/OpenAI4S

    Predict ranked products from reactants and reagents with ReactionT5v2-forward; use for outcome prediction or round-trip recovery.

    620 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Reaction Yield Estimation

    PKU-YuanGroup/OpenAI4S

    Estimate yield for a fully specified reactant/reagent/product record with ReactionT5v2-yield.

    620 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Rfdiffusion

    PKU-YuanGroup/OpenAI4S

    Generate de novo protein backbones with RFdiffusion for protein-target binders, hotspot-conditioned interfaces, motif scaffolding, partial diffusion, or symmetric assemblies.

    620 GitHub stars~2.2k tokensUpdated today
    Auto-check passed

Questions about Evaluate Model

What does Evaluate Model do?

Evaluate binary classification or regression models with confusion-matrix metrics, tie-aware ROC AUC, regression errors, and deterministic bootstrap confidence intervals; emphasizes held-out data…. Evaluate Model is an agent skill from PKU-YuanGroup/OpenAI4S. Evaluate binary classification or regression models with confusion-matrix metrics, tie-aware ROC AUC, regression errors, and deterministic bootstrap confidence intervals; emphasizes held-out data, uncertainty, baselines, and subgroup checks.

When should I use Evaluate Model?

Evaluate Model fits situations like: tasks that involve Machine learning; tasks that involve Statistics.

How do I install Evaluate Model in Claude Code?

Run `npx skills add PKU-YuanGroup/OpenAI4S --skill evaluate-model -a claude-code`. Or copy the skill folder (skills/evaluate-model in PKU-YuanGroup/OpenAI4S) into .claude/skills/evaluate-model in your project. Claude Code loads it when a task matches its description.

How do I install Evaluate Model in Codex?

Run `npx skills add PKU-YuanGroup/OpenAI4S --skill evaluate-model -a codex`. Or copy the skill folder (skills/evaluate-model in PKU-YuanGroup/OpenAI4S) into .agents/skills/evaluate-model in your project. Codex loads it when a task matches its description.

Can I use Evaluate Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PKU-YuanGroup/OpenAI4S --skill evaluate-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-model, .gemini/skills/evaluate-model, .github/skills/evaluate-model and .opencode/skills/evaluate-model in your project.

What does Evaluate Model need to run?

Going by SKILL.md and its folder, Evaluate Model needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Evaluate Model access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluate Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluate Model use?

Evaluate Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluate Model use?

About 583 tokens (SKILL.md is roughly 2.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluate Model?

Skills that share tags, products or a category with Evaluate Model: scikit-survival Time-to-Event Modeling (davila7/claude-code-templates, 32k stars), Data Scientist (davila7/claude-code-templates, 32k stars), Scientific Toolkit Skill (zLanqing/codex-claude-academic-skills, 4.7k stars) and Data Science (travisjneuman/.claude, 101 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluate Model?

PKU-YuanGroup (a GitHub organization) maintains it in PKU-YuanGroup/OpenAI4S, which has 620 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 9, 2026.

Source: PKU-YuanGroup/OpenAI4S on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.