Agent skill

Evaluating Machine Learning Models

by foryourhealth111-pixel in foryourhealth111-pixel/Vibe-Skills

Evaluate trained machine learning models with the right metrics and comparison logic.

MITAuto-check passedData & Analytics

Install Evaluating Machine Learning Models

skills CLI
$ npx skills add foryourhealth111-pixel/Vibe-Skills --skill evaluating-machine-learning-models -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install foryourhealth111-pixel/Vibe-Skills evaluating-machine-learning-models --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/foryourhealth111-pixel/Vibe-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/bundled/skills/evaluating-machine-learning-models .claude/skills/evaluating-machine-learning-models && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluating-machine-learning-models
GitHub stars
3.6k
Token cost
~390 tokens
SKILL.md length
128 words
Files
8 (incl. scripts, references, assets)
Skills in repo
81
Repo updated
First seen
Licence
MIT

At a glance

Evaluate trained machine learning models with the right metrics and comparison logic.

  • Benchmark review
  • SKILL.md covers Overview, When to Use This Skill, Not For / Boundaries and Typical Outputs, plus 1 more section
  • Runs Python scripts from its folder
  • Threshold selection

What it does

Evaluating Machine Learning Models is an agent skill from foryourhealth111-pixel/Vibe-Skills. Evaluate trained machine learning models with the right metrics and comparison logic. Use for benchmark review, threshold selection, calibration, validation, and model comparison; not for feature engineering or leakage auditing.

Its SKILL.md is about 390 tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts, reference files and assets (for example `assets/README.md`, `assets/visualization_script.py` and `references/README.md`).

It sits in Data & Analytics, covering Machine learning and Performance reviews. The repository describes itself as: Intelligent Skill routing and workflow orchestration for AI agents — +21.12 pp reward, −29.6% tokens on SkillsBench with DeepSeekV4Flash-VE. The licence is MIT.

When your agent uses it

  • Benchmark review
  • Threshold selection
  • Model comparison
  • Not for feature engineering

Example prompts

  • “/evaluating-machine-learning-models”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Grep, Glob, Bash(cmd:*)

What it can do on your machine

Read from SKILL.md and the folder at commit ddcaa2a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Grep
    • Glob
    • Bash(cmd:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluating Machine Learning Models loads about 390 tokens when it runs, and up to ~490 if it reads all its reference files. Until then it costs about 66 tokens; SKILL.md has 128 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~390
With references · SKILL.md plus every file in references/, read only if the agent opens them
~490

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from foryourhealth111-pixel/Vibe-Skills at commit ddcaa2a, republished under its MIT licence (© foryourhealth111-pixel). 128 words, ~390 tokens.

Download SKILL.mdSave it as .claude/skills/evaluating-machine-learning-models/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
evaluating-machine-learning-models
description
Evaluate trained machine learning models with the right metrics and comparison logic. Use for benchmark review, threshold selection, calibration, validation, and model comparison; not for feature engineering or leakage auditing.
allowed-tools
Read, Write, Edit, Grep, Glob, Bash(cmd:*)
version
1.0.0
author
Jeremy Longshore <jeremy@intentsolutions.io>
license
MIT

Model Evaluation Suite

Use this skill when the model exists and the question is whether it is good enough.

Overview

This skill focuses on choosing and interpreting the right evaluation metrics for the problem, then comparing candidate models or thresholds.

When to Use This Skill

  • Comparing candidate models with consistent metrics
  • Reviewing precision/recall/F1/AUC, regression error, calibration, or ranking quality
  • Stress-testing validation strategy before deployment or publication

Not For / Boundaries

  • Building the training pipeline itself: use scikit-learn for classical modeling or ml-pipeline-workflow for end-to-end workflow ownership
  • Engineering features: use preprocessing-data-with-automated-pipelines
  • Checking train/test contamination: use ml-data-leakage-guard

Typical Outputs

  • Metric suite recommendations
  • Model comparison tables
  • Notes on threshold tradeoffs, calibration, and validation weaknesses
  • scikit-learn for class-level error breakdowns and confusion matrices
  • scientific-reporting when the evaluation must become a deliverable

© foryourhealth111-pixel, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references, assets) in bundled/skills/evaluating-machine-learning-models of foryourhealth111-pixel/Vibe-Skills.

  • SKILL.md
  • assets/README.md
  • assets/visualization_script.py
  • references/README.md
  • scripts/README.md
  • scripts/data_loader.py
  • scripts/evaluate_model.py
  • scripts/metrics_calculator.py

Open the folder on GitHubat commit ddcaa2a

Compare with similar skills

Evaluating Machine Learning Models next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluating Machine Learning Models compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluating Machine Learning Models this skillforyourhealth111-pixel/Vibe-Skills3.6k—~390Automated safety check: PassMIT
Geomlitalo-goncalves/geoML109—~4.6kAutomated safety check: PassGPL-3.0
Univariate Multivariable Cox Regressionaipoch/medical-research-skills2k—~2.7kAutomated safety check: PassMIT
Bio Machine Learning Survival AnalysisGPTomics/bioSkills1.2k1 repos~4.6kAutomated safety check: PassMIT
Non Tumor Mechanism Guided Diagnostic ML Research Planneraipoch/medical-research-skills2k—~4.8kAutomated safety check: PassMIT
Bio Clip Seq M6a ClipGPTomics/bioSkills1.2k2 repos~5.7kAutomated safety check: PassMIT

Similar skills

  • Geoml

    italo-goncalves/geoML

    Working knowledge of the geoML Python package (github.com/italo-goncalves/geoML): variational Gaussian processes for spatial data, implicit geological modelling, block models, drillhole data…

    109 GitHub stars~4.6k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Univariate Multivariable Cox Regression

    aipoch/medical-research-skills

    A skill your agent uses when running prognostic survival analysis on a clinical cohort with time-to-event data to estimate univariate and multivariable Cox proportional hazards models, export result…

    2k GitHub stars~2.7k tokensUpdated 22 days ago
    Data & AnalyticsAuto-check passed
  • Builds and validates predictive time-to-event models on clinical and omics data with penalized Cox, random survival forests, gradient-boosted and deep survival models, and prediction-grade…

    1.2k GitHub starsUsed in 1 repo~4.6k tokens
    Data & AnalyticsAuto-check passed
  • Generates complete conventional non-oncology diagnostic machine-learning research designs from a user-provided disease context, optional mechanism theme, and validation direction.

    2k GitHub stars~4.8k tokensUpdated 22 days ago
    Data & AnalyticsAuto-check passed
  • Bio Clip Seq M6a Clip

    GPTomics/bioSkills

    Map N6-methyladenosine (m6A) RNA modifications at single-nucleotide resolution using miCLIP (Linder 2015), miCLIP2 + m6Aboost machine learning (Kortel 2021), GLORI (Liu 2023, antibody-free chemical…

    1.2k GitHub starsUsed in 2 repos~5.7k tokens
    Research & ScienceAuto-check passed
  • External Model Validation

    aipoch/medical-research-skills

    A skill your agent uses when validating an existing prognostic risk signature on an external bulk expression cohort with survival outcomes, producing risk scores, Kaplan-Meier curves, risk…

    2k GitHub stars~3.2k tokensUpdated 22 days ago
    Research & ScienceAuto-check passed

More from foryourhealth111-pixel/Vibe-Skills

All 81 skills in this repo
  • Market Research Reports

    foryourhealth111-pixel/Vibe-Skills

    Produces long consulting-style market research and industry reports covering market sizing, competitive landscape, market entry and investment theses.

    3.6k GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check: notes
  • Academic Venue Templates

    foryourhealth111-pixel/Vibe-Skills

    Supplies venue-specific LaTeX templates and formatting rules for journals, conferences and posters, and checks a manuscript against page limits and submission requirements.

    3.6k GitHub stars~3.9k tokensUpdated 1 mo ago
    Auto-check: notes
  • Digital Brain

    foryourhealth111-pixel/Vibe-Skills

    This skill should be used when the user asks to "write a post", "check my voice", "look up contact", "prepare for meeting", "weekly review", "track goals", or mentions personal brand, content…

    3.6k GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Smart File Writer

    foryourhealth111-pixel/Vibe-Skills

    Diagnoses why a file write failed (permissions, disk space, path length, locks, read-only mounts) before retrying, instead of repeating the same call blindly.

    3.6k GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Automated Video Studio

    foryourhealth111-pixel/Vibe-Skills

    Turns footage, audio and a storyboard plan into a finished short video with FFmpeg jump-cuts, subtitle burn-in and a final polish pass.

    3.6k GitHub stars~838 tokensUpdated 1 mo ago
    Auto-check passed
  • Citation Management

    foryourhealth111-pixel/Vibe-Skills

    Turns DOIs, PMIDs and arXiv IDs into clean BibTeX, searches Google Scholar and PubMed, and checks and deduplicates a reference list.

    3.6k GitHub stars~7.6k tokensUpdated 1 mo ago
    Auto-check: notes

Questions about Evaluating Machine Learning Models

What does Evaluating Machine Learning Models do?

Evaluate trained machine learning models with the right metrics and comparison logic. Evaluating Machine Learning Models is an agent skill from foryourhealth111-pixel/Vibe-Skills. Evaluate trained machine learning models with the right metrics and comparison logic.

When should I use Evaluating Machine Learning Models?

Evaluating Machine Learning Models fits situations like: benchmark review; threshold selection; model comparison; not for feature engineering.

How do I install Evaluating Machine Learning Models in Claude Code?

Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill evaluating-machine-learning-models -a claude-code`. Or copy the skill folder (bundled/skills/evaluating-machine-learning-models in foryourhealth111-pixel/Vibe-Skills) into .claude/skills/evaluating-machine-learning-models in your project. Claude Code loads it when a task matches its description.

How do I install Evaluating Machine Learning Models in Codex?

Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill evaluating-machine-learning-models -a codex`. Or copy the skill folder (bundled/skills/evaluating-machine-learning-models in foryourhealth111-pixel/Vibe-Skills) into .agents/skills/evaluating-machine-learning-models in your project. Codex loads it when a task matches its description.

Can I use Evaluating Machine Learning Models in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill evaluating-machine-learning-models -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluating-machine-learning-models, .gemini/skills/evaluating-machine-learning-models, .github/skills/evaluating-machine-learning-models and .opencode/skills/evaluating-machine-learning-models in your project.

What does Evaluating Machine Learning Models need to run?

Going by SKILL.md and its folder, Evaluating Machine Learning Models needs Python for the scripts in its folder. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Grep, Glob, Bash(cmd:*).

Does Evaluating Machine Learning Models access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluating Machine Learning Models safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Evaluating Machine Learning Models use?

Evaluating Machine Learning Models is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluating Machine Learning Models use?

About 390 tokens (SKILL.md is roughly 1.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 100 tokens, read only when the agent opens those files.

What are the alternatives to Evaluating Machine Learning Models?

Skills that share tags, products or a category with Evaluating Machine Learning Models: Geoml (italo-goncalves/geoML, 109 stars), Univariate Multivariable Cox Regression (aipoch/medical-research-skills, 2k stars), Bio Machine Learning Survival Analysis (GPTomics/bioSkills, 1.2k stars) and Non Tumor Mechanism Guided Diagnostic ML Research Planner (aipoch/medical-research-skills, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluating Machine Learning Models?

foryourhealth111-pixel (a GitHub user) maintains it in foryourhealth111-pixel/Vibe-Skills, which has 3,612 GitHub stars. The repository holds 81 skills in this directory. The repository was last updated on August 31, 2026.

Source: foryourhealth111-pixel/Vibe-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.