Agent skill

Evaluating Machine Learning Models

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Build this skill allows AI assistant to evaluate machine learning models using a comprehensive suite of metrics.

MITAuto-check passedData & Analytics

Install Evaluating Machine Learning Models

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill evaluating-machine-learning-models -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace evaluating-machine-learning-models --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/evaluating-machine-learning-models .claude/skills/evaluating-machine-learning-models && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluating-machine-learning-models
GitHub stars
2.8k
Token cost
~914 tokens
SKILL.md length
400 words
Files
5 (incl. scripts, references, assets)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Build this skill allows AI assistant to evaluate machine learning models using a comprehensive suite of metrics.

  • Works in 3 steps: Analyzing Context: Claude analyzes the… → Executing Evaluation: Claude uses the… → Presenting Results: Claude presents the…
  • Requests model performance analysis
  • SKILL.md covers Overview, How It Works, When to Use This Skill and Examples, plus 7 more sections
  • Runs Python scripts from its folder

What it does

Evaluating Machine Learning Models is an agent skill from jeremylongshore/tons-of-skills-marketplace. Build this skill allows AI assistant to evaluate machine learning models using a comprehensive suite of metrics. it should be used when the user requests model performance analysis, validation, or testing. AI assistant can use this skill to assess model accuracy, p... Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

Its SKILL.md is about 910 tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts, reference files and assets (for example `assets/README.md`, `assets/visualization_script.py` and `references/README.md`). Compatibility notes: Designed for Claude Code

It sits in Data & Analytics, covering Machine learning. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Requests model performance analysis
  • Assess model accuracy
  • Appropriate context detected
  • With relevant phrases based on skill purpose

Example prompts

  • “/evaluating-machine-learning-models”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Grep, Glob, Bash(cmd:*)

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Analyzing Context: Claude analyzes the user's request to identify the model to be evaluated and any specific metrics of interest.
  2. Executing Evaluation: Claude uses the /eval-model command to initiate the model evaluation process within the model-evaluation-suite plugin.
  3. Presenting Results: Claude presents the generated metrics and insights to the user, highlighting key performance indicators and potential…

What it can do on your machine

Read from SKILL.md and the folder at commit 80f86df. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Grep
    • Glob
    • Bash(cmd:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Evaluating Machine Learning Models loads about 914 tokens when it runs, and up to ~931 if it reads all its reference files. Until then it costs about 99 tokens; SKILL.md has 400 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~914
With references · SKILL.md plus every file in references/, read only if the agent opens them
~931

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit 80f86df, republished under its MIT licence (© jeremylongshore). 400 words, ~914 tokens.

Download SKILL.mdSave it as .claude/skills/evaluating-machine-learning-models/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
evaluating-machine-learning-models
description
Build this skill allows AI assistant to evaluate machine learning models using a comprehensive suite of metrics. it should be used when the user requests model performance analysis, validation, or testing. AI assistant can use this skill to assess model accuracy, p... Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.
allowed-tools
Read, Write, Edit, Grep, Glob, Bash(cmd:*)
compatibility
Designed for Claude Code
version
1.22.0
author
Jeremy Longshore <jeremy@intentsolutions.io>
license
MIT
tags
ai, testing, ml, performance

Model Evaluation Suite

Evaluate machine learning models using a comprehensive suite of metrics including accuracy, precision, recall, F1-score, and custom KPIs.

Overview

This skill empowers Claude to perform thorough evaluations of machine learning models, providing detailed performance insights. It leverages the model-evaluation-suite plugin to generate a range of metrics, enabling informed decisions about model selection and optimization.

How It Works

  1. Analyzing Context: Claude analyzes the user's request to identify the model to be evaluated and any specific metrics of interest.
  2. Executing Evaluation: Claude uses the /eval-model command to initiate the model evaluation process within the model-evaluation-suite plugin.
  3. Presenting Results: Claude presents the generated metrics and insights to the user, highlighting key performance indicators and potential areas for improvement.

When to Use This Skill

This skill activates when you need to:

  • Assess the performance of a machine learning model.
  • Compare the performance of multiple models.
  • Identify areas where a model can be improved.
  • Validate a model's performance before deployment.

Examples

Example 1: Evaluating Model Accuracy

User request: "Evaluate the accuracy of my image classification model."

The skill will:

  1. Invoke the /eval-model command.
  2. Analyze the model's performance on a held-out dataset.
  3. Report the accuracy score and other relevant metrics.
Example 2: Comparing Model Performance

User request: "Compare the F1-score of model A and model B."

The skill will:

  1. Invoke the /eval-model command for both models.
  2. Extract the F1-score from the evaluation results.
  3. Present a comparison of the F1-scores for model A and model B.
Show full SKILL.md (149 more words)Show less

Best Practices

  • Specify Metrics: Clearly define the specific metrics of interest for the evaluation.
  • Data Validation: Ensure the data used for evaluation is representative of the real-world data the model will encounter.
  • Interpret Results: Provide context and interpretation of the evaluation results to facilitate informed decision-making.

Integration

This skill integrates seamlessly with the model-evaluation-suite plugin, providing a comprehensive solution for model evaluation within the Claude Code environment. It can be combined with other skills to build automated machine learning workflows.

Prerequisites

  • Appropriate file access permissions
  • Required dependencies installed

Instructions

  1. Invoke this skill when the trigger conditions are met
  2. Provide necessary context and parameters
  3. Review the generated output
  4. Apply modifications as needed

Output

The skill produces structured output relevant to the task.

Error Handling

  • Invalid input: Prompts for correction
  • Missing dependencies: Lists required components
  • Permission errors: Suggests remediation steps

Resources

  • Project documentation
  • Related skills and commands

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references, assets) in skills/.curated/evaluating-machine-learning-models of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • assets/README.md
  • assets/visualization_script.py
  • references/README.md
  • scripts/README.md

Open the folder on GitHubat commit 80f86df

Compare with similar skills

Evaluating Machine Learning Models next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluating Machine Learning Models compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluating Machine Learning Models this skilljeremylongshore/tons-of-skills-marketplace2.8k—~914Automated safety check: PassMIT
Scikit LearnzLanqing/codex-claude-academic-skills4.7k16 repos~3.9kAutomated safety check: PassBSD-3-Clause
Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill188—~4kAutomated safety check: PassMIT
Senior Data ScientistRaidriar7170/hermes-skilleval1255 repos~1.4kAutomated safety check: PassMIT
Geomlitalo-goncalves/geoML109—~4.6kAutomated safety check: PassGPL-3.0
QuantMind Training Config Generatorqusong0627/QuantMind1.7k—~1.5kAutomated safety check: PassAGPL-3.0

Similar skills

  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 16 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Agentic Kaggle Workflow

    FrankS-IntelLab/agentic-kaggle-skill

    Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

    188 GitHub stars~4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 5 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Geoml

    italo-goncalves/geoML

    Working knowledge of the geoML Python package (github.com/italo-goncalves/geoML): variational Gaussian processes for spatial data, implicit geological modelling, block models, drillhole data…

    109 GitHub stars~4.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Turns a plain-language model training request into a validated QuantMind training config file that can be imported from the Model Training page.

    1.7k GitHub stars~1.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Retention Analysis

    liangdabiao/claude-data-analysis-ultra-main

    Analyze user retention and churn using survival analysis, cohort analysis, and machine learning.

    290 GitHub stars~1.3k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Evaluating Machine Learning Models

What does Evaluating Machine Learning Models do?

Build this skill allows AI assistant to evaluate machine learning models using a comprehensive suite of metrics. Evaluating Machine Learning Models is an agent skill from jeremylongshore/tons-of-skills-marketplace. Build this skill allows AI assistant to evaluate machine learning models using a comprehensive suite of metrics.

When should I use Evaluating Machine Learning Models?

Evaluating Machine Learning Models fits situations like: requests model performance analysis; assess model accuracy; appropriate context detected; with relevant phrases based on skill purpose.

How do I install Evaluating Machine Learning Models in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill evaluating-machine-learning-models -a claude-code`. Or copy the skill folder (skills/.curated/evaluating-machine-learning-models in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/evaluating-machine-learning-models in your project. Claude Code loads it when a task matches its description.

How do I install Evaluating Machine Learning Models in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill evaluating-machine-learning-models -a codex`. Or copy the skill folder (skills/.curated/evaluating-machine-learning-models in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/evaluating-machine-learning-models in your project. Codex loads it when a task matches its description.

Can I use Evaluating Machine Learning Models in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill evaluating-machine-learning-models -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluating-machine-learning-models, .gemini/skills/evaluating-machine-learning-models, .github/skills/evaluating-machine-learning-models and .opencode/skills/evaluating-machine-learning-models in your project.

What does Evaluating Machine Learning Models need to run?

Going by SKILL.md and its folder, Evaluating Machine Learning Models needs Python for the scripts in its folder. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Grep, Glob, Bash(cmd:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Evaluating Machine Learning Models access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluating Machine Learning Models safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Evaluating Machine Learning Models use?

Evaluating Machine Learning Models is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluating Machine Learning Models use?

About 914 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 17 tokens, read only when the agent opens those files.

What are the alternatives to Evaluating Machine Learning Models?

Skills that share tags, products or a category with Evaluating Machine Learning Models: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.7k stars), Agentic Kaggle Workflow (FrankS-IntelLab/agentic-kaggle-skill, 188 stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars) and Geoml (italo-goncalves/geoML, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluating Machine Learning Models?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,825 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 9, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.