Agent skill

Model Evaluator

by FerroxLabs in FerroxLabs/wayland

ML model assessment covering classification metrics (precision, recall, F1, AUC-ROC), regression metrics (MAE, RMSE, R2), confusion matrix analysis, cross-validation strategies, bias detection…

Apache-2.0Auto-check passedData & Analytics

Install Model Evaluator

skills CLI
$ npx skills add FerroxLabs/wayland --skill model-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install FerroxLabs/wayland model-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/ai-machine-learning/model-evaluator .claude/skills/model-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-evaluator
GitHub stars
608
Token cost
~4.1k tokens
SKILL.md length
565 words
Files
1
Skills in repo
1,194
Repo updated
First seen
Licence
Apache-2.0

At a glance

ML model assessment covering classification metrics (precision, recall, F1, AUC-ROC), regression metrics (MAE, RMSE, R2), confusion matrix analysis, cross-validation strategies, bias detection…

  • The user asks about model evaluator
  • SKILL.md covers Overview, Classification Metrics, Confusion Matrix Analysis and Regression Metrics, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Model evaluator best practices

What it does

Model Evaluator is an agent skill from FerroxLabs/wayland. ML model assessment covering classification metrics (precision, recall, F1, AUC-ROC), regression metrics (MAE, RMSE, R2), confusion matrix analysis, cross-validation strategies, bias detection, fairness metrics, and A/B testing for models. Use when the user asks about model evaluator, model evaluator best practices, or needs guidance on model evaluator implementation. Do NOT use when the user needs a different specialized skill or is asking about an unrelated technology domain.

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Machine learning and A/B testing. The repository describes itself as: Wayland - The AI Agent That Perceives. Reasons. Acts. Evolves. The licence is Apache-2.0.

When your agent uses it

  • The user asks about model evaluator
  • Model evaluator best practices
  • Needs guidance on model evaluator implementation
  • The user needs a different specialized skill

Example prompts

  • “/model-evaluator”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 4c030c7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Evaluator loads about 4.1k tokens when it runs. Until then it costs about 125 tokens; SKILL.md has 565 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~125
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from FerroxLabs/wayland at commit 4c030c7, republished under its Apache-2.0 licence (© FerroxLabs). 565 words, ~4,054 tokens.

Download SKILL.mdSave it as .claude/skills/model-evaluator/SKILL.md (or your agent's skills folder).
name
model-evaluator
description
ML model assessment covering classification metrics (precision, recall, F1, AUC-ROC), regression metrics (MAE, RMSE, R2), confusion matrix analysis, cross-validation strategies, bias detection, fairness metrics, and A/B testing for models. Use when the user asks about model evaluator, model evaluator best practices, or needs guidance on model evaluator implementation. Do NOT use when the user needs a different specialized skill or is asking about an unrelated technology domain.
license
Apache-2.0
metadata.author
foundry-skills
metadata.version
1.0.0
metadata.tags
ai-ml testing guide
metadata.category
ai-machine-learning
metadata.subcategory
ml-fundamentals
metadata.disclaimer
none
metadata.difficulty
intermediate

Model Evaluator

Overview

Rigorous model assessment is essential to deploying trustworthy ML systems. This skill covers comprehensive metrics for classification and regression, strategies for cross-validation, statistical testing, bias and fairness auditing, and A/B testing frameworks for comparing models in production.

Classification Metrics

Core Metrics
python
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score, f1_score,
    roc_auc_score, average_precision_score, classification_report,
    confusion_matrix,
)
import numpy as np

def classification_report_full(y_true, y_pred, y_prob=None) -> dict:
    """Comprehensive classification metrics."""
    metrics = {
        "accuracy": accuracy_score(y_true, y_pred),
        "precision_macro": precision_score(y_true, y_pred, average="macro"),
        "recall_macro": recall_score(y_true, y_pred, average="macro"),
        "f1_macro": f1_score(y_true, y_pred, average="macro"),
        # ... (condensed) ...
            metrics["auc_roc_ovr"] = roc_auc_score(
                y_true, y_prob, multi_class="ovr", average="macro"
            )

    return metrics
Metric Selection Guide
MetricWhen to UseSensitive To
AccuracyBalanced classes onlyClass imbalance
PrecisionCost of false positives is highThreshold selection
RecallCost of false negatives is highThreshold selection
F1 ScoreBalance precision and recallThreshold selection
AUC-ROCOverall ranking abilityNot threshold-dependent
Average PrecisionImbalanced classes, rankingClass distribution
Cohen's KappaAgreement beyond chanceNone
Decision Framework
Is your dataset balanced (classes within 2x of each other)?
  YES -> Accuracy is meaningful, but also report F1
  NO  -> DO NOT rely on accuracy. Use these instead:
         - F1 (balanced view)
         - Average Precision (best for heavy imbalance)
         - AUC-ROC (threshold-independent ranking)

What is more costly?
  False positives (spam filter, fraud alert):
    -> Optimize for PRECISION
  False negatives (cancer screening, security):
    -> Optimize for RECALL
  Both equally bad:
    -> Optimize for F1 score

Confusion Matrix Analysis

Visualization and Interpretation
python
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay

def plot_confusion_matrix(
    y_true, y_pred,
    class_names: list[str] = None,
    normalize: str = None,
    figsize: tuple = (8, 6),
) -> plt.Figure:
    """Plot confusion matrix with detailed annotations."""
    cm = confusion_matrix(y_true, y_pred, normalize=normalize)

    fig, ax = plt.subplots(figsize=figsize)
    # ... (condensed) ...
                })
        confused_with.sort(key=lambda x: x["count"], reverse=True)
        analysis[cls]["most_confused_with"] = confused_with[:3]

    return analysis
ROC and Precision-Recall Curves
python
from sklearn.metrics import roc_curve, precision_recall_curve, auc

def plot_roc_pr_curves(y_true, y_prob, figsize=(14, 5)):
    """Plot ROC and Precision-Recall curves side by side."""
    fig, (ax1, ax2) = plt.subplots(1, 2, figsize=figsize)

    # ROC Curve
    fpr, tpr, roc_thresholds = roc_curve(y_true, y_prob)
    roc_auc = auc(fpr, tpr)
    ax1.plot(fpr, tpr, label=f"AUC = {roc_auc:.3f}")
    ax1.plot([0, 1], [0, 1], "k--", alpha=0.3)
    ax1.set_xlabel("False Positive Rate")
    ax1.set_ylabel("True Positive Rate")
    ax1.set_title("ROC Curve")
    # ... (condensed) ...
    ax2.set_title("Precision-Recall Curve")
    ax2.legend()

    plt.tight_layout()
    return fig
Threshold Optimization
python
from sklearn.metrics import balanced_accuracy_score

def find_optimal_threshold(
    y_true, y_prob,
    metric: str = "f1",
) -> tuple[float, float]:
    """Find the threshold that maximizes a given metric."""

    thresholds = np.arange(0.1, 0.95, 0.01)
    best_threshold = 0.5
    best_score = 0

    for threshold in thresholds:
        y_pred = (y_prob >= threshold).astype(int)
# ... (condensed) ...
        if score > best_score:
            best_score = score
            best_threshold = threshold

    return best_threshold, best_score

Regression Metrics

Core Regression Metrics
python
from sklearn.metrics import (
    mean_absolute_error, mean_squared_error, r2_score,
    mean_absolute_percentage_error, median_absolute_error,
)

def regression_report(y_true, y_pred) -> dict:
    """Comprehensive regression metrics."""
    return {
        "mae": mean_absolute_error(y_true, y_pred),
        "rmse": np.sqrt(mean_squared_error(y_true, y_pred)),
        "mse": mean_squared_error(y_true, y_pred),
        "r2": r2_score(y_true, y_pred),
        "mape": mean_absolute_percentage_error(y_true, y_pred),
        "median_ae": median_absolute_error(y_true, y_pred),
        "max_error": float(np.max(np.abs(y_true - y_pred))),
    }
Metric Interpretation
MetricRangeInterpretation
MAE[0, inf)Average absolute error in original units
RMSE[0, inf)Penalizes large errors more than MAE
R2(-inf, 1]1 = perfect; 0 = predicts mean; <0 = worse than mean
MAPE[0, inf)Percentage error (avoid when y has zeros)
Median AE[0, inf)Robust to outliers
Residual Analysis
python
from scipy import stats as sp_stats

def plot_residual_analysis(y_true, y_pred, figsize=(14, 10)):
    """Comprehensive residual analysis plots."""
    residuals = y_true - y_pred

    fig, axes = plt.subplots(2, 2, figsize=figsize)

    # Predicted vs Actual
    axes[0, 0].scatter(y_pred, y_true, alpha=0.5, s=10)
    min_val, max_val = min(y_true.min(), y_pred.min()), max(y_true.max(), y_pred.max())
    axes[0, 0].plot([min_val, max_val], [min_val, max_val], "r--")
    axes[0, 0].set_xlabel("Predicted")
    axes[0, 0].set_ylabel("Actual")
    # ... (condensed) ...
    sp_stats.probplot(residuals, dist="norm", plot=axes[1, 1])
    axes[1, 1].set_title("QQ Plot")

    plt.tight_layout()
    return fig

Cross-Validation Strategies

Choosing the Right Strategy
python
from sklearn.model_selection import (
    KFold, StratifiedKFold, TimeSeriesSplit,
    GroupKFold, RepeatedStratifiedKFold,
    cross_val_score,
)

def get_cv_strategy(
    task_type: str,
    data_type: str = "standard",
    n_splits: int = 5,
    groups=None,
):
    """Select appropriate cross-validation strategy."""

    # ... (condensed) ...

    if task_type == "classification":
        return StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=42)

    return KFold(n_splits=n_splits, shuffle=True, random_state=42)
Cross-Validation with Confidence Intervals
python
def cv_with_confidence(model, X, y, cv=5, scoring="f1") -> dict:
    """Cross-validation with confidence interval."""
    scores = cross_val_score(model, X, y, cv=cv, scoring=scoring)

    mean = scores.mean()
    std = scores.std()
    n = len(scores)
    se = std / np.sqrt(n)

    # 95% confidence interval
    ci_low = mean - 1.96 * se
    ci_high = mean + 1.96 * se

    return {
        "mean": mean,
        "std": std,
        "scores": scores.tolist(),
        "ci_95": (ci_low, ci_high),
        "n_folds": n,
    }

Statistical Model Comparison

Paired t-Test for Model Comparison
python
from scipy import stats

def compare_models_statistical(
    model_a_scores: list[float],
    model_b_scores: list[float],
    alpha: float = 0.05,
) -> dict:
    """Statistical comparison of two models using paired t-test."""

    t_stat, p_value = stats.ttest_rel(model_a_scores, model_b_scores)
    mean_diff = np.mean(model_a_scores) - np.mean(model_b_scores)

    return {
        "model_a_mean": np.mean(model_a_scores),
        "model_b_mean": np.mean(model_b_scores),
        "mean_difference": mean_diff,
        "t_statistic": t_stat,
        "p_value": p_value,
        "significant": p_value < alpha,
        "better_model": "A" if mean_diff > 0 else "B",
    }
McNemar's Test (for Classification)
python
def mcnemar_test(y_true, y_pred_a, y_pred_b) -> dict:
    """McNemar's test: are two classifiers significantly different?"""
    from statsmodels.stats.contingency_tables import mcnemar as mcnemar_fn

    correct_a = (y_pred_a == y_true)
    correct_b = (y_pred_b == y_true)

    n01 = ((~correct_a) & correct_b).sum()  # A wrong, B right
    n10 = (correct_a & (~correct_b)).sum()   # A right, B wrong

    table = [[0, n01], [n10, 0]]
    result = mcnemar_fn(table, exact=True)

    return {
        "a_right_b_wrong": int(n10),
        "a_wrong_b_right": int(n01),
        "p_value": result.pvalue,
        "significant": result.pvalue < 0.05,
    }

Bias and Fairness Assessment

Fairness Metrics
python
def compute_fairness_metrics(
    y_true: np.ndarray,
    y_pred: np.ndarray,
    sensitive_attr: np.ndarray,
    privileged_value=1,
    unprivileged_value=0,
) -> dict:
    """Compute fairness metrics across a sensitive attribute."""

    priv_mask = sensitive_attr == privileged_value
    unpriv_mask = sensitive_attr == unprivileged_value

    # Demographic parity
    rate_priv = y_pred[priv_mask].mean()
    # ... (condensed) ...
        "tpr_privileged": round(tpr_priv, 4),
        "tpr_unprivileged": round(tpr_unpriv, 4),
        "fpr_privileged": round(fpr_priv, 4),
        "fpr_unprivileged": round(fpr_unpriv, 4),
    }
Fairness Metric Definitions
MetricDefinitionFair When
Demographic ParitySelection rate ratio across groupsRatio between 0.8-1.25
Equal OpportunityTrue positive rate ratioRatio between 0.8-1.25
Equalized OddsTPR and FPR equal across groupsBoth ratios near 1.0
Predictive ParityPositive predictive value equalRatio between 0.8-1.25
CalibrationP(Y=1 given score=s) same across groupsCalibration curves overlap
Subgroup Analysis
python
import pandas as pd

def subgroup_performance(
    y_true, y_pred, y_prob,
    group_column: np.ndarray,
    group_names: dict,
) -> pd.DataFrame:
    """Compute performance metrics per subgroup."""
    results = []

    for group_val, group_name in group_names.items():
        mask = group_column == group_val
        if mask.sum() < 10:
            continue
# ... (condensed) ...
                metrics["auc_roc"] = float(roc_auc_score(y_true[mask], prob))

        results.append(metrics)

    return pd.DataFrame(results)

A/B Testing for Models

Online A/B Test Framework
python
import hashlib
from dataclasses import dataclass, field
from scipy import stats

@dataclass
class ModelABTest:
    """A/B test framework for comparing models in production."""
    name: str
    model_a_name: str
    model_b_name: str
    traffic_split: float = 0.5
    results_a: list = field(default_factory=list)
    results_b: list = field(default_factory=list)

    # ... (condensed) ...
            "effect_size": "small" if abs(cohens_d) < 0.5 else "medium" if abs(cohens_d) < 0.8 else "large",
            "recommendation": "B" if p_value < 0.05 and mean_b > mean_a else "A" if p_value < 0.05 else "continue_testing",
            "samples_a": len(self.results_a),
            "samples_b": len(self.results_b),
        }
Sample Size Planning
python
from scipy.stats import norm

def required_sample_size(
    baseline_metric: float,
    minimum_detectable_effect: float,
    alpha: float = 0.05,
    power: float = 0.8,
    metric_std: float = None,
) -> int:
    """Calculate required sample size per group for A/B test."""
    if metric_std is None:
        metric_std = np.sqrt(baseline_metric * (1 - baseline_metric))

    z_alpha = norm.ppf(1 - alpha / 2)
    z_beta = norm.ppf(power)

    n = (2 * metric_std**2 * (z_alpha + z_beta)**2) / minimum_detectable_effect**2
    return int(np.ceil(n))

Comprehensive Scoring Report

python
def generate_scoring_report(
    model_name: str,
    y_true, y_pred, y_prob=None,
    sensitive_attrs: dict = None,
) -> dict:
    """Generate comprehensive scoring report."""

    report = {
        "model": model_name,
        "dataset_size": len(y_true),
        "class_distribution": {int(k): int(v) for k, v in zip(*np.unique(y_true, return_counts=True))},
    }

    report["metrics"] = classification_report_full(y_true, y_pred, y_prob)
# ... (condensed) ...
            report["fairness"][attr_name] = compute_fairness_metrics(
                y_true, y_pred, attr_values
            )

    return report

Checklist

  • Select metrics appropriate to the task and class distribution
  • Report multiple metrics (never rely on accuracy alone)
  • Analyze the confusion matrix for systematic error patterns
  • Optimize the classification threshold if using probabilities
  • Use proper cross-validation (stratified for classification, temporal for time series)
  • Compute confidence intervals on performance estimates
  • Perform statistical tests when comparing models
  • Audit for bias across sensitive attributes (gender, race, age)
  • Check fairness metrics (demographic parity, equal opportunity)
  • Plan A/B tests with proper sample size calculations
  • Document all metrics, thresholds, and decisions in a scoring report
Show full SKILL.md (198 more words)Show less

When to Use

Use this skill when:

  • Designing or implementing model evaluator solutions
  • Reviewing or improving existing model evaluator approaches
  • Making architectural or implementation decisions about model evaluator
  • Learning model evaluator patterns and best practices
  • Troubleshooting model evaluator-related issues

Do NOT use this skill when:

  • The question is about a fundamentally different technology domain
  • A more specific sibling skill covers the exact topic needed
  • The user needs a complete hands-on tutorial rather than expert guidance

Output Format

markdown
# Model Evaluator Analysis

## Context Assessment
[Situation summary and constraints]

## Recommended Approach
[Primary recommendation with rationale]

## Implementation Steps
1. [Step with specific details]
2. [Step with specific details]
3. [Step with specific details]

## Trade-offs and Considerations
- [Key trade-off 1]
- [Key trade-off 2]

## Next Steps
- [Immediate action item]
- [Follow-up action item]

Example

Input: "Help me implement model evaluator for a medium-scale production application"

Output: A structured analysis covering current state assessment, recommended model evaluator approach with specific patterns, implementation roadmap with milestones, and risk mitigation strategies tailored to the application scale and constraints.

Edge Cases

  • Legacy system integration: When model evaluator must coexist with legacy approaches, provide a gradual migration path rather than a complete rewrite
  • Scale mismatch: When the solution complexity exceeds the project scale, recommend a simpler approach and note when to revisit
  • Team skill gaps: When the team lacks experience with the recommended approach, include learning resources and simpler alternatives
  • Conflicting requirements: When constraints conflict (e.g., performance vs. maintainability), explicitly state the trade-off and recommend based on stated priorities

© FerroxLabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in src/process/resources/skills-library/bodies/skills/ai-machine-learning/model-evaluator of FerroxLabs/wayland.

Open the folder on GitHubat commit 4c030c7

Compare with similar skills

Model Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Evaluator this skillFerroxLabs/wayland608—~4.1kAutomated safety check: PassApache-2.0
Senior Data ScientistRaidriar7170/hermes-skilleval1256 repos~1.4kAutomated safety check: PassMIT
Data Scientistborghei/Claude-Skills874—~3.3kAutomated safety check: PassMIT
Automl SkillLeoYeAI/openclaw-master-skills2.2k—~3.6kAutomated safety check: PassMIT
Senior Data Scientistborghei/Claude-Skills874—~1.7kAutomated safety check: PassMIT
Data Sciencemajiayu000/claude-skill-registry6661 repos~4.3kAutomated safety check: PassMIT

Similar skills

  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 6 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Data Scientist

    borghei/Claude-Skills

    Data science across machine learning, statistical modeling, and experimentation.

    874 GitHub stars~3.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Automl Skill

    LeoYeAI/openclaw-master-skills

    AutoML 自动化机器学习技能 | Automated Machine Learning Skill. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~3.6k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    borghei/Claude-Skills

    A skill your agent uses when the user asks to "design an experiment", "build a predictive model", "run A/B test analysis", "perform causal inference", "engineer features", "evaluate model…

    874 GitHub stars~1.7k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Data Science

    majiayu000/claude-skill-registry

    A skill your agent uses when performing exploratory data analysis, statistical testing, data visualization, or building predictive models.

    666 GitHub starsUsed in 1 repo~4.3k tokens
    Data & AnalyticsAuto-check passed
  • ML Experiment Evaluation

    hashgraph-online/awesome-codex-plugins

    Plan evaluation strategies for machine-learning product changes.

    1.2k GitHub stars~801 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from FerroxLabs/wayland

All 1,194 skills in this repo
  • Star Office Helper

    FerroxLabs/wayland

    Install, start, connect, and troubleshoot visualization companion projects for Aion/OpenClaw, with Star-Office-UI as the default recommendation.

    608 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check: notes
  • Openclaw Setup

    FerroxLabs/wayland

    OpenClaw usage expert: Helps you install, deploy, configure, and use OpenClaw personal AI assistant.

    608 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Tvcontrol Setup

    FerroxLabs/wayland

    Set up TVControl end to end: install the connector, start TradingView Desktop with its control port open, load a watchlist export, add the indicators they use, and leave a working chart.

    608 GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed
  • Ab Testing Specialist

    FerroxLabs/wayland

    End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.

    608 GitHub stars~3.7k tokensUpdated yesterday
    Auto-check passed
  • Academic Writer

    FerroxLabs/wayland

    Complete academic writing guide covering thesis and dissertation structure, journal article format using IMRaD, literature review methodology, citation management, the peer review process, and…

    608 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Accessibility Auditor

    FerroxLabs/wayland

    Web accessibility expertise covering WCAG 2.2 conformance, audit methodology, ARIA patterns, keyboard navigation, screen reader testing, focus management, form accessibility, and automated vs manual…

    608 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed

Questions about Model Evaluator

What does Model Evaluator do?

ML model assessment covering classification metrics (precision, recall, F1, AUC-ROC), regression metrics (MAE, RMSE, R2), confusion matrix analysis, cross-validation strategies, bias detection…. Model Evaluator is an agent skill from FerroxLabs/wayland. ML model assessment covering classification metrics (precision, recall, F1, AUC-ROC), regression metrics (MAE, RMSE, R2), confusion matrix analysis, cross-validation strategies, bias detection, fairness metrics, and A/B testing for models.

When should I use Model Evaluator?

Model Evaluator fits situations like: the user asks about model evaluator; model evaluator best practices; needs guidance on model evaluator implementation; the user needs a different specialized skill.

How do I install Model Evaluator in Claude Code?

Run `npx skills add FerroxLabs/wayland --skill model-evaluator -a claude-code`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/ai-machine-learning/model-evaluator in FerroxLabs/wayland) into .claude/skills/model-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Model Evaluator in Codex?

Run `npx skills add FerroxLabs/wayland --skill model-evaluator -a codex`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/ai-machine-learning/model-evaluator in FerroxLabs/wayland) into .agents/skills/model-evaluator in your project. Codex loads it when a task matches its description.

Can I use Model Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FerroxLabs/wayland --skill model-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-evaluator, .gemini/skills/model-evaluator, .github/skills/model-evaluator and .opencode/skills/model-evaluator in your project.

What does Model Evaluator need to run?

SKILL.md names no scripts, command-line tools or credentials: Model Evaluator is instructions for the agent only. Our summary lists: Python 3.

Does Model Evaluator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Model Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Model Evaluator use?

Model Evaluator is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Evaluator use?

About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Model Evaluator?

Skills that share tags, products or a category with Model Evaluator: Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Data Scientist (borghei/Claude-Skills, 874 stars), Automl Skill (LeoYeAI/openclaw-master-skills, 2.2k stars) and Senior Data Scientist (borghei/Claude-Skills, 874 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Evaluator?

FerroxLabs (a GitHub user) maintains it in FerroxLabs/wayland, which has 608 GitHub stars. The repository holds 1,194 skills in this directory. The repository was last updated on October 6, 2026.

Source: FerroxLabs/wayland on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.