Automated pipeline for retraining ML models with new construction data.

MITAuto-check passedData & Analytics

Install ML Model Retrainer

skills CLI
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill ml-model-retrainer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction ml-model-retrainer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .claude/skills && cp -r skills-src/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer .claude/skills/ml-model-retrainer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-model-retrainer
GitHub stars
345
Token cost
~4.6k tokens
SKILL.md length
64 words
Files
3
Skills in repo
36
Repo updated
First seen
Licence
MIT

At a glance

Automated pipeline for retraining ML models with new construction data.

  • Validate model performance
  • SKILL.md covers Overview, Business Case, Technical Implementation and Quick Start, plus 1 more section
  • Calls pip
  • Tasks that involve Machine learning

What it does

ML Model Retrainer is an agent skill from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction. Automated pipeline for retraining ML models with new construction data. Monitor model drift, trigger retraining, and validate model performance.

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `claw.json` and `instructions.md`).

It sits in Data & Analytics, covering Machine learning. The repository describes itself as: 221 AI skills for construction: BIM analysis, cost estimation, scheduling, document control, and automation with Claude Code. The licence is MIT.

When your agent uses it

  • Validate model performance
  • Tasks that involve Machine learning

Example prompts

  • “/ml-model-retrainer”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit ce45bbf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Model Retrainer loads about 4.6k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 64 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~41
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction at commit ce45bbf, republished under its MIT licence (© datadrivenconstruction). 64 words, ~4,601 tokens.

Download SKILL.mdSave it as .claude/skills/ml-model-retrainer/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
ml-model-retrainer
description
Automated pipeline for retraining ML models with new construction data. Monitor model drift, trigger retraining, and validate model performance.
homepage
https://datadrivenconstruction.io

ML Model Retrainer for Construction

Overview

Automated pipeline for keeping construction ML models up-to-date. Monitor for data drift, trigger retraining when needed, validate performance, and manage model versions.

Business Case

ML models degrade over time as:

  • Market conditions change (material prices, labor rates)
  • New construction methods emerge
  • Project complexity evolves
  • Regional factors shift

Continuous retraining ensures predictions remain accurate.

Technical Implementation

python
from dataclasses import dataclass, field
from typing import List, Dict, Any, Optional, Callable
from datetime import datetime, timedelta
import pandas as pd
import numpy as np
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
from sklearn.model_selection import cross_val_score
import pickle
import hashlib
import json
import os

@dataclass
class ModelVersion:
    version_id: str
    model_name: str
    created_at: datetime
    training_samples: int
    metrics: Dict[str, float]
    feature_columns: List[str]
    hyperparameters: Dict[str, Any]
    data_hash: str
    is_active: bool = False

@dataclass
class DriftReport:
    checked_at: datetime
    data_drift_detected: bool
    performance_drift_detected: bool
    drift_score: float
    affected_features: List[str]
    recommendation: str

@dataclass
class RetrainingResult:
    success: bool
    new_version: Optional[ModelVersion]
    old_metrics: Dict[str, float]
    new_metrics: Dict[str, float]
    improvement: Dict[str, float]
    validation_passed: bool
    notes: List[str]

class MLModelRetrainer:
    """Automated ML model retraining for construction predictions."""

    def __init__(self, model_dir: str = "./models"):
        self.model_dir = model_dir
        self.models: Dict[str, Any] = {}
        self.versions: Dict[str, List[ModelVersion]] = {}
        self.active_versions: Dict[str, ModelVersion] = {}
        self.drift_thresholds = {
            'performance_degradation': 0.15,  # 15% degradation triggers retrain
            'data_drift_score': 0.3,
            'min_new_samples': 50,
        }

        os.makedirs(model_dir, exist_ok=True)

    def register_model(self, model_name: str, model: Any,
                       feature_columns: List[str],
                       hyperparameters: Dict = None) -> ModelVersion:
        """Register a new model for management."""
        version = ModelVersion(
            version_id=f"{model_name}-v{datetime.now().strftime('%Y%m%d%H%M%S')}",
            model_name=model_name,
            created_at=datetime.now(),
            training_samples=0,
            metrics={},
            feature_columns=feature_columns,
            hyperparameters=hyperparameters or {},
            data_hash="",
            is_active=True
        )

        self.models[model_name] = model
        if model_name not in self.versions:
            self.versions[model_name] = []
        self.versions[model_name].append(version)
        self.active_versions[model_name] = version

        return version

    def calculate_data_hash(self, data: pd.DataFrame) -> str:
        """Calculate hash of training data for change detection."""
        data_str = data.to_json()
        return hashlib.md5(data_str.encode()).hexdigest()

    def detect_data_drift(self, model_name: str,
                          reference_data: pd.DataFrame,
                          current_data: pd.DataFrame) -> DriftReport:
        """Detect data drift between reference and current data."""

        drift_scores = {}
        affected_features = []

        # Compare distributions for each feature
        for col in reference_data.select_dtypes(include=[np.number]).columns:
            if col in current_data.columns:
                ref_mean = reference_data[col].mean()
                ref_std = reference_data[col].std()
                cur_mean = current_data[col].mean()
                cur_std = current_data[col].std()

                # Normalized difference
                if ref_std > 0:
                    mean_drift = abs(cur_mean - ref_mean) / ref_std
                    std_drift = abs(cur_std - ref_std) / ref_std
                    drift_scores[col] = (mean_drift + std_drift) / 2

                    if drift_scores[col] > 0.5:
                        affected_features.append(f"{col} (drift: {drift_scores[col]:.2f})")

        avg_drift = np.mean(list(drift_scores.values())) if drift_scores else 0
        data_drift_detected = avg_drift > self.drift_thresholds['data_drift_score']

        recommendation = "No action needed"
        if data_drift_detected:
            recommendation = "Data drift detected - consider retraining"
        elif avg_drift > self.drift_thresholds['data_drift_score'] * 0.7:
            recommendation = "Minor drift detected - monitor closely"

        return DriftReport(
            checked_at=datetime.now(),
            data_drift_detected=data_drift_detected,
            performance_drift_detected=False,
            drift_score=avg_drift,
            affected_features=affected_features,
            recommendation=recommendation
        )

    def evaluate_model_performance(self, model_name: str,
                                   test_data: pd.DataFrame,
                                   target_col: str) -> Dict[str, float]:
        """Evaluate current model performance on new data."""

        if model_name not in self.models:
            raise ValueError(f"Model {model_name} not registered")

        model = self.models[model_name]
        version = self.active_versions[model_name]

        # Prepare features
        X = test_data[version.feature_columns].fillna(0)
        y = test_data[target_col]

        # Predict
        y_pred = model.predict(X)

        # Calculate metrics
        metrics = {
            'mae': mean_absolute_error(y, y_pred),
            'rmse': np.sqrt(mean_squared_error(y, y_pred)),
            'r2': r2_score(y, y_pred),
            'mape': np.mean(np.abs((y - y_pred) / y.replace(0, 1))) * 100,
        }

        return metrics

    def check_performance_drift(self, model_name: str,
                                 baseline_metrics: Dict[str, float],
                                 current_metrics: Dict[str, float]) -> DriftReport:
        """Check if model performance has degraded."""

        # Calculate degradation for each metric
        degradation = {}
        for metric in ['mae', 'rmse']:
            if metric in baseline_metrics and metric in current_metrics:
                # Higher is worse for these metrics
                change = (current_metrics[metric] - baseline_metrics[metric]) / baseline_metrics[metric]
                degradation[metric] = change

        for metric in ['r2']:
            if metric in baseline_metrics and metric in current_metrics:
                # Lower is worse for R2
                change = (baseline_metrics[metric] - current_metrics[metric]) / abs(baseline_metrics[metric])
                degradation[metric] = change

        avg_degradation = np.mean(list(degradation.values())) if degradation else 0
        performance_drift = avg_degradation > self.drift_thresholds['performance_degradation']

        affected = [f"{m}: {d:+.1%}" for m, d in degradation.items() if d > 0.1]

        recommendation = "No action needed"
        if performance_drift:
            recommendation = "Performance degraded - retraining recommended"
        elif avg_degradation > self.drift_thresholds['performance_degradation'] * 0.5:
            recommendation = "Performance declining - monitor closely"

        return DriftReport(
            checked_at=datetime.now(),
            data_drift_detected=False,
            performance_drift_detected=performance_drift,
            drift_score=avg_degradation,
            affected_features=affected,
            recommendation=recommendation
        )

    def retrain_model(self, model_name: str,
                      training_data: pd.DataFrame,
                      target_col: str,
                      model_class: type,
                      hyperparameters: Dict = None,
                      validation_data: pd.DataFrame = None) -> RetrainingResult:
        """Retrain model with new data."""

        if model_name not in self.active_versions:
            raise ValueError(f"Model {model_name} not found")

        old_version = self.active_versions[model_name]
        old_metrics = old_version.metrics.copy()

        notes = []

        # Check minimum samples
        if len(training_data) < self.drift_thresholds['min_new_samples']:
            notes.append(f"Warning: Only {len(training_data)} samples (minimum: {self.drift_thresholds['min_new_samples']})")

        # Prepare data
        X = training_data[old_version.feature_columns].fillna(0)
        y = training_data[target_col]

        # Train new model
        hyperparams = hyperparameters or old_version.hyperparameters
        new_model = model_class(**hyperparams)
        new_model.fit(X, y)

        # Evaluate on validation data
        if validation_data is not None:
            X_val = validation_data[old_version.feature_columns].fillna(0)
            y_val = validation_data[target_col]
            y_pred = new_model.predict(X_val)

            new_metrics = {
                'mae': mean_absolute_error(y_val, y_pred),
                'rmse': np.sqrt(mean_squared_error(y_val, y_pred)),
                'r2': r2_score(y_val, y_pred),
            }
        else:
            # Cross-validation
            cv_scores = cross_val_score(new_model, X, y, cv=5, scoring='neg_mean_absolute_error')
            new_metrics = {
                'mae': -cv_scores.mean(),
                'mae_std': cv_scores.std(),
            }
            new_model.fit(X, y)  # Refit on full data

        # Calculate improvement
        improvement = {}
        for metric in new_metrics:
            if metric in old_metrics:
                if metric in ['mae', 'rmse']:
                    imp = (old_metrics[metric] - new_metrics[metric]) / old_metrics[metric]
                else:
                    imp = (new_metrics[metric] - old_metrics[metric]) / abs(old_metrics[metric])
                improvement[metric] = imp

        # Validation check
        validation_passed = True
        if 'mae' in improvement and improvement['mae'] < -0.1:
            validation_passed = False
            notes.append("New model performs worse - not deploying")

        if validation_passed:
            # Create new version
            new_version = ModelVersion(
                version_id=f"{model_name}-v{datetime.now().strftime('%Y%m%d%H%M%S')}",
                model_name=model_name,
                created_at=datetime.now(),
                training_samples=len(training_data),
                metrics=new_metrics,
                feature_columns=old_version.feature_columns,
                hyperparameters=hyperparams,
                data_hash=self.calculate_data_hash(training_data),
                is_active=True
            )

            # Deactivate old version
            old_version.is_active = False

            # Update registries
            self.models[model_name] = new_model
            self.versions[model_name].append(new_version)
            self.active_versions[model_name] = new_version

            notes.append(f"Model updated: {old_version.version_id} -> {new_version.version_id}")

            return RetrainingResult(
                success=True,
                new_version=new_version,
                old_metrics=old_metrics,
                new_metrics=new_metrics,
                improvement=improvement,
                validation_passed=True,
                notes=notes
            )
        else:
            return RetrainingResult(
                success=False,
                new_version=None,
                old_metrics=old_metrics,
                new_metrics=new_metrics,
                improvement=improvement,
                validation_passed=False,
                notes=notes
            )

    def save_model(self, model_name: str, path: str = None):
        """Save model to disk."""
        if path is None:
            version = self.active_versions[model_name]
            path = os.path.join(self.model_dir, f"{version.version_id}.pkl")

        model_data = {
            'model': self.models[model_name],
            'version': self.active_versions[model_name],
        }

        with open(path, 'wb') as f:
            pickle.dump(model_data, f)

        return path

    def load_model(self, path: str) -> str:
        """Load model from disk."""
        with open(path, 'rb') as f:
            model_data = pickle.load(f)

        model_name = model_data['version'].model_name
        self.models[model_name] = model_data['model']
        self.active_versions[model_name] = model_data['version']

        if model_name not in self.versions:
            self.versions[model_name] = []
        self.versions[model_name].append(model_data['version'])

        return model_name

    def get_model_history(self, model_name: str) -> pd.DataFrame:
        """Get version history for a model."""
        if model_name not in self.versions:
            return pd.DataFrame()

        history = []
        for v in self.versions[model_name]:
            history.append({
                'version_id': v.version_id,
                'created_at': v.created_at,
                'training_samples': v.training_samples,
                'mae': v.metrics.get('mae'),
                'r2': v.metrics.get('r2'),
                'is_active': v.is_active
            })

        return pd.DataFrame(history)

    def run_maintenance_check(self, model_name: str,
                               reference_data: pd.DataFrame,
                               current_data: pd.DataFrame,
                               target_col: str) -> Dict:
        """Run complete maintenance check for a model."""

        results = {
            'model_name': model_name,
            'checked_at': datetime.now(),
            'actions_needed': []
        }

        # Check data drift
        data_drift = self.detect_data_drift(model_name, reference_data, current_data)
        results['data_drift'] = {
            'detected': data_drift.data_drift_detected,
            'score': data_drift.drift_score,
            'affected_features': data_drift.affected_features
        }

        if data_drift.data_drift_detected:
            results['actions_needed'].append("Retrain due to data drift")

        # Check performance
        baseline_metrics = self.active_versions[model_name].metrics
        current_metrics = self.evaluate_model_performance(model_name, current_data, target_col)

        perf_drift = self.check_performance_drift(model_name, baseline_metrics, current_metrics)
        results['performance_drift'] = {
            'detected': perf_drift.performance_drift_detected,
            'score': perf_drift.drift_score,
            'metrics_affected': perf_drift.affected_features
        }

        if perf_drift.performance_drift_detected:
            results['actions_needed'].append("Retrain due to performance degradation")

        # Overall recommendation
        if results['actions_needed']:
            results['recommendation'] = "Retraining recommended"
        else:
            results['recommendation'] = "Model performing well - no action needed"

        return results

    def generate_report(self, model_name: str) -> str:
        """Generate model status report."""
        lines = [f"# Model Status Report: {model_name}", ""]
        lines.append(f"**Generated:** {datetime.now().strftime('%Y-%m-%d %H:%M')}")

        if model_name in self.active_versions:
            version = self.active_versions[model_name]
            lines.append("")
            lines.append("## Active Version")
            lines.append(f"- **Version:** {version.version_id}")
            lines.append(f"- **Created:** {version.created_at.strftime('%Y-%m-%d')}")
            lines.append(f"- **Training Samples:** {version.training_samples:,}")
            lines.append("")
            lines.append("## Performance Metrics")
            for metric, value in version.metrics.items():
                lines.append(f"- **{metric}:** {value:.4f}")

        # Version history
        lines.append("")
        lines.append("## Version History")
        history = self.get_model_history(model_name)
        if not history.empty:
            lines.append(history.to_markdown(index=False))

        return "\n".join(lines)

Quick Start

python
from sklearn.ensemble import GradientBoostingRegressor
import pandas as pd

# Load data
historical = pd.read_excel("historical_projects.xlsx")
new_data = pd.read_excel("recent_projects.xlsx")

# Initialize retrainer
retrainer = MLModelRetrainer("./models")

# Train initial model
features = ['gross_area', 'contract_value', 'planned_duration', 'num_subcontractors']
X = historical[features]
y = historical['delay_days']

model = GradientBoostingRegressor(n_estimators=100)
model.fit(X, y)

# Register model
version = retrainer.register_model(
    "schedule_delay",
    model,
    features,
    {'n_estimators': 100}
)

# Run maintenance check
check = retrainer.run_maintenance_check(
    "schedule_delay",
    historical,
    new_data,
    "delay_days"
)
print(f"Recommendation: {check['recommendation']}")

# Retrain if needed
if check['actions_needed']:
    result = retrainer.retrain_model(
        "schedule_delay",
        pd.concat([historical, new_data]),
        "delay_days",
        GradientBoostingRegressor,
        {'n_estimators': 100}
    )
    print(f"Retraining {'successful' if result.success else 'failed'}")
    for note in result.notes:
        print(f"  - {note}")

# Save model
retrainer.save_model("schedule_delay")

# Generate report
print(retrainer.generate_report("schedule_delay"))

Dependencies

bash
pip install pandas numpy scikit-learn

© datadrivenconstruction, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in 2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer of datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.

  • SKILL.md
  • claw.json
  • instructions.md

Open the folder on GitHubat commit ce45bbf

Compare with similar skills

ML Model Retrainer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Model Retrainer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Model Retrainer this skilldatadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction345—~4.6kAutomated safety check: PassMIT
Scikit LearnzLanqing/codex-claude-academic-skills4.7k16 repos~3.9kAutomated safety check: PassBSD-3-Clause
Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill188—~4kAutomated safety check: PassMIT
Senior Data ScientistRaidriar7170/hermes-skilleval1255 repos~1.4kAutomated safety check: PassMIT
Geomlitalo-goncalves/geoML109—~4.6kAutomated safety check: PassGPL-3.0
QuantMind Training Config Generatorqusong0627/QuantMind1.7k—~1.5kAutomated safety check: PassAGPL-3.0

Similar skills

  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 16 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Agentic Kaggle Workflow

    FrankS-IntelLab/agentic-kaggle-skill

    Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

    188 GitHub stars~4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 5 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Geoml

    italo-goncalves/geoML

    Working knowledge of the geoML Python package (github.com/italo-goncalves/geoML): variational Gaussian processes for spatial data, implicit geological modelling, block models, drillhole data…

    109 GitHub stars~4.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Turns a plain-language model training request into a validated QuantMind training config file that can be imported from the Model Training page.

    1.7k GitHub stars~1.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Retention Analysis

    liangdabiao/claude-data-analysis-ultra-main

    Analyze user retention and churn using survival analysis, cohort analysis, and machine learning.

    290 GitHub stars~1.3k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction

All 36 skills in this repo
  • AI Agent Orchestration

    datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction

    Orchestrate multiple AI agents for construction workflows: estimator, scheduler, document, QA and safety agents coordinated by a supervisor agent, with human checkpoints.

    345 GitHub stars~679 tokensUpdated 1 mo ago
    Auto-check passed
  • Embodied Carbon Esg

    datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction

    Estimate embodied carbon and produce ESG/climate reporting for construction: LCA per work item, material-based carbon factors, EU taxonomy and CSRD alignment.

    345 GitHub stars~664 tokensUpdated 1 mo ago
    Auto-check passed
  • Generative AI Design

    datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction

    Generative design for construction: text-to-BIM concepts, option generation, and AI-assisted design iteration with cost and carbon feedback.

    345 GitHub stars~593 tokensUpdated 1 mo ago
    Auto-check passed
  • Material Passports Circular

    datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction

    Material passports and circular construction: generate per-element material inventories from BOQ/BIM, mark reuse potential and recycled content, and prepare deconstruction data.

    345 GitHub stars~634 tokensUpdated 1 mo ago
    Auto-check passed
  • Oce Cost Browser

    datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction

    Browse and search the OpenConstructionERP cost database: classification tree, SQL and semantic search, autocomplete, certainty badges, and the resource catalog.

    345 GitHub stars~637 tokensUpdated 1 mo ago
    Auto-check passed
  • Oce Estimate Boq

    datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction

    Create bills of quantities and estimates in OpenConstructionERP: search cost items, build BOQ sections, link BIM elements in bulk, validate the BOQ, and export GAEB/XLSX/JSON.

    345 GitHub stars~763 tokensUpdated 1 mo ago
    Auto-check passed

Questions about ML Model Retrainer

What does ML Model Retrainer do?

Automated pipeline for retraining ML models with new construction data. ML Model Retrainer is an agent skill from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction. Automated pipeline for retraining ML models with new construction data.

When should I use ML Model Retrainer?

ML Model Retrainer fits situations like: validate model performance; tasks that involve Machine learning.

How do I install ML Model Retrainer in Claude Code?

Run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill ml-model-retrainer -a claude-code`. Or copy the skill folder (2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction) into .claude/skills/ml-model-retrainer in your project. Claude Code loads it when a task matches its description.

How do I install ML Model Retrainer in Codex?

Run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill ml-model-retrainer -a codex`. Or copy the skill folder (2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction) into .agents/skills/ml-model-retrainer in your project. Codex loads it when a task matches its description.

Can I use ML Model Retrainer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill ml-model-retrainer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-model-retrainer, .gemini/skills/ml-model-retrainer, .github/skills/ml-model-retrainer and .opencode/skills/ml-model-retrainer in your project.

What does ML Model Retrainer need to run?

Going by SKILL.md and its folder, ML Model Retrainer needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does ML Model Retrainer access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is ML Model Retrainer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does ML Model Retrainer use?

ML Model Retrainer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Model Retrainer use?

About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to ML Model Retrainer?

Skills that share tags, products or a category with ML Model Retrainer: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.7k stars), Agentic Kaggle Workflow (FrankS-IntelLab/agentic-kaggle-skill, 188 stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars) and Geoml (italo-goncalves/geoML, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Model Retrainer?

datadrivenconstruction (a GitHub user) maintains it in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction, which has 345 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on August 22, 2026.

Source: datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.