Scikit Learn
zLanqing/codex-claude-academic-skills
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
Agent skill
by datadrivenconstruction in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Automated pipeline for retraining ML models with new construction data.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill ml-model-retrainer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction ml-model-retrainer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .claude/skills && cp -r skills-src/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer .claude/skills/ml-model-retrainer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ml-model-retrainer" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer into .claude/skills/ml-model-retrainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-model-retrainer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill ml-model-retrainer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction ml-model-retrainer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .agents/skills && cp -r skills-src/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer .agents/skills/ml-model-retrainer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ml-model-retrainer" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer into .agents/skills/ml-model-retrainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-model-retrainer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill ml-model-retrainer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction ml-model-retrainer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer .cursor/skills/ml-model-retrainer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ml-model-retrainer" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer into .cursor/skills/ml-model-retrainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-model-retrainer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git --path 2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill ml-model-retrainer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction ml-model-retrainer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer .gemini/skills/ml-model-retrainer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ml-model-retrainer" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer into .gemini/skills/ml-model-retrainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-model-retrainer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction ml-model-retrainerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill ml-model-retrainer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .github/skills && cp -r skills-src/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer .github/skills/ml-model-retrainer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ml-model-retrainer" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer into .github/skills/ml-model-retrainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-model-retrainer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill ml-model-retrainer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction ml-model-retrainer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer .opencode/skills/ml-model-retrainer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ml-model-retrainer" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer into .opencode/skills/ml-model-retrainer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-model-retrainer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ml-model-retrainerAutomated pipeline for retraining ML models with new construction data.
ML Model Retrainer is an agent skill from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction. Automated pipeline for retraining ML models with new construction data. Monitor model drift, trigger retraining, and validate model performance.
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `claw.json` and `instructions.md`).
It sits in Data & Analytics, covering Machine learning. The repository describes itself as: 221 AI skills for construction: BIM analysis, cost estimation, scheduling, document control, and automation with Claude Code. The licence is MIT.
Read from SKILL.md and the folder at commit ce45bbf. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
ML Model Retrainer loads about 4.6k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 64 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction at commit ce45bbf, republished under its MIT licence (© datadrivenconstruction). 64 words, ~4,601 tokens.
.claude/skills/ml-model-retrainer/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Automated pipeline for keeping construction ML models up-to-date. Monitor for data drift, trigger retraining when needed, validate performance, and manage model versions.
ML models degrade over time as:
Continuous retraining ensures predictions remain accurate.
from dataclasses import dataclass, field
from typing import List, Dict, Any, Optional, Callable
from datetime import datetime, timedelta
import pandas as pd
import numpy as np
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
from sklearn.model_selection import cross_val_score
import pickle
import hashlib
import json
import os
@dataclass
class ModelVersion:
version_id: str
model_name: str
created_at: datetime
training_samples: int
metrics: Dict[str, float]
feature_columns: List[str]
hyperparameters: Dict[str, Any]
data_hash: str
is_active: bool = False
@dataclass
class DriftReport:
checked_at: datetime
data_drift_detected: bool
performance_drift_detected: bool
drift_score: float
affected_features: List[str]
recommendation: str
@dataclass
class RetrainingResult:
success: bool
new_version: Optional[ModelVersion]
old_metrics: Dict[str, float]
new_metrics: Dict[str, float]
improvement: Dict[str, float]
validation_passed: bool
notes: List[str]
class MLModelRetrainer:
"""Automated ML model retraining for construction predictions."""
def __init__(self, model_dir: str = "./models"):
self.model_dir = model_dir
self.models: Dict[str, Any] = {}
self.versions: Dict[str, List[ModelVersion]] = {}
self.active_versions: Dict[str, ModelVersion] = {}
self.drift_thresholds = {
'performance_degradation': 0.15, # 15% degradation triggers retrain
'data_drift_score': 0.3,
'min_new_samples': 50,
}
os.makedirs(model_dir, exist_ok=True)
def register_model(self, model_name: str, model: Any,
feature_columns: List[str],
hyperparameters: Dict = None) -> ModelVersion:
"""Register a new model for management."""
version = ModelVersion(
version_id=f"{model_name}-v{datetime.now().strftime('%Y%m%d%H%M%S')}",
model_name=model_name,
created_at=datetime.now(),
training_samples=0,
metrics={},
feature_columns=feature_columns,
hyperparameters=hyperparameters or {},
data_hash="",
is_active=True
)
self.models[model_name] = model
if model_name not in self.versions:
self.versions[model_name] = []
self.versions[model_name].append(version)
self.active_versions[model_name] = version
return version
def calculate_data_hash(self, data: pd.DataFrame) -> str:
"""Calculate hash of training data for change detection."""
data_str = data.to_json()
return hashlib.md5(data_str.encode()).hexdigest()
def detect_data_drift(self, model_name: str,
reference_data: pd.DataFrame,
current_data: pd.DataFrame) -> DriftReport:
"""Detect data drift between reference and current data."""
drift_scores = {}
affected_features = []
# Compare distributions for each feature
for col in reference_data.select_dtypes(include=[np.number]).columns:
if col in current_data.columns:
ref_mean = reference_data[col].mean()
ref_std = reference_data[col].std()
cur_mean = current_data[col].mean()
cur_std = current_data[col].std()
# Normalized difference
if ref_std > 0:
mean_drift = abs(cur_mean - ref_mean) / ref_std
std_drift = abs(cur_std - ref_std) / ref_std
drift_scores[col] = (mean_drift + std_drift) / 2
if drift_scores[col] > 0.5:
affected_features.append(f"{col} (drift: {drift_scores[col]:.2f})")
avg_drift = np.mean(list(drift_scores.values())) if drift_scores else 0
data_drift_detected = avg_drift > self.drift_thresholds['data_drift_score']
recommendation = "No action needed"
if data_drift_detected:
recommendation = "Data drift detected - consider retraining"
elif avg_drift > self.drift_thresholds['data_drift_score'] * 0.7:
recommendation = "Minor drift detected - monitor closely"
return DriftReport(
checked_at=datetime.now(),
data_drift_detected=data_drift_detected,
performance_drift_detected=False,
drift_score=avg_drift,
affected_features=affected_features,
recommendation=recommendation
)
def evaluate_model_performance(self, model_name: str,
test_data: pd.DataFrame,
target_col: str) -> Dict[str, float]:
"""Evaluate current model performance on new data."""
if model_name not in self.models:
raise ValueError(f"Model {model_name} not registered")
model = self.models[model_name]
version = self.active_versions[model_name]
# Prepare features
X = test_data[version.feature_columns].fillna(0)
y = test_data[target_col]
# Predict
y_pred = model.predict(X)
# Calculate metrics
metrics = {
'mae': mean_absolute_error(y, y_pred),
'rmse': np.sqrt(mean_squared_error(y, y_pred)),
'r2': r2_score(y, y_pred),
'mape': np.mean(np.abs((y - y_pred) / y.replace(0, 1))) * 100,
}
return metrics
def check_performance_drift(self, model_name: str,
baseline_metrics: Dict[str, float],
current_metrics: Dict[str, float]) -> DriftReport:
"""Check if model performance has degraded."""
# Calculate degradation for each metric
degradation = {}
for metric in ['mae', 'rmse']:
if metric in baseline_metrics and metric in current_metrics:
# Higher is worse for these metrics
change = (current_metrics[metric] - baseline_metrics[metric]) / baseline_metrics[metric]
degradation[metric] = change
for metric in ['r2']:
if metric in baseline_metrics and metric in current_metrics:
# Lower is worse for R2
change = (baseline_metrics[metric] - current_metrics[metric]) / abs(baseline_metrics[metric])
degradation[metric] = change
avg_degradation = np.mean(list(degradation.values())) if degradation else 0
performance_drift = avg_degradation > self.drift_thresholds['performance_degradation']
affected = [f"{m}: {d:+.1%}" for m, d in degradation.items() if d > 0.1]
recommendation = "No action needed"
if performance_drift:
recommendation = "Performance degraded - retraining recommended"
elif avg_degradation > self.drift_thresholds['performance_degradation'] * 0.5:
recommendation = "Performance declining - monitor closely"
return DriftReport(
checked_at=datetime.now(),
data_drift_detected=False,
performance_drift_detected=performance_drift,
drift_score=avg_degradation,
affected_features=affected,
recommendation=recommendation
)
def retrain_model(self, model_name: str,
training_data: pd.DataFrame,
target_col: str,
model_class: type,
hyperparameters: Dict = None,
validation_data: pd.DataFrame = None) -> RetrainingResult:
"""Retrain model with new data."""
if model_name not in self.active_versions:
raise ValueError(f"Model {model_name} not found")
old_version = self.active_versions[model_name]
old_metrics = old_version.metrics.copy()
notes = []
# Check minimum samples
if len(training_data) < self.drift_thresholds['min_new_samples']:
notes.append(f"Warning: Only {len(training_data)} samples (minimum: {self.drift_thresholds['min_new_samples']})")
# Prepare data
X = training_data[old_version.feature_columns].fillna(0)
y = training_data[target_col]
# Train new model
hyperparams = hyperparameters or old_version.hyperparameters
new_model = model_class(**hyperparams)
new_model.fit(X, y)
# Evaluate on validation data
if validation_data is not None:
X_val = validation_data[old_version.feature_columns].fillna(0)
y_val = validation_data[target_col]
y_pred = new_model.predict(X_val)
new_metrics = {
'mae': mean_absolute_error(y_val, y_pred),
'rmse': np.sqrt(mean_squared_error(y_val, y_pred)),
'r2': r2_score(y_val, y_pred),
}
else:
# Cross-validation
cv_scores = cross_val_score(new_model, X, y, cv=5, scoring='neg_mean_absolute_error')
new_metrics = {
'mae': -cv_scores.mean(),
'mae_std': cv_scores.std(),
}
new_model.fit(X, y) # Refit on full data
# Calculate improvement
improvement = {}
for metric in new_metrics:
if metric in old_metrics:
if metric in ['mae', 'rmse']:
imp = (old_metrics[metric] - new_metrics[metric]) / old_metrics[metric]
else:
imp = (new_metrics[metric] - old_metrics[metric]) / abs(old_metrics[metric])
improvement[metric] = imp
# Validation check
validation_passed = True
if 'mae' in improvement and improvement['mae'] < -0.1:
validation_passed = False
notes.append("New model performs worse - not deploying")
if validation_passed:
# Create new version
new_version = ModelVersion(
version_id=f"{model_name}-v{datetime.now().strftime('%Y%m%d%H%M%S')}",
model_name=model_name,
created_at=datetime.now(),
training_samples=len(training_data),
metrics=new_metrics,
feature_columns=old_version.feature_columns,
hyperparameters=hyperparams,
data_hash=self.calculate_data_hash(training_data),
is_active=True
)
# Deactivate old version
old_version.is_active = False
# Update registries
self.models[model_name] = new_model
self.versions[model_name].append(new_version)
self.active_versions[model_name] = new_version
notes.append(f"Model updated: {old_version.version_id} -> {new_version.version_id}")
return RetrainingResult(
success=True,
new_version=new_version,
old_metrics=old_metrics,
new_metrics=new_metrics,
improvement=improvement,
validation_passed=True,
notes=notes
)
else:
return RetrainingResult(
success=False,
new_version=None,
old_metrics=old_metrics,
new_metrics=new_metrics,
improvement=improvement,
validation_passed=False,
notes=notes
)
def save_model(self, model_name: str, path: str = None):
"""Save model to disk."""
if path is None:
version = self.active_versions[model_name]
path = os.path.join(self.model_dir, f"{version.version_id}.pkl")
model_data = {
'model': self.models[model_name],
'version': self.active_versions[model_name],
}
with open(path, 'wb') as f:
pickle.dump(model_data, f)
return path
def load_model(self, path: str) -> str:
"""Load model from disk."""
with open(path, 'rb') as f:
model_data = pickle.load(f)
model_name = model_data['version'].model_name
self.models[model_name] = model_data['model']
self.active_versions[model_name] = model_data['version']
if model_name not in self.versions:
self.versions[model_name] = []
self.versions[model_name].append(model_data['version'])
return model_name
def get_model_history(self, model_name: str) -> pd.DataFrame:
"""Get version history for a model."""
if model_name not in self.versions:
return pd.DataFrame()
history = []
for v in self.versions[model_name]:
history.append({
'version_id': v.version_id,
'created_at': v.created_at,
'training_samples': v.training_samples,
'mae': v.metrics.get('mae'),
'r2': v.metrics.get('r2'),
'is_active': v.is_active
})
return pd.DataFrame(history)
def run_maintenance_check(self, model_name: str,
reference_data: pd.DataFrame,
current_data: pd.DataFrame,
target_col: str) -> Dict:
"""Run complete maintenance check for a model."""
results = {
'model_name': model_name,
'checked_at': datetime.now(),
'actions_needed': []
}
# Check data drift
data_drift = self.detect_data_drift(model_name, reference_data, current_data)
results['data_drift'] = {
'detected': data_drift.data_drift_detected,
'score': data_drift.drift_score,
'affected_features': data_drift.affected_features
}
if data_drift.data_drift_detected:
results['actions_needed'].append("Retrain due to data drift")
# Check performance
baseline_metrics = self.active_versions[model_name].metrics
current_metrics = self.evaluate_model_performance(model_name, current_data, target_col)
perf_drift = self.check_performance_drift(model_name, baseline_metrics, current_metrics)
results['performance_drift'] = {
'detected': perf_drift.performance_drift_detected,
'score': perf_drift.drift_score,
'metrics_affected': perf_drift.affected_features
}
if perf_drift.performance_drift_detected:
results['actions_needed'].append("Retrain due to performance degradation")
# Overall recommendation
if results['actions_needed']:
results['recommendation'] = "Retraining recommended"
else:
results['recommendation'] = "Model performing well - no action needed"
return results
def generate_report(self, model_name: str) -> str:
"""Generate model status report."""
lines = [f"# Model Status Report: {model_name}", ""]
lines.append(f"**Generated:** {datetime.now().strftime('%Y-%m-%d %H:%M')}")
if model_name in self.active_versions:
version = self.active_versions[model_name]
lines.append("")
lines.append("## Active Version")
lines.append(f"- **Version:** {version.version_id}")
lines.append(f"- **Created:** {version.created_at.strftime('%Y-%m-%d')}")
lines.append(f"- **Training Samples:** {version.training_samples:,}")
lines.append("")
lines.append("## Performance Metrics")
for metric, value in version.metrics.items():
lines.append(f"- **{metric}:** {value:.4f}")
# Version history
lines.append("")
lines.append("## Version History")
history = self.get_model_history(model_name)
if not history.empty:
lines.append(history.to_markdown(index=False))
return "\n".join(lines)from sklearn.ensemble import GradientBoostingRegressor
import pandas as pd
# Load data
historical = pd.read_excel("historical_projects.xlsx")
new_data = pd.read_excel("recent_projects.xlsx")
# Initialize retrainer
retrainer = MLModelRetrainer("./models")
# Train initial model
features = ['gross_area', 'contract_value', 'planned_duration', 'num_subcontractors']
X = historical[features]
y = historical['delay_days']
model = GradientBoostingRegressor(n_estimators=100)
model.fit(X, y)
# Register model
version = retrainer.register_model(
"schedule_delay",
model,
features,
{'n_estimators': 100}
)
# Run maintenance check
check = retrainer.run_maintenance_check(
"schedule_delay",
historical,
new_data,
"delay_days"
)
print(f"Recommendation: {check['recommendation']}")
# Retrain if needed
if check['actions_needed']:
result = retrainer.retrain_model(
"schedule_delay",
pd.concat([historical, new_data]),
"delay_days",
GradientBoostingRegressor,
{'n_estimators': 100}
)
print(f"Retraining {'successful' if result.success else 'failed'}")
for note in result.notes:
print(f" - {note}")
# Save model
retrainer.save_model("schedule_delay")
# Generate report
print(retrainer.generate_report("schedule_delay"))pip install pandas numpy scikit-learn© datadrivenconstruction, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in 2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer of datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.
Open the folder on GitHubat commit ce45bbf
ML Model Retrainer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| ML Model Retrainer this skilldatadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction | 345 | — | ~4.6k | Automated safety check: Pass | MIT | |
| Scikit LearnzLanqing/codex-claude-academic-skills | 4.7k | 16 repos | ~3.9k | Automated safety check: Pass | BSD-3-Clause | |
| Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill | 188 | — | ~4k | Automated safety check: Pass | MIT | |
| Senior Data ScientistRaidriar7170/hermes-skilleval | 125 | 5 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Geomlitalo-goncalves/geoML | 109 | — | ~4.6k | Automated safety check: Pass | GPL-3.0 | |
| QuantMind Training Config Generatorqusong0627/QuantMind | 1.7k | — | ~1.5k | Automated safety check: Pass | AGPL-3.0 |
zLanqing/codex-claude-academic-skills
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
FrankS-IntelLab/agentic-kaggle-skill
Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.
Raidriar7170/hermes-skilleval
World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.
italo-goncalves/geoML
Working knowledge of the geoML Python package (github.com/italo-goncalves/geoML): variational Gaussian processes for spatial data, implicit geological modelling, block models, drillhole data…
qusong0627/QuantMind
Turns a plain-language model training request into a validated QuantMind training config file that can be imported from the Model Training page.
liangdabiao/claude-data-analysis-ultra-main
Analyze user retention and churn using survival analysis, cohort analysis, and machine learning.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Orchestrate multiple AI agents for construction workflows: estimator, scheduler, document, QA and safety agents coordinated by a supervisor agent, with human checkpoints.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Estimate embodied carbon and produce ESG/climate reporting for construction: LCA per work item, material-based carbon factors, EU taxonomy and CSRD alignment.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Generative design for construction: text-to-BIM concepts, option generation, and AI-assisted design iteration with cost and carbon feedback.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Material passports and circular construction: generate per-element material inventories from BOQ/BIM, mark reuse potential and recycled content, and prepare deconstruction data.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Browse and search the OpenConstructionERP cost database: classification tree, SQL and semantic search, autocomplete, certainty badges, and the resource catalog.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Create bills of quantities and estimates in OpenConstructionERP: search cost items, build BOQ sections, link BIM elements in bulk, validate the BOQ, and export GAEB/XLSX/JSON.
Categories
Automated pipeline for retraining ML models with new construction data. ML Model Retrainer is an agent skill from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction. Automated pipeline for retraining ML models with new construction data.
ML Model Retrainer fits situations like: validate model performance; tasks that involve Machine learning.
Run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill ml-model-retrainer -a claude-code`. Or copy the skill folder (2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction) into .claude/skills/ml-model-retrainer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill ml-model-retrainer -a codex`. Or copy the skill folder (2_DDC_Book/4.5-ML-Cost-Prediction/ml-model-retrainer in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction) into .agents/skills/ml-model-retrainer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill ml-model-retrainer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-model-retrainer, .gemini/skills/ml-model-retrainer, .github/skills/ml-model-retrainer and .opencode/skills/ml-model-retrainer in your project.
Going by SKILL.md and its folder, ML Model Retrainer needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
ML Model Retrainer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with ML Model Retrainer: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.7k stars), Agentic Kaggle Workflow (FrankS-IntelLab/agentic-kaggle-skill, 188 stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars) and Geoml (italo-goncalves/geoML, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
datadrivenconstruction (a GitHub user) maintains it in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction, which has 345 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on August 22, 2026.
Source: datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.