Statistical Analysis
majiayu000/claude-skill-registry
Apply statistical methods including descriptive stats, trend analysis, outlier detection, and hypothesis testing.
Agent skill
by datadrivenconstruction in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Detect anomalies and outliers in construction data: unusual costs, schedule variances, productivity spikes.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-anomaly-detector -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-anomaly-detector --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .claude/skills && cp -r skills-src/2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector .claude/skills/data-anomaly-detector && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-anomaly-detector" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector into .claude/skills/data-anomaly-detector/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-anomaly-detector", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detectorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-anomaly-detector -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-anomaly-detector --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .agents/skills && cp -r skills-src/2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector .agents/skills/data-anomaly-detector && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-anomaly-detector" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector into .agents/skills/data-anomaly-detector/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-anomaly-detector", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-anomaly-detector -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-anomaly-detector --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector .cursor/skills/data-anomaly-detector && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-anomaly-detector" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector into .cursor/skills/data-anomaly-detector/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-anomaly-detector", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git --path 2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-anomaly-detector -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-anomaly-detector --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector .gemini/skills/data-anomaly-detector && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-anomaly-detector" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector into .gemini/skills/data-anomaly-detector/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-anomaly-detector", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-anomaly-detectorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-anomaly-detector -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .github/skills && cp -r skills-src/2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector .github/skills/data-anomaly-detector && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-anomaly-detector" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector into .github/skills/data-anomaly-detector/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-anomaly-detector", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-anomaly-detector -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-anomaly-detector --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector .opencode/skills/data-anomaly-detector && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-anomaly-detector" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector into .opencode/skills/data-anomaly-detector/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-anomaly-detector", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-anomaly-detectorDetect anomalies and outliers in construction data: unusual costs, schedule variances, productivity spikes.
Data Anomaly Detector is an agent skill from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction. Detect anomalies and outliers in construction data: unusual costs, schedule variances, productivity spikes. Statistical and ML-based detection methods.
Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `claw.json` and `instructions.md`).
It sits in Data & Analytics, covering Anomaly detection and Data cleaning. The repository describes itself as: 221 AI skills for construction: BIM analysis, cost estimation, scheduling, document control, and automation with Claude Code. The licence is MIT.
Read from SKILL.md and the folder at commit ce45bbf. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Anomaly Detector loads about 4.8k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 81 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction at commit ce45bbf, republished under its MIT licence (© datadrivenconstruction). 81 words, ~4,757 tokens.
.claude/skills/data-anomaly-detector/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Detect unusual patterns, outliers, and anomalies in construction data. Identify cost overruns, schedule delays, productivity issues, and data quality problems before they impact projects.
Construction data often contains anomalies that indicate:
Early detection prevents costly corrections and project delays.
from dataclasses import dataclass, field
from typing import List, Dict, Any, Optional, Tuple
from enum import Enum
import pandas as pd
import numpy as np
from datetime import datetime
from scipy import stats
class AnomalyType(Enum):
OUTLIER = "outlier"
PATTERN_BREAK = "pattern_break"
MISSING_SEQUENCE = "missing_sequence"
DUPLICATE = "duplicate"
IMPOSSIBLE_VALUE = "impossible_value"
TREND_DEVIATION = "trend_deviation"
class AnomalySeverity(Enum):
CRITICAL = "critical"
HIGH = "high"
MEDIUM = "medium"
LOW = "low"
@dataclass
class Anomaly:
id: str
anomaly_type: AnomalyType
severity: AnomalySeverity
field: str
value: Any
expected_range: Optional[Tuple[float, float]] = None
description: str = ""
row_index: Optional[int] = None
detection_method: str = ""
confidence: float = 0.0
suggested_action: str = ""
@dataclass
class AnomalyReport:
source: str
detected_at: datetime
total_records: int
anomalies: List[Anomaly]
summary: Dict[str, int]
class ConstructionAnomalyDetector:
"""Detect anomalies in construction data."""
# Construction-specific thresholds
COST_THRESHOLDS = {
'concrete_per_cy': (200, 800),
'steel_per_ton': (1500, 4000),
'labor_per_hour': (25, 150),
'overhead_percentage': (5, 25),
'contingency_percentage': (3, 20),
}
SCHEDULE_THRESHOLDS = {
'max_activity_duration': 365, # days
'max_lag': 30, # days
'min_productivity': 0.1,
'max_productivity': 10.0,
}
def __init__(self):
self.anomalies: List[Anomaly] = []
self.detection_history: List[AnomalyReport] = []
def detect_cost_anomalies(self, df: pd.DataFrame, cost_column: str,
group_by: str = None) -> List[Anomaly]:
"""Detect anomalies in cost data."""
anomalies = []
# Statistical outlier detection (IQR method)
Q1 = df[cost_column].quantile(0.25)
Q3 = df[cost_column].quantile(0.75)
IQR = Q3 - Q1
lower_bound = Q1 - 1.5 * IQR
upper_bound = Q3 + 1.5 * IQR
outliers = df[(df[cost_column] < lower_bound) | (df[cost_column] > upper_bound)]
for idx, row in outliers.iterrows():
value = row[cost_column]
severity = AnomalySeverity.HIGH if abs(value - df[cost_column].median()) > 3 * IQR else AnomalySeverity.MEDIUM
anomalies.append(Anomaly(
id=f"COST-{idx}",
anomaly_type=AnomalyType.OUTLIER,
severity=severity,
field=cost_column,
value=value,
expected_range=(lower_bound, upper_bound),
description=f"Cost value {value:,.2f} outside expected range",
row_index=idx,
detection_method="IQR",
confidence=0.95,
suggested_action="Review cost estimate for errors"
))
# Negative cost check
negatives = df[df[cost_column] < 0]
for idx, row in negatives.iterrows():
anomalies.append(Anomaly(
id=f"COST-NEG-{idx}",
anomaly_type=AnomalyType.IMPOSSIBLE_VALUE,
severity=AnomalySeverity.CRITICAL,
field=cost_column,
value=row[cost_column],
expected_range=(0, None),
description="Negative cost value detected",
row_index=idx,
detection_method="Business Rule",
confidence=1.0,
suggested_action="Correct data entry error or investigate credit"
))
# Group-based anomalies (if grouped)
if group_by and group_by in df.columns:
group_stats = df.groupby(group_by)[cost_column].agg(['mean', 'std'])
for group_name, stats in group_stats.iterrows():
group_data = df[df[group_by] == group_name]
z_scores = np.abs((group_data[cost_column] - stats['mean']) / stats['std'])
for idx, z in z_scores.items():
if z > 3:
anomalies.append(Anomaly(
id=f"COST-GROUP-{idx}",
anomaly_type=AnomalyType.OUTLIER,
severity=AnomalySeverity.MEDIUM,
field=cost_column,
value=df.loc[idx, cost_column],
description=f"Unusual cost for group {group_name} (z-score: {z:.2f})",
row_index=idx,
detection_method="Z-Score by Group",
confidence=min(z / 5, 1.0)
))
return anomalies
def detect_schedule_anomalies(self, df: pd.DataFrame) -> List[Anomaly]:
"""Detect anomalies in schedule data."""
anomalies = []
# Check for required columns
required = ['start_date', 'end_date']
if not all(col in df.columns for col in required):
return anomalies
# Convert dates
df['start_date'] = pd.to_datetime(df['start_date'])
df['end_date'] = pd.to_datetime(df['end_date'])
# Calculate duration
df['duration'] = (df['end_date'] - df['start_date']).dt.days
# Negative duration (end before start)
negative_duration = df[df['duration'] < 0]
for idx, row in negative_duration.iterrows():
anomalies.append(Anomaly(
id=f"SCHED-NEG-{idx}",
anomaly_type=AnomalyType.IMPOSSIBLE_VALUE,
severity=AnomalySeverity.CRITICAL,
field="duration",
value=row['duration'],
description="End date before start date",
row_index=idx,
detection_method="Business Rule",
confidence=1.0,
suggested_action="Correct dates"
))
# Extremely long durations
long_tasks = df[df['duration'] > self.SCHEDULE_THRESHOLDS['max_activity_duration']]
for idx, row in long_tasks.iterrows():
anomalies.append(Anomaly(
id=f"SCHED-LONG-{idx}",
anomaly_type=AnomalyType.OUTLIER,
severity=AnomalySeverity.MEDIUM,
field="duration",
value=row['duration'],
expected_range=(0, self.SCHEDULE_THRESHOLDS['max_activity_duration']),
description=f"Task duration {row['duration']} days exceeds threshold",
row_index=idx,
detection_method="Threshold",
confidence=0.9,
suggested_action="Review if task should be broken down"
))
# Zero duration non-milestones
if 'is_milestone' in df.columns:
zero_duration = df[(df['duration'] == 0) & (~df['is_milestone'])]
for idx, row in zero_duration.iterrows():
anomalies.append(Anomaly(
id=f"SCHED-ZERO-{idx}",
anomaly_type=AnomalyType.IMPOSSIBLE_VALUE,
severity=AnomalySeverity.HIGH,
field="duration",
value=0,
description="Zero duration task that is not a milestone",
row_index=idx,
detection_method="Business Rule",
confidence=1.0,
suggested_action="Add duration or mark as milestone"
))
return anomalies
def detect_productivity_anomalies(self, df: pd.DataFrame,
quantity_col: str,
hours_col: str) -> List[Anomaly]:
"""Detect productivity anomalies."""
anomalies = []
# Calculate productivity
df['productivity'] = df[quantity_col] / df[hours_col].replace(0, np.nan)
# Use Modified Z-Score (more robust for skewed data)
median = df['productivity'].median()
mad = np.abs(df['productivity'] - median).median()
modified_z = 0.6745 * (df['productivity'] - median) / mad
outliers = df[np.abs(modified_z) > 3.5]
for idx, row in outliers.iterrows():
prod = row['productivity']
z = modified_z.loc[idx]
severity = AnomalySeverity.HIGH if abs(z) > 5 else AnomalySeverity.MEDIUM
direction = "high" if z > 0 else "low"
anomalies.append(Anomaly(
id=f"PROD-{idx}",
anomaly_type=AnomalyType.OUTLIER,
severity=severity,
field="productivity",
value=prod,
description=f"Unusually {direction} productivity: {prod:.2f} units/hour",
row_index=idx,
detection_method="Modified Z-Score",
confidence=min(abs(z) / 7, 1.0),
suggested_action=f"Investigate {direction} productivity cause"
))
return anomalies
def detect_time_series_anomalies(self, df: pd.DataFrame,
date_col: str,
value_col: str,
window: int = 7) -> List[Anomaly]:
"""Detect anomalies in time series data (e.g., daily costs, progress)."""
anomalies = []
df = df.sort_values(date_col).copy()
df['rolling_mean'] = df[value_col].rolling(window=window, center=True).mean()
df['rolling_std'] = df[value_col].rolling(window=window, center=True).std()
# Points outside 2 standard deviations from rolling mean
df['z_score'] = (df[value_col] - df['rolling_mean']) / df['rolling_std']
outliers = df[np.abs(df['z_score']) > 2].dropna()
for idx, row in outliers.iterrows():
anomalies.append(Anomaly(
id=f"TS-{idx}",
anomaly_type=AnomalyType.TREND_DEVIATION,
severity=AnomalySeverity.MEDIUM if abs(row['z_score']) < 3 else AnomalySeverity.HIGH,
field=value_col,
value=row[value_col],
expected_range=(
row['rolling_mean'] - 2 * row['rolling_std'],
row['rolling_mean'] + 2 * row['rolling_std']
),
description=f"Value deviates from {window}-day trend",
row_index=idx,
detection_method="Rolling Z-Score",
confidence=min(abs(row['z_score']) / 4, 1.0)
))
return anomalies
def detect_duplicate_anomalies(self, df: pd.DataFrame,
key_columns: List[str]) -> List[Anomaly]:
"""Detect duplicate records."""
anomalies = []
duplicates = df[df.duplicated(subset=key_columns, keep=False)]
if len(duplicates) > 0:
dup_groups = duplicates.groupby(key_columns).size()
for keys, count in dup_groups.items():
anomalies.append(Anomaly(
id=f"DUP-{hash(str(keys)) % 10000}",
anomaly_type=AnomalyType.DUPLICATE,
severity=AnomalySeverity.HIGH,
field=str(key_columns),
value=keys,
description=f"Found {count} duplicate records for {keys}",
detection_method="Exact Match",
confidence=1.0,
suggested_action="Review and remove duplicates"
))
return anomalies
def detect_sequence_gaps(self, df: pd.DataFrame, sequence_col: str) -> List[Anomaly]:
"""Detect gaps in sequential data (invoice numbers, PO numbers, etc.)."""
anomalies = []
# Extract numeric part if mixed format
df['seq_num'] = pd.to_numeric(
df[sequence_col].astype(str).str.extract(r'(\d+)')[0],
errors='coerce'
)
sorted_seq = df['seq_num'].dropna().sort_values()
expected = range(int(sorted_seq.min()), int(sorted_seq.max()) + 1)
actual = set(sorted_seq.astype(int))
missing = set(expected) - actual
if missing:
# Group consecutive missing numbers
missing_ranges = []
sorted_missing = sorted(missing)
start = sorted_missing[0]
end = start
for num in sorted_missing[1:]:
if num == end + 1:
end = num
else:
missing_ranges.append((start, end))
start = num
end = num
missing_ranges.append((start, end))
for start, end in missing_ranges:
range_str = str(start) if start == end else f"{start}-{end}"
anomalies.append(Anomaly(
id=f"SEQ-{start}",
anomaly_type=AnomalyType.MISSING_SEQUENCE,
severity=AnomalySeverity.MEDIUM,
field=sequence_col,
value=range_str,
description=f"Missing sequence number(s): {range_str}",
detection_method="Sequence Analysis",
confidence=1.0,
suggested_action="Investigate missing numbers"
))
return anomalies
def run_full_detection(self, df: pd.DataFrame, config: Dict) -> AnomalyReport:
"""Run all applicable anomaly detection methods."""
all_anomalies = []
# Cost anomalies
if 'cost_columns' in config:
for col in config['cost_columns']:
if col in df.columns:
all_anomalies.extend(
self.detect_cost_anomalies(df, col, config.get('group_by'))
)
# Schedule anomalies
if 'start_date' in df.columns and 'end_date' in df.columns:
all_anomalies.extend(self.detect_schedule_anomalies(df))
# Productivity
if 'quantity_col' in config and 'hours_col' in config:
all_anomalies.extend(
self.detect_productivity_anomalies(
df, config['quantity_col'], config['hours_col']
)
)
# Duplicates
if 'key_columns' in config:
all_anomalies.extend(
self.detect_duplicate_anomalies(df, config['key_columns'])
)
# Sequence gaps
if 'sequence_column' in config:
all_anomalies.extend(
self.detect_sequence_gaps(df, config['sequence_column'])
)
# Create summary
summary = {}
for a in all_anomalies:
key = f"{a.anomaly_type.value}_{a.severity.value}"
summary[key] = summary.get(key, 0) + 1
report = AnomalyReport(
source=config.get('source_name', 'Unknown'),
detected_at=datetime.now(),
total_records=len(df),
anomalies=all_anomalies,
summary=summary
)
self.detection_history.append(report)
return report
def generate_report(self, report: AnomalyReport) -> str:
"""Generate markdown anomaly report."""
lines = [f"# Anomaly Detection Report", ""]
lines.append(f"**Source:** {report.source}")
lines.append(f"**Detected At:** {report.detected_at.strftime('%Y-%m-%d %H:%M')}")
lines.append(f"**Total Records:** {report.total_records:,}")
lines.append(f"**Anomalies Found:** {len(report.anomalies)}")
lines.append("")
# Summary by severity
lines.append("## Summary by Severity")
for severity in AnomalySeverity:
count = sum(1 for a in report.anomalies if a.severity == severity)
if count > 0:
lines.append(f"- **{severity.value.upper()}:** {count}")
lines.append("")
# Critical anomalies first
critical = [a for a in report.anomalies if a.severity == AnomalySeverity.CRITICAL]
if critical:
lines.append("## Critical Anomalies")
for a in critical:
lines.append(f"\n### {a.id}")
lines.append(f"- **Type:** {a.anomaly_type.value}")
lines.append(f"- **Field:** {a.field}")
lines.append(f"- **Value:** {a.value}")
lines.append(f"- **Description:** {a.description}")
lines.append(f"- **Action:** {a.suggested_action}")
# All anomalies table
lines.append("\n## All Anomalies")
lines.append("| ID | Type | Severity | Field | Description |")
lines.append("|-----|------|----------|-------|-------------|")
for a in report.anomalies[:50]:
lines.append(f"| {a.id} | {a.anomaly_type.value} | {a.severity.value} | {a.field} | {a.description[:50]} |")
if len(report.anomalies) > 50:
lines.append(f"\n*... and {len(report.anomalies) - 50} more anomalies*")
return "\n".join(lines)import pandas as pd
# Load data
df = pd.read_excel("project_costs.xlsx")
# Initialize detector
detector = ConstructionAnomalyDetector()
# Run detection
config = {
'source_name': 'Project Costs Q1 2026',
'cost_columns': ['total_cost', 'labor_cost', 'material_cost'],
'group_by': 'cost_code',
'key_columns': ['project_id', 'cost_code', 'date'],
'sequence_column': 'invoice_number'
}
report = detector.run_full_detection(df, config)
# Generate report
print(detector.generate_report(report))
# Get critical anomalies for immediate action
critical = [a for a in report.anomalies if a.severity == AnomalySeverity.CRITICAL]
print(f"\n{len(critical)} critical anomalies require immediate attention")pip install pandas numpy scipy© datadrivenconstruction, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in 2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector of datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.
Open the folder on GitHubat commit ce45bbf
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction, which our catalogue first saw on October 7, 2026.
Data Anomaly Detector next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Anomaly Detector this skilldatadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction | 344 | 1 repos | ~4.8k | Automated safety check: Pass | MIT | |
| Statistical Analysismajiayu000/claude-skill-registry | 666 | 2 repos | ~2.6k | Automated safety check: Pass | MIT | |
| Stat Edaasgard-ai-platform/skills | 241 | — | ~954 | Automated safety check: Pass | MIT | |
| TimesFM Forecastinggoogle-research/timesfm | 34k | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | |
| Question2reportrefraction-ray/xalpha | 2.7k | — | ~3.2k | Automated safety check: Pass | MIT | |
| Dingo VerifyMigoXLab/dingo | 757 | — | ~741 | Automated safety check: Notes | Apache-2.0 |
majiayu000/claude-skill-registry
Apply statistical methods including descriptive stats, trend analysis, outlier detection, and hypothesis testing.
asgard-ai-platform/skills
Conduct Exploratory Data Analysis (EDA) using descriptive statistics, visualizations, and data quality checks.
google-research/timesfm
Forecasts any univariate time series zero-shot with Google's TimesFM model, returning point forecasts and calibrated prediction intervals without training.
refraction-ray/xalpha
Turn a natural-language financial question into a polished, self-contained HTML report.
MigoXLab/dingo
A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.
platonai/Browser4
Validates data against common and custom rules (required fields, formats, ranges).
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Orchestrate multiple AI agents for construction workflows: estimator, scheduler, document, QA and safety agents coordinated by a supervisor agent, with human checkpoints.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Estimate embodied carbon and produce ESG/climate reporting for construction: LCA per work item, material-based carbon factors, EU taxonomy and CSRD alignment.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Material passports and circular construction: generate per-element material inventories from BOQ/BIM, mark reuse potential and recycled content, and prepare deconstruction data.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Browse and search the OpenConstructionERP cost database: classification tree, SQL and semantic search, autocomplete, certainty badges, and the resource catalog.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Create bills of quantities and estimates in OpenConstructionERP: search cost items, build BOQ sections, link BIM elements in bulk, validate the BOQ, and export GAEB/XLSX/JSON.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Field operations in OpenConstructionERP: punch list, daily diary, HSE observations and task tracking on site.
Categories
Detect anomalies and outliers in construction data: unusual costs, schedule variances, productivity spikes. Data Anomaly Detector is an agent skill from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction. Detect anomalies and outliers in construction data: unusual costs, schedule variances, productivity spikes.
Data Anomaly Detector fits situations like: tasks that involve Anomaly detection; tasks that involve Data cleaning.
Run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-anomaly-detector -a claude-code`. Or copy the skill folder (2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction) into .claude/skills/data-anomaly-detector in your project. Claude Code loads it when a task matches its description.
Run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-anomaly-detector -a codex`. Or copy the skill folder (2_DDC_Book/2.6-Data-Quality-Validation/data-anomaly-detector in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction) into .agents/skills/data-anomaly-detector in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-anomaly-detector -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-anomaly-detector, .gemini/skills/data-anomaly-detector, .github/skills/data-anomaly-detector and .opencode/skills/data-anomaly-detector in your project.
Going by SKILL.md and its folder, Data Anomaly Detector needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Data Anomaly Detector is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Data Anomaly Detector: Statistical Analysis (majiayu000/claude-skill-registry, 666 stars), Stat Eda (asgard-ai-platform/skills, 241 stars), TimesFM Forecasting (google-research/timesfm, 34k stars) and Question2report (refraction-ray/xalpha, 2.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
datadrivenconstruction (a GitHub user) maintains it in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction, which has 344 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on August 22, 2026.
Source: datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.