Exploratory Data Analysis
spacering-net/codeg
Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.
Agent skill
by datadrivenconstruction in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Classify construction data by type (structured, unstructured, semi-structured).
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-type-classifier -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-type-classifier --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .claude/skills && cp -r skills-src/2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier .claude/skills/data-type-classifier && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-type-classifier" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier into .claude/skills/data-type-classifier/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-type-classifier", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.1-Data-Types-Classification/data-type-classifierType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-type-classifier -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-type-classifier --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .agents/skills && cp -r skills-src/2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier .agents/skills/data-type-classifier && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-type-classifier" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier into .agents/skills/data-type-classifier/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-type-classifier", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-type-classifier -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-type-classifier --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier .cursor/skills/data-type-classifier && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-type-classifier" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier into .cursor/skills/data-type-classifier/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-type-classifier", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git --path 2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-type-classifier -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-type-classifier --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier .gemini/skills/data-type-classifier && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-type-classifier" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier into .gemini/skills/data-type-classifier/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-type-classifier", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-type-classifierInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-type-classifier -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .github/skills && cp -r skills-src/2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier .github/skills/data-type-classifier && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-type-classifier" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier into .github/skills/data-type-classifier/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-type-classifier", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-type-classifier -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-type-classifier --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier .opencode/skills/data-type-classifier && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-type-classifier" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier into .opencode/skills/data-type-classifier/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-type-classifier", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-type-classifierClassify construction data by type (structured, unstructured, semi-structured).
Data Type Classifier is an agent skill from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction. Classify construction data by type (structured, unstructured, semi-structured). Analyze data sources and recommend appropriate storage/processing methods
Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `claw.json` and `instructions.md`).
It sits in Data & Analytics, covering Data analysis. The repository describes itself as: 221 AI skills for construction: BIM analysis, cost estimation, scheduling, document control, and automation with Claude Code. The licence is MIT.
Read from SKILL.md and the folder at commit ce45bbf. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
datadrivenconstruction.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Type Classifier loads about 6.1k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 110 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction at commit ce45bbf, republished under its MIT licence (© datadrivenconstruction). 110 words, ~6,060 tokens.
.claude/skills/data-type-classifier/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Based on DDC methodology (Chapter 2.1), this skill classifies construction data by type, analyzes data sources, and recommends appropriate storage, processing, and integration methods.
Book Reference: "Типы данных в строительстве" / "Data Types in Construction"
from dataclasses import dataclass, field
from enum import Enum
from typing import List, Dict, Optional, Any, Tuple
from datetime import datetime
import json
import re
import mimetypes
class DataStructure(Enum):
"""Data structure classification"""
STRUCTURED = "structured" # Tables, databases, spreadsheets
SEMI_STRUCTURED = "semi_structured" # JSON, XML, IFC
UNSTRUCTURED = "unstructured" # Documents, images, videos
GEOMETRIC = "geometric" # CAD, BIM geometry
TEMPORAL = "temporal" # Time-series, schedules
SPATIAL = "spatial" # GIS, coordinates
class DataFormat(Enum):
"""Common construction data formats"""
# Structured
CSV = "csv"
EXCEL = "excel"
SQL = "sql"
PARQUET = "parquet"
# Semi-structured
JSON = "json"
XML = "xml"
IFC = "ifc"
BCF = "bcf"
# Unstructured
PDF = "pdf"
DOCX = "docx"
IMAGE = "image"
VIDEO = "video"
# Geometric
DWG = "dwg"
DXF = "dxf"
RVT = "rvt"
NWD = "nwd"
OBJ = "obj"
STL = "stl"
# Schedule
MPP = "mpp"
P6 = "p6"
XER = "xer"
class StorageRecommendation(Enum):
"""Storage system recommendations"""
RELATIONAL_DB = "relational_database"
DOCUMENT_DB = "document_database"
OBJECT_STORAGE = "object_storage"
GRAPH_DB = "graph_database"
TIME_SERIES_DB = "time_series_database"
VECTOR_DB = "vector_database"
FILE_SYSTEM = "file_system"
DATA_LAKE = "data_lake"
@dataclass
class DataCharacteristics:
"""Characteristics of a data source"""
has_schema: bool
has_relationships: bool
is_queryable: bool
is_binary: bool
has_geometry: bool
has_temporal: bool
has_text_content: bool
avg_record_size: Optional[int] = None # bytes
estimated_volume: Optional[str] = None # small/medium/large/huge
update_frequency: Optional[str] = None
@dataclass
class DataClassification:
"""Classification result for a data source"""
source_name: str
source_type: str
detected_format: DataFormat
structure: DataStructure
characteristics: DataCharacteristics
storage_recommendation: StorageRecommendation
processing_tools: List[str]
integration_options: List[str]
quality_considerations: List[str]
confidence: float
@dataclass
class ClassificationReport:
"""Complete classification report"""
total_sources: int
classifications: List[DataClassification]
summary_by_structure: Dict[str, int]
summary_by_format: Dict[str, int]
storage_recommendations: Dict[str, List[str]]
integration_strategy: Dict[str, str]
class DataTypeClassifier:
"""
Classify construction data by type and recommend processing methods.
Based on DDC methodology Chapter 2.1.
"""
def __init__(self):
self.format_signatures = self._define_format_signatures()
self.structure_mapping = self._define_structure_mapping()
self.storage_mapping = self._define_storage_mapping()
self.processing_tools = self._define_processing_tools()
def _define_format_signatures(self) -> Dict[str, Dict]:
"""Define format detection signatures"""
return {
# File extensions
".csv": {"format": DataFormat.CSV, "structure": DataStructure.STRUCTURED},
".xlsx": {"format": DataFormat.EXCEL, "structure": DataStructure.STRUCTURED},
".xls": {"format": DataFormat.EXCEL, "structure": DataStructure.STRUCTURED},
".json": {"format": DataFormat.JSON, "structure": DataStructure.SEMI_STRUCTURED},
".xml": {"format": DataFormat.XML, "structure": DataStructure.SEMI_STRUCTURED},
".ifc": {"format": DataFormat.IFC, "structure": DataStructure.SEMI_STRUCTURED},
".bcf": {"format": DataFormat.BCF, "structure": DataStructure.SEMI_STRUCTURED},
".pdf": {"format": DataFormat.PDF, "structure": DataStructure.UNSTRUCTURED},
".docx": {"format": DataFormat.DOCX, "structure": DataStructure.UNSTRUCTURED},
".dwg": {"format": DataFormat.DWG, "structure": DataStructure.GEOMETRIC},
".dxf": {"format": DataFormat.DXF, "structure": DataStructure.GEOMETRIC},
".rvt": {"format": DataFormat.RVT, "structure": DataStructure.GEOMETRIC},
".nwd": {"format": DataFormat.NWD, "structure": DataStructure.GEOMETRIC},
".mpp": {"format": DataFormat.MPP, "structure": DataStructure.TEMPORAL},
".xer": {"format": DataFormat.XER, "structure": DataStructure.TEMPORAL},
".parquet": {"format": DataFormat.PARQUET, "structure": DataStructure.STRUCTURED},
".jpg": {"format": DataFormat.IMAGE, "structure": DataStructure.UNSTRUCTURED},
".png": {"format": DataFormat.IMAGE, "structure": DataStructure.UNSTRUCTURED},
".mp4": {"format": DataFormat.VIDEO, "structure": DataStructure.UNSTRUCTURED}
}
def _define_structure_mapping(self) -> Dict[DataStructure, Dict]:
"""Define characteristics for each structure type"""
return {
DataStructure.STRUCTURED: {
"description": "Tabular data with fixed schema",
"examples": ["Cost databases", "Material lists", "Vendor records"],
"query_support": True,
"schema_required": True
},
DataStructure.SEMI_STRUCTURED: {
"description": "Hierarchical data with flexible schema",
"examples": ["BIM models (IFC)", "API responses", "Configuration files"],
"query_support": True,
"schema_required": False
},
DataStructure.UNSTRUCTURED: {
"description": "No predefined schema or format",
"examples": ["Contracts", "Photos", "Emails", "Meeting notes"],
"query_support": False,
"schema_required": False
},
DataStructure.GEOMETRIC: {
"description": "3D/2D geometric and spatial data",
"examples": ["CAD drawings", "BIM geometry", "Point clouds"],
"query_support": True,
"schema_required": True
},
DataStructure.TEMPORAL: {
"description": "Time-based sequential data",
"examples": ["Schedules", "Progress data", "Sensor readings"],
"query_support": True,
"schema_required": True
},
DataStructure.SPATIAL: {
"description": "Geographic and location data",
"examples": ["Site maps", "GPS tracks", "GIS layers"],
"query_support": True,
"schema_required": True
}
}
def _define_storage_mapping(self) -> Dict[DataStructure, StorageRecommendation]:
"""Map data structures to storage recommendations"""
return {
DataStructure.STRUCTURED: StorageRecommendation.RELATIONAL_DB,
DataStructure.SEMI_STRUCTURED: StorageRecommendation.DOCUMENT_DB,
DataStructure.UNSTRUCTURED: StorageRecommendation.OBJECT_STORAGE,
DataStructure.GEOMETRIC: StorageRecommendation.FILE_SYSTEM,
DataStructure.TEMPORAL: StorageRecommendation.TIME_SERIES_DB,
DataStructure.SPATIAL: StorageRecommendation.RELATIONAL_DB
}
def _define_processing_tools(self) -> Dict[DataFormat, List[str]]:
"""Define processing tools for each format"""
return {
DataFormat.CSV: ["pandas", "polars", "duckdb"],
DataFormat.EXCEL: ["pandas", "openpyxl", "xlrd"],
DataFormat.JSON: ["json", "pandas", "jq"],
DataFormat.XML: ["lxml", "ElementTree", "BeautifulSoup"],
DataFormat.IFC: ["ifcopenshell", "IfcOpenShell", "xBIM"],
DataFormat.BCF: ["bcfpython", "ifcopenshell"],
DataFormat.PDF: ["pdfplumber", "PyPDF2", "pdf2image"],
DataFormat.DOCX: ["python-docx", "mammoth"],
DataFormat.DWG: ["ezdxf", "Teigha", "ODA SDK"],
DataFormat.DXF: ["ezdxf", "dxfgrabber"],
DataFormat.RVT: ["Revit API", "pyRevit", "Dynamo"],
DataFormat.NWD: ["Navisworks API", "NW API"],
DataFormat.MPP: ["mpxj", "Project API"],
DataFormat.XER: ["xerparser", "P6 API"],
DataFormat.PARQUET: ["pandas", "pyarrow", "polars"],
DataFormat.IMAGE: ["PIL", "opencv", "scikit-image"],
DataFormat.VIDEO: ["opencv", "ffmpeg", "moviepy"]
}
def classify_source(
self,
source_name: str,
source_type: str,
file_extension: Optional[str] = None,
sample_data: Optional[Any] = None,
metadata: Optional[Dict] = None
) -> DataClassification:
"""
Classify a single data source.
Args:
source_name: Name of the data source
source_type: Type (file, database, api, etc.)
file_extension: File extension if applicable
sample_data: Sample of the data for analysis
metadata: Additional metadata
Returns:
Classification result
"""
# Detect format
detected_format, structure = self._detect_format(
file_extension, source_type, sample_data
)
# Analyze characteristics
characteristics = self._analyze_characteristics(
detected_format, structure, sample_data, metadata
)
# Determine storage recommendation
storage = self._recommend_storage(structure, characteristics)
# Get processing tools
tools = self.processing_tools.get(detected_format, [])
# Determine integration options
integration = self._get_integration_options(detected_format, structure)
# Quality considerations
quality = self._get_quality_considerations(detected_format, structure)
# Calculate confidence
confidence = self._calculate_confidence(
file_extension, sample_data, metadata
)
return DataClassification(
source_name=source_name,
source_type=source_type,
detected_format=detected_format,
structure=structure,
characteristics=characteristics,
storage_recommendation=storage,
processing_tools=tools,
integration_options=integration,
quality_considerations=quality,
confidence=confidence
)
def _detect_format(
self,
extension: Optional[str],
source_type: str,
sample: Optional[Any]
) -> Tuple[DataFormat, DataStructure]:
"""Detect data format and structure"""
# Check file extension
if extension:
ext = extension.lower() if extension.startswith('.') else f".{extension.lower()}"
if ext in self.format_signatures:
sig = self.format_signatures[ext]
return sig["format"], sig["structure"]
# Check source type
if source_type == "database":
return DataFormat.SQL, DataStructure.STRUCTURED
elif source_type == "api":
return DataFormat.JSON, DataStructure.SEMI_STRUCTURED
# Analyze sample data
if sample:
if isinstance(sample, dict):
return DataFormat.JSON, DataStructure.SEMI_STRUCTURED
elif isinstance(sample, list) and all(isinstance(x, dict) for x in sample):
return DataFormat.JSON, DataStructure.STRUCTURED
elif isinstance(sample, str):
if sample.strip().startswith('<'):
return DataFormat.XML, DataStructure.SEMI_STRUCTURED
elif sample.strip().startswith('{'):
return DataFormat.JSON, DataStructure.SEMI_STRUCTURED
# Default
return DataFormat.JSON, DataStructure.SEMI_STRUCTURED
def _analyze_characteristics(
self,
format: DataFormat,
structure: DataStructure,
sample: Optional[Any],
metadata: Optional[Dict]
) -> DataCharacteristics:
"""Analyze data characteristics"""
return DataCharacteristics(
has_schema=structure in [DataStructure.STRUCTURED, DataStructure.TEMPORAL],
has_relationships=format in [DataFormat.IFC, DataFormat.SQL],
is_queryable=structure != DataStructure.UNSTRUCTURED,
is_binary=format in [
DataFormat.DWG, DataFormat.RVT, DataFormat.NWD,
DataFormat.IMAGE, DataFormat.VIDEO, DataFormat.PDF
],
has_geometry=structure == DataStructure.GEOMETRIC or format == DataFormat.IFC,
has_temporal=structure == DataStructure.TEMPORAL,
has_text_content=format in [
DataFormat.PDF, DataFormat.DOCX, DataFormat.CSV
],
estimated_volume=metadata.get("volume") if metadata else None,
update_frequency=metadata.get("update_frequency") if metadata else None
)
def _recommend_storage(
self,
structure: DataStructure,
characteristics: DataCharacteristics
) -> StorageRecommendation:
"""Recommend storage solution"""
# Special cases
if characteristics.has_text_content and not characteristics.has_schema:
return StorageRecommendation.VECTOR_DB
if characteristics.is_binary and characteristics.estimated_volume == "huge":
return StorageRecommendation.OBJECT_STORAGE
if characteristics.has_relationships:
return StorageRecommendation.GRAPH_DB
# Default mapping
return self.storage_mapping.get(structure, StorageRecommendation.FILE_SYSTEM)
def _get_integration_options(
self,
format: DataFormat,
structure: DataStructure
) -> List[str]:
"""Get integration options for the data"""
options = []
if structure == DataStructure.STRUCTURED:
options.extend(["Direct SQL queries", "ETL pipelines", "API export"])
elif structure == DataStructure.SEMI_STRUCTURED:
options.extend(["JSON/XML parsing", "Schema validation", "API integration"])
elif structure == DataStructure.UNSTRUCTURED:
options.extend(["OCR extraction", "NLP processing", "ML classification"])
elif structure == DataStructure.GEOMETRIC:
options.extend(["IFC export", "Geometry extraction", "Clash detection"])
# Format-specific options
if format == DataFormat.IFC:
options.append("IFC import/export via IfcOpenShell")
elif format == DataFormat.EXCEL:
options.append("Pandas DataFrame conversion")
elif format == DataFormat.PDF:
options.append("PDF text/table extraction")
return options
def _get_quality_considerations(
self,
format: DataFormat,
structure: DataStructure
) -> List[str]:
"""Get quality considerations"""
considerations = []
if structure == DataStructure.STRUCTURED:
considerations.extend([
"Validate schema consistency",
"Check for null/missing values",
"Verify data types"
])
elif structure == DataStructure.UNSTRUCTURED:
considerations.extend([
"OCR accuracy verification",
"Text encoding issues",
"Content extraction completeness"
])
elif structure == DataStructure.GEOMETRIC:
considerations.extend([
"Model validity (closed solids)",
"Coordinate system consistency",
"Unit verification"
])
# Format-specific
if format == DataFormat.IFC:
considerations.append("IFC schema version compatibility")
elif format == DataFormat.EXCEL:
considerations.append("Formula vs value extraction")
return considerations
def _calculate_confidence(
self,
extension: Optional[str],
sample: Optional[Any],
metadata: Optional[Dict]
) -> float:
"""Calculate classification confidence"""
confidence = 0.5 # Base confidence
if extension:
confidence += 0.3 # Extension provides good hint
if sample:
confidence += 0.15 # Sample data helps
if metadata:
confidence += 0.05 # Metadata adds context
return min(1.0, confidence)
def classify_multiple(
self,
sources: List[Dict]
) -> ClassificationReport:
"""
Classify multiple data sources.
Args:
sources: List of source definitions
Returns:
Complete classification report
"""
classifications = []
for source in sources:
classification = self.classify_source(
source_name=source["name"],
source_type=source.get("type", "file"),
file_extension=source.get("extension"),
sample_data=source.get("sample"),
metadata=source.get("metadata")
)
classifications.append(classification)
# Generate summaries
summary_structure = {}
summary_format = {}
storage_recs = {}
for c in classifications:
# Structure summary
struct = c.structure.value
summary_structure[struct] = summary_structure.get(struct, 0) + 1
# Format summary
fmt = c.detected_format.value
summary_format[fmt] = summary_format.get(fmt, 0) + 1
# Storage recommendations
storage = c.storage_recommendation.value
if storage not in storage_recs:
storage_recs[storage] = []
storage_recs[storage].append(c.source_name)
# Integration strategy
strategy = self._generate_integration_strategy(classifications)
return ClassificationReport(
total_sources=len(sources),
classifications=classifications,
summary_by_structure=summary_structure,
summary_by_format=summary_format,
storage_recommendations=storage_recs,
integration_strategy=strategy
)
def _generate_integration_strategy(
self,
classifications: List[DataClassification]
) -> Dict[str, str]:
"""Generate integration strategy"""
strategy = {}
# Group by structure
structured = [c for c in classifications if c.structure == DataStructure.STRUCTURED]
semi = [c for c in classifications if c.structure == DataStructure.SEMI_STRUCTURED]
unstructured = [c for c in classifications if c.structure == DataStructure.UNSTRUCTURED]
geometric = [c for c in classifications if c.structure == DataStructure.GEOMETRIC]
if structured:
strategy["structured_data"] = (
"Use ETL pipeline to consolidate into central data warehouse. "
"Implement SQL-based querying and reporting."
)
if semi:
strategy["semi_structured_data"] = (
"Use document database for flexible storage. "
"Implement schema validation at ingestion."
)
if unstructured:
strategy["unstructured_data"] = (
"Extract text content using OCR/NLP. "
"Store in vector database for semantic search."
)
if geometric:
strategy["geometric_data"] = (
"Standardize on IFC format for exchange. "
"Maintain native formats for editing."
)
return strategy
def generate_report(self, report: ClassificationReport) -> str:
"""Generate classification report"""
output = f"""
# Data Classification Report
**Total Sources Analyzed:** {report.total_sources}
## Summary by Structure
"""
for struct, count in report.summary_by_structure.items():
output += f"- **{struct.title()}**: {count} sources\n"
output += "\n## Summary by Format\n\n"
for fmt, count in report.summary_by_format.items():
output += f"- **{fmt.upper()}**: {count} sources\n"
output += "\n## Storage Recommendations\n\n"
for storage, sources in report.storage_recommendations.items():
output += f"### {storage.replace('_', ' ').title()}\n"
for src in sources:
output += f"- {src}\n"
output += "\n"
output += "## Integration Strategy\n\n"
for category, strategy in report.integration_strategy.items():
output += f"### {category.replace('_', ' ').title()}\n{strategy}\n\n"
output += "## Detailed Classifications\n\n"
for c in report.classifications[:10]:
output += f"""
### {c.source_name}
- **Format:** {c.detected_format.value}
- **Structure:** {c.structure.value}
- **Storage:** {c.storage_recommendation.value}
- **Tools:** {', '.join(c.processing_tools[:3])}
- **Confidence:** {c.confidence:.0%}
"""
return outputclassifier = DataTypeClassifier()
# Classify a BIM model
classification = classifier.classify_source(
source_name="Building Model",
source_type="file",
file_extension=".ifc",
metadata={"volume": "large"}
)
print(f"Format: {classification.detected_format.value}")
print(f"Structure: {classification.structure.value}")
print(f"Storage: {classification.storage_recommendation.value}")
print(f"Tools: {classification.processing_tools}")sources = [
{"name": "Cost Database", "type": "database", "extension": ".sql"},
{"name": "Building Model", "type": "file", "extension": ".ifc"},
{"name": "Contract PDFs", "type": "file", "extension": ".pdf"},
{"name": "Site Photos", "type": "file", "extension": ".jpg"},
{"name": "Schedule", "type": "file", "extension": ".mpp"}
]
report = classifier.classify_multiple(sources)
print(f"Total: {report.total_sources}")
print(f"By structure: {report.summary_by_structure}")report_text = classifier.generate_report(report)
print(report_text)
# Save to file
with open("classification_report.md", "w") as f:
f.write(report_text)| Component | Purpose |
|---|---|
DataTypeClassifier | Main classification engine |
DataStructure | Structure types (structured, semi, unstructured) |
DataFormat | File format detection |
StorageRecommendation | Storage system recommendations |
DataClassification | Classification result |
ClassificationReport | Multi-source report |
© datadrivenconstruction, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in 2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier of datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.
Open the folder on GitHubat commit ce45bbf
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction, which our catalogue first saw on October 7, 2026.
Data Type Classifier next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Type Classifier this skilldatadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction | 344 | 1 repos | ~6.1k | Automated safety check: Pass | MIT | |
| Exploratory Data Analysisspacering-net/codeg | 3.8k | 15 repos | ~3.6k | Automated safety check: Pass | MIT | |
| Excel and CSV Data Analysisbytedance/deer-flow | 83k | 4 repos | ~2.2k | Automated safety check: Pass | MIT | |
| Exploratory Data AnalysisOleafly/Oleafly | 205 | 2 repos | ~3.4k | Automated safety check: Notes | MIT | |
| Python Executorcortega26/chile-hub | 113 | 2 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill | 188 | — | ~4k | Automated safety check: Pass | MIT |
spacering-net/codeg
Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
Oleafly/Oleafly
Perform bounded, local exploratory analysis of explicitly supported scientific files.
cortega26/chile-hub
Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).
FrankS-IntelLab/agentic-kaggle-skill
Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.
mcncarl/yichen-skills
Read, decrypt, query, search, and export local WeCom/企业微信 5.x desktop databases on macOS into a private read-only vault.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Orchestrate multiple AI agents for construction workflows: estimator, scheduler, document, QA and safety agents coordinated by a supervisor agent, with human checkpoints.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Estimate embodied carbon and produce ESG/climate reporting for construction: LCA per work item, material-based carbon factors, EU taxonomy and CSRD alignment.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Material passports and circular construction: generate per-element material inventories from BOQ/BIM, mark reuse potential and recycled content, and prepare deconstruction data.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Browse and search the OpenConstructionERP cost database: classification tree, SQL and semantic search, autocomplete, certainty badges, and the resource catalog.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Create bills of quantities and estimates in OpenConstructionERP: search cost items, build BOQ sections, link BIM elements in bulk, validate the BOQ, and export GAEB/XLSX/JSON.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Field operations in OpenConstructionERP: punch list, daily diary, HSE observations and task tracking on site.
Categories
Classify construction data by type (structured, unstructured, semi-structured). Data Type Classifier is an agent skill from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction. Classify construction data by type (structured, unstructured, semi-structured).
Data Type Classifier fits situations like: tasks that involve Data analysis.
Run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-type-classifier -a claude-code`. Or copy the skill folder (2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction) into .claude/skills/data-type-classifier in your project. Claude Code loads it when a task matches its description.
Run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-type-classifier -a codex`. Or copy the skill folder (2_DDC_Book/2.1-Data-Types-Classification/data-type-classifier in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction) into .agents/skills/data-type-classifier in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-type-classifier -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-type-classifier, .gemini/skills/data-type-classifier, .github/skills/data-type-classifier and .opencode/skills/data-type-classifier in your project.
SKILL.md names no scripts, command-line tools or credentials: Data Type Classifier is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: datadrivenconstruction.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Data Type Classifier is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.1k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Data Type Classifier: Exploratory Data Analysis (spacering-net/codeg, 3.8k stars), Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), Exploratory Data Analysis (Oleafly/Oleafly, 205 stars) and Python Executor (cortega26/chile-hub, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
datadrivenconstruction (a GitHub user) maintains it in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction, which has 344 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on August 22, 2026.
Source: datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.