Data Quality Frameworks
wshobson/agents
Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.
Agent skill
by datadrivenconstruction in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Track data origin, transformations, and flow through construction systems.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-lineage-tracker -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-lineage-tracker --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .claude/skills && cp -r skills-src/2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker .claude/skills/data-lineage-tracker && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-lineage-tracker" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker into .claude/skills/data-lineage-tracker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-lineage-tracker", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-trackerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-lineage-tracker -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-lineage-tracker --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .agents/skills && cp -r skills-src/2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker .agents/skills/data-lineage-tracker && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-lineage-tracker" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker into .agents/skills/data-lineage-tracker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-lineage-tracker", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-lineage-tracker -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-lineage-tracker --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker .cursor/skills/data-lineage-tracker && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-lineage-tracker" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker into .cursor/skills/data-lineage-tracker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-lineage-tracker", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git --path 2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-lineage-tracker -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-lineage-tracker --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker .gemini/skills/data-lineage-tracker && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-lineage-tracker" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker into .gemini/skills/data-lineage-tracker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-lineage-tracker", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-lineage-trackerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-lineage-tracker -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .github/skills && cp -r skills-src/2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker .github/skills/data-lineage-tracker && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-lineage-tracker" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker into .github/skills/data-lineage-tracker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-lineage-tracker", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-lineage-tracker -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction data-lineage-tracker --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker .opencode/skills/data-lineage-tracker && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-lineage-tracker" agent skill from https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction/tree/main/2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker into .opencode/skills/data-lineage-tracker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-lineage-tracker", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-lineage-trackerTrack data origin, transformations, and flow through construction systems.
Data Lineage Tracker is an agent skill from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction. Track data origin, transformations, and flow through construction systems. Essential for audit trails, compliance, and debugging data issues.
Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `claw.json` and `instructions.md`).
It sits in Data & Analytics, covering Data governance and Data cleaning. The repository describes itself as: 221 AI skills for construction: BIM analysis, cost estimation, scheduling, document control, and automation with Claude Code. The licence is MIT.
Read from SKILL.md and the folder at commit ce45bbf. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Lineage Tracker loads about 4.4k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 84 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction at commit ce45bbf, republished under its MIT licence (© datadrivenconstruction). 84 words, ~4,433 tokens.
.claude/skills/data-lineage-tracker/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Track the origin, transformations, and flow of construction data through systems. Provides audit trails for compliance, helps debug data issues, and ensures data governance.
Construction projects require data accountability:
from dataclasses import dataclass, field
from typing import List, Dict, Any, Optional, Set
from datetime import datetime
from enum import Enum
import json
import hashlib
import uuid
class TransformationType(Enum):
EXTRACT = "extract"
TRANSFORM = "transform"
LOAD = "load"
AGGREGATE = "aggregate"
JOIN = "join"
FILTER = "filter"
CALCULATE = "calculate"
MANUAL_EDIT = "manual_edit"
IMPORT = "import"
EXPORT = "export"
@dataclass
class DataSource:
id: str
name: str
system: str
location: str
owner: str
created_at: datetime
@dataclass
class TransformationStep:
id: str
transformation_type: TransformationType
description: str
input_entities: List[str]
output_entities: List[str]
logic: str # SQL, Python, or description
performed_by: str # user or system
performed_at: datetime
parameters: Dict[str, Any] = field(default_factory=dict)
@dataclass
class DataEntity:
id: str
name: str
source_id: str
entity_type: str # table, file, field, record
created_at: datetime
version: int = 1
checksum: Optional[str] = None
parent_entities: List[str] = field(default_factory=list)
metadata: Dict[str, Any] = field(default_factory=dict)
@dataclass
class LineageRecord:
id: str
entity_id: str
transformation_id: str
upstream_entities: List[str]
downstream_entities: List[str]
recorded_at: datetime
class ConstructionDataLineageTracker:
"""Track data lineage for construction data flows."""
def __init__(self, project_id: str):
self.project_id = project_id
self.sources: Dict[str, DataSource] = {}
self.entities: Dict[str, DataEntity] = {}
self.transformations: Dict[str, TransformationStep] = {}
self.lineage_records: List[LineageRecord] = []
def register_source(self, name: str, system: str, location: str, owner: str) -> DataSource:
"""Register a new data source."""
source = DataSource(
id=f"SRC-{uuid.uuid4().hex[:8]}",
name=name,
system=system,
location=location,
owner=owner,
created_at=datetime.now()
)
self.sources[source.id] = source
return source
def register_entity(self, name: str, source_id: str, entity_type: str,
parent_entities: List[str] = None,
metadata: Dict = None) -> DataEntity:
"""Register a data entity (table, file, field)."""
entity = DataEntity(
id=f"ENT-{uuid.uuid4().hex[:8]}",
name=name,
source_id=source_id,
entity_type=entity_type,
created_at=datetime.now(),
parent_entities=parent_entities or [],
metadata=metadata or {}
)
self.entities[entity.id] = entity
return entity
def calculate_checksum(self, data: Any) -> str:
"""Calculate checksum for data verification."""
if isinstance(data, str):
content = data
else:
content = json.dumps(data, sort_keys=True, default=str)
return hashlib.sha256(content.encode()).hexdigest()[:16]
def record_transformation(self,
transformation_type: TransformationType,
description: str,
input_entities: List[str],
output_entities: List[str],
logic: str,
performed_by: str,
parameters: Dict = None) -> TransformationStep:
"""Record a data transformation."""
transformation = TransformationStep(
id=f"TRF-{uuid.uuid4().hex[:8]}",
transformation_type=transformation_type,
description=description,
input_entities=input_entities,
output_entities=output_entities,
logic=logic,
performed_by=performed_by,
performed_at=datetime.now(),
parameters=parameters or {}
)
self.transformations[transformation.id] = transformation
# Create lineage records
for output_id in output_entities:
record = LineageRecord(
id=f"LIN-{uuid.uuid4().hex[:8]}",
entity_id=output_id,
transformation_id=transformation.id,
upstream_entities=input_entities,
downstream_entities=[],
recorded_at=datetime.now()
)
self.lineage_records.append(record)
# Update downstream references for input entities
for input_id in input_entities:
for existing_record in self.lineage_records:
if existing_record.entity_id == input_id:
existing_record.downstream_entities.append(output_id)
return transformation
def trace_upstream(self, entity_id: str, depth: int = None) -> List[Dict]:
"""Trace all upstream sources of an entity."""
visited = set()
lineage = []
def trace(eid: str, current_depth: int):
if eid in visited:
return
if depth is not None and current_depth > depth:
return
visited.add(eid)
entity = self.entities.get(eid)
if not entity:
return
# Find transformations that produced this entity
for record in self.lineage_records:
if record.entity_id == eid:
transformation = self.transformations.get(record.transformation_id)
if transformation:
lineage.append({
'entity': entity.name,
'entity_id': eid,
'depth': current_depth,
'transformation': transformation.description,
'transformation_type': transformation.transformation_type.value,
'performed_at': transformation.performed_at.isoformat(),
'performed_by': transformation.performed_by,
'upstream': record.upstream_entities
})
for upstream_id in record.upstream_entities:
trace(upstream_id, current_depth + 1)
trace(entity_id, 0)
return sorted(lineage, key=lambda x: x['depth'])
def trace_downstream(self, entity_id: str, depth: int = None) -> List[Dict]:
"""Trace all downstream dependencies of an entity."""
visited = set()
dependencies = []
def trace(eid: str, current_depth: int):
if eid in visited:
return
if depth is not None and current_depth > depth:
return
visited.add(eid)
entity = self.entities.get(eid)
if not entity:
return
# Find entities that use this entity
for record in self.lineage_records:
if eid in record.upstream_entities:
transformation = self.transformations.get(record.transformation_id)
if transformation:
dependencies.append({
'entity': self.entities[record.entity_id].name if record.entity_id in self.entities else record.entity_id,
'entity_id': record.entity_id,
'depth': current_depth,
'transformation': transformation.description,
'transformation_type': transformation.transformation_type.value
})
trace(record.entity_id, current_depth + 1)
trace(entity_id, 0)
return sorted(dependencies, key=lambda x: x['depth'])
def get_entity_history(self, entity_id: str) -> List[Dict]:
"""Get complete history of changes to an entity."""
history = []
for record in self.lineage_records:
if record.entity_id == entity_id:
transformation = self.transformations.get(record.transformation_id)
if transformation:
history.append({
'timestamp': transformation.performed_at.isoformat(),
'action': transformation.transformation_type.value,
'description': transformation.description,
'performed_by': transformation.performed_by,
'inputs': [
self.entities[eid].name if eid in self.entities else eid
for eid in record.upstream_entities
]
})
return sorted(history, key=lambda x: x['timestamp'])
def impact_analysis(self, entity_id: str) -> Dict:
"""Analyze impact of changes to an entity."""
downstream = self.trace_downstream(entity_id)
impact = {
'entity': self.entities[entity_id].name if entity_id in self.entities else entity_id,
'total_affected': len(downstream),
'affected_by_depth': {},
'affected_entities': downstream
}
for dep in downstream:
depth = dep['depth']
impact['affected_by_depth'][depth] = impact['affected_by_depth'].get(depth, 0) + 1
return impact
def validate_lineage(self) -> List[str]:
"""Validate lineage for completeness and consistency."""
issues = []
# Check for orphan entities (no source or transformation)
for eid, entity in self.entities.items():
has_lineage = any(r.entity_id == eid for r in self.lineage_records)
if not has_lineage and entity.entity_type != 'source':
issues.append(f"Entity '{entity.name}' has no lineage record")
# Check for broken references
all_entity_ids = set(self.entities.keys())
for record in self.lineage_records:
for upstream_id in record.upstream_entities:
if upstream_id not in all_entity_ids:
issues.append(f"Lineage references unknown entity: {upstream_id}")
# Check for circular dependencies
for eid in self.entities:
upstream = set()
to_check = [eid]
while to_check:
current = to_check.pop()
if current in upstream:
issues.append(f"Circular dependency detected involving entity: {self.entities[eid].name}")
break
upstream.add(current)
for record in self.lineage_records:
if record.entity_id == current:
to_check.extend(record.upstream_entities)
return issues
def generate_lineage_graph(self, entity_id: str) -> str:
"""Generate Mermaid diagram of lineage."""
lines = ["```mermaid", "graph LR"]
upstream = self.trace_upstream(entity_id, depth=5)
downstream = self.trace_downstream(entity_id, depth=5)
# Add nodes
added_nodes = set()
for item in upstream + downstream:
node_id = item['entity_id'].replace('-', '_')
if node_id not in added_nodes:
entity = self.entities.get(item['entity_id'])
name = entity.name if entity else item['entity_id']
lines.append(f" {node_id}[{name}]")
added_nodes.add(node_id)
# Add target node
target_node = entity_id.replace('-', '_')
if target_node not in added_nodes:
entity = self.entities.get(entity_id)
name = entity.name if entity else entity_id
lines.append(f" {target_node}[{name}]:::target")
# Add edges
for item in upstream:
for upstream_id in item.get('upstream', []):
from_node = upstream_id.replace('-', '_')
to_node = item['entity_id'].replace('-', '_')
lines.append(f" {from_node} --> {to_node}")
for item in downstream:
from_node = entity_id.replace('-', '_')
to_node = item['entity_id'].replace('-', '_')
if to_node != from_node:
lines.append(f" {from_node} --> {to_node}")
lines.append(" classDef target fill:#f96")
lines.append("```")
return "\n".join(lines)
def export_lineage(self) -> Dict:
"""Export complete lineage data."""
return {
'project_id': self.project_id,
'exported_at': datetime.now().isoformat(),
'sources': {k: {
'id': v.id,
'name': v.name,
'system': v.system,
'location': v.location,
'owner': v.owner
} for k, v in self.sources.items()},
'entities': {k: {
'id': v.id,
'name': v.name,
'source_id': v.source_id,
'entity_type': v.entity_type,
'parent_entities': v.parent_entities
} for k, v in self.entities.items()},
'transformations': {k: {
'id': v.id,
'type': v.transformation_type.value,
'description': v.description,
'input_entities': v.input_entities,
'output_entities': v.output_entities,
'performed_by': v.performed_by,
'performed_at': v.performed_at.isoformat()
} for k, v in self.transformations.items()},
'lineage_records': [{
'id': r.id,
'entity_id': r.entity_id,
'transformation_id': r.transformation_id,
'upstream_entities': r.upstream_entities
} for r in self.lineage_records]
}
def generate_report(self) -> str:
"""Generate lineage report."""
lines = [f"# Data Lineage Report: {self.project_id}", ""]
lines.append(f"**Generated:** {datetime.now().strftime('%Y-%m-%d %H:%M')}")
lines.append(f"**Sources:** {len(self.sources)}")
lines.append(f"**Entities:** {len(self.entities)}")
lines.append(f"**Transformations:** {len(self.transformations)}")
lines.append("")
# Sources
lines.append("## Data Sources")
for source in self.sources.values():
lines.append(f"- **{source.name}** ({source.system})")
lines.append(f" - Location: {source.location}")
lines.append(f" - Owner: {source.owner}")
lines.append("")
# Validation
issues = self.validate_lineage()
if issues:
lines.append("## Lineage Issues")
for issue in issues:
lines.append(f"- ⚠️ {issue}")
lines.append("")
# Transformation summary
lines.append("## Transformation Summary")
type_counts = {}
for t in self.transformations.values():
type_counts[t.transformation_type.value] = type_counts.get(t.transformation_type.value, 0) + 1
for t_type, count in sorted(type_counts.items()):
lines.append(f"- {t_type}: {count}")
return "\n".join(lines)# Initialize tracker
tracker = ConstructionDataLineageTracker("PROJECT-001")
# Register sources
procore = tracker.register_source("Procore", "SaaS", "cloud", "PM Team")
sage = tracker.register_source("Sage 300", "Database", "on-prem", "Finance")
# Register entities
budget = tracker.register_entity("Project Budget", procore.id, "table")
costs = tracker.register_entity("Job Costs", sage.id, "table")
report = tracker.register_entity("Cost Variance Report", procore.id, "file")
# Record transformation
tracker.record_transformation(
transformation_type=TransformationType.JOIN,
description="Join budget and actual costs for variance calculation",
input_entities=[budget.id, costs.id],
output_entities=[report.id],
logic="SELECT b.*, c.actual, (b.budget - c.actual) as variance FROM budget b JOIN costs c ON b.cost_code = c.cost_code",
performed_by="ETL Pipeline"
)
# Trace lineage
upstream = tracker.trace_upstream(report.id)
print("Upstream lineage:", upstream)
# Generate graph
print(tracker.generate_lineage_graph(report.id))
# Export for audit
lineage_data = tracker.export_lineage()© datadrivenconstruction, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in 2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker of datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction.
Open the folder on GitHubat commit ce45bbf
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction, which our catalogue first saw on October 7, 2026.
Data Lineage Tracker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Lineage Tracker this skilldatadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction | 344 | 1 repos | ~4.4k | Automated safety check: Pass | MIT | |
| Data Quality Frameworkswshobson/agents | 40k | 11 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Research Data Feasibility and Leakage ChecksLight0305/Light-skills | 641 | — | ~4.9k | Automated safety check: Pass | MIT | |
| Datalineage Summarygoogle/skills | 21k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Monte Carlo Context Detectionsickn33/agentic-awesome-skills | 47k | 1 repos | ~2.6k | Automated safety check: Warn | MIT | |
| Querying AWS Sagemaker Catalogaws/agent-toolkit-for-aws | 2.8k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 |
wshobson/agents
Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.
Light0305/Light-skills
Finds usable public datasets, judges whether the data can support a research idea, and checks train and test splits for leakage before results are trusted.
google/skills
Summarizes Google Cloud Data Lineage graphs to help users debug data quality issues and understand data provenance for BQ/GCS.
sickn33/agentic-awesome-skills
Route data-related requests to the right Monte Carlo skill or workflow.
aws/agent-toolkit-for-aws
Runs SQL analytics on SageMaker Catalog asset metadata tables exported as Apache Iceberg in S3 Tables.
cbrock84/headcount
Establishes ownership, definitions, quality, access, and lineage for the organization's data.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Orchestrate multiple AI agents for construction workflows: estimator, scheduler, document, QA and safety agents coordinated by a supervisor agent, with human checkpoints.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Estimate embodied carbon and produce ESG/climate reporting for construction: LCA per work item, material-based carbon factors, EU taxonomy and CSRD alignment.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Material passports and circular construction: generate per-element material inventories from BOQ/BIM, mark reuse potential and recycled content, and prepare deconstruction data.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Browse and search the OpenConstructionERP cost database: classification tree, SQL and semantic search, autocomplete, certainty badges, and the resource catalog.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Create bills of quantities and estimates in OpenConstructionERP: search cost items, build BOQ sections, link BIM elements in bulk, validate the BOQ, and export GAEB/XLSX/JSON.
datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction
Field operations in OpenConstructionERP: punch list, daily diary, HSE observations and task tracking on site.
Categories
Track data origin, transformations, and flow through construction systems. Data Lineage Tracker is an agent skill from datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction. Track data origin, transformations, and flow through construction systems.
Data Lineage Tracker fits situations like: tasks that involve Data governance; tasks that involve Data cleaning.
Run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-lineage-tracker -a claude-code`. Or copy the skill folder (2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction) into .claude/skills/data-lineage-tracker in your project. Claude Code loads it when a task matches its description.
Run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-lineage-tracker -a codex`. Or copy the skill folder (2_DDC_Book/2.6-Data-Quality-Validation/data-lineage-tracker in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction) into .agents/skills/data-lineage-tracker in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill data-lineage-tracker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-lineage-tracker, .gemini/skills/data-lineage-tracker, .github/skills/data-lineage-tracker and .opencode/skills/data-lineage-tracker in your project.
SKILL.md names no scripts, command-line tools or credentials: Data Lineage Tracker is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Data Lineage Tracker is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Data Lineage Tracker: Data Quality Frameworks (wshobson/agents, 40k stars), Research Data Feasibility and Leakage Checks (Light0305/Light-skills, 641 stars), Datalineage Summary (google/skills, 21k stars) and Monte Carlo Context Detection (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
datadrivenconstruction (a GitHub user) maintains it in datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction, which has 344 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on August 22, 2026.
Source: datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.