Data Lakehouse Architect
FerroxLabs/wayland
Architecture expertise for data lakehouse platforms covering Delta Lake, Apache Iceberg, Apache Hudi, medallion architecture design, table format selection, storage optimization, schema evolution…
Strategic guidance for designing modern data platforms, covering storage paradigms (data lake, warehouse, lakehouse), modeling approaches (dimensional, normalized, data vault, wide tables), data…
$ npx skills add ancoleman/ai-design-components --skill architecting-data -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ancoleman/ai-design-components architecting-data --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/architecting-data .claude/skills/architecting-data && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "architecting-data" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/architecting-data into .claude/skills/architecting-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "architecting-data", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ancoleman/ai-design-components/tree/main/skills/architecting-dataType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ancoleman/ai-design-components --skill architecting-data -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ancoleman/ai-design-components architecting-data --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/architecting-data .agents/skills/architecting-data && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "architecting-data" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/architecting-data into .agents/skills/architecting-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "architecting-data", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ancoleman/ai-design-components --skill architecting-data -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ancoleman/ai-design-components architecting-data --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/architecting-data .cursor/skills/architecting-data && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "architecting-data" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/architecting-data into .cursor/skills/architecting-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "architecting-data", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ancoleman/ai-design-components.git --path skills/architecting-data--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ancoleman/ai-design-components --skill architecting-data -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ancoleman/ai-design-components architecting-data --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/architecting-data .gemini/skills/architecting-data && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "architecting-data" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/architecting-data into .gemini/skills/architecting-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "architecting-data", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ancoleman/ai-design-components architecting-dataInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ancoleman/ai-design-components --skill architecting-data -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/architecting-data .github/skills/architecting-data && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "architecting-data" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/architecting-data into .github/skills/architecting-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "architecting-data", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ancoleman/ai-design-components --skill architecting-data -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ancoleman/ai-design-components architecting-data --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/architecting-data .opencode/skills/architecting-data && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "architecting-data" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/architecting-data into .opencode/skills/architecting-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "architecting-data", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
architecting-dataStrategic guidance for designing modern data platforms, covering storage paradigms (data lake, warehouse, lakehouse), modeling approaches (dimensional, normalized, data vault, wide tables), data…
Architecting Data is an agent skill from ancoleman/ai-design-components. Strategic guidance for designing modern data platforms, covering storage paradigms (data lake, warehouse, lakehouse), modeling approaches (dimensional, normalized, data vault, wide tables), data mesh principles, and medallion architecture patterns. Use when architecting data platforms, choosing between centralized vs decentralized patterns, selecting table formats (Iceberg, Delta Lake), or designing data governance frameworks.
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 16 other files, including reference files (for example `examples/dbt-project/README.md`, `outputs.yaml` and `references/data-mesh-guide.md`).
It sits in Databases, covering Data warehousing, Data governance and Software architecture. It works with Databricks. The repository describes itself as: Comprehensive UI/UX and Backend component design skills for AI-assisted development with Claude. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 76551b7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are sql).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
iceberg.apache.orgdocs.getdbt.comdatamesh-architecture.comdatabricks.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Architecting Data loads about 3.6k tokens when it runs, and up to ~23k if it reads all its reference files. Until then it costs about 112 tokens; SKILL.md has 1,276 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ancoleman/ai-design-components at commit 76551b7, republished under its MIT licence (© ancoleman). 1,276 words, ~3,588 tokens.
.claude/skills/architecting-data/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.Guide architects and platform engineers through strategic data architecture decisions for modern cloud-native data platforms.
Invoke this skill when:
Three primary patterns for analytical data storage:
Data Lake: Centralized repository for raw data at scale
Data Warehouse: Structured repository optimized for BI
Data Lakehouse: Hybrid combining lake flexibility with warehouse reliability
Decision Framework:
For detailed comparison, see references/storage-paradigms.md.
Four primary modeling patterns:
Dimensional (Kimball): Star/snowflake schemas for BI
Normalized (3NF): Eliminate redundancy for transactional systems
Data Vault 2.0: Flexible model with complete audit trail
Wide Tables: Denormalized, optimized for columnar storage
Decision Framework:
For detailed patterns, see references/modeling-approaches.md.
Decentralized architecture for large organizations (>500 people).
Four Core Principles:
Readiness Assessment (Score 1-5 each):
Scoring: 24-30: Strong candidate | 18-23: Hybrid | 12-17: Build foundation first | 6-11: Centralized
Red Flags: Small org (<100 people), unclear domains, no platform team, weak governance
For full guide, see references/data-mesh-guide.md.
Standard lakehouse pattern: Bronze (raw) → Silver (cleaned) → Gold (business-level)
Bronze Layer: Exact copy of source data, immutable, append-only
Silver Layer: Validated, deduplicated, typed data
Gold Layer: Business logic, aggregates, dimensional models, ML features
Data Quality by Layer:
For patterns, see references/medallion-pattern.md.
Enable ACID transactions on data lakes:
Apache Iceberg: Multi-engine, vendor-neutral (Context7: 79.7 score)
Delta Lake: Databricks ecosystem, Spark-optimized
Apache Hudi: Optimized for CDC and frequent upserts
Recommendation: Apache Iceberg for new projects (vendor-neutral, broadest support)
For comparison, see references/table-formats.md.
Standard Layers:
Tool Selection:
For detailed recommendations, see references/tool-recommendations.md and references/modern-data-stack.md.
Data Catalog: Searchable inventory (DataHub, Alation, Collibra)
Data Lineage: Track data flow (OpenLineage, Marquez)
Data Quality: Validation and testing (Great Expectations, Soda, dbt tests)
Access Control:
For governance patterns, see references/governance-patterns.md.
Step 1: Identify Primary Use Case
Step 2: Evaluate Budget
Recommendation by Org Size:
See references/decision-frameworks.md.
Decision Tree:
See references/decision-frameworks.md.
Use 6-factor assessment. Score interpretation:
See references/decision-frameworks.md.
Decision Tree:
Recommendation: Apache Iceberg for new projects
Context: 50-person startup, PostgreSQL + MongoDB + Stripe
Recommendation:
Context: Legacy Oracle warehouse, need cloud migration
Recommendation:
Context: 200-person company, 5-person central data team
Recommendation: NOT YET. Build foundation first.
dbt: Score 87.0, 3,532+ code snippets
Apache Iceberg: Score 79.7, 832+ code snippets
Tool Stack by Use Case:
Startup: BigQuery + Airbyte + dbt + Metabase (<$1K/month)
Growth: Snowflake + Fivetran + dbt + Airflow + Tableau ($10K-50K/month)
Enterprise: Snowflake + Databricks + Fivetran + Kafka + dbt + Airflow + Alation ($50K-500K/month)
See references/tool-recommendations.md.
-- Bronze: Raw ingestion
CREATE TABLE bronze.raw_customers (_ingested_at TIMESTAMP, _raw_data STRING);
-- Silver: Cleaned
CREATE TABLE silver.customers AS
SELECT json_extract(_raw_data, '$.id') AS customer_id, ...
FROM bronze.raw_customers
QUALIFY ROW_NUMBER() OVER (PARTITION BY customer_id ORDER BY _ingested_at DESC) = 1;
-- Gold: Business-level
CREATE TABLE gold.fact_sales AS
SELECT s.order_id, d.date_key, c.customer_key, ...
FROM silver.sales s
JOIN gold.dim_date d ON s.order_date = d.date;CREATE TABLE catalog.db.sales (order_id BIGINT, amount DECIMAL(10,2))
USING iceberg
PARTITIONED BY (days(order_date));
-- Time travel
SELECT * FROM catalog.db.sales TIMESTAMP AS OF '2025-01-01';-- models/staging/stg_customers.sql
WITH source AS (SELECT * FROM {{ source('raw', 'customers') }}),
cleaned AS (
SELECT customer_id, UPPER(customer_name) AS customer_name
FROM source WHERE customer_id IS NOT NULL
)
SELECT * FROM cleanedFor complete examples, see examples/.
Direct Dependencies:
Complementary:
Downstream:
Common Workflows:
End-to-End Analytics:
data-architecture (warehouse) → ingesting-data (Fivetran) →
data-transformation (dbt) → visualizing-data (Tableau)Data Platform for AI/ML:
data-architecture (lakehouse) → ingesting-data (Kafka) →
data-transformation (dbt features) → ai-data-engineering (feature store)Reference Files:
Examples:
External Resources:
© ancoleman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 13 other files (references) in skills/architecting-data of ancoleman/ai-design-components.
Open the folder on GitHubat commit 76551b7
Architecting Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Architecting Data this skillancoleman/ai-design-components | 526 | — | ~3.6k | Automated safety check: Pass | MIT | |
| Data Lakehouse ArchitectFerroxLabs/wayland | 608 | — | ~3k | Automated safety check: Pass | Apache-2.0 | |
| Databricks Dbsqldatabricks/databricks-agent-skills | 345 | 1 repos | ~2.8k | Automated safety check: Pass | Custom licence | |
| Altimate Data Warehouse DelegateAltimateAI/data-engineering-skills | 127 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Rocky New Adapterrocky-data/rocky | 304 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| SQL Queriesw95/awesome-claude-corporate-skills | 235 | 3 repos | ~2.8k | Automated safety check: Pass | MIT |
FerroxLabs/wayland
Architecture expertise for data lakehouse platforms covering Delta Lake, Apache Iceberg, Apache Hudi, medallion architecture design, table format selection, storage optimization, schema evolution…
databricks/databricks-agent-skills
Databricks SQL (DBSQL) advanced features and SQL warehouse capabilities.
AltimateAI/data-engineering-skills
Delegates dbt and warehouse tasks such as lineage, migrations and cost attribution to the altimate-code CLI agent and relays its answer back.
rocky-data/rocky
Adding a new warehouse or source adapter crate to the Rocky engine.
w95/awesome-claude-corporate-skills
Write correct, performant SQL across all major data warehouse dialects (Snowflake, BigQuery, Databricks, PostgreSQL, etc.).
databricks/databricks-agent-skills
Apache Iceberg tables on Databricks — Managed Iceberg tables, External Iceberg Reads (fka Uniform), Compatibility Mode, Iceberg REST Catalog (IRC), Iceberg v3, Snowflake interop, PyIceberg, OSS…
ancoleman/ai-design-components
Builds AI chat interfaces and conversational UI with streaming responses, context management, and multi-modal support.
ancoleman/ai-design-components
Builds form components and data collection interfaces including contact forms, registration flows, checkout processes, surveys, and settings pages.
ancoleman/ai-design-components
Builds tables and data grids for displaying tabular information, from simple HTML tables to complex enterprise data grids.
ancoleman/ai-design-components
Creates comprehensive dashboard and analytics interfaces that combine data visualization, KPI cards, real-time updates, and interactive layouts.
ancoleman/ai-design-components
Designs layout systems and responsive interfaces including grid systems, flexbox patterns, sidebar layouts, and responsive breakpoints.
ancoleman/ai-design-components
Displays chronological events and activity through timelines, activity feeds, Gantt charts, and calendar interfaces.
Works with
Categories
Strategic guidance for designing modern data platforms, covering storage paradigms (data lake, warehouse, lakehouse), modeling approaches (dimensional, normalized, data vault, wide tables), data…. Architecting Data is an agent skill from ancoleman/ai-design-components. Strategic guidance for designing modern data platforms, covering storage paradigms (data lake, warehouse, lakehouse), modeling approaches (dimensional, normalized, data vault, wide tables), data mesh principles, and medallion architecture patterns.
Architecting Data fits situations like: architecting data platforms; choosing between centralized vs decentralized patterns; selecting table formats (Iceberg; designing data governance frameworks.
Run `npx skills add ancoleman/ai-design-components --skill architecting-data -a claude-code`. Or copy the skill folder (skills/architecting-data in ancoleman/ai-design-components) into .claude/skills/architecting-data in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ancoleman/ai-design-components --skill architecting-data -a codex`. Or copy the skill folder (skills/architecting-data in ancoleman/ai-design-components) into .agents/skills/architecting-data in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ancoleman/ai-design-components --skill architecting-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/architecting-data, .gemini/skills/architecting-data, .github/skills/architecting-data and .opencode/skills/architecting-data in your project.
SKILL.md names no scripts, command-line tools or credentials: Architecting Data is instructions for the agent only.
SKILL.md names 4 domains. As links in the text: iceberg.apache.org, docs.getdbt.com, datamesh-architecture.com and databricks.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Architecting Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Architecting Data: Data Lakehouse Architect (FerroxLabs/wayland, 608 stars), Databricks Dbsql (databricks/databricks-agent-skills, 345 stars), Altimate Data Warehouse Delegate (AltimateAI/data-engineering-skills, 127 stars) and Rocky New Adapter (rocky-data/rocky, 304 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ancoleman (a GitHub user) maintains it in ancoleman/ai-design-components, which has 526 GitHub stars. The repository holds 75 skills in this directory. The repository was last updated on December 11, 2025.
Source: ancoleman/ai-design-components on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.